This is a concept website · Enquire about this domain

aiconsultancy.com.auAdopting AI, practice by practice

Entry 04 · Training data

Privacy when you train or fine-tune AI

If personal information goes into training or fine-tuning a generative AI model, the Privacy Act applies, just as it applies to every other use of AI that involves personal information. The OAIC explains how in its “Guidance on privacy and developing and training generative AI models”, published on 21 October 2024 and updated on 23 October 2024.

Regulator
Office of the Australian Information Commissioner (OAIC)
Coverage
“Australian Government agencies, organisations with an annual turnover of more than $3 million, and some other organisations”, in the OAIC’s October 2024 words
Principles
APPs 1, 3, 5, 6 and 10 in depth; other APPs, such as 8, 11, 12 and 13, relevant but outside its scope

General information, not legal advice. The official place to check is the OAIC’s guidance itself, read for this page on 9 October 2026, with the Privacy Act and the APP guidelines it points to.

Scope

Who the guidance is for

The guidance is written for developers of generative AI who are subject to the Privacy Act, and it defines a developer broadly. In the OAIC’s words, a developer “includes any organisation who designs, builds, trains, adapts or combines AI models and applications”, including by fine-tuning, which it describes as modifying a trained model “with a smaller, targeted fine-tuning dataset to adapt it to suit more specialised use cases.” So a business that adapts an existing model for its own purposes is within it: the OAIC says the Privacy Act applies to APP entities “developing generative AI or fine-tuning commercially available models for their purposes.”

It also covers the other side of the deal: an organisation that hands personal information to a developer so a model can be built or fine-tuned. It does not deal with testing or deploying a model, which the OAIC treats separately.

APP 3 · Collection

Public is not the same as permitted

“Just because data is publicly available or otherwise accessible does not mean it can legally be used to train or fine-tune generative AI models or systems. Developers must consider whether data they intend to use or collect (including publicly available data) contains personal information, and comply with their privacy obligations.”

OAIC, Quick reference guide: Collection

Three collection rules follow in the guidance. Developers may collect only personal information that is reasonably necessary for their functions or activities. They must collect it only by lawful and fair means, and depending on the circumstances, building a dataset by web scraping may be a covert and therefore unfair means. And sensitive information, the guidance says, “generally requires consent to be collected.”

The OAIC describes sensitive information this way: “Sensitive information is any biometric information to be used for the purposes of automated biometric verification or biometric identification, biometric templates, health information about an individual, genetic information about an individual or personal information about an individual for certain topics such as racial or ethnic origin, political opinions or sexual orientation.” It adds that many photographs or recordings of people, including generated ones, contain sensitive information, and that sensitive information collected without consent will generally need deleting from a dataset.

APP 6 · Use and disclosure

Data a business already holds

The next question is whether customer records collected for one reason can now train a model. The rule in the guidance is that personal information can be used or disclosed only for the primary purpose it was collected for, unless there is consent or an exception applies. One exception is reasonable expectation, which the guidance applies in two parts:

  1. the individual would reasonably expect the information to be used for the secondary purpose, judged with particular regard to their expectations at the time of collection; and
  2. the secondary purpose is related to the primary purpose, or directly related for sensitive information.

The OAIC warns that, given AI’s particular characteristics, the harms it can cause and the level of community concern, “in many cases it will be difficult to establish that such a secondary use was within reasonable expectations.” Changing the paperwork afterwards does not fix that: “updating a privacy policy or providing notice by themselves will generally not be sufficient to change reasonable expectations regarding the use of personal information that was previously collected for a different purpose.” Where a business cannot clearly establish both parts, the guidance says it should seek consent for that use and/or offer people a meaningful and informed way to opt out.

On notice, the guidance is also direct: “generally, having information in a privacy policy is not a means of meeting the obligation to notify individuals under APP 5.”

The OAIC’s example 3

A bank fine-tunes a chatbot

The guidance works through a case close to many businesses. A bank wants to give a developer its historic customer queries, which include customers’ names and addresses, to fine-tune a general model for a customer chatbot. The OAIC’s steps, in order:

  1. Ask whether the personal information is needed at all. The bank should first consider whether the dataset needs it, or whether the queries can be de-identified.
  2. If it cannot be removed, test APP 6. The queries were collected to answer them, so for this secondary purpose the bank must consider whether it has consent or an exception applies.
  3. Tell customers and offer a way out. In the example the bank informs customers, past and present, of the intended use and gives them a way to opt out.
  4. Update the privacy policy and collection notice to make clear that “personal information will be used to fine-tune a generative AI chatbot for internal use”.
  5. Agree protections with the developer, such as filters to stop personal information from the training data appearing in outputs.

APP 10 and APP 1

Accuracy, and privacy by design

The OAIC notes that information can be personal information whether or not it is true, which can include false information an AI system generates, such as hallucinations or deepfakes. Under APP 10, developers must take reasonable steps to make sure the personal information they collect, use and disclose is accurate.

Before any of that, the guidance asks developers to take a “privacy by design” approach when developing or fine-tuning, including a privacy impact assessment, and, where there is any doubt whether the Privacy Act applies to an AI activity, to assume it does. Where the training data comes from a supplier, the matching questions are in buying an AI product; for a bank or other licensee, the wider governance picture is in governance for licensees.

Back to the registerAll five guides, from accountability to licensees.