Scope
Who the guidance is for
The guidance is written for developers of generative AI who are subject to the Privacy Act, and it defines a developer broadly. In the OAIC’s words, a developer “includes any organisation who designs, builds, trains, adapts or combines AI models and applications”, including by fine-tuning, which it describes as modifying a trained model “with a smaller, targeted fine-tuning dataset to adapt it to suit more specialised use cases.” So a business that adapts an existing model for its own purposes is within it: the OAIC says the Privacy Act applies to APP entities “developing generative AI or fine-tuning commercially available models for their purposes.”
It also covers the other side of the deal: an organisation that hands personal information to a developer so a model can be built or fine-tuned. It does not deal with testing or deploying a model, which the OAIC treats separately.
APP 3 · Collection
Public is not the same as permitted
“Just because data is publicly available or otherwise accessible does not mean it can legally be used to train or fine-tune generative AI models or systems. Developers must consider whether data they intend to use or collect (including publicly available data) contains personal information, and comply with their privacy obligations.”
Three collection rules follow in the guidance. Developers may collect only personal information that is reasonably necessary for their functions or activities. They must collect it only by lawful and fair means, and depending on the circumstances, building a dataset by web scraping may be a covert and therefore unfair means. And sensitive information, the guidance says, “generally requires consent to be collected.”
The OAIC describes sensitive information this way: “Sensitive information is any biometric information to be used for the purposes of automated biometric verification or biometric identification, biometric templates, health information about an individual, genetic information about an individual or personal information about an individual for certain topics such as racial or ethnic origin, political opinions or sexual orientation.” It adds that many photographs or recordings of people, including generated ones, contain sensitive information, and that sensitive information collected without consent will generally need deleting from a dataset.
APP 6 · Use and disclosure
Data a business already holds
The next question is whether customer records collected for one reason can now train a model. The rule in the guidance is that personal information can be used or disclosed only for the primary purpose it was collected for, unless there is consent or an exception applies. One exception is reasonable expectation, which the guidance applies in two parts:
- the individual would reasonably expect the information to be used for the secondary purpose, judged with particular regard to their expectations at the time of collection; and
- the secondary purpose is related to the primary purpose, or directly related for sensitive information.
The OAIC warns that, given AI’s particular characteristics, the harms it can cause and the level of community concern, “in many cases it will be difficult to establish that such a secondary use was within reasonable expectations.” Changing the paperwork afterwards does not fix that: “updating a privacy policy or providing notice by themselves will generally not be sufficient to change reasonable expectations regarding the use of personal information that was previously collected for a different purpose.” Where a business cannot clearly establish both parts, the guidance says it should seek consent for that use and/or offer people a meaningful and informed way to opt out.
On notice, the guidance is also direct: “generally, having information in a privacy policy is not a means of meeting the obligation to notify individuals under APP 5.”
The OAIC’s example 3
A bank fine-tunes a chatbot
The guidance works through a case close to many businesses. A bank wants to give a developer its historic customer queries, which include customers’ names and addresses, to fine-tune a general model for a customer chatbot. The OAIC’s steps, in order:
- Ask whether the personal information is needed at all. The bank should first consider whether the dataset needs it, or whether the queries can be de-identified.
- If it cannot be removed, test APP 6. The queries were collected to answer them, so for this secondary purpose the bank must consider whether it has consent or an exception applies.
- Tell customers and offer a way out. In the example the bank informs customers, past and present, of the intended use and gives them a way to opt out.
- Update the privacy policy and collection notice to make clear that “personal information will be used to fine-tune a generative AI chatbot for internal use”.
- Agree protections with the developer, such as filters to stop personal information from the training data appearing in outputs.
APP 10 and APP 1
Accuracy, and privacy by design
The OAIC notes that information can be personal information whether or not it is true, which can include false information an AI system generates, such as hallucinations or deepfakes. Under APP 10, developers must take reasonable steps to make sure the personal information they collect, use and disclose is accurate.
Before any of that, the guidance asks developers to take a “privacy by design” approach when developing or fine-tuning, including a privacy impact assessment, and, where there is any doubt whether the Privacy Act applies to an AI activity, to assume it does. Where the training data comes from a supplier, the matching questions are in buying an AI product; for a bank or other licensee, the wider governance picture is in governance for licensees.
Back to the registerAll five guides, from accountability to licensees.