Legal Analysis: Does US Law Permit AI to be Trained Upon Australia Without First Seeking Permission as OpenAI Claims?

OpenAI is half right, and the half that is right is the part that feels wrong to us Aussies. US law can let a company copy Australian work inside the United States without asking, even when Australian law would not. That is not a loophole invented for chatbots. It is how copyright has worked for a century. It is also not the blank cheque the hearing made it sound like.

At the Joint Select Committee this week, Google and OpenAI told senators they train in the United States on material that is freely available online, lawfully accessed, and not opted out of. Google's Roxanne Carter said that under American fair use "we will train on that content," and that seeking permission first would take something like 200 years. OpenAI's Adam Cohen said fair use meant there was no obligation to negotiate over publicly available data. Microsoft added that Australian law does not give developers the certainty they want. Senator David Pocock's question was the obvious one: why should every Australian author have to attach a notice to stop a US company using their work? The instinct behind that question is sound. The legal answer is less satisfying.

Copyright is territorial. The Berne Convention gives an Australian book national treatment in the United States, meaning it is protected there as if it were American. It does not export the Australian Copyright Act to California. If the copying that builds the training set happens on US servers, the law that decides whether that copying is excused is US law. Australia can forbid the same act in Sydney. It cannot, by itself, forbid it in a US data centre. That is why the companies keep saying they train in America. They are choosing the forum as much as they are describing the engineering.

"Publicly available" is doing too much work in their answers. A novel posted by its publisher, a news article, a song lyric on a lyric site: available is not the same as free of copyright. Public domain is a different category. Most of what a crawler finds is owned by someone. Fair use is the claim that the copying is nevertheless permitted.

Fair use, in section 107 of the US Copyright Act, is not a right to take whatever is useful. It is a defence, weighed case by case on four factors: the purpose and character of the use, the nature of the work, the amount taken, and the effect on the market. The industry's best precedent is still Authors Guild v. Google, where the Second Circuit held that scanning millions of books to make them searchable was transformative and fair, because the use was not a substitute for reading the books. The industry's worst recent precedent is Thomson Reuters v. Ross Intelligence, where a Delaware court rejected fair use for an AI tool trained on Westlaw headnotes to compete in the same market. Other training cases, including suits against OpenAI, Meta and Anthropic, are split, settled, or unfinished. Outputs that regurgitate a chapter are a separate infringement even if the training copy were excused. So, when Cohen says fair use means there is no obligation to negotiate, he is stating the company's litigation position, not a settled rule. A court may agree. A court may not. "Enables us" is a prediction.

That is the counter-intuitive core. American copyright is mainly an incentive system, not a natural right to control every use. Fair use exists so that criticism, search, scholarship and, on the industry's argument, machine learning can proceed without a licence when the use does not replace the original. Australian law is closer to the permission model Pocock was reaching for. It has fair dealing for named purposes, not an open fairness defence, and it has moral rights. The Albanese government has already said no company should use Australian books, music, art or news to train a model without the artist's control over price and value. Labor has since floated a narrower deal: training allowed if creators can opt out, or if the firms strike enough local content agreements. OpenAI and Anthropic have answered by asking for a clear local permission to train, tied to data-centre investment. The two systems are not confused. They are in conflict, and the companies have noticed which one they can litigate under today.

The practical holes in the US-law answer are larger than the hearing suggested. First, opt-out is not fair use. Google-Extended and OpenAI's bot controls are product choices. If training is fair, a missing robots.txt does not create liability. If training is not fair, an opt-out tool does not cure the copies already made. Putting the burden on every Australian to find the right setting reverses the ordinary rule that the user of a work seeks the licence. It may be the only scalable system. It is not what the statute requires.

Second, "lawfully accessed" hides a chain of copies. A crawler downloads. A pipeline deduplicates, filters and stores. A training run copies again. Some of those copies may be fair and others not. Memorisation is the awkward fact: if a model can emit a substantial part of an Australian article on request, the market-effect factor gets much harder for the company.

Third, territoriality cuts both ways. Training in the US does not immunise a product served in Australia, a download onto an Australian server, or a local fine-tune. Choice of law is messy once the act and the harm sit in different countries. Australia cannot rewrite section 107. It can regulate the service offered to Australians, the data centres it is being asked to host, and the imports of a model trained on local work. That is the bargain now on the table, and it is why the companies are in the room at all. If US fair use already covered everything they want to do, they would not be asking Canberra for a regime that "legally permits AI training, and does so very clearly."

Fourth, the tax exchange Pocock raised is a separate argument and a weaker one. Paying the tax that is due is not a copyright licence. A data-centre investment is not a royalty. Both can be reasons for a government to deal. Neither converts an unlicensed copy into a licensed one.

So, is OpenAI right? On the narrow claim, yes. Copying done in the United States is judged by US law, Australian authorship does not change that, and fair use is a real defence that does not require permission in advance. On the claim as it landed in the committee, no. Fair use is not a standing permission to train on the Australian internet. It has not been finally won in the cases that matter. Opt-out is not the legal onus. And the request for an Australian training exception is an admission that the US theory does not, by itself, secure the right to do the same thing here.

The intuition that this is taking without asking is not a mistake about the facts. It is a disagreement about which tradition should govern. The US tradition says some large-scale copying is allowed because the new use is different and the market for the original survives. The Australian tradition, as the government has stated it, says the creator decides. OpenAI is betting that the first tradition applies to its servers, and that the second can be negotiated. Both bets are still open.

https://www.theepochtimes.com/world/google-openai-say-us-law-lets-them-train-ai-on-australian-content-without-seeking-permission-6100350