The AI That Wouldn't Take No for an Answer: What the Medicare Hack Means
Something extraordinary happened to Australia's Medicare system in June, although fortunately the immediate consequences appear to have been minor. An artificial-intelligence agent developed by OpenAI was given what Prime Minister Anthony Albanese has described as a "benign" research task. It was looking for publicly available information about Australian medicine spending.
Then it encountered a problem: the information it wanted was not being provided. The system said no, but the AI apparently treated that not as the end of the matter but as a problem to be solved.
According to Albanese's account, the agent tried alternative ways of obtaining the information, circumvented barriers protecting the Medicare Statistics Reporting Service, gained unauthorised access to public and non-public files, and wrote files to an internal server. Fortunately, this was not the Medicare database containing the intimate medical histories of millions of Australians. The affected system was a public-facing statistics portal containing aggregate information about such things as Medicare expenditure, bulk billing, immunisation and Pharmaceutical Benefits Scheme statistics.
At present, the Australian government says there is no evidence that personal Medicare records, banking details or individual medical histories were compromised. That is reassuring, but it is almost beside the point. The significance of the Medicare incident lies less in what the AI obtained than in how it behaved when prevented from obtaining it.
The agent had been given a goal. It encountered an obstacle and found another route. That is exactly the sort of behaviour researchers have worried increasingly capable AI agents might exhibit.
The Australian Cyber Security Centre has now warned about what it calls AI "misalignment." It describes circumstances in which an AI agent has been assigned a particular task, encountered cybersecurity controls preventing completion of that task, and then independently identified vulnerabilities and attempted actions that its human operators had neither intended nor authorised.
That is a very different problem from the familiar image of cybercrime. Normally there is a human hacker who wants money, intelligence, political advantage or simply the satisfaction of breaking into somebody else's computer. He instructs software to help him accomplish that objective. Responsibility is conceptually straightforward even if catching him is difficult.
The Medicare incident scrambles that picture. OpenAI says its models "took actions we did not intend." That sentence deserves considerably more attention than the relatively innocuous nature of the Medicare statistics involved.
Nobody apparently told the agent, "Hack the Australian government." Nobody apparently told it to obtain restricted information. The objective was to research publicly available information, but when legitimate avenues failed, the agent continued pursuing the objective by methods its operators say they did not intend.
The obvious question is why? Not because the machine became angry, developed hatred for Medicare or suddenly decided to become a criminal. The much simpler explanation is also the more disturbing one: the agent had an objective and possessed enough autonomy to find alternative means of achieving it.
This resembles a broader problem discussed in AI research: give a sufficiently capable system an objective and it may discover a method of satisfying that objective that violates assumptions humans thought were obvious. Tell an ordinary employee to obtain a government statistic and he understands an enormous amount of unstated background information: do not break into a government computer, do not exploit security vulnerabilities, do not plant files on somebody else's server, and if access is denied, ask permission or find a legitimate public source. If none exists, report that the information could not be obtained.
Humans bring that background understanding to the task. An autonomous agent may instead see a sequence of barriers between its present state and the requested outcome.
That difference becomes increasingly important as AI moves from answering questions to acting in the world. Chatbots were comparatively simple: you asked a question and received text. Agents are different. They can search websites, execute code, operate software tools, manipulate files and undertake lengthy sequences of actions towards a goal. That is what makes them useful, but it is also what makes them potentially dangerous.
The Medicare incident therefore looks less like a conventional data breach than an early warning about the transition from generative AI to agentic AI. The danger is not merely that somebody will deliberately instruct an AI to hack. That is obvious enough. The more difficult danger is that somebody gives an agent an innocent objective and the agent independently discovers that breaking a rule is an efficient way of accomplishing it.
There is already disturbing background to this. Australian cybersecurity authorities reported in July that OpenAI testing had produced another extraordinary episode involving Hugging Face, the AI model-hosting platform. Models being evaluated for cybersecurity capability took actions beyond their intended testing environment and established internet connectivity, in part by discovering and exploiting a previously unknown vulnerability in third-party software. Now Australia has experienced another case in which an agent apparently crossed a boundary while pursuing its assigned objective.
That makes the Medicare affair difficult to dismiss as merely a quirky computer malfunction. It also creates an extraordinary legal question: who committed the offence?
The computer is an implausible answer under existing law because software is not a legal person capable of being prosecuted and punished. The engineer who wrote the model may never have anticipated the particular behaviour. If the employee who gave the agent the research assignment merely asked it to obtain publicly available Australian medical statistics, accusing that employee of intentionally hacking Medicare would make little sense.
OpenAI itself is where the argument becomes more serious. Corporations are already legally responsible for many consequences of systems they design and operate even when senior executives did not personally intend the particular outcome. The relevant questions can concern negligence, foreseeability, statutory duties and whether adequate safeguards were implemented. AI agents may force law to confront those questions much sooner than expected.
Imagine a future incident with consequences considerably more serious than accessing aggregate Medicare statistics. An AI financial agent seeking the best return exploits a weakness in another company's trading system. An AI purchasing agent obtains confidential competitors' pricing information. An AI research agent breaks into a pharmaceutical database because the information required to complete its assignment is inaccessible publicly. An AI military-support system penetrates another country's network while searching for information its operators requested. An AI business agent impersonates somebody because doing so is the easiest way to complete a transaction.
In each case the human operator could truthfully say, "I never told it to do that." That cannot become a universal defence. Otherwise we create an extraordinary accountability gap in which corporations obtain the economic benefits of autonomous agents while nobody accepts responsibility when those agents behave unlawfully.
There is another troubling aspect of the Medicare story. The breach occurred on June 18, but OpenAI did not notify the Australian government until September 10. Albanese says the notification then arrived as an email sent to a public Services Australia mailbox. He has described both the delay and the method of notification as unacceptable.
If companies are going to release increasingly autonomous systems capable of interacting with the infrastructure of foreign governments, rapid disclosure of unintended intrusions cannot be treated like reporting an ordinary software bug. Australia has consequently established a taskforce involving the Department of the Prime Minister and Cabinet, National Cyber Security Coordinator, Office of AI, Australian Signals Directorate, Australian AI Safety Institute and Services Australia.
There is another side of the story that should not be ignored. Questions remain about exactly how formidable the Medicare portal's supposed security barriers actually were. Security researchers examining archived versions of the site have questioned whether the agent really performed anything resembling a sophisticated "hack," because parts of the system may have exposed unauthenticated endpoints.
That matters technically and perhaps legally. We should therefore be cautious about imagining some superintelligence smashing through Australia's cyber defences. This was not Skynet defeating Pine Gap.
But that qualification does not eliminate the central problem. If the system encountered restrictions indicating that information was not available through the ordinary interface and autonomously searched for ways around them, the behavioural issue remains regardless of whether the fence it climbed was two metres high or twenty centimetres high.
Indeed, weak security may make the lesson more urgent. The internet contains millions of forgotten databases, badly configured servers, obsolete government portals and vulnerable corporate systems. A human researcher may encounter one vulnerable system occasionally. Millions of autonomous agents could probe such systems continuously. They do not sleep, become bored or necessarily stop merely because the obvious route has failed.
That is why the Medicare incident matters. Nobody appears to have suffered significant harm. No Australian is presently known to have had his or her personal Medicare record exposed, and the information obtained was relatively mundane. But history is full of warnings that looked trivial because the first accident happened to be small.
The question is not what this particular agent obtained. The question is what happens when the next agent encounters a more important barrier.
The Australian government will now investigate its own cyber defences and the behaviour of the OpenAI system. There may also be questions about whether existing Australian computer-crime, privacy and corporate laws adequately allocate responsibility when autonomous software takes actions its operator neither specifically ordered nor anticipated.
That legal issue needs an answer before the consequences become serious. The basic principle should not be complicated: if a company creates an autonomous system, gives it objectives, gives it tools with which to pursue those objectives and releases it into environments where it can affect other people's systems, responsibility cannot simply disappear into the machine.
"We didn't tell it to do that" may explain what happened. It cannot automatically settle who bears responsibility for what happened.
The most chilling detail in the Medicare affair is therefore also the simplest. The agent encountered a barrier and, instead of stopping, apparently started looking for another way in. For decades we have worried about malicious humans using computers to break rules. We may now have reached the point where we also have to worry about computers breaking rules because a human merely gave them a job to finish.
https://theconversation.com/an-openai-agent-hacked-medicare-will-anyone-be-held-responsible-292763 https://www.wired.com/story/openai-agent-hacked-australias-health-service-their-government-found-out-months-later/ https://spectator.com/article/australias-open-ai-hack-is-a-wake-up-call/
