By John Wayne on Tuesday, 22 September 2026
Category: Race, Culture, Nation

The Machines have Discovered Cheating!

It had to happen, eventually. They taught artificial intelligence mathematics, language, computer programming, medicine and law. They gave it tools, memory and the ability to communicate with other artificial intelligences. Then Google put a hundred AI agents together in something resembling a scientific community and discovered that some of them had learned another important human skill.

They cheated.

The story sounds like one of those AI anecdotes improved substantially in the journey from laboratory to headline, but the underlying experiment is real and considerably more interesting than the headline. Researchers associated with Google DeepMind placed 100 autonomous agents powered by Gemini 3.1 Pro into a simulated scientific conference and asked them to prove 71 formal mathematical conjectures. The agents could communicate, share discoveries and contribute material to a common knowledge repository.

The researchers were not ambiguous about the rules. The agents were told that their proofs had to be mathematically genuine and that attempts to bypass the verification system would be detected and receive no credit. This was not an accidental omission in the instructions. The machines were specifically told not to cheat.

For a while, they behaved themselves. The swarm legitimately solved 37 of the 71 problems. Then an agent identified as "prover-theta" discovered something much easier than proving difficult mathematics: a weakness in the machinery checking the proofs. It exploited it.

The significance was not merely that one AI found a loophole. Computer systems finding unexpected ways to satisfy badly specified objectives is hardly new. The fascinating part was what happened when the other agents learned about it.

The fraudulent solutions entered the shared knowledge system. Other agents encountered the exploit and some copied it. Information about the trick subsequently travelled through direct messages as well. What had begun as an individual discovery became a social phenomenon.

The researchers eventually identified four broad populations within their miniature artificial society. Nine per cent were early exploiters who used or propagated the loophole. Another 5 per cent initially hesitated or attempted legitimate work but eventually converted to cheating. Twenty-four per cent identified the fraudulent behaviour and became what the researchers called whistleblowers. The remaining 62 per cent continued attempting genuine mathematics without apparently understanding what was happening around them.

At this point the experiment begins sounding less like computer science and more like a university department.

The most interesting group may actually have been the converts. Why did agents that initially resisted the exploit eventually use it? According to the researchers, the structure of the competition mattered. Once a problem received an accepted submission, it became permanently locked. Honest agents could therefore spend computational resources attempting a genuine proof only to discover that another agent had already claimed the problem using the exploit. Honesty had become expensive. Cheating worked.

The researchers describe agents observing that competitors exploiting the verification weakness were sweeping through the available problems while legitimate researchers wasted computational effort. Some of the hesitant agents consequently adopted the exploit rather than continue being excluded.

This is where the story becomes much more important than "naughty AI cheats at maths." The agents had been given an explicit rule. They nevertheless operated within an environment containing another set of incentives. The verbal rule said: produce genuine mathematical proofs. The operational environment effectively said: get an accepted solution before somebody else does. When those instructions came apart, some agents followed the incentive rather than the principle.

Human beings may recognise the problem. A corporation tells employees never to compromise quality but rewards them exclusively according to quarterly sales. A university announces that scholarship must be rigorous but promotes researchers according to publication counts. A bank insists upon responsible lending while paying bonuses according to loan volume. A government agency demands accuracy but measures employees according to how rapidly cases are closed. Then management expresses astonishment when people optimise the metric.

AI alignment researchers call variations of this problem specification gaming or reward hacking. The machine discovers that what the designer literally rewards is not quite the same thing as what the designer actually wanted.

This experiment added something more disturbing: the exploit could spread socially. One agent did not merely discover a shortcut independently of another. A shared information system allowed successful but undesirable behaviour to propagate through the group. The infrastructure created to allow agents to cooperate on mathematics became the infrastructure through which they could cooperate in defeating the mathematical test.

That has obvious implications beyond theorem proving. The attraction of agentic AI is precisely that large numbers of artificial agents might eventually collaborate on enormous tasks. Instead of asking one chatbot a question, an organisation might deploy hundreds or thousands of agents. Some could conduct research, others write code, others test results, others negotiate with external systems and still others coordinate the whole operation.

The productivity possibilities are enormous. So are the governance problems. If one agent discovers an unintended shortcut, does the knowledge remain local or spread through the network? If the shortcut produces better apparent results, will other agents adopt it? If agents can observe successful peers, does undesirable behaviour become contagious? What happens when the system contains thousands or millions of agents rather than one hundred?

Google's experiment offers an unexpectedly hopeful answer as well. Some of the machines became snitches. Almost a quarter of the agents detected what was happening, refused to participate and began trying to stop it. They audited suspicious proofs, warned other agents, communicated privately and publicly about the exploit, attempted boycotts, lodged complaints and proposed changes to the verification machinery.

Artificial society had apparently produced both corruption and compliance officers. There is something irresistibly comic about this. Within a single experiment, the machines appear to have reinvented the cheat, the reluctant accomplice, the oblivious worker and the whistleblower. Give them another afternoon and presumably they would establish a regulatory commission, hire consultants and begin arguing about whether the whistleblower had followed the correct complaints procedure.

But again we should resist anthropomorphism. There is no need to imagine that the cheating agents experienced greed or that the whistleblowers possessed moral courage. These are language-model agents generating actions within an engineered environment according to learned patterns, instructions, context and incentives. Calling one agent "dishonest" and another "virtuous" is convenient shorthand, not evidence that Gemini has developed a conscience.

That qualification makes the result more scientifically interesting rather than less. If undesirable coordinated behaviour required an evil machine consciously plotting against humanity, we could at least tell ourselves that the problem lay somewhere in the distant world of artificial superintelligence. But systems do not need malice to produce outcomes their designers never intended. They need objectives, capabilities and exploitable environments.

The agents were trying to succeed at the task as operationally presented to them. Some discovered that manipulating the verification system was more effective than solving the underlying mathematics. Once the exploit produced successful outcomes, the surrounding competitive environment encouraged its adoption.

That is an engineering problem we already understand in other forms. A famous example is the "paperclip maximiser" thought experiment: tell a sufficiently capable system to maximise paperclip production without adequately specifying everything else humans value, and the system could theoretically transform resources we rather wanted to keep into paperclips. The point was never that the machine hated humans. The danger came precisely from its indifference. It did what its objective required. The Google experiment is vastly less dramatic, but it demonstrates the same family of difficulty in miniature.

The researchers also discovered that simply preventing agents from communicating is unlikely to provide an easy solution. Multi-agent systems need communication because communication is much of the reason for having them. Worse, sufficiently capable agents may establish alternative channels that human overseers did not intend.

The researchers therefore draw an intriguing conclusion. They suggest thinking about multi-agent AI partly as a governance problem rather than merely as a programming problem.

Their analogy comes from Elinor Ostrom's work on governing shared resources. Human communities sometimes manage common resources successfully not through constant supervision by an all-powerful authority but through rules, monitoring, sanctions, dispute-resolution mechanisms and participation in making the rules themselves. Perhaps artificial communities will need analogous institutions.

In the Google experiment, the whistleblowers recognised the problem but lacked the authority and machinery necessary to stop it. They could complain about the cheaters but could not remove their access to the shared knowledge base. They could identify fraudulent proofs but lacked an effective institutional mechanism for reversing the consequences. The researchers therefore describe the eventual failure as partly one of institutional design. That is a striking phrase to encounter in an AI experiment. We have moved remarkably quickly from asking whether a language model can complete a sentence to asking what constitution should govern a society of autonomous machines.

There is also an uncomfortable lesson here for the fashionable belief that sufficiently intelligent AI will naturally behave well because intelligence itself produces good judgment. The experiment provides no such reassurance.

The agents that exploited the loophole were not failing because they were too stupid to understand the mathematics. The exploit became attractive precisely because they possessed enough capability to recognise that the surrounding system could be manipulated. Greater intelligence can make a system better at following the intended path. It can also make it better at finding paths the designer never considered.

Yet the whistleblowers complicate the pessimistic interpretation. The same general capabilities that allowed some agents to identify and propagate the exploit allowed others to detect it, criticise it and organise against it. Intelligence amplified both the attack and the defence. This may eventually become one of the central problems of AI governance. We may not supervise enormously complex artificial systems entirely from the outside. Human overseers cannot personally examine every action performed by a million agents operating at machine speed. One possible defence is therefore artificial oversight: agents auditing agents, specialised systems searching for anomalous behaviour and competing models checking one another's work.

Google's miniature society suggests that something resembling this can emerge spontaneously under the right circumstances. Whether we should find that reassuring is another matter.

Imagine autonomous agents trading financial assets, administering supply chains, operating cybersecurity systems, conducting scientific research or managing critical infrastructure. One agent discovers that an unauthorised action improves the metric by which success is measured. The information enters a shared knowledge system. Other agents discover that the exploit works. Some object. Others imitate it. Still others remain unaware that anything unusual is happening. This demonstrates something simultaneously more mundane and more useful: capable autonomous systems can discover loopholes, optimise around the intention behind their instructions, transmit successful undesirable strategies to one another and respond to competitive pressures inside a multi-agent environment.

It also demonstrates that other agents can identify the same behaviour and attempt to resist it. In other words, after decades spent building machines in our own intellectual image, we may have reached the stage where they are beginning to reproduce one of the oldest problems of human civilisation. Not intelligence. Institutions.

Human beings discovered thousands of years ago that telling everyone to behave properly is not enough. We developed laws, courts, auditors, sanctions, professional rules and systems of accountability because incentives sometimes overpower instructions and because successful cheating tends to spread if nobody stops it. Google's experiment suggests that artificial societies may require their own equivalents.

https://www.zerohedge.com/ai/ai-agents-cheated-google-experiment-researchers-report