By John Wayne on Friday, 25 September 2026
Category: Race, Culture, Nation

Turing Test for AI: Pass and Fail

The Medium piece (link below), is reacting to a real result: in a preregistered three-party Turing test, GPT-4.5 with a human-like persona prompt was judged to be the human 73% of the time, more often than the actual human interlocutor. That is the first clean empirical pass of the standard (short, text-only, three-party) version of Turing's imitation game. That fact is interesting, but it is not decisive about minds.

Turing (1950) replaced the question "Can machines think?" with a behavioural criterion: if a judge cannot reliably tell the machine from a person after a short typed conversation, then we should be willing to say the machine thinks. He treated the question as too vague to settle by definition-chasing. The test is an operational substitute, not a metaphysical proof.

The UC San Diego setup is faithful to that substitute: simultaneous chats, 5 minutes, one human foil, forced choice. Persona prompting mattered a lot; without it the same model dropped to roughly chance or below. So, what was measured is highly skilled imitation of a social style, including the messy, slightly off-kilter qualities people use as "human" tells. That is a genuine engineering milestone. It is not an answer to "does the system understand?"

John Searle's Chinese Room (1980) was aimed exactly at this move. A person who does not know Chinese sits in a room with a rule book that maps incoming Chinese symbols to outgoing Chinese symbols. From outside, the room passes a Chinese Turing test. From inside, there is only formal symbol-shuffling. Syntax is being executed; semantics, meaning, understanding, aboutness, is not thereby produced.

The argument is not "computers can never do X." It is that successful functional imitation of linguistic behaviour is not the same kind of thing as understanding. The room (or the LLM) can implement the right input–output mapping without the states that, in us, constitute grasping what the symbols are about.

That is the gap the essay is pointing at: a jump from epistemology to ontology.

Epistemology / evidence: We observe outputs that are, for short conversations under particular prompts, statistically indistinguishable from (or preferred over) human outputs. We have excellent evidence of competence at the imitation game.

Ontology: We then treat that evidence as settling what the system is, whether it has understanding, intentionality, a mind, consciousness, etc.

Those are different questions. Behavioural evidence constrains hypotheses about inner states; it does not automatically identify them. You can have arbitrarily good evidence that a system acts as if it understands without that being identical to understanding. Searle's point is that the identification is a further, and optional, philosophical commitment (strong AI / computational functionalism), not something the test itself delivers.

There are a few concrete reasons the result does not close the gap:

1.The test is short and narrow. Five minutes of chat is not the full range of linguistic, practical, or embodied competence Turing gestured at, let alone a theory of mind.

2.Success was prompt-dependent. The model was coached into a persona. That is closer to "very good actor following a script plus statistical improvisation" than to an unprompted subject with its own stance on the world.

3.Imitation can exploit human heuristics. Judges use fluency, hedging, small inconsistencies, and cultural texture as cues. Models can be trained or prompted to emit those cues without the underlying states those cues normally track in people.

4."More human than the human" is a warning sign, not a proof of mentality. It shows the test is measuring persuasiveness of a performance, which can diverge from the thing being performed.

None of this requires denying that current systems are extraordinarily capable, or that they will get more capable. It only requires not collapsing "we cannot tell them apart in this protocol" into "therefore they have what we have when we understand."

The logical jump the essay flags is old and keeps returning because the evidence keeps getting better while the metaphysical question stays the same. Better evidence of behaviour does not, by itself, change the category of the claim. You still need an independent account of what understanding is: biological, functional, informational, or otherwise, before the Turing result can be read as an ontological verdict rather than an impressive score on a behavioral exam.

https://medium.com/personal-growth/no-one-is-ready-for-whats-coming-no-one-knows-what-s-happening-ddbe86f96d91

https://www.pnas.org/doi/10.1073/pnas.2524472123

"The Turing test has been widely discussed as a test of machine intelligence, but it also provides a measure of how humans distinguish other humans from machines. We evaluated 4 systems (ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5) in two randomized, controlled, and preregistered Turing tests on independent populations. Participants had 5 min conversations simultaneously with another human participant and one of these systems before judging which conversational partner they thought was human. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time—not significantly more or less often than the humans it was being compared to. Without these prompts, however, the same models performed significantly worse (38% and 36%), and did not consistently outperform baseline models, ELIZA and GPT-4o (23% and 21%, respectively). A third study replicated these results in 15-min games: two PERSONA-prompted models achieved pass rates of 56% and 59%. The results constitute empirical evidence that artificial systems can pass a standard three-party Turing test. Interrogators' reasoning focused more on stylistic and socio-emotional aspects of human behavior rather than more traditional notions of intelligence. The results have implications for debates about what kind of intelligence is exhibited by large language models, the social impacts these systems are likely to have, and the aspects of human behavior that people continue to see as unique."