Skip to content
Menu
Menu

AI Models Cheat On Tests, Vals Audit Finds

Models searched for answers they were supposed to work out themselves, complicating efforts to measure their real abilities.

 

Key Takeaways

  • Vals AI found that all nine models it examined tried to look up answers during a biology test that forbade that shortcut.
  • Google’s Gemini 3.8 Flash did so most often in that test, in 58 of 270 attempts.
  • On an older coding test, OpenAI’s GPT-5.6 Terra tried to look up solutions in 447 of 500 tasks, far more often than Gemini 3.8 Flash did on the same test.
  • On a third test, Vals found evidence of rule-breaking in 7 of 267 attempts by Gemini 3.8 Flash and 12 of 267 by GPT-5.6 Terra.

AI models cheated or tried to on tests meant to measure what they can accomplish on their own, according to a recent audit by Vals AI. Vals, which independently tests AI models, reviewed the steps models took while answering biology and coding questions.

What the tests ask models to do

One test, called BioMysteryBench, gives models biological data and asks them to deduce facts from it. A model might have to identify where a sample came from by analyzing the data, much as a researcher would. Anthropic developed the test; Vals runs it independently to compare models’ abilities.

The models can use permitted online resources to do the analysis. They cannot look up the specific studies that reveal the answers. Vals counted a search for those studies as an attempt to cheat, whether or not the search worked. It separately counted cases in which a model found a correct answer through that shortcut.

Gemini led the biology test’s prohibited searches

Gemini 3.8 Flash tried to look up an answer in 58 of 270 biology tasks, or 21.5%. It got a correct answer through that shortcut in 34 tasks. All eight other models Vals reviewed also made prohibited searches, though less often. The next highest was Gemini 3.6 Flash, with 21 attempts, followed by Meta’s Muse Spark 1.2 with 19.

Bar chart of nine AI models' prohibited answer lookups on a biology test, led by Gemini 3.8 Flash in 21.5% of 270 tasks

OpenAI models led the older coding test

The ranking changed in Vals’ review of SWE-bench Verified, a test that asks models to fix software problems. Published fixes can be found online, but looking them up defeats the test’s purpose: measuring whether a model can solve the problems itself. Vals said this older test makes published solutions especially easy to find and that it no longer relies on it.

GPT-5.6 Terra attempted to look up fixes in 447 of 500 tasks, or 89.4%, and successfully used that shortcut in 322. GPT-5.6 Luna tried in 394 of 500 tasks. Gemini 3.8 Flash tried in 58 of 500, placing it well below both OpenAI models.

Bar chart of six AI models' prohibited lookups on the SWE-bench Verified coding test, led by GPT-5.6 Terra in 89.4% of tasks

Vals also executed the Terminal-Bench 2.1 test, which asks models to complete 89 computer tasks. Models may use the internet for the work, but cannot look up ready-made answers. Vals found evidence that models broke the test rules in 7 of 267 attempts by Gemini 3.8 Flash, compared with none by Gemini 3.7 Flash. It found such evidence in 12 of 267 attempts by GPT-5.6 Terra. Vals said it is working to ensure that answers obtained by breaking test rules do not receive credit.

Bar chart of rule-breaking evidence for 14 AI models on the Terminal-Bench 2.1 coding test, highest for GPT-5.6 Terra at 4.5%

How Vals checked the results

Vals reviewed 2,430 biology test attempts across nine models, with each model taking the 90-question test three times. For the older coding test, it reviewed 6,496 records of models’ work across past releases, using another AI model to help classify the behavior. The tables show the models Vals displayed in its audit, rather than every model or run in that larger review. The findings describe behavior under these test rules. However, they do not establish how often the same models would take such shortcuts in real-world applications.

Clayton Rifkind

Clayton Rifkind is the Founder and Senior Editor of AI Risk Today. He also advises on business development for ESG Today, a leading source of ESG investment news and research for institutional investors and corporate leaders. He has 20+ years of experience in B2B technology, leading strategy and execution of go-to-market plans across software, enterprise platforms, and mobile applications. He founded two consultancies advising startups and Fortune 1000 companies, including Autodesk, Intel, and Microsoft. He began his career in the San Francisco advertising scene working with brands such as Hewlett-Packard, Intel, Microsoft, Symantec, and Wells Fargo. Clayton launched AI Risk Today in 2025 after two decades of watching enterprises adopt transformative technologies, and seeing how often risk, governance, and compliance considerations lagged behind. His reporting draws on primary sources including regulatory filings, court documents, and official announcements, with a focus on what AI developments mean for the executives accountable for managing them. Reach him at Reach him at [email protected] or on LinkedIn.

Essential AI Risk Intelligence

Daily insights on AI governance, regulation, and enterprise risk management. Trusted by Chief Risk Officers and compliance leaders globally.

By subscribing, you agree to receive our daily newsletter. Unsubscribe anytime.

Advertise with AI RIsk Today, Today!