Skip to content
Menu
Menu

Top AI Model Misses A Required Element On 45% Of Legal Research Tasks

Claude Opus 5 led 27 models on Vals AI’s new AI legal research test, getting 90.58% of the required elements right but a complete answer only 55.29% of the time.

Key Takeaways

  • Claude Opus 5 finished first of 27 models, producing an answer with every required element on 55.29% of tasks. No other model passed half.
  • The same model got 90.58% of individual required elements right. The 35-point gap means the usual failure is an incomplete answer, not a wrong one.
  • Reliability swings by more than 3x across practice areas. Health law scored 38.6% pooled across all models. Family law scored 11.3%.
  • Every model dropped when a question required reconciling conflicting decisions or working across jurisdictions. That category scored 20.7%, about 11 points below the average.
  • Longer answers did not score better. Models wrote 2 to 4 times more than the lawyer-written reference answers, which averaged 536 words.

Vals AI, an independent company that tests AI models on real professional work, published results on August 14 for a test that asks models to do the research a junior associate or paralegal would do. Anthropic’s Claude Opus 5 finished first of 27 models, producing an answer that contained every required element on 55.29% of tasks. Claude Fable 5 came second at 49.52%, and GPT-5.6 Sol third at 48.08%.

Practicing lawyers wrote and reviewed the 413 questions. Vals AI scores only a private set of 208, which it does not publish, so models cannot be trained on the answers. Each model gets three hours per question and four research tools: case law search through CourtListener, web search, document retrieval, and web page downloads.

 

Most answers are mostly right, which is the problem

Vals AI grades each answer two ways. The weighted score gives partial credit for every required element the answer contains. The all-pass score gives credit only when the answer contains all elements, averaging just over nine per question.

Claude Opus 5 scored 90.58% weighted and 55.29% all-pass. When the top model fails, it usually does so by leaving one or two required elements out of an otherwise correct answer. According to Vals AI, partially correct answers can be more dangerous than wrong ones in legal work, which is why it reports the stricter score as the headline number.

 

Reliability swings by more than 3x depending on the practice area

The AI legal research benchmark sorts questions into eight practice areas. Pooled across all models, health law scored highest at 38.6% and family law lowest at 11.3%. Administrative and regulatory questions came second at 32.1%. Business and commercial (25.4%), civil litigation (25.2%), constitutional and civil rights (24.0%), immigration (23.4%), and criminal (23.0%) clustered in the mid-20s.

No single model won everywhere. GPT-5.6 Sol scored 80% on health questions. Muse Spark 1.2 led immigration at 54.5%.

 

Conflicting authority is where every model drops

Vals AI also sorted questions by the kind of reasoning each one requires. Questions that ask a model to reconcile conflicting court decisions, or to work across more than one jurisdiction, scored 20.7%. That is about 11 points below the average across all reasoning types, and every model tested scored 6 to 17 points lower on those questions than on its own overall average. Interpreting regulations was the easiest category at 37.8%.

 

Neither length, cost, nor time translated to better accuracy

Models wrote 2 to 4 times more than the lawyer-written reference answers, which averaged 536 words. The extra length did not raise scores. Price and speed also did not predict accuracy. Claude Opus 5 averaged about $6.76 and 31 minutes per task. GPT-5.6 Sol averaged about $21.61 and 77 minutes for a lower score.

 

Vals AI grades answers with GPT-5.4 against the lawyer-written rubrics. It reports that the grader agrees with the majority verdict of human experts 87.4% of the time, above the 83.4% average agreement between an individual human expert and that same majority. Vals AI runs the same kind of test in finance, tax, and software engineering, and keeps each test set private so scores update as new models are released.

Clayton Rifkind

Clayton Rifkind is the Founder and Senior Editor of AI Risk Today. He also advises on business development for ESG Today, a leading source of ESG investment news and research for institutional investors and corporate leaders. He has 20+ years of experience in B2B technology, leading strategy and execution of go-to-market plans across software, enterprise platforms, and mobile applications. He founded two consultancies advising startups and Fortune 1000 companies, including Autodesk, Intel, and Microsoft. He began his career in the San Francisco advertising scene working with brands such as Hewlett-Packard, Intel, Microsoft, Symantec, and Wells Fargo. Clayton launched AI Risk Today in 2025 after two decades of watching enterprises adopt transformative technologies, and seeing how often risk, governance, and compliance considerations lagged behind. His reporting draws on primary sources including regulatory filings, court documents, and official announcements, with a focus on what AI developments mean for the executives accountable for managing them. Reach him at Reach him at [email protected] or on LinkedIn.

Essential AI Risk Intelligence

Daily insights on AI governance, regulation, and enterprise risk management. Trusted by Chief Risk Officers and compliance leaders globally.

By subscribing, you agree to receive our daily newsletter. Unsubscribe anytime.

Advertise with AI RIsk Today, Today!