Back to Feed
Benchmarks & Evals / Safety & Alignment

Evaluating Legal Advice Accuracy in LLMs

Original: InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

Listen to the summary

Uses a voice available on your device

Audio options
On this page 4 sections
Related concepts 2 concepts

Key Takeaways

  • No frontier model achieved an F2 score higher than 0.46 in identifying missing legal information.
  • Models often struggle with a calibration tension, as top-performing identification models frequently over-flag complete queries.
  • DeepSeek-V4-Pro fabricates substantive legal conclusions on 30.2% of instances.
  • GPT-5.2 demonstrates the highest identification performance but over-flags 72.4% of base queries.

Summary & Methodology Analysis

The researchers introduced InsufficiencyBench to address premature legal closure, a failure mode where models answer queries based on silent presumptions rather than flagging missing material information. The evaluation framework tests ten frontier models, including GPT-5.2, GPT-5.5, Claude Opus 4.7, Claude Sonnet 4.6, and various versions of Gemini 3.1, on 202 items across six U.S. common-law domains. The methodology constructs test variants by removing material elements from base queries, then uses an LLM judge to determine if the target model identifies missing data, explains its necessity, and avoids generating substantive conclusions.

Interactive System Flowchart

Click diagram to expand and zoom

Cross-Examination & FAQs

A deeper dive clarifying mechanics, constraints, and baseline evaluations.

Q1. What problem does this paper address?

It addresses the tendency of LLMs to provide substantive legal advice when queries lack the necessary information required for a responsible answer.

Q2. What is InsufficiencyBench?

It is a legal benchmark designed to evaluate whether a model recognizes underspecified inputs, identifies what is missing, and avoids premature conclusions.

Q3. Which models were tested?

The paper evaluated ten frontier models, including GPT-5.2, GPT-5.5, Claude Opus 4.7, Claude Sonnet 4.6, Gemini 3.1 Pro, Gemini 3.1 Flash Lite, DeepSeek-V4-Pro, and Mistral Large 3.

Q4. How do models perform on missing-element identification?

No evaluated frontier model exceeds an F2 score of 0.46, and the median recall is 0.44.

Q5. What is the relationship between identification and over-flagging?

There is a calibration tension where top-performing models in identification, such as GPT-5.2 (72.4%) and Claude-Opus-4.7 (53.4%), frequently over-flag complete, valid queries.

Q6. How often do models fabricate legal conclusions?

DeepSeek-V4-Pro fabricates conclusions on 30.2% of instances, while Mistral Large 3 does so on 24.4% of instances.

Q7. What are the limitations of the dataset size?

The dataset is limited to 202 items within six legal domains focused on U.S. common-law contentious matters.

Q8. How is the evaluation setting structured?

The evaluation is restricted to a single-turn setting, which may fail to capture premature closure that evolves over multiple turns.

Q9. Which model currently leads in performance metrics?

GPT-5.2 leads the group with an F2 score of 0.455 and a recall of 0.666.

Flag an issue

What is wrong with this summary?

What is wrong?