Back to Feed
Benchmarks & Evals

Benchmarking Cognitive Bias in LLMs

Original: AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

Listen to the summary

Uses a voice available on your device

Audio options
On this page 4 sections
Related concepts 4 concepts

Key Takeaways

  • AnchorBench measures the anchoring effect using synthetic numeric judgment items across six business domains.
  • The benchmark tests five distinct pathway suites including External, History, In-Context Learning, Retrieval-Augmented Generation, and Tool.
  • Anchoring is strongly pathway-dependent, with External and RAG showing the broadest positive effects.
  • Task accuracy and anchoring discrimination are only weakly correlated at r = -0.24.

Summary & Methodology Analysis

Large language models, which are deep learning systems that process sequential data via attention mechanisms, suffer from cognitive biases similar to humans. Specifically, the anchoring effect occurs when an initial reference value shifts subsequent judgments toward itself. Existing work typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant anchors from plausible ones. To address this, the researchers built AnchorBench, a multi-pathway benchmark consisting of synthetic numeric judgment items across six business domains with fixed evidence and deterministic gold answers. The paper delivers the anchor through five distinct pathway suites: External, History, In-Context Learning, Retrieval-Augmented Generation, and Tool. It varies anchor relevance across control, irrelevant, and plausible conditions using identical numeric anchor values, and evaluates fourteen models using standardized inference and deterministic hierarchical parsing.

The evaluation covers ten open-weight models and four frontier API models, specifically including Llama 3.1, Llama 3.2, Llama 3.3, Qwen2.5, Gemma 3, OLMo 2, OpenAI GPT-5.4-mini, OpenAI GPT-5.4, Anthropic Claude Haiku 4.5, Anthropic Claude Sonnet 4.6, Google Gemini 2.5-Flash, xAI Grok-3-mini-beta, and xAI Grok-4.3. The methodology measures task accuracy on control responses, anchor susceptibility using Unified Anchor Influence and Toward-Anchor Rate, and relevance discrimination using the discrimination gap. The key results show that anchoring is strongly pathway-dependent, with External and RAG showing the broadest positive effects, In-Context Learning near zero, and History and Tool varying by model. Additionally, task accuracy and anchoring discrimination are only weakly correlated, with r = -0.24 and a bootstrap 95 percent confidence interval of [-0.43, -0.00].

Despite these findings, the study has notable limitations. AnchorBench studies anchoring in a controlled setting using synthetic deterministic-aggregation tasks, which clean up comparison but do not cover the complexity of real judgment. Furthermore, the History suite uses a different interaction structure from the control, the Tool suite is affected by cross-family format differences, and small effects are sensitive to the Unified Anchor Influence exclusion rule, making a few cross-suite comparisons less precise. The paper does not specify compute costs, execution runtimes, or hardware requirements for running these evaluations.

Interactive System Flowchart

Click diagram to expand and zoom

Cross-Examination & FAQs

A deeper dive clarifying mechanics, constraints, and baseline evaluations.

Q1. What is the main problem addressed by the paper?

The paper addresses the anchoring effect in large language models, where an initial reference value shifts a subsequent judgment toward itself, noting that prior work evaluated a narrow set of anchor pathways and rarely distinguished irrelevant from plausible anchors.

Q2. What is AnchorBench?

AnchorBench is a multi-pathway benchmark consisting of synthetic numeric judgment items across six business domains with fixed evidence and deterministic gold answers.

Q3. Which models were evaluated in the study?

The study evaluated fourteen models, comprising ten open-weight models and four frontier API models including Llama 3.1, Llama 3.2, Llama 3.3, Qwen2.5, Gemma 3, OLMo 2, OpenAI GPT-5.4-mini, OpenAI GPT-5.4, Anthropic Claude Haiku 4.5, Anthropic Claude Sonnet 4.6, Google Gemini 2.5-Flash, xAI Grok-3-mini-beta, and xAI Grok-4.3.

Q4. What are the five pathway suites used to deliver the anchor?

The five pathway suites are External, History, In-Context Learning, Retrieval-Augmented Generation, and Tool.

Q5. How were anchor relevance conditions varied in the experiments?

The experiments varied anchor relevance across control, irrelevant, and plausible conditions using identical numeric anchor values.

Q6. What metrics were used to measure anchor susceptibility and relevance discrimination?

Anchor susceptibility was measured using Unified Anchor Influence and Toward-Anchor Rate, while relevance discrimination was measured using the discrimination gap.

Q7. What correlation was found between task accuracy and anchoring discrimination?

Task accuracy and anchoring discrimination are only weakly correlated, with r = -0.24 and a bootstrap 95 percent confidence interval of [-0.43, -0.00].

Q8. What are the main limitations regarding the benchmark design?

AnchorBench uses a controlled setting with synthetic deterministic-aggregation tasks that do not cover the complexity of real judgment, the History suite uses a different interaction structure from the control, and the Tool suite is affected by cross-family format differences.

Q9. How do data exclusions affect cross-suite comparisons?

Small effects are sensitive to the Unified Anchor Influence exclusion rule, which makes a few cross-suite comparisons less precise.

Flag an issue

What is wrong with this summary?

What is wrong?