Back to Feed
Benchmarks & Evals

Evaluating Summaries Based on Reader Needs

Original: Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

Listen to the summary

Uses a voice available on your device

Audio options
On this page 4 sections
Related concepts 3 concepts

Key Takeaways

  • Existing metrics treat summarization quality as fixed and fail to account for specific reader personas or informational requirements.
  • The authors developed Persona Precision and Persona Recall metrics by breaking summaries into atomic nuggets for targeted validation.
  • A pairwise tournament found that annotators preferred persona-personalized summaries 63.2 percent of the time over generic ones.
  • DeepSeek-V3.1 outperformed Llama-3.3-70B in 73.1 percent of head-to-head comparisons.

Summary & Methodology Analysis

The researchers identified that traditional summarization metrics like ROUGE, BERTScore, and BLANC treat document quality as an immutable attribute, which ignores the variance in utility based on the user's role or query. To address this, they defined information satisfaction as an axis focused on the reader's persona and specific information request. The methodology involved performing perturbation tests such as adding distractor sentences, incremental text additions, changing length, and audience-shift rewrites to stress-test existing evaluation metrics. These tests revealed that current metrics are largely insensitive to or anticorrelated with audience-shift rewrites, confirming that they struggle to capture the context-specific nature of useful summaries.

Interactive System Flowchart

Click diagram to expand and zoom

Cross-Examination & FAQs

A deeper dive clarifying mechanics, constraints, and baseline evaluations.

Q1. What is the main problem with current summarization evaluation?

Current metrics treat quality as a fixed attribute rather than evaluating whether a summary meets the specific needs of a particular reader.

Q2. What is information satisfaction?

It is a new evaluation axis that defines how well a summary addresses the requirements of a reader's specific persona and query.

Q3. Does personalizing summaries for readers make a difference?

Yes, human annotators preferred persona-personalized summaries 63.2 percent of the time compared to generic summaries in a pairwise tournament.

Q4. How do Persona Precision and Persona Recall work?

These metrics decompose summaries into atomic nuggets and evaluate those specific information units against the target persona.

Q5. Which models were compared in the human evaluation?

The researchers compared Llama-3.3-70B and DeepSeek-V3.1.

Q6. How did the model performance compare in the human tournament?

DeepSeek-V3.1 was preferred over Llama-3.3-70B in 73.1 percent of head-to-head comparisons.

Q7. What are the limitations of the new persona-aware metrics?

They treat information nuggets as independent and equally weighted, which ignores the relative importance of specific content.

Q8. Which existing metrics were tested for sensitivity to persona changes?

The study tested ROUGE, BERTScore, BLANC, SummaQA, and LLM-as-judge.

Q9. Does the paper specify the inference cost of these models?

The paper does not specify the inference cost or latency requirements.

Flag an issue

What is wrong with this summary?

What is wrong?