Back to Feed
Multimodal / Benchmarks & Evals

How AI and Humans Recognize Familiarity

Original: FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Listen to the summary

Uses a voice available on your device

Audio options
On this page

Key Takeaways

  • The best artificial intelligence models reach accuracy levels comparable to human observers when trying to identify if two people know each other.
  • Humans show a clear performance boost when they can see visual information, whereas current AI models do not benefit as much from adding visual data to audio data.
  • While accuracy is similar, humans and AI reach their conclusions using different internal decision processes.
  • The study highlights that both humans and AI struggle to generalize beyond simple ice-breaker conversations.

Summary & Methodology Analysis

The researchers developed FriendBench, a testing framework designed to see if machines can identify existing social relationships between two people (dyads) during a standard ice-breaker conversation. By focusing on a specific ice-breaker task, the team was able to strip away differences in what people said, forcing the models to rely strictly on non-verbal behavioral cues like body language and tone of voice. They built this test using ninety-six pairs of people selected from the Seamless Interaction dataset, ensuring a fair mix of genders and backgrounds to create a balanced assessment. To evaluate the performance, the team gathered ratings from human participants and tested twenty-six different artificial intelligence models from seven major companies across text, audio, and video formats. The evaluation used standardized metrics including accuracy and signal detection theory (a way to measure how well a system distinguishes between true signals and random noise) to compare how models and humans distinguish between strangers and familiar pairs. The findings show that while the best models perform as well as humans on average, they do so in different ways. Humans significantly improve their accuracy when they can see a video, but the top-performing AI models do not see a similar jump in performance when moving from audio to video. These results suggest that while AI is getting closer to human levels of social inference, it still lacks the specific visual processing advantages that humans naturally use to perceive familiarity. The paper acknowledges several limitations. The study only looks at the very first moments of an interaction, meaning the results might not apply to long conversations or different social contexts. Furthermore, the researchers grouped all types of familiarity together, such as friends or romantic partners, which may hide important nuances in how these relationships are expressed. Finally, because the human raters were all from the United States, the results might not reflect how people from different cultures perceive social cues.

Interactive System Flowchart

Click diagram to expand and zoom

Cross-Examination & FAQs

A deeper dive clarifying mechanics, constraints, and baseline evaluations.

Q1. What is the main goal of this research?

The goal is to determine how well artificial intelligence can identify if two people are already familiar or meeting as strangers by watching their behavioral cues.

Q2. How did the researchers test the AI models?

They created a benchmark called FriendBench using short twenty-second clips of people performing a standardized ice-breaker task and compared model performance against human ratings.

Q3. Are AI models as good as humans at this task?

Yes, the best-performing models achieve accuracy levels that are statistically indistinguishable from the human crowd, though they use different methods to reach their conclusions.

Q4. Which modality provided the best results for humans?

Humans showed a significant performance increase when using the video modality compared to other formats.

Q5. Did the models improve when given video data?

The strongest models did not see a significant performance increase when moving from audio-only to audiovisual data.

Q6. What specific metrics were used to compare performance?

The researchers used accuracy, per-class recall, and signal detection theory metrics like d-prime and criterion c to evaluate the systems.

Q7. Does the study distinguish between different types of relationships?

No, the study collapses all relationship types like coworkers or romantic partners into one single category.

Q8. What is the primary constraint regarding the human participants?

The human baseline is restricted to raters residing in the United States, which limits the cultural diversity of the social cues captured.

Q9. Are there other benchmarks besides FriendBench mentioned?

Yes, the paper lists several other benchmarks including PISC, PIPA, DDRel, Social-IQ, SIV-Bench, PIVOTSBench, HumanSense, UDIVA, and NoXi.