ArXiv

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Authors
Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino...
Categories
cs.CL, cs.AI, cs.CV, cs.HC
arXiv
https://arxiv.org/abs/2607.29602v1
PDF
https://arxiv.org/pdf/2607.29602v1

Brief

FriendBench evaluates whether two people in a 20-second ice-breaker are previously familiar or strangers, comparing 26 multimodal models (seven companies) to human panels across text, audio, and video on 96 balanced dyads. The best model matches human accuracy across modalities but leans toward “stranger” labels (an effective prior difference); only humans gain from added visible behavior. The paper releases data and predictions.

Why it matters

FriendBench introduces a dyadic familiarity benchmark using 20-second ice-breaker clips across text, audio, and video, evaluating 26 models from seven companies against matched human panels on 96 balanced dyads (Girard et al., arXiv 2026-07-31).

Key details

  • The top model and the human crowd are statistically indistinguishable in accuracy in every modality, but strongest models show a systematic bias toward labeling pairs as “strangers” (an effective prior difference); only humans benefit from visible behavior beyond speech. The authors release stimuli, human ratings, and model predictions.
Source evidence

Abstract

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward "stranger"---a difference in effective prior, not discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.

Comment: 15 pages, 3 figures