ArXiv

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

Authors
Yang Wang, Yanan Ma, Yiqi Liu...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2608.04939v1
PDF
https://arxiv.org/pdf/2608.04939v1

Brief

DrivelHub+ introduces a diagnostic benchmark of 1,000 social-media videos with human-written implicit narrative explanations to test models' ability to infer non-literal pragmatic meanings (e.g., humor, irony, satire) that arise from multimodal cues and cultural context. The authors evaluate current video-language models on an Explanation task (generate pragmatic explanations) and a Representation task (reasoning-as-retrieval for video↔text retrieval), highlighting a measured gap between perception-focused description and pragmatic comprehension. Summary based on the abstract; full text was not reviewed.

Why it matters

DrivelHub+ is a new benchmark of 1,000 social-media videos, each annotated with a human-written implicit narrative explanation, targeting implicit, non-linear, rhetorically layered meanings such as humor, irony, and satire.

Key details

  • Evaluation covers two tasks: (1) Explanation — models must produce natural-language pragmatic explanations of a video's implied meaning; (2) Representation — a reasoning-as-retrieval protocol testing video-to-text and text-to-video alignment between videos and their implicit narratives.
  • Paper: Yang Wang et al., arXiv:2608.04939v1 (published 2026-08-05); benchmark is positioned as a diagnostic to measure the gap between conventional recognition/description-focused video understanding and deeper pragmatic multimodal comprehension.
Source evidence

Abstract

Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models. In this paper, we introduce \textit{DrivelHub+}, a benchmark for evaluating whether models can infer the implicit, non-linear, and rhetorically layered meanings of social media videos that appear nonsensical on the surface but convey deliberate pragmatic meanings. DrivelHub+ consists of 1,000 videos collected from social media, each annotated with a human-written implicit narrative explanation. Unlike conventional video understanding tasks focused on recognition or description, we present a benchmark that targets contextual multimodal reasoning. We evaluate current video-language models from two perspectives: explanation, where models must explain the pragmatic comprehension of a video in natural language; and representation, where we adapt reasoning-as-retrieval to test whether model representations align videos with their corresponding implicit narratives in both video-to-text and text-to-video retrieval. Our benchmark provides a diagnostic setting for measuring the gap between multimodal perception and pragmatic comprehension, asking whether current models can move beyond describing what is shown to inferring what is meant.