ArXiv

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

Authors
Kevin Du, Clara Kümpel, Michelle Wastl...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2607.18232v1
PDF
https://arxiv.org/pdf/2607.18232v1

Brief

The paper develops a typology of user expressions of belief (EoBs) across four linguistic dimensions with 17 fine-grained types, generates controlled EoB–query pairs anchored to world-knowledge facts, and evaluates 16 LLMs (Llama3, Qwen3, Gemma3; 1B–30B; base vs instruct). Results show model size and instruction tuning affect tendency to adopt user beliefs, and some EoBs more reliably persuade models.

Why it matters

Introduces a linguistically grounded typology of expressions of belief (EoBs) along four dimensions—form, evidentiality, epistemic stance, and tone—covering 17 fine-grained EoB types and uses controlled EoB–query pairs paired with world-knowledge facts.

Key details

  • Evaluates 16 LLMs (including Llama3, Qwen3, Gemma3) across scales from 1B–30B parameters and training stages (base vs instruct) and finds that larger models and instruction-tuned models tend to be less likely to follow contextual user beliefs than smaller and base models.
  • Finds specific EoB types that statistically significantly persuade models more consistently, revealing systematic patterns in how linguistic framing affects LLM context integration (published at ACL 2026).
Source evidence

Abstract

Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true. In others, they should stick to their prior knowledge. Notably, users' expressions of belief (EoBs) can take linguistically diverse forms - using presuppositions, evidential and certainty markers, or varied tones - each of which may have a different persuasiveness over the LLMs. We introduce a typology to systematically evaluate how different EoBs affect whether models follow context versus prior knowledge. The typology is grounded in four linguistically motivated dimensions: form, evidentiality, epistemic stance, and tone, spanning 17 fine-grained types. By pairing these EoBs with world knowledge facts, we generate controlled EoB-query pairs that isolate the effect of linguistic variation. Using this benchmark, we evaluate 16 LLMs that differ in architecture (Llama3, Qwen3, Gemma3), scale (1B-30B parameters), and training stages (base vs instruct). We identify meaningful variations in response behavior across these axes, e.g., that bigger models and instruction models tend to be less context-following than smaller models and base models. We further identify specific EoBs that statistically significantly persuade LMs more consistently than others. Our work reveals systematic patterns in how linguistic framing affects LLM context integration, with implications for prompt engineering and model robustness.

Comment: Published at ACL 2026