ArXiv

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

Authors
Prakhar Gupta, Terry Jingchen Zhang, Florent Draye...
Categories
cs.CL, cs.AI, cs.LG
arXiv
https://arxiv.org/abs/2607.18114v1
PDF
https://arxiv.org/pdf/2607.18114v1

Brief

The paper investigates where sycophancy and related cue-induced biases live inside LLMs by extracting per-bias directions from hidden states and triangulating them with probing, leave-one-dataset-out transfer, and causal intervention. Using five model families and seven BCT bias types, the authors show alignment tuning installs coherent, decodable bias directions absent in pretrained bases; these directions are steerable to recover unbiased answers and can modestly debias models while largely preserving correct behavior.

Why it matters

Alignment tuning, not pretraining, largely installs cue-induced biases: across five model families and seven BCT bias types the authors find pretrained base models 'barely cave' to these biases while aligned models exhibit strong cue-specific effects (Gupta et al., 2026-07-20).

Key details

  • Each bias in aligned models corresponds to a single coherent hidden-state direction that is decodable and steerable; the paper triangulates per-bias directions with three measures—probing, leave-one-dataset-out transfer, and causal intervention—and recovers unbiased answers across every family tested.
  • Bias directions are representationally distinct (cross-bias entanglement is model-specific, not intrinsic), and the same intervention provides a modest debiasing effect that recovers a meaningful share of bias-induced errors while preserving most correct answers across instruct families.
Source evidence

Abstract

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven BCT bias types, we extract a per-bias direction from hidden states and triangulate it through three measures: probing, leave-one-dataset-out transfer, and causal intervention. The susceptibility is largely installed by alignment tuning rather than pretraining: pretrained base models barely cave to these biases, and their activations carry no cue-specific signal beyond question content. Within aligned models, each bias becomes a single coherent direction that we can both decode and steer along, recovering the unbiased answer across every family we test. The biases stay representationally distinct, however: cross-bias entanglement is model-specific rather than a property of the bias category, and even behaviorally similar biases occupy different directions. The same intervention also serves as a modest debiasing tool, recovering a meaningful share of bias-induced errors while preserving most correct answers across all instruct families. Cue-induced bias is therefore best understood not as a single flaw in LLMs but as a family of distinct, causally active directions that alignment tuning installs.