ArXiv

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

Authors
Yongli Xiang, Zhifang Zhang, Bojun Yang...
Categories
cs.CR, cs.CL, cs.CY
arXiv
https://arxiv.org/abs/2608.03700v1
PDF
https://arxiv.org/pdf/2608.03700v1

Brief

AntiSkillBench studies safety risks introduced when personal interaction histories are distilled into reusable persona skills. The benchmark bundles 7,500 persona-grounded dialogues from 50 profiles, an evaluation suite for privacy leakage, attribute disclosure, and behavioral impersonation across three distillation strategies, and four defense configurations. Results on three frontier agents show persistent leakage—from explicit attributes to communication style—and that current defenses are limited and dependent on the distillation protocol, making AntiSkillBench a challenging testbed for privacy-preserving, authenticity-aware persona skills.

Why it matters

AntiSkillBench provides a benchmark dataset of 7,500 persona-grounded dialogue traces derived from 50 behaviorally rich profiles and an evaluation suite measuring skill-level privacy leakage, agent-level attribute disclosure, and behavioral impersonation across three skill-distillation strategies.

Key details

  • Experiments on three frontier agents show persona-skill risks persist across model backbones and distillation protocols—leaking explicit attributes, communication styles, and personality traits—and four tested defense configurations (online vs post-hoc; active risk suppression vs passive provenance protection) exhibit limited, distillation-dependent effectiveness and fail to generalize.
Source evidence

Abstract

Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.

Comment: Project page: https://yonglixiang.github.io/AntiSkillBench