ArXiv

ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors

Authors
Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho...
Categories
cs.CV
arXiv
https://arxiv.org/abs/2607.15246v1
PDF
https://arxiv.org/pdf/2607.15246v1

Brief

ARMOR++ is an agentic multi-agent framework that improves black-box, no-query transfer attacks on deepfake detectors by combining Qwen2.5-VL spatial priors and a Qwen3 orchestrator to mix five complementary primitives (dense, saliency, spatial, frequency, block) with adaptive hyperparameter reparameterization and entropy-regularized mixing. Evaluated on AADD-2025, it significantly outperforms prior agentic and non-agentic baselines, revealing a persistent reliability gap in deployed detectors.

Why it matters

ARMOR++ (2026-07-16) achieves a statistically confirmed, substantial gain in blind-target Attack Success Rate (ASR) over the state-of-the-art agentic baseline on the AADD-2025 benchmark, and also outperforms non-agentic baselines and defended detectors.

Key details

  • The framework uses Qwen2.5-VL (vision-language model) to provide spatial semantic priors and Qwen3 (LLM) to orchestrate primitive selection, adaptive hyperparameter reparameterization, and entropy-regularized perturbation mixing.
  • ARMOR++ composes five complementary primitives—dense optimization, saliency-based methods, spatial transformations, frequency-domain perturbations, and block-structured modifications—to improve black-box, no-query transferability from convolutional surrogates to transformer-based deepfake detectors across low- and high-quality images.
Source evidence

Abstract

The reliability of deepfake detectors frequently degrades under black-box adversarial transfer, as these models often rely on fragile, architecture-dependent forensic cues. Existing transfer attacks often lack semantic awareness and struggle to maintain effectiveness under strict no-query constraints, particularly when perturbations are transferred from convolutional surrogates to transformer-based targets. To address these limitations, this paper introduces ARMOR++, a robust multi-agent framework designed for high-transferability deepfake evasion. The framework leverages the Qwen2.5-VL Vision-Language Model (VLM) to supply spatial semantic priors, while the Qwen3 Large Language Model (LLM) orchestrates primitive selection, adaptive hyperparameter reparameterization, and entropy-regularized perturbation mixing. By integrating five complementary primitives, spanning dense optimization, saliency-based methods, spatial transformations, frequency-domain perturbations, and block-structured modifications, ARMOR++ effectively targets heterogeneous inductive biases. Rigorous evaluation on the AADD-2025 benchmark demonstrates that ARMOR++ significantly outperforms existing agentic and non-agentic baselines across both low- and high-quality image regimes. Statistical analysis confirms a substantial gain in blind-target Attack Success Rate (ASR) over the state-of-the-art agentic baseline, with further performance advantages evidenced against non-agentic benchmarks and under robust defensive configurations. These findings highlight a significant residual reliability gap in current deepfake detector deployments and demonstrate the efficacy of agentic orchestration in identifying latent vulnerabilities.