ArXiv

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Authors
Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar...
Categories
cs.CV
arXiv
https://arxiv.org/abs/2606.27376v1
PDF
https://arxiv.org/pdf/2606.27376v1

Brief

The paper presents Ask–Solve–Generate, a self-evolving framework that lets unified large multimodal models improve vision understanding and image generation using only unlabeled images. A Proposer, Solver and Generator interact via self-consistency rewards; Solver Token Entropy (STE) stabilizes training, and a multi-scale QA+cycle-caption evaluator assesses generations. The approach yields consistent improvements (e.g., BAGEL +3.5% MMMU, GenEval 82%→85%) and generalizes across BLIP3o, BAGEL, and VARGPT-v1.1. Full text and code were released on arXiv.

Why it matters

Proposes a self-evolving training framework (published 2026-06-25) that decomposes a unified LMM into three internal roles—Proposer (generates visual questions), Solver (answers and evaluates), and Generator (synthesizes images)—and trains using only unlabeled images with no human annotations, preference labels, or external reward/judge models.

Key details

  • Introduces Solver Token Entropy (STE), a continuous token-level difficulty signal to stabilize learning when sample-level consistency is unreliable, and a multi-scale internal evaluation for generation combining question–answer fidelity with cycle-consistent captioning to couple understanding and generation.
  • Reports consistent gains across eight understanding metrics and across BLIP3o, rectified-flow BAGEL, and autoregressive VARGPT-v1.1; on BAGEL achieves +3.5% absolute on MMMU and improves GenEval image-generation score from 82% to 85%. Code and models are publicly released.
Source evidence

Abstract

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. We ask whether a unified LMM can improve both abilities autonomously using only unlabeled images. We propose a self-evolving training framework with three internal roles: a Proposer that generates visual questions, a Solver that answers and evaluates them, and a Generator that synthesizes images. Training uses only self-derived consistency signals, without human annotations, preference labels, or task-trained external reward/judge models. To stabilize learning, we introduce Solver Token Entropy (STE), a continuous difficulty signal based on token-level prediction uncertainty that remains useful even when sample-level consistency becomes unreliable. For image generation, we design a multi-scale internal evaluation scheme that combines question-answer fidelity scoring with cycle-consistent captioning. This creates a solver-mediated coupling, where better visual understanding enables more reliable generation assessment and stronger internal training signals. The framework preserves the same role decomposition, reward logic, and training schedule across diffusion-based BLIP3o, rectified-flow BAGEL, and autoregressive VARGPT-v1.1 architectures, requiring only each backbone's native prompting and generation interface. Across eight understanding metrics, our method consistently improves over the corresponding base models. On BAGEL, it achieves a $+3.5\%$ absolute gain on MMMU and improves GenEval image generation performance from $82\%$ to $85\%$. Code and models are publicly released.