ArXiv

Introduces a pixel-pair temporal warped flow field that generates corresponding…

Authors
Dingyun Zhang, Lixue Gong, Wei Liu
Categories
cs.CV
arXiv
https://arxiv.org/abs/2607.18227v1
PDF
https://arxiv.org/pdf/2607.18227v1

Brief

The paper addresses costly, mask-dependent video-editing data collection by introducing a pixel-pair temporal warped flow field that converts image editing samples into video editing samples in real time. Combined with modality-mimic generation/editing losses and sense-related tasks (referring expression segmentation) plus latent and attention-region losses, the approach trains video-editing models from generated data without external masks or auxiliary MLLMs, broadening scalable editing tasks.

Why it matters

Introduces a pixel-pair temporal warped flow field that generates corresponding video editing samples in real time from image editing samples, enabling training of video-editing models using only such generated data (Zhang, Gong, Liu; arXiv:2607.18227v1, 2026-07-20).

Key details

  • Proposes two losses—modality mimic generation loss and modality mimic editing loss—to align image and video modalities via mutual imitation, treating the image modality as a special case of video to equalize their output distributions.
  • Adds sense-related tasks (e.g., referring expression segmentation) plus editing-region-aware latent-level and attention-level losses so the model internalizes instruction-conditioned localization and region-only modification, removing reliance on external masks, MLLM fine-tuning, I2V pair synthesis, or ControlNet-like guidance.
Source evidence

Abstract

In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited task scalability. As a result, the diversity of editing tasks remains substantially narrower than that available for image editing models. We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate across multiple levels of video editing tasks that a model can learn video editing using only such data. We regard the image modality as a particular form of the video modality. Accordingly, we design a modality mimic generation loss and a modality mimic editing loss to relatively align the capabilities--and thereby the output distributions--of the two modalities through mutual imitation. Moreover, language-based visual editing entails the comprehension of the editing instruction and the reference visual content, the localization of the region corresponding to that instruction within the reference visual contents, and the modification of that region alone. Existing approaches predominantly rely on external aids, such as fine-tuning an additional MLLM or explicitly supplying a mask sequence as auxiliary input during inference. In contrast, we aspire for the model to internalize this capability. To that end, we introduce sense-related tasks--for instance, referring expression segmentation--along with corresponding editing-region-aware latent-level loss and attention-level loss.

Comment: Due to file size constraints, the figures in the arXiv file have been heavily lossy-compressed. Please visit the uncompressed file at: https://huggingface.co/datasets/FlowMimic/Uncompressed/blob/main/main.pdf