Twitter/X

Cambrian-P (announced 2026-05-27) is the newest model in the Cambrian series; the…

Brief

sainingxie introduced Cambrian-P (May 27, 2026), arguing that camera pose is an easy-to-obtain, minimal sufficient 3D signal that—when modeled jointly with frames—provides global grounding for video multimodal models. Jihan Yang echoes this, saying camera pose fills a missing piece that lets MLLMs move beyond activity recognition to understand ego/object dynamics and scene geometry.

Why it matters

Cambrian-P (announced 2026-05-27) is the newest model in the Cambrian series; the 'P' stands for pose and the author claims pose is the minimal sufficient 3D signal needed for robust video multimodal models.

Key details

  • The approach jointly models raw video frames and camera/pose data to turn image sequences into a globally grounded structure, addressing gaps in MLLMs' handling of scene geometry, ego-motion, and object dynamics noted by Jihan Yang.
Source evidence

📸latest in our cambrian series: cambrian-p, p for pose.
i think pose is probably the minimal sufficient 3d signal (and it’s easy to get!) that we need for robust video multimodal models -- jointly modeling frames and pose turns image sequences into a globally grounded structure.

Jihan Yang (@jihanyang13)

Camera pose matters for video understanding!

Today's MLLMs excel at recognizing activities, but still struggle with the underlying space and ego/object dynamics in video. We trace this gap to a missing piece: camera pose.

Introducing Cambrian-P: a multimodal LLM natively grounded in camera pose. (1/n)

— https://nitter.net/jihanyang13/status/2059413001778720870#m