Twitter/X

MiniCPM-V 4.6 is a 1.3B-parameter multimodal model from OpenBMB that emphasizes…

Brief

MiniCPM-V 4.6 is a 1.3B-parameter multimodal model from OpenBMB that emphasizes edge/mobile practicality: it claims 75.7 ms TTFT, a 55% reduction in vision-encoding cost via LLaVA-UHD v4, and 3136x3136 image support. OpenBMB alleges it beats Gemma4-E2B-it and Qwen3.5-0.8B on benchmarks while using far fewer tokens and running faster on a single RTX 4090; code and demos are available on GitHub and Hugging Face.

Source evidence

75.7ms TTFT
55% lower vision encoding cost
3136² high-res support
1.3B params

MiniCPM-V 4.6 is pushing multimodal AI
toward something far more practical for edge devices and mobile hardware

Very impressive engineering work from OpenBMB

⭐ Star the repo & download the model:
GitHub: github.com/OpenBMB/MiniCPM-V
HF: huggingface.co/openbmb/MiniC…

OpenBMB (@OpenBMB)

1/5 MiniCPM-V 4.6 (1.3B) is now live 🚀🚀
High-res visual processing, optimized for consumer-grade and mobile hardware. We’ve leveraged the latest LLaVA-UHD v4 technique to cut vision encoding costs by 55%, enabling native edge deployment with extreme efficiency.
🔥 Beats Gemma4-E2B-it and Qwen3.5-0.8B across key multimodal and Artificial Analysis benchmarks — scoring higher than Qwen3.5-0.8B using just 2.5% of its token budget.
⚡ TTFT (75.7ms) 2.2x Faster than Qwen3.5-0.8B even with 3136² high-res images.
🏗️ ~1.5x Token Throughput compared with Qwen3.5-0.8B on a single RTX 4090.
Try the model here:
🤗 Hugging Face:
huggingface.co/openbmb/MiniC…
💻 GitHub:
github.com/OpenBMB/MiniCPM-V
🔭 Modelscope:
modelscope.cn/models/OpenBMB…
🌐 Web Demo:
huggingface.co/spaces/openbm…
📱 App Demo:
github.com/OpenBMB/MiniCPM-V…

Video

— https://nitter.net/OpenBMB/status/2053857238805311824#m