MiniMax H3 just became the best video editor in AI.
1 on Artificial Analysis's Video Editing Leaderboard, #2 in Text to Video, #3 in Image to Video. And unlike every model above it, H3 can take a video as input, not just generate one from scratch.
That changes what video AI actually means in production.
Right now, the tools people use to package video for advertising and brand content are Adobe After Effects, DaVinci Resolve, and manual compositing workflows. They're powerful, but they require technical operators and hours of iteration to get timing, text overlays, and motion graphics aligned.
H3 takes text, images, existing video clips, and audio together as inputs, generates 5 to 15 second clips at 24fps with native audio embedded. The instruction-based editing is what makes it different. You're not starting from scratch with a prompt. You're directing an existing clip.
For creative fields specifically, this is where it gets interesting.
Advertising teams spend enormous effort on visual packaging: supers (on-screen text), brand color grading, transitions, and audio sync. These are all instructable edits, and H3's multimodal architecture is built to handle exactly that layer. The model has a strong enough aesthetic sense to respond to directorial intent, not just literal prompts.
Pricing is also positioned to win. At $7.80 per minute for 2K video with audio, H3 significantly undercuts Kling 3.0 ($20.16/min) and Dreamina Seedance 2.0 ($22.45/min).
And MiniMax plans to release the weights under a Community License that permits commercial use for companies under $20M revenue, with attribution. If that ships, H3 becomes the strongest open-weights video model by a wide margin, ahead of LTX-2.3 which held that position before.
Most video AI is still about generation. H3 is the first model that starts to feel like post-production.
Video