Skip to main content
Back to Blog
MiniMax H3Hailuo 3.0AI video generationopen weightsSeedance 2.5Kling 3.0

MiniMax H3: The Open-Weights Video Model That Just Undercut Everyone by 3x

ghosty
Founder, SaaSCity
MiniMax H3: The Open-Weights Video Model That Just Undercut Everyone by 3x

A Chinese startup just shipped the #1-ranked video editing model on Artificial Analysis, priced it at roughly a third of what its closest rivals charge, and announced it's giving away the weights.

That's the whole story in one sentence. The details are more interesting.

MiniMax released H3 on July 31, 2026 — a general-purpose omni-modal generation model that takes text, images, video, and audio in a single context and returns 2K video with native stereo sound. The Hong Kong market noticed immediately: shares of MiniMax (0100.HK) climbed roughly 13% by midday, touching HK$231.6 intraday, with nearly HK$1 billion in turnover on the session.

But the stock move isn't the point. The point is that video generation has been the last major modality where closed models had a comfortable moat, and MiniMax just announced it's walking away from that moat on purpose.


What MiniMax H3 Actually Does

H3 is the successor to the Hailuo line (Hailuo 01 → Hailuo 02 → Hailuo 2.3 → H3, also branded Hailuo 3.0). Shanghai-based MiniMax, founded in 2022 and one of China's so-called "AI tigers," listed in Hong Kong in January — the second of that group to go public after Z.AI. It's the same lab behind MiniMax M3, the open-weight coding model that landed earlier this year.

Here's what H3 ships with:

  • Native 2K output (2560×1440) at 24fps, offered by default rather than as a premium upsell
  • 5 to 15 second clips
  • Native stereo audio generated jointly with the video — dialogue, score, foley, room tone, all timed to the cut
  • Multimodal input in one call: up to 9 reference images, 3 video clips, and 3 audio tracks, capped at 12 files total
  • Prompts up to 7,000 characters, so a full shot list fits in a single request
  • Instruction-based video editing — swap a product, rewrite signage, replace a line of dialogue, relight day to night
  • V2V motion transfer, first-and-last-frame control, and reference-to-video

The commercial targeting is explicit. MiniMax is aiming this at advertising, e-commerce, branding, product design, UI/UX, and gaming — not at people making surreal TikToks. Film title sequences, game menu animations, motion posters, product reveals.


The Omni-Reference Trick Is the Real Unlock

Most video pipelines today are Frankenstein jobs. You generate a clip with one model, clone a voice with another, add music with a third, then stitch it in an editor. Every handoff loses fidelity and adds latency.

H3 collapses that into one call. The example MiniMax used in its own launch post says everything about the design intent:

"Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3."

You describe the relationship between your references and the output in plain language. No control nets, no separate conditioning modules, no per-asset config. The model reads identity from one input, camera movement from another, and vocal performance from a third, then carries all three through to one coherent result.

MiniMax's framing is that prior generative models were fragmented by task — separate expert models for text-to-image, editing, subject reference, style transfer — and that fragmentation was an artifact of how the field trained models, not a law of physics. H3 folds text-to-image, text-to-video, image-to-video, multi-shot, editing, and audio into one pretraining objective.


Under the Hood: Four Pieces Worth Knowing

MiniMax named four technologies in its launch post. The full technical report is "coming soon," so treat these as company claims until it lands. Still, they're specific enough to be useful.

Contextual Omni Representation. This is the data-and-annotation layer. It compresses multimodal inputs from an average of roughly 100,000 tokens down to about 4,000, with language acting as the bridge that makes heterogeneous inputs mutually intelligible. Language as the generalization substrate — same bet the LLM world made, applied to pixels and waveforms.

H3-VAE. A full tokenizer rewrite with a much higher compression rate, which MiniMax says bought a 4× effective sequence length gain. That's the load-bearing piece behind offering 2K by default: without it, native 2K at these prices doesn't pencil out.

H3-Omni Transformer. A training architecture that separates understanding and generation workloads and tunes hardware utilization for each, while balancing per-sample heterogeneous compute against cross-sample load. End-to-end training throughput went up nearly 30%. Notably, MiniMax says it threw out the Hailuo-02 architecture entirely — a team abandoning its own proven design because clever task-specific tricks don't generalize.

In-Context Regeneration. Instead of bolting on a dedicated super-resolution module for 2K, the base model regenerates its own low-res output in-context. Two consequences: it reuses generative capability the model already has, and it can look back at the original multimodal context during upscaling. Traditional super-res has to guess at small text and fine detail. This doesn't.

One more thing that matters strategically: MiniMax says chip compatibility was a design constraint from the earliest stages, with H3 built to run on several Chinese-made accelerators. In a market shaped by export controls, that's not a footnote.


Benchmarks: Where H3 Actually Lands

Artificial Analysis has H3 at #1 in Video Editing, #2 in Text-to-Video, and #3 in Image-to-Video on its leaderboards, with the evaluation run on the 2K tier.

The video editing crown is the least surprising and most defensible result. Video input support is rare — most models can't take an existing clip and apply instruction-based edits at all — so the competitive field there is thin. H3 owning it says less about raw generation quality than about the fact that it's playing a game few others have entered.

The top-3 finishes in T2V and I2V are the ones that should make competitors uncomfortable, because those are crowded categories.

Artificial Analysis also noted that if the weights land as promised, H3 would become the strongest open-weights video model by a wide margin — well ahead of the previous open leader, LTX-2.3.


MiniMax H3 Pricing

TierPer secondPer minuteStatus
2K (default)$0.13$7.80Live
768p$0.09$5.40Closed beta / listed as coming soon

First five reference images are free; each additional image runs $0.04. Reference audio is free. One gotcha worth budgeting for: reference video seconds are added to your billed output duration — a 5-second clip with a 3-second reference video bills as 8 seconds.

The comparison Artificial Analysis ran on per-minute cost with audio:

  • MiniMax H3 (2K): $7.80/min
  • HappyHorse-1.1: $9.90/min
  • Kling 3.0 (1080p): $20.16/min
  • Dreamina Seedance 2.0 (1080p): $22.45/min
  • Google Gemini Omni Flash: $6.00/min

So H3 is roughly a third the cost of Seedance 2.0 and Kling 3.0 at comparable-or-better resolution — MiniMax's "less than one-third of mainstream models" claim holds up against those two. It's not the cheapest thing available; Gemini Omni Flash still undercuts it. But H3 is delivering 2K where those rivals are delivering 1080p.

H3 is live now in the Hailuo AI app and via the MiniMax API under the model ID MiniMax-H3. Day-zero integrations landed across fal.ai, Atlas Cloud, PixVerse, EvoLink, OpenRouter, Topview, Leonardo.Ai, and others — an unusually broad simultaneous rollout that suggests partners had access well before Friday.


Launch Your Video Tool on SaaSCity

A 3x price cut is a product opportunity. Every video SaaS that was uneconomical at $22/minute becomes viable at $7.80 — and the ones getting built this quarter will be built on H3.

SaaSCity is a free startup directory where your product gets a building on an interactive 3D city map instead of a row in a spreadsheet.

  • 🆓 Free to list: submit in under 2 minutes, no credit card
  • 📈 Dofollow backlink that compounds your domain rating
  • 🗺️ 3D map visibility — a permanent, indexed page for your tool
  • 🎯 Builder audience: founders and engineers picking their video stack right now

Submit your product for free →


The Open-Weights Question

MiniMax says it plans to publish H3's weights "in the coming days, subject to applicable laws and regulations." Reuters confirmed the timeline. Artificial Analysis reports the release will land under the MiniMax Community License, which permits commercial use for organizations under $20M in revenue, with prominent attribution required.

That's not OSI-approved open source. It's a source-available commercial license with a revenue ceiling — the same shape Meta used for Llama. For an indie studio or a Series A startup, it's functionally free. For a company doing $50M in ad production, it isn't. (Compare Ideogram 4's release, where the license terms shaped who could actually build on it far more than the benchmarks did.)

As of publication, the weights are not downloadable. Community reports point to early August. Until files are on ModelScope or Hugging Face, this is a promise, not a product.

The other unanswered question is size. MiniMax hasn't disclosed a parameter count, and the top-voted reply under its launch tweet was a creator asking whether they'd need a supercomputer to run it. That answer determines whether "open weights" means "the community can fine-tune this" or "three labs with H100 clusters can fine-tune this."


H3 vs. Seedance 2.0 and Kling 3.0

The China video-model race got loud fast this year. ByteDance's Seedance 2.0 drew serious attention for combining text, image, audio, and video inputs. Kuaishou shipped Kling 3.0. Seedance 2.5 reached consumers on Jimeng the same day H3 launched — but its API doesn't open until August 7.

That week-long gap is the whole competitive story right now. Seedance 2.5 promises up to 30 seconds and a larger reference budget. If you're shipping before August 7, none of that is callable from code, and H3 is.

Where H3 wins: cost per second at 2K, video editing, native audio in one pass, and — pending release — weights you can run yourself.

Where it doesn't: 15 seconds is a hard ceiling. If your brief needs a single continuous scene longer than that, H3 can't do it, and no amount of price advantage fixes that. Seedance 2.5's longer unit is a real differentiator for narrative work.


What to Watch Next

Three things will tell you whether the launch-day narrative survives contact with reality.

The weights. Date, file size, parameter count, license text. Everything about H3's ecosystem impact depends on this.

The technical report. Contextual Omni Representation and In-Context Regeneration are interesting claims. Reproducible numbers would make them important claims.

Independent side-by-sides. Vendor showcase clips are marketing. The comparisons worth reading are prompt-matched runs from creators with access to H3, Seedance 2.5, and Kling 3.0 — and those are only starting to appear now.


Quick Answers

Is MiniMax H3 open source? Not yet, and "open source" is the wrong word. MiniMax has committed to publishing downloadable weights within days under the MiniMax Community License — commercial use allowed for organizations under $20M in revenue, with attribution. That's source-available, not OSI open source. As of July 31 the files aren't public.

What's the maximum clip length? 15 seconds, with a 5-second floor. There's no stitching mode that extends a single continuous shot beyond that.

Does it generate audio? Yes, natively and in the same pass — stereo, including dialogue, music, foley, and ambience synced to the cut. You can also hand it a reference recording and have it transfer that voice onto your character.

What does a 15-second 2K clip cost? About $1.95 at the listed $0.13/second, before any reference-video seconds get added to the bill.

Is Hailuo 3.0 the same thing as H3? Yes. H3 is the model name; Hailuo is the consumer product brand it ships under. You'll see both, plus MiniMax-H3 as the API model ID.

Can it edit video I already have? That's its strongest category. Instruction-based edits — object swaps, signage rewrites, relighting, dialogue replacement — with the rest of the shot held stable. It's #1 on Artificial Analysis for video editing.

What hardware will I need to run the weights locally? Unknown. MiniMax hasn't published a parameter count. Wait for the model card.


The Takeaway

Video generation spent two years as the most expensive, most closed, most vertically-integrated corner of generative AI. H3 attacks all three at once: it's cheaper by a multiple, it's about to be downloadable, and it replaces a four-tool pipeline with a single API call.

Whether it holds the technical lead is genuinely unclear — Seedance 2.5's API opens in a week, and one week is a long time in this market. What's clearer is that the pricing floor and the openness floor both just moved, and moved in the direction that favors whoever is building on top rather than whoever is renting out capacity. It's the same move DeepSeek made on the text side the same day.

If you're running a video pipeline right now, the useful next step isn't reading more launch coverage. It's a matched test: same brief, same references, H3 against whatever you're currently paying for. At $7.80 a minute, that experiment costs less than lunch.

What would you actually build if state-of-the-art video generation cost you nothing but electricity? That question stops being hypothetical the day those weights go live.


Advertise Your Startup on SaaSCity

Already shipping something in this space — a video production micro-SaaS, an editing tool, a generation wrapper? Don't let it launch into silence.

SaaSCity gets your product a building on the map, a permanent indexed page, and a launch slot in front of builders instead of an archive nobody reads.

Submit your startup → · Browse the directory →


Specs, pricing, and benchmark placements verified against MiniMax's July 31, 2026 launch materials, the MiniMax API pricing page, Artificial Analysis leaderboards, and Reuters reporting, as of July 31, 2026. Weight availability and license terms are company statements, not shipped artifacts — confirm before you build on them.

SaaSCity.io covers AI model launches and the tooling decisions behind them. Explore the directory or list your own product.

Get your SaaS in front of founders

List your product on the SaaSCity live city map - a permanent listing, real discovery, and a backlink from a high-DR directory. Free to start; upgrade for a dofollow link and a building on the map.