Project MonetRequest demo
Home/Blog/MiniMax H3 Max: API, Pricing, Features & Reference Video

AI · Project Monet Briefing

MiniMax H3 Max: Fast AI Video, API, Pricing and Reference-to-Video

H3 Max is fal Research's speed-optimized post-training of MiniMax H3, with 5–15 second 480p/768p video, synchronized audio and live text, image and reference-to-video routes.

Published 2026-09-07 · Updated 2026-09-07 · By Project Monet Editorial Team

MiniMax H3 Max fast AI video workflow from text image and references to synchronized video and audio

01

What is MiniMax H3 Max?

MiniMax H3 Max is a video-generation model that fal says it post-trained on top of the open-weight MiniMax H3 base model. fal describes the derivative as tuned for stronger prompt adherence and aesthetics while being co-optimized with its inference stack for higher throughput.

That attribution matters: H3 Max is not simply a faster hosted copy of standard H3, and it should not be described as a MiniMax-post-trained release. The base model comes from MiniMax; the H3 Max post-training and serving optimization described here come from fal.

02

Current H3 Max capabilities

  • Text-to-video
  • Image-to-video, including first/last-frame workflows
  • Reference-to-video with images, video and audio references
  • Synchronized audio generated with the video
  • 5–15 second clips
  • 480p and 768p output

fal's September 1 guide says the reference-to-video route is now live and can condition on up to 12 supplied reference files across the supported reference types. This is newer than launch copy that described reference video as a follow-up capability.

03

How fast is H3 Max?

fal reports that a five-second 768p clip can complete in under three seconds of backend inference on its optimized stack. Treat that as a vendor-reported systems result rather than a universal end-to-end latency guarantee.

MiniMax Design publishes a different product-level expectation: about 15 seconds for a five-second video and about 40 seconds for a 15-second video. Queueing, prompt expansion, uploads, orchestration, network transfer and service load can all make user-perceived latency longer than backend inference.

04

MiniMax H3 Max pricing

As rechecked on September 7, 2026, fal's current standard output pricing is $0.05 per generated second at 480p and $0.08 per generated second at 768p. A five-second clip therefore costs $0.25 or $0.40, and a 15-second clip costs $0.75 or $1.20, before any reference-input charges.

Reference-to-video has separate conditioning costs after a free token allowance. fal's current guide says the first 4,096 reference tokens are free and additional reference tokens are billed separately, so a reference workflow can cost more than the output-video price alone.

05

Where H3 Max is available

Developers can use H3 Max through fal's API/playground routes for text-to-video, image-to-video and reference-to-video. MiniMax Design also exposes H3 Max in a creator-facing workflow for text and image generation.

fal's current guide advertises a small daily free-generation allowance on its product experience. Free quotas are operational details and can change, so check the live product page before relying on them for a production workflow.

06

H3 Max vs standard MiniMax H3

H3 Max prioritizes rapid iteration at 480p/768p. Standard MiniMax H3 remains the broader choice when a workflow needs 2K output or instruction-based video editing.

Reference conditioning is no longer a reason by itself to choose standard H3, because fal now documents reference-to-video for H3 Max as well. Choose based on the output and workflow capabilities you actually need rather than assuming Max replaces the base model everywhere.

07

Where H3 Max fits best

The strongest fit is iteration-heavy short-form video: ad concepts, social clips, product motion tests, storyboard exploration, character or scene variations and applications where waiting for every generation slows the creative loop.

Synchronized audio is especially useful when a brief includes dialogue, ambience, effects or music and the team would otherwise need a separate audio-generation pass.

08

Important limitations

  • H3 Max currently tops out at 768p rather than standard H3's 2K option.
  • Separate calls are not guaranteed to preserve identity or scene continuity deterministically.
  • Reference conditioning helps but does not remove the need for careful prompting and review.
  • fal's speed and preference-evaluation claims are vendor-reported, not independent universal benchmarks.
  • Provider pricing and free allowances can change after publication.

09

MiniMax H3 Max FAQ

Who built H3 Max? MiniMax built the H3 base model; fal says it post-trained H3 Max on top of that open-weight base and optimized it with fal's inference stack.

Does H3 Max generate audio? Yes. fal's current documentation says synchronized audio is generated alongside the video.

Is reference-to-video live? Yes. fal now publishes a live reference-to-video route and current setup documentation.

What does it cost? The standard rates verified September 7 are $0.05/s at 480p and $0.08/s at 768p, with separate conditioning charges possible for reference inputs.

10

Continue reading

For implementation details, endpoint selection, queue handling, server-side authentication and cost controls, use the dedicated H3 Max API and pricing guide.

Sources

Primary and supporting sources

Facts were rechecked against the linked sources immediately before publication. Pricing, product availability and rollout status can change.

Project Monet

Useful signals. Clear decisions. Better digital work.

Project Monet turns relevant shifts in AI, creator tools and the web into practical context—and builds focused websites for businesses ready to grow.

Request a free homepage concept