Technical Specifications of MiniMax H3 Max
| Specification | MiniMax H3 Max |
|---|---|
| Model ID | minimax-h3-max |
| Model family | MiniMax H3 |
| Model type | Generative video model |
| Model origin | Post-trained from MiniMax H3 open weights |
| Post-training | fal Research |
| Primary optimization | Prompt adherence, aesthetics, inference throughput |
| Input | Text, image; reference-to-video support is available through the H3 Max family |
| Output | Video with synchronized/native audio |
| Duration | 5–15 seconds |
| Resolution | 480p and 768p; 1080p refinement availability depends on the current endpoint |
| Frame rate | 24 FPS |
| Audio | Synchronized stereo audio |
| Text-to-video | Yes |
| Image-to-video | Yes |
| First-to-last frame | Supported through image-to-video with an end frame |
| Reference-to-video | Available through the H3 Max family/API |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Main API endpoints | api.cometapi.com/v1/videos |
| Base model | MiniMax H3 |
MiniMax H3 Max is a post-trained version of MiniMax H3 rather than a separately developed successor with a larger capability ceiling. fal Research states that it post-trained H3's open weights with additional data focused on prompt adherence and aesthetics, while its inference team co-designed the serving stack around the resulting model.
The distinction matters when comparing H3 Max with standard MiniMax H3. H3 itself is a general-purpose omni-modal video system supporting text, image, video, and audio understanding and video generation with native stereo audio. MiniMax's open-source H3 documentation specifies output durations of 4–15 seconds and resolutions up to 2K through its H3 variants.
H3 Max instead focuses on a different operating point: faster generation, strong prompt following, and visual quality at 480p/768p output. fal reports that a 5-second 768p clip can be generated in under three seconds on its infrastructure.
What is MiniMax H3 Max?
MiniMax H3 Max is an AI video generation model created through post-training of the open-weight MiniMax H3 model. fal Research describes the model as specifically tuned for stronger prompt adherence and aesthetics, while its inference stack is optimized for high throughput.
The model supports text-to-video and image-to-video generation, with the image-to-video endpoint also supporting first-to-last-frame generation when an end image is supplied. The H3 Max API family additionally exposes reference-to-video functionality.
Its most distinctive characteristic is therefore not a larger parameter count or a longer context window. It is the combination of video quality, prompt adherence, and low generation latency.
Main Features of MiniMax H3 Max
- Post-trained MiniMax H3 foundation: H3 Max is derived from MiniMax H3's open weights rather than being an unrelated video model. fal Research added post-training data targeted at prompt adherence and aesthetics.
- Fast video generation: fal reports generating a 5-second 768p clip in approximately 2.78 seconds on its infrastructure, substantially reducing the waiting time associated with iterative video generation.
- Strong prompt adherence: The post-training process specifically targets better adherence to detailed instructions, including scene composition, actions, camera movement, and visual style.
- Text-to-video and image-to-video: Developers can create video directly from a text prompt or use an image as the starting frame. The image endpoint can also accept an end frame for first-to-last-frame generation.
- Native synchronized audio: H3 Max generates video with synchronized audio, making it suitable for short-form scenes that require dialogue, environmental sound, or audiovisual coordination.
- Multiple aspect ratios: The model supports common landscape, portrait, square, and cinematic formats, including 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.
Benchmark Performance of MiniMax H3 Max
| Evaluation | MiniMax H3 Max Result | Evaluation Type | Sample Size / Confidence | What It Measures |
|---|---|---|---|---|
| Design Arena – Image-to-Video | 1,341 Elo | Human preference Elo | Arena-based comparison | Overall user preference for generated video quality |
| Artificial Analysis – Image-to-Video + Audio | 1,201 Elo | Human preference Elo | 2,177 samples; ±11 at 95% CI | Preference for image-to-video generation with audio |
| fal Human Preference Evaluation – Overall Quality | #1 among evaluated models | Head-to-head human evaluation | Compared with 12 leading video models | Overall visual quality |
| fal Human Preference Evaluation – Prompt Understanding | #1 among evaluated models | Head-to-head human evaluation | Compared with 12 leading video models | How accurately the generated video follows the prompt |
| fal Human Preference Evaluation – Aesthetics | #1 among evaluated models | Head-to-head human evaluation | Compared with 12 leading video models | Visual aesthetics and perceived quality |
| Generation Speed – 5-second 768p video | <3 seconds | Inference-speed measurement | fal infrastructure | End-to-end generation speed |
| Generation Speed vs Official MiniMax H3 | ~35× higher throughput | Provider-reported throughput comparison | fal infrastructure | Relative inference throughput |
The strongest quantitative public results are the 1,341 Elo Design Arena score and the 1,201 Elo Artificial Analysis score. However, both are human-preference Elo measurements, not conventional accuracy benchmarks. The Artificial Analysis result reports a 95% confidence interval of ±11 over 2,177 samples.
For CometAPI content, the safest formulation is therefore to distinguish:
Capability evidence: prompt adherence, aesthetics, audiovisual generation and image/video conditioning.
Performance evidence: provider-reported generation latency and third-party preference evaluations.
Avoid presenting "#1" human-preference result as an objective overall leaderboard position.
MiniMax H3 Max vs MiniMax H3
| Capability | MiniMax H3 Max | MiniMax H3 |
|---|---|---|
| Origin | H3 open weights + fal post-training | Original MiniMax H3 |
| Primary focus | Prompt adherence + aesthetics + speed | General-purpose omni-modal video generation |
| Typical output duration | 5–15 sec | 4–15 sec |
| Standard output | 480p / 768p | Up to 2K through documented H3 variants |
| Native audio | Yes | Yes |
| Text-to-video | Yes | Yes |
| Image-to-video | Yes | Yes |
| First/last-frame workflow | Supported | Supported |
| Reference workflows | H3 Max reference endpoint available | Broad reference-to-video capabilities |
| Generation speed | Optimized for high throughput | Standard H3 inference |
| Best documented distinction | Fast iterative generation | Broader capability/resolution ceiling |
The most important difference is operating point rather than model generation. Standard H3 is designed as a broader omni-modal system and supports output up to 2K in MiniMax's documentation, whereas H3 Max was developed around faster generation and prompt adherence at its supported output resolutions.
Representative Use Cases of MiniMax H3 Max
Rapid creative iteration
H3 Max is well suited to workflows where creators repeatedly generate, inspect, and revise short clips. Lower latency makes it practical to test several prompt variations rather than waiting several minutes for each iteration.
Social media video
Its short-duration format, portrait aspect ratio support, synchronized audio, and fast generation make H3 Max suitable for short-form social content and advertising concepts.
Product visualization
Image-to-video generation can animate product stills into short promotional sequences while allowing the source image to guide composition and appearance.
Cinematic concept development
Detailed prompts describing camera movement, subject action, environment, lighting, and visual style can be used to prototype cinematic shots before committing to more expensive production workflows.
First-to-last-frame transitions
When an opening and ending image are supplied, H3 Max can be used for controlled transition sequences rather than relying entirely on unconstrained text-to-video generation.
How to Access MiniMax H3 Max API with CometAPI
CometAPI can be used as a unified API layer for integrating supported AI models without maintaining separate provider-specific integrations for every model.
Step 1: Create a CometAPI API key
Sign in to CometAPI and create an API token from the developer console.
Step 2: Select minimax-h3-max
Use minimax-h3-max as the model identifier when the model is available in your CometAPI account.
Step 3: Send a video-generation request
Submit your prompt and the supported generation parameters through CometAPI's video API interface. Depending on the currently exposed CometAPI schema, parameters may include duration, resolution, aspect ratio, image inputs, and other video-generation controls.
Before publishing an integration example, verify the current CometAPI model page and API documentation for the exact endpoint, parameter names, supported resolutions, and asynchronous task-handling behavior. These details can change independently of the underlying H3 Max model.