GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now live on CometAPI →
new/CometAPI research

What is MiniMax M3.1-Flash-Preview

MiniMax M3.1-Flash is a native multimodal, 1M-context Frontier Coding model with tunable thinking depth

CometAPI
AnnaAI model and API research team
Updated Sep 28, 2026 7 min read
What is MiniMax M3.1-Flash-Preview
Use this pattern

Make the first API call.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_COMETAPI_KEY",
    base_url="https://api.cometapi.com/v1",
)

response = client.chat.completions.create(
    model="gpt-5-mini",
    messages=[{"role": "user", "content": "Build this workflow."}],
)

print(response.choices[0].message.content)

TLDR On September 27, 2026, MiniMax quietly but officially launched M3.1-Flash-Preview inside MiniMax Code (and via Token Plan). It is a frontier multimodal coding model with a full 1-million-token context window, native support for text/image/video input, and a new five-level reasoning-effort control (low → max). Positioned explicitly for everyday development workflows—bug fixing through implementation, testing, and delivery—it prioritizes speed and cost-efficiency while keeping thinking always on.

Availability is currently limited to Token Plan + MiniMax Code (with API support via Subscription Key); full public pricing and open weights are not yet published. CometAPI already surfaces the model (minimax-m3.1-flash-preview) at competitive rates, making unified access straightforward.

Key Takeaways

  • 1M-token context + native multimodality — Handles large codebases, long agent sessions, documents, images, and video in one window.
  • Five-level reasoning effort (low / medium / high / xhigh / max, default max) — Trade latency and token usage for deeper deliberation on demand. Thinking cannot be disabled.
  • Everyday coding focus — Explicitly optimized for the closed loop of bug localization → implementation → testing → delivery inside MiniMax Code.
  • Preview limitations — Confirmed in MiniMax Code and Token Plan; API calls require a Subscription Key. No independent model card, public benchmarks, open weights, or standalone pay-as-you-go pricing announced at launch.
  • CometAPI recommendation — Access via a single OpenAI-compatible endpoint (model ID: minimax-m3.1-flash-preview) with discounted pricing, ideal for teams already using CometAPI’s 500+ model catalog.

What Is MiniMax M3.1-Flash-Preview?

MiniMax M3.1-Flash-Preview is the latest entry in MiniMax’s M-series of large language models. Official documentation describes it as a “Frontier multimodal coding model with 1M context window and tunable thinking depth.” It was first spotted by developers in the MiniMax Code model picker and formally confirmed by the official @MiniMaxAgent account on X on September 27, 2026:

“MiniMax’s latest text model, M3.1-Flash-Preview, debuts today on MiniMax Code. Built for everyday development, it’s fast, reliable, and ready for real work, from quick bug fixes to full features.”

Unlike a general chat model, it is purpose-built for software engineering and agentic workflows. MiniMax’s product messaging emphasizes real-world developer tasks: quick bug fixes, feature implementation, boundary-case handling, regression testing, and change verification. The “Flash” designation signals a focus on speed and everyday practicality relative to the heavier MiniMax-M3 flagship, while still inheriting the large context and coding strengths of the M3 family.

The model supports:

  • Up to 1,000,000 tokens of context — suitable for entire codebases, long documents, and multi-step agent sessions.
  • Native multimodal inputs (text, images, and video).
  • Agentic reasoning, tool calling, and structured task execution.
  • Always-on deep thinking that can be tuned via a five-level effort slider.

Thinking content is returned separately (as reasoning_content in OpenAI-compatible mode or thinking blocks in Anthropic-compatible mode), keeping the final answer clean for display or downstream use.

1M-Token Context + Coding Orientation

The standout technical specification is the 1,000,000-token context window. This matches MiniMax-M3 and is far larger than most competing coding models. Official docs highlight its suitability for:

  • Entire code repositories
  • Long multi-step agent sessions
  • Large documents and knowledge bases
  • Multimodal inputs (text + images + video) within the same context

This long-context capability, combined with native multimodality, enables workflows that previously required heavy context management or external retrieval systems. Developers can keep large codebases, design mockups, error screenshots, and related documentation live in a single conversation.

The model is explicitly optimized for Agent and coding scenarios: tool calling, structured output, and multi-turn engineering tasks. Output-speed figures for related models (M3 ≈ 100+ TPS, M2.7-highspeed ≈ 100 TPS) suggest Flash variants target the higher end of the latency/throughput spectrum, though exact TPS numbers for M3.1-Flash-Preview itself have not been published.

Five-Level Reasoning Effort

One of the most practical new features is the tunable thinking depth controlled by the effort (Anthropic-compatible) or reasoning_effort (OpenAI-compatible) parameter. The five levels are:

LevelTypical Use CaseTrade-off
lowSimple syntax fixes, quick lookupsLowest latency & token usage
mediumRoutine refactors, small featuresBalanced
highComplex logic, multi-file changesDeeper analysis
xhighHard debugging, architectural decisionsHigher compute & latency
maxHighest-stakes or ambiguous problems (default)Maximum deliberation

Thinking is always on for M3.1-Flash-Preview; attempting to disable it returns a 400 error. Higher effort levels produce more internal reasoning tokens and increase latency, giving developers fine-grained control that was less granular on earlier M-series models.

In MiniMax Code the control appears as a simple slider or selector; via API it is passed in the request body. This design mirrors the growing industry trend of exposing “reasoning budget” knobs so users can match model behavior to task difficulty and cost constraints.

Everyday Development Workflow

MiniMax positions M3.1-Flash-Preview squarely inside the daily developer loop rather than as a pure research or long-horizon agent model. The official messaging and product placement emphasize a closed-loop experience inside MiniMax Code, Bug Fixing → Implementation → Testing → Delivery:

  1. Bug localization — Paste stack traces, error screenshots, or relevant code sections; the model can reason over large context to identify root causes.
  2. Implementation — Generate or edit code across multiple files while retaining project-wide context.
  3. Testing & verification — Suggest or write tests, run regression checks, and analyze impact.
  4. Delivery — Produce clean diffs, commit messages, or documentation ready for review.

The combination of 1M context, multimodal input (e.g., screenshots of failing UIs), and adjustable effort makes the model particularly effective for the high-frequency, latency-sensitive tasks that dominate most engineering days. Limited-time promotions accompanying the launch (double daily check-in credits September 28–October 7 UTC+8 and Token Plan quota resets) further encourage real-world testing.

Current Availability and Preview Limitations

Confirmed as of late September 2026:

  • Selectable inside MiniMax Code (model picker alongside M3, M2.7, etc.).
  • Available via Token Plan / Subscription Key.
  • Documented in official MiniMax API reference pages with OpenAI-compatible and Anthropic-compatible examples.
  • Thinking depth controllable via reasoning_effort or equivalent fields.
  • Promotional activity: Token Plan quota resets and double check-in points (September 28 – October 7) usable on the new model.

Confirmed API path: Official docs list the model name MiniMax-M3.1-Flash-Preview and provide both Anthropic-compatible (https://api.minimax.io/anthropic) and OpenAI-compatible (https://api.minimax.io/v1) endpoints. A Subscription Key (Token Plan) is required. Example parameters include output_config.effort or reasoning_effort. Thinking content is returned separately (thinking blocks or reasoning_content).

Still limited or unconfirmed:

  • Full independent public benchmark suite and model card with parameter counts.
  • Standalone pay-as-you-go pricing page with the same transparency as MiniMax-M3.
  • Open weights (not announced for this Preview).

Preview status means features, quotas, and pricing can still change. Developers should treat current performance as indicative rather than final.

How to Access MiniMax M3.1-Flash-Preview

1. Inside MiniMax Code (easiest for interactive use)

Download or open MiniMax Code (desktop apps available for macOS and Windows). Select M3.1-Flash-Preview from the model picker and adjust the reasoning level. Daily check-in promotions can double free credits through early October 2026.

2. Official MiniMax API (Token Plan)

  • Obtain a Subscription Key from the MiniMax console.
  • Use either the Anthropic or OpenAI SDK by changing base_url and setting model="MiniMax-M3.1-Flash-Preview".
  • Pass the desired effort level.

CometAPI already lists minimax-m3.1-flash-preview in its catalog (with 20% off pricing visible at the time of writing). Because CometAPI exposes an OpenAI-compatible endpoint (https://api.cometapi.com/v1), existing codebases require only a change of base_url and API key. This is especially convenient for teams that already route GPT, Claude, Gemini, MiniMax-M3, and other models through a single key and want consistent observability, billing, and fallback logic.

Example (Python / OpenAI SDK via CometAPI):

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_COMETAPI_KEY",
    base_url="https://api.cometapi.com/v1",
)

response = client.chat.completions.create(
    model="minimax-m3.1-flash-preview",
    messages=[{"role": "user", "content": "Refactor this function for clarity..."}],
    # reasoning_effort or equivalent may be supported depending on gateway mapping
)

CometAPI’s unified catalog also includes MiniMax-M3, M2.7 series, and the new H3 video models, allowing seamless switching between text and multimodal generation under one billing relationship.

Continue learning

Connect this article to the next decision.

View all topics
Published on Sep 28, 2026
Last updated Sep 28, 2026
0 views
Reviewed for clarity, source attribution and current API terminology.

Ready to cut AI development costs by 20%?

Start free in minutes. Free trial credits included. No credit card required.

Read More