GPT-6.1 Sol are now live on CometAPI →
guide/CometAPI research

How to Use MiMo-V2.5 API Complete Developer Guide

Learn Xiaomi API specifications, benchmarks, pricing, multimodal requests, streaming, Python integration, and production best practices.

CometAPI
Deon GoodwinAI model and API research team
Updated Oct 7, 2026 10 min read
How to Use MiMo-V2.5 API Complete Developer Guide
Use this pattern

Make the first API call.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_COMETAPI_KEY",
    base_url="https://api.cometapi.com/v1",
)

response = client.chat.completions.create(
    model="gpt-5-mini",
    messages=[{"role": "user", "content": "Build this workflow."}],
)

print(response.choices[0].message.content)

TL;DR

MiMo-V2.5 API in CometAPI provides access to Xiaomi's sparse mixture-of-experts model for coding, multimodal understanding, and agentic workloads. The MiMo-V2.5 API is the practical default when a project needs a 1-million-token context window, native multimodal input, and low token cost; MiMo-V2.5 Pro is the better fit when difficult coding, reasoning, or long-horizon tool use matters more than price. This guide shows how to call the API through CometAPI, stream responses, structure multimodal requests, estimate cost, and harden a production integration.

Key Takeaways

  • The MiMo-V2.5 model uses 310B total parameters with 15B active parameters per token and supports a 1M-token context window.
  • Use the MiMo-V2.5 route for multimodal work, large-context analysis, and cost-sensitive agents; choose MiMo-V2.5 Pro for the hardest coding and reasoning tasks.
  • The OpenAI-compatible endpoint makes migration straightforward, but production systems still need timeouts, retry logic, token budgets, and provider-side capability checks.
  • At the stated CometAPI rates, a request with 10,000 input tokens and 2,000 output tokens costs about $0.001568 on the MiMo-V2.5 route.

MiMo-V2.5 Model Overview

Xiaomi introduced the model on April 22, 2026. MiMo-V2.5 API in CometAPI exposes it through an OpenAI-compatible chat-completions interface, so existing clients can usually migrate by changing the base URL, API key, and model identifier.

According to the official model card, the model combines a sparse MoE language backbone with dedicated vision and audio encoders. It accepts text, image, video, and audio inputs while producing text output. Its large context window is useful for repository-scale code review, document sets, long transcripts, and multi-step agents.

SpecificationValueWhy it matters
ArchitectureSparse mixture of expertsActivates a limited portion of the network for each token.
Total parameters310BProvides broad capacity without using every parameter per token.
Active parameters15BSupports efficient inference relative to total model size.
Context window1,000,000 tokensHandles large repositories and long document collections.
Maximum outputUp to 128K tokens through the current routeSupports long reports and extended code generation.
Native inputsText, image, video, audioEnables unified multimodal workflows.
Output modalityTextResponses are returned as generated text.
Vision encoder729M-parameter ViTProcesses visual content for image and video tasks.
Audio encoder261M-parameter TransformerProcesses speech and other audio inputs.
LicenseMITPermits broad commercial and research use under the license terms.

The official multimodal benchmark image below reports results across image understanding, multimodal-agent, and video-understanding tasks.

How to Use MiMo-V2.5 API Complete Developer Guide

Official Xiaomi MiMo-V2.5 multimodal benchmark comparison

Provider routes can expose a narrower set of media formats than the underlying model supports. Validate the exact image, audio, and video payload format against the active endpoint before shipping a multimodal feature.

MiMo-V2.5 Benchmark Performance

Xiaomi reports strong results in coding, terminal use, and multimodal-agent evaluation. These numbers are useful for positioning the model, but they are not substitutes for tests on your own prompts, tools, latency targets, and failure conditions.

BenchmarkReported scoreEvaluation focus
SWE-Bench Pro56.1Repository-level software engineering
Terminal-Bench 2.065.8Terminal-based agent tasks
Claw-Eval General62.1 Pass@3General agent capability
Claw-Eval Multimodal23.8 Pass@3Multimodal agent tasks
Claw-Eval Multi-Turn63.2 Pass@3Multi-turn agent behavior
ResearchClawBench16.91Research-oriented agent workflows

The official coding comparison shows 65.8 on Terminal-Bench 2.0 and 56.1 on SWE-Bench Pro, while also placing the model in a broader coding-agent comparison.

How to Use MiMo-V2.5 API Complete Developer Guide

Official Xiaomi MiMo-V2.5 coding benchmark comparison

What the results suggest: the model is especially attractive for applications that combine code, tools, and mixed media. For deterministic production decisions, run an internal evaluation that measures task success, token use, end-to-end latency, retry frequency, and human correction rate.

MiMo-V2.5 vs MiMo-V2.5 Pro: A Multi-Dimensional Comparison

DimensionMiMo-V2.5MiMo-V2.5-Pro
Total parameters310B1.02T
Active parameters15B42B
Context window1M tokens1M tokens
Primary strengthOmnimodal and general agent workloadsComplex agents and demanding coding
Input price per 1M tokens$0.112$0.348
Output price per 1M tokens$0.224$0.696
Best fitMultimodal, large-context, cost-sensitive systemsHard reasoning, coding, and long-horizon execution
Official API input modalitiesText, image, video, audioText
Official API output modalityTextText

Comparison result: start with the MiMo-V2.5 route when cost, throughput, or multimodal coverage drives the decision. Move selected requests to MiMo-V2.5 Pro when internal evaluations show a meaningful accuracy gain on difficult coding, reasoning, or tool-use cases. A tiered router often delivers a better cost-quality balance than sending every request to MiMo-V2.5 Pro.

How to Call the MiMo-V2.5 API

The API follows the familiar chat-completions schema. Keep the key on your server, load it from an environment variable, and never embed it in browser or mobile code.

Set the API key

export COMETAPI_KEY="your_api_key_here"

$env:COMETAPI_KEY = "your_api_key_here"


### Send a cURL request

curl https://api.cometapi.com/v1/chat/completions
-H "Authorization: Bearer $COMETAPI_KEY"
-H "Content-Type: application/json"
-d '{
"model": "mimo-v2.5",
"messages": [
{"role": "system", "content": "You are a careful coding assistant."},
{"role": "user", "content": "Explain the bug and propose a minimal patch."}
],
"temperature": 0.2
}'


### Use Python with the OpenAI SDK

import os
from openai import OpenAI

client = OpenAI(
api_key=os.environ["COMETAPI_KEY"],
base_url="https://api.cometapi.com/v1",
)

response = client.chat.completions.create(
model="mimo-v2.5",
messages=[
{"role": "system", "content": "You are a precise engineering assistant."},
{"role": "user", "content": "Review this migration plan for hidden risks."},
],
temperature=0.2,
)

print(response.choices[0].message.content)


> Do not log API keys, full authorization headers, or sensitive prompt content. Use a secrets manager in production and rotate credentials on a defined schedule.

## How to Stream MiMo-V2.5 API Responses

Streaming reduces perceived latency for long answers. The service returns incremental events, which the SDK exposes as chunks.

stream = client.chat.completions.create(
model="mimo-v2.5",
messages=[{"role": "user", "content": "Draft a safe database migration checklist."}],
stream=True,
)

for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)


In a web service, forward deltas with Server-Sent Events or WebSocket messages, handle client disconnects, and stop upstream generation when the user cancels.

## MiMo-V2.5 API Multimodal and Structured Workflows

For image-aware tasks, send a text instruction plus an image object in the same user message. The exact accepted media representation can vary by provider route, so treat this as a payload pattern and confirm current route support.

{
"model": "mimo-v2.5",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Extract the visible error and suggest the next diagnostic step."},
{"type": "image_url", "image_url": {"url": "https://example.com/error.png"}}
]
}
]
}


For machine-readable output, request a compact JSON object and validate it against your own schema. Never assume generated JSON is safe merely because it parses.

import json

raw = response.choices[0].message.content
result = json.loads(raw)

required = {"summary", "risks", "next_action"}
missing = required - result.keys()
if missing:
raise ValueError(f"Missing fields: {sorted(missing)}")


## MiMo-V2.5 API Pricing on CometAPI and Cost Control

| Route                          | Input per 1M tokens | Output per 1M tokens | 100-request example |
| ------------------------------ | ------------------- | -------------------- | ------------------- |
| MiMo-V2.5                      | $0.112              | $0.224               | $0.1568             |
| MiMo-V2.5 Pro                  | $0.348              | $0.696               | $0.4872             |
| Xiaomi pay-as-you-go reference | $0.14               | $0.28                | $0.1960             |

The example assumes 100 requests, each with 10,000 input tokens and 2,000 output tokens. Under those assumptions, the MiMo-V2.5 route is about 68% cheaper than MiMo-V2.5 Pro. Xiaomi's [published pay-as-you-go rates](https://platform.xiaomimimo.com/docs/zh-CN/integration/roocode) provide a useful reference point, while actual bills should be verified against the active provider price before deployment.

MiMo-V2.5 input: 100 × 10,000 / 1,000,000 × $0.112 = $0.1120
MiMo-V2.5 output: 100 × 2,000 / 1,000,000 × $0.224 = $0.0448
MiMo-V2.5 total: $0.1568

Pro input: 100 × 10,000 / 1,000,000 × $0.348 = $0.3480
Pro output: 100 × 2,000 / 1,000,000 × $0.696 = $0.1392
Pro total: $0.4872


Control cost with maximum-output limits, prompt compression, cached context, request classification, and a router that escalates only difficult tasks. Track cost per successful task rather than cost per request; a cheaper request that fails repeatedly can be more expensive overall.

## How to Design Long-Context MiMo-V2.5 API Requests

A 1M-token window makes larger inputs possible, but it does not remove retrieval and attention tradeoffs. Build the prompt so the model can locate the relevant evidence quickly.

* Place the task, output contract, and decision criteria near the beginning.
* Separate documents with stable identifiers and clear delimiters.
* Include only evidence that can affect the answer.
* Ask the model to cite document identifiers or line references in its result.
* Measure accuracy as context grows; do not treat maximum context as the ideal context.

> Long context increases both token cost and the chance that irrelevant material distracts the model. Retrieval, summarization, and hierarchical prompting remain important even when the full corpus fits.

## Production Reliability

### Retry only transient failures

Retry rate limits and temporary server failures with exponential backoff and jitter. Do not automatically retry invalid authentication, malformed payloads, or policy errors.

import random
import time

def delay_for(attempt: int) -> float:
return min(30.0, (2 ** attempt) + random.random())

for attempt in range(5):
try:
result = call_model()
break
except TransientAPIError:
if attempt == 4:
raise
time.sleep(delay_for(attempt))


### Set operational guardrails

* Use connection and overall request timeouts.
* Attach an idempotency key to workflows with external side effects.
* Record model ID, latency, input and output tokens, retry count, and final status.
* Redact secrets and personal data before logging.
* Require tool allowlists and confirmation for destructive actions.
* Maintain a fallback model or queue for provider outages.

## MiMo-V2.5 API Common Errors and Solutions

| Symptom       | Likely cause                                   | Recommended action                                               |
| ------------- | ---------------------------------------------- | ---------------------------------------------------------------- |
| 401 or 403    | Missing, invalid, or unauthorized API key      | Verify the server-side secret and project permissions.           |
| 400           | Malformed messages or unsupported media object | Validate JSON and confirm route-specific payload support.        |
| 413           | Request body is too large                      | Reduce media size or split the input.                            |
| 429           | Rate or quota limit                            | Back off with jitter and enforce client-side concurrency limits. |
| 5xx           | Temporary provider failure                     | Retry within a bounded budget or fail over.                      |
| Slow response | Large context, long output, or tool latency    | Stream output, cap tokens, and profile every stage.              |
| Invalid JSON  | Generation drift                               | Validate, repair once if safe, and reject persistent failures.   |

## Conclusion

The MiMo-V2.5 model is the strongest starting point for teams that need multimodal input, large context, and aggressive cost control. Pro should be a deliberate escalation path for tasks where measured quality gains justify the higher price. Begin with a small evaluation set, instrument every request, and promote the integration only after testing real payloads, long-context behavior, retry handling, and failure recovery.

## FAQ

### What endpoint should I use?

Use `/api.cometapi.com/v1/chat/completions`, or set `/api.cometapi.com/v1` as the base URL in an OpenAI-compatible SDK.

### What model identifier should I send?

Use `mimo-v2.5` for the MiMo-V2.5 route. Confirm the current identifier before launch if your account uses a provider-specific alias.

### When should I choose MiMo-V2.5 Pro?

Choose MiMo-V2.5 Pro when your evaluation shows a material advantage on difficult coding, reasoning, or long-running tool tasks. Keep routine or multimodal traffic on the MiMo-V2.5 route when its quality is sufficient.

### Does a 1M-token context mean I should send everything?

No. Large context is a capacity limit, not a prompt-design target. Retrieve and organize the evidence that can affect the answer, then measure quality and cost as the input grows.

### Can I call the API directly from a browser?

Do not expose a secret API key in frontend code. Call the provider from your backend and apply authentication, rate limits, logging, and data controls there.

### How do I verify multimodal support?

Test the exact media payload against the active provider route. The underlying model's native modalities do not guarantee that every route exposes every format in the same way.
Continue learning

Connect this article to the next decision.

View all topics
Published on Oct 7, 2026
Last updated Oct 7, 2026
0 views
Reviewed for clarity, source attribution and current API terminology.

Read More