Minimax
MiniMax M2.5
Total Context
204.8K
Max Output
131.1K
Released
N/A
MiniMax-M3MiniMax M3 is a frontier multimodal MiniMax model with a 1M-token context window, built for long-horizon agent workflows, coding, and tool use. It uses MiniMax Sparse Attention to reduce the cost of large-context workloads compared with earlier generations. Use the MiniMax API through TokenHub to review the current MiniMax M3 price, model ID, supported endpoints, context limits, and benchmark results before deploying it for production software tasks or collaborative agent workflows.
Context Window
512K tokens
Maximum Output
128K tokens
Release Date
Jun 1, 2026
Modalities
| Input Price | Output Price | Cache Read |
|---|---|---|
| $0.6/M | $2.4/M | $0.12/M |
Reasoning
Tool calling
Temperature parameter
Attachments
Knowledge Base
Endpoint Protocols
MiniMax M3 combines advanced software engineering, a million-token context window and native visual understanding for long-running development and agent tasks.
Handles bug fixing, frontend and backend development, performance optimization and repeated collaboration across extended engineering sessions.
MiniMax Sparse Attention enables context windows up to one million tokens while reducing the computational burden of long-context processing.
Processes text, images and video through a model trained with mixed modalities from the beginning, supporting visually grounded reasoning.
MiniMax M3 suits repository-scale engineering, long-running technical research and visual analysis that combines documents, code and experimental results.
Analyze a large codebase, discuss evolving requirements and implement coordinated frontend, backend or performance changes over multiple rounds.
Keep a paper, source code and experiment logs in one context, then implement experiments, inspect figures and compare results with the publication.
Interpret diagrams, charts, screenshots or video together with technical text and code to identify issues and propose implementation changes.
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.TOKENHUB_API_KEY,
baseURL: "https://us-api.tokenhub.com/v1",
})
const result = await client.chat.completions.create({})
console.log(result.choices[0]?.message?.content)MiniMax-M3
| Index score | ||
|---|---|---|
| Artificial Analysis Intelligence Index | Artificial Analysis broad capability aggregate | 44.4 |
| Artificial Analysis Coding Index | Artificial Analysis software task aggregate | 43.4 |
| Knowledge & Reasoning | ||
| GPQA | Advanced science problem solving | 92.9% |
| HLE | Broad expert-level exam set | 37.1% |
| Coding & Engineering | ||
| SciCode | Scientific coding challenges | 45.4% |
| Terminal-Bench Hard | Hard terminal task execution | 42.4% |
| Instruction Following & Agent Tasks | ||
| IFBench | Prompt constraint adherence | 82.9% |
| AA-LCR | Long-context reasoning | 74% |
| τ²-Bench | Agent workflow tasks | 88.9% |
Metrics sourced from Artificial Analysis
Answers to common questions about the MiniMax M3 API, pricing, coding and agent workflows, long-context limits, and model comparisons on TokenHub.
MiniMax M3 is a MiniMax multimodal model for long-horizon agent work, coding, and tool use. Its model profile emphasizes large-context workloads and production-oriented software tasks where a model needs to work across substantial project information.
MiniMax M3 is worth evaluating for coding and agent workflows that need long context, repeated tool use, or collaboration across multiple task steps. Before adopting it in production, test it with representative repository tasks, tool calls, and completion criteria from your own workflow.
Create a TokenHub API key, copy the exact MiniMax M3 model ID shown on this page, and choose a supported endpoint in the API section. Use the generated request example as the starting point for your OpenAI-compatible client or other supported integration.
Review the current input, output, and cache-read prices on this page, then estimate cost using your typical prompt size, response length, request volume, and cache-hit rate. Long-context and agent workflows can have a very different cost profile from a short text-only request.
Check the current context window, maximum output, input modalities, supported endpoints, and pricing shown in the model specifications. A large context limit does not remove the need to structure documents, control tool loops, and test how the model performs with your real project data.
Compare MiniMax M3 and Kimi K3 against the workload you need to ship. Review the current API price, context and output limits, modalities, supported endpoints, and benchmark data, then run a small evaluation using your own coding, reasoning, or agent tasks before committing to either model.
Create an API key in the TokenHub workspace, then copy the exact MiniMax M3 model ID displayed on this page. Confirm the supported endpoint and current specifications here before sending production traffic, because availability and model details can change.
Use one API key to access MiniMax M3 and more AI models through TokenHub.
Media and Discussions
Selected public videos and posts related to this model.
X (Twitter)
Reddit
YouTube