Minimax logoMinimax

    minimax-m3

    minimax/minimax-m3

    ChatVisionFilesToolsSystem promptReasoning
    Use this model

    Context

    1.0M

    Max output

    524K

    Input / 1M

    $0.30

    Output / 1M

    $1.20

    Cached / 1M

    $0.06

    Cutoff

    2026

    About this model

    MiniMax-M3 is an open-weight, natively multimodal Mixture-of-Experts model with ~428B total parameters (~23B activated), built on MiniMax Sparse Attention (MSA) for efficient long-context processing. Accepts text, image, video, and PDF inputs and produces text outputs, with a deep thinking mode for complex reasoning.

    Best suited for

    • Long-horizon agentic workflows, coding, tool/function calling, multi-turn cowork, RAG over large documents, and high-throughput pipelines requiring extended context.

    Built-in tools

    Gtwy web search

    Hosted by the gateway — enable them per request without wiring your own endpoint.

    Capabilities

    Vision

    Accepts images alongside text in the same message.

    Files

    Accepts file attachments — PDFs, transcripts, spreadsheets.

    Tools

    Native function calling, so agents can invoke your endpoints.

    System prompt

    Honours a dedicated system role, separate from the user turn.

    Reasoning

    Emits a separate thinking pass before the answer.

    Supported parameters

    creativity_level

    Controls randomness (temperature). Higher values (0.8) make output more random, lower values (0.2) more focused and deterministic.

    probability_cutoff

    Nucleus sampling parameter (top_p). The model considers tokens with top_p probability mass. Default 0.95 for M3.

    max_tokensMax Tokens Limit

    Maps to max_completion_tokens. Max response length in tokens. M3 max 524288 (512K), recommended 131072 (128K).

    tools

    Lists tool definitions available to the model. Function tools are supported.

    reasoning

    Controls the level of reasoning used by the model.

    reasoning_split

    Separates thinking into reasoning_content instead of embedding <think> tags in content. Keep true to keep content clean.

    stream

    Sends the response in real-time as it's being generated.

    service_tier

    Use this model

    OpenAI-SDK compatible, with gateway fallback and routing across providers.

    Pricing & provider details

    Minimax logo

    Minimax

    minimax/minimax-m3

    Input · per 1M
    $0.30
    Output · per 1M
    $1.20
    Cached input · per 1M
    $0.06
    Context window
    1,048,576 tokens
    Max output
    524,288 tokens
    Knowledge cutoff
    2026
    Auto-router
    Not supported

    Start building today

    Route minimax-m3 — and every other model in the catalogue — through one endpoint, with failover built in.