magistral-small-latest
Auto-routermistral/magistral-small-latest
Context
128K
Max output
40K
Input / 1M
$0.50
Output / 1M
$1.50
Cached / 1M
Free
Cutoff
01 June, 2025
About this model
Lightweight, fast, and proficient in over 80 programming languages.
Best suited for
- Cost-efficient reasoning for internal tools, automated compliance checks, complex data extraction, and problem-solving with speed and auditability.
Built-in tools
Hosted by the gateway — enable them per request without wiring your own endpoint.
Capabilities
Vision
Accepts images alongside text in the same message.
Tools
Native function calling, so agents can invoke your endpoints.
System prompt
Honours a dedicated system role, separate from the user turn.
Supported parameters
creativity_levelControls the creativity of responses. Higher values (e.g., 0.7) increase creativity; lower values (e.g., 0.2) make responses more predictable.
max_tokensMax Tokens LimitSpecifies the maximum number of text units (tokens) allowed in a response, limiting its length.
probability_cutoffProbability Cutoff (Top P)Focuses on the most likely words based on a percentage of probability.
repetition_penaltyThe `frequency_penalty` controls how often the model repeats itself, with higher positive values reducing repetition and negative values encouraging it.
novelty_penaltyDiscourages responses that are too similar to previous ones.
stopThis parameter tells the model to stop generating text when it reaches any of the specified sequences (like a word or punctuation)
toolsLists tool definitions or capabilities available to the model.
tool_choiceDecides whether to use tools or just the model for generating responses.
response_typeDefines the format or type of the generated response.
parallel_tool_callsEnables parallel execution of tools, allowing multiple tools to run simultaneously.
streamSends the response in real-time as it's being generated.
Use this model
OpenAI-SDK compatible, with gateway fallback and routing across providers.
Other Mistral models
| Model | Context | Max output | Input | Output | Capabilities |
|---|---|---|---|---|---|
| mistral-medium-latest Chat | 128K | 128K | $0.40 | $2.00 | Vision, Tools, System prompt |
| magistral-medium-latest Chat | 41K | 40K | $0.40 | $2.00 | Vision, Tools, System prompt |
| codestral-latest Chat | 256K | 40K | $0.30 | $0.90 | Vision, Tools, System prompt |
| mistral-small-latest Chat | 128K | 128K | $0.10 | $0.30 | Vision, Tools, System prompt |
Pricing & provider details
Mistral
mistral/magistral-small-latest
- Input · per 1M
- $0.50
- Output · per 1M
- $1.50
- Cached input · per 1M
- Free
- Context window
- 128,000 tokens
- Max output
- 40,000 tokens
- Knowledge cutoff
- 01 June, 2025
- Auto-router
- Supported
Start building today
Route magistral-small-latest — and every other model in the catalogue — through one endpoint, with failover built in.