Context
8K
Max output
8K
Input / 1M
$0.20
Output / 1M
$2.00
Cached / 1M
Free
Cutoff
—
About this model
Moonshot V1 8K is Moonshot AI's generation model with an 8,192-token context window, suited for short-context chat, generation, and tool-calling tasks. Supports ToolCalls, JSON Mode, and Partial Mode.
Best suited for
- Short-context conversational AI and chat
- Text generation and summarization within 8K tokens
- Tool-calling and function-calling agents
- JSON/structured output generation
- Lightweight, cost-efficient inference tasks
Built-in tools
Hosted by the gateway — enable them per request without wiring your own endpoint.
Capabilities
Files
Accepts file attachments — PDFs, transcripts, spreadsheets.
Tools
Native function calling, so agents can invoke your endpoints.
System prompt
Honours a dedicated system role, separate from the user turn.
Supported parameters
max_tokensThe maximum number of tokens to generate for the chat completion. Maps to max_completion_tokens.
creativity_levelTemperatureControls the creativity of responses. Higher values (e.g., 0.7) increase randomness; lower values (e.g., 0.2) make responses more focused and deterministic. Maps to temperature.
probability_cutoffTop PNucleus sampling. The model considers tokens with a cumulative probability mass of top_p. Maps to top_p.
response_countNThe number of results to generate for each input message. Must not exceed 5. Maps to n.
novelty_penaltyPresence PenaltyPositive values penalize new tokens based on whether they appear in the text so far, increasing the likelihood of new topics. Maps to presence_penalty.
repetition_penaltyFrequency PenaltyPositive values penalize new tokens based on their existing frequency in the text, reducing verbatim repetition. Maps to frequency_penalty.
toolsLists tool definitions or capabilities available to the model.
tool_choiceDecides whether to use tools or just the model for generating responses.
response_typeadditional_stop_sequencesStop SequencesUp to 5 stop sequences that halt generation when fully matched. Maps to stop.
streamSends the response in real-time as it's being generated.
Use this model
OpenAI-SDK compatible, with gateway fallback and routing across providers.
Other Moonshot models
| Model | Context | Max output | Input | Output | Capabilities |
|---|---|---|---|---|---|
| kimi-k2.6 Chat | 262K | 262K | $0.95 | $4.00 | Vision, Files, Tools, System prompt |
| moonshot-v1-32k Chat | 33K | 33K | $1.00 | $3.00 | Files, Tools, System prompt |
| moonshot-v1-128k Chat | 131K | 131K | $2.00 | $5.00 | Files, Tools, System prompt |
| moonshot-v1-8k-vision-preview Chat | 8K | 8K | $0.20 | $2.00 | Vision, Files, Tools, System prompt |
| moonshot-v1-32k-vision-preview Chat | 33K | 33K | $1.00 | $3.00 | Vision, Files, Tools, System prompt |
| moonshot-v1-128k-vision-preview Chat | 131K | 131K | $2.00 | $5.00 | Vision, Files, Tools, System prompt |
| kimi-k2.5 Chat | 262K | 262K | $0.60 | $3.00 | Vision, Files, Tools, System prompt |
| kimi-k2.7-code Chat | 262K | 262K | $0.95 | $4.00 | Vision, Files, Tools, System prompt |
| kimi-k2.7-code-highspeed Chat | 262K | 262K | $1.90 | $8.00 | Vision, Files, Tools, System prompt |
Pricing & provider details
Moonshot
moonshot/moonshot-v1-8k
- Input · per 1M
- $0.20
- Output · per 1M
- $2.00
- Cached input · per 1M
- Free
- Context window
- 8,192 tokens
- Max output
- 8,192 tokens
- Knowledge cutoff
- —
- Auto-router
- Not supported
Start building today
Route moonshot-v1-8k — and every other model in the catalogue — through one endpoint, with failover built in.