deepseek-v4-flash
Auto-routerdeepseek/deepseek-v4-flash
Context
1M
Max output
348K
Input / 1M
$0.14
Output / 1M
$0.28
Cached / 1M
$0.0028
Cutoff
May 2024
About this model
DeepSeek-V4-Flash is a fast, capable model supporting thinking and non-thinking modes with 1M context. Accepts text inputs and produces text outputs.
Best suited for
- Multi-turn conversation, reasoning tasks, JSON output, function/tool calling, and high-throughput pipelines.
Built-in tools
Hosted by the gateway — enable them per request without wiring your own endpoint.
Capabilities
Tools
Native function calling, so agents can invoke your endpoints.
System prompt
Honours a dedicated system role, separate from the user turn.
Reasoning
Emits a separate thinking pass before the answer.
Supported parameters
stopThis parameter tells the model to stop generating text when it reaches any of the specified sequences (like a word or punctuation)
creativity_levelControls randomness. Higher values (0.8) make output more random, lower values (0.2) more focused and deterministic.
probability_cutoffNucleus sampling parameter. The model considers tokens with top_p probability mass.
log_probabilityLog ProbabilitiesWhether to return log probabilities of the output tokens.
max_tokensMax Tokens LimitSpecifies the maximum number of text units (tokens) allowed in a response, limiting its length.
toolsLists tool definitions or capabilities available to the model.
tool_choiceDecides whether to use tools or just the model for generating responses.
response_typeDefines the format or type of the generated response.
parallel_tool_callsEnables parallel execution of tools, allowing multiple tools to run simultaneously.
reasoningControls the level of reasoning used by the model.
streamSends the response in real-time as it's being generated.
service_tierUse this model
OpenAI-SDK compatible, with gateway fallback and routing across providers.
Other Deepseek models
| Model | Context | Max output | Input | Output | Capabilities |
|---|---|---|---|---|---|
| deepseek-v4-pro Chat | 1M | 348K | $0.435 | $0.87 | Tools, System prompt, Reasoning |
Pricing & provider details
Deepseek
deepseek/deepseek-v4-flash
- Input · per 1M
- $0.14
- Output · per 1M
- $0.28
- Cached input · per 1M
- $0.0028
- Context window
- 1,000,000 tokens
- Max output
- 348,000 tokens
- Knowledge cutoff
- May 2024
- Auto-router
- Supported
Start building today
Route deepseek-v4-flash — and every other model in the catalogue — through one endpoint, with failover built in.