Context
1.0M
Max output
66K
Input / 1M
$0.30
Output / 1M
$2.50
Modality
Chat
Cutoff
Jan 2025
About this model
best model in terms of price-performance, offering well-rounded capabilities. 2.5 Flash is best for large scale processing, low-latency, high volume tasks that require thinking, and agentic use cases.
Best suited for
- High-volume chat, real-time summarization, content moderation, simple coding, and low-latency applications needing speed and cost efficiency.
Built-in tools
Hosted by the gateway — enable them per request without wiring your own endpoint.
Capabilities
Files
Accepts file attachments — PDFs, transcripts, spreadsheets.
Tools
Native function calling, so agents can invoke your endpoints.
System prompt
Honours a dedicated system role, separate from the user turn.
Supported parameters
creativity_levelControls the creativity of responses. Higher values (e.g., 0.7) increase creativity; lower values (e.g., 0.2) make responses more predictable.
max_tokensMax Tokens LimitSpecifies the maximum number of text units (tokens) allowed in a response, limiting its length.
probability_cutoffProbability Cutoff (Top P)Focuses on the most likely words based on a percentage of probability.
log_probabilityIf true, returns the log probabilities of each output token returned in the content of message.
stopThis parameter tells the model to stop generating text when it reaches any of the specified sequences (like a word or punctuation)
toolsLists tool definitions or capabilities available to the model.
tool_choiceDecides whether to use tools or just the model for generating responses.
response_typeDefines the format or type of the generated response.
parallel_tool_callsEnables parallel execution of tools, allowing multiple tools to run simultaneously.
streamSends the response in real-time as it's being generated.
Use this model
OpenAI-SDK compatible, with gateway fallback and routing across providers.
Other Gemini models
| Model | Context | Max output | Input | Output | Capabilities |
|---|---|---|---|---|---|
| gemini-2.5-pro Chat | 1.0M | 66K | $1.25 | $10.00 | Vision, Tools, System prompt |
| gemini-2.5-flash-lite Chat | 1.0M | 66K | $0.10 | $0.40 | Vision, Tools, System prompt |
| gemini-3.1-pro-preview Chat | 1.0M | 66K | $2.00 | $12.00 | Vision, Tools, System prompt, Reasoning |
| gemini-3-flash-preview Chat | 1.0M | 66K | $0.50 | $3.00 | Vision, Tools, System prompt |
| gemini-3-pro-preview Chat | 1.0M | 66K | $2.00 | $12.00 | Vision, Tools, System prompt |
| gemini-3.5-flash Chat | 1.0M | 66K | $15.00 | $35.00 | Vision, Tools, System prompt |
| gemini-3.7-flash Chat | 1M | 66K | $0.75 | $3.75 | Vision, Files, Tools, System prompt, Reasoning |
| gemini-2.5-flash-image Image | — | — | — | — | Vision |
| gemini-3-pro-image-preview Image | — | — | — | — | Vision, Tools |
| imagen-4.0-generate-001 Image | — | — | — | — | — |
| imagen-4.0-fast-generate-001 Image | — | — | — | — | — |
| imagen-4.0-ultra-generate-001 Image | — | — | — | — | — |
Pricing & provider details
Gemini
gemini/gemini-2.5-flash
- Input · per 1M
- $0.30
- Output · per 1M
- $2.50
- Cached input · per 1M
- —
- Context window
- 1,048,576 tokens
- Max output
- 65,536 tokens
- Knowledge cutoff
- Jan 2025
- Auto-router
- Supported
Start building today
Route gemini-2.5-flash — and every other model in the catalogue — through one endpoint, with failover built in.