Gemini logoGemini

    gemini-2.5-flash-lite

    Auto-router

    gemini/gemini-2.5-flash-lite

    ChatVisionToolsSystem prompt
    Use this model

    Context

    1.0M

    Max output

    66K

    Input / 1M

    $0.10

    Output / 1M

    $0.40

    Modality

    Chat

    Cutoff

    Jan 2025

    About this model

    The most cost-efficient and fastest model in the 2.5 family, optimized for extreme low-latency and high-throughput scenarios.

    Best suited for

    • Ultra-low-cost, high-throughput tasks, classification, simple translation, customer inquiry routing, and budget-constrained workflows.

    Built-in tools

    Gtwy web search

    Hosted by the gateway — enable them per request without wiring your own endpoint.

    Capabilities

    Vision

    Accepts images alongside text in the same message.

    Tools

    Native function calling, so agents can invoke your endpoints.

    System prompt

    Honours a dedicated system role, separate from the user turn.

    Supported parameters

    creativity_level

    Controls the creativity of responses. Higher values (e.g., 0.7) increase creativity; lower values (e.g., 0.2) make responses more predictable.

    max_tokensMax Tokens Limit

    Specifies the maximum number of text units (tokens) allowed in a response, limiting its length.

    probability_cutoffProbability Cutoff (Top P)

    Focuses on the most likely words based on a percentage of probability.

    log_probability

    If true, returns the log probabilities of each output token returned in the content of message.

    stop

    This parameter tells the model to stop generating text when it reaches any of the specified sequences (like a word or punctuation)

    tools

    Lists tool definitions or capabilities available to the model.

    tool_choice

    Decides whether to use tools or just the model for generating responses.

    response_type

    Defines the format or type of the generated response.

    parallel_tool_calls

    Enables parallel execution of tools, allowing multiple tools to run simultaneously.

    stream

    Sends the response in real-time as it's being generated.

    Use this model

    OpenAI-SDK compatible, with gateway fallback and routing across providers.

    Other Gemini models

    ModelContextMax outputInputOutputCapabilities
    gemini-2.5-flash

    Chat

    1.0M66K$0.30$2.50Files, Tools, System prompt
    gemini-2.5-pro

    Chat

    1.0M66K$1.25$10.00Vision, Tools, System prompt
    gemini-3.1-pro-preview

    Chat

    1.0M66K$2.00$12.00Vision, Tools, System prompt, Reasoning
    gemini-3-flash-preview

    Chat

    1.0M66K$0.50$3.00Vision, Tools, System prompt
    gemini-3-pro-preview

    Chat

    1.0M66K$2.00$12.00Vision, Tools, System prompt
    gemini-3.5-flash

    Chat

    1.0M66K$15.00$35.00Vision, Tools, System prompt
    gemini-3.7-flash

    Chat

    1M66K$0.75$3.75Vision, Files, Tools, System prompt, Reasoning
    gemini-2.5-flash-image

    Image

    Vision
    gemini-3-pro-image-preview

    Image

    Vision, Tools
    imagen-4.0-generate-001

    Image

    imagen-4.0-fast-generate-001

    Image

    imagen-4.0-ultra-generate-001

    Image

    Pricing & provider details

    Gemini logo

    Gemini

    gemini/gemini-2.5-flash-lite

    Input · per 1M
    $0.10
    Output · per 1M
    $0.40
    Cached input · per 1M
    Context window
    1,048,576 tokens
    Max output
    65,536 tokens
    Knowledge cutoff
    Jan 2025
    Auto-router
    Supported

    Start building today

    Route gemini-2.5-flash-lite — and every other model in the catalogue — through one endpoint, with failover built in.