Context
—
Max output
131K
Input / 1M
$7.14
Output / 1M
$28.56
Modality
Chat
Cutoff
Unknown
About this model
Compact open-weight Mixture-of-Experts model optimized for cost-efficient deployment and agentic workflows.
Best suited for
- Cost-efficient AI deployment
- Agentic workflows
- Tool-calling assistants
- Long-context conversations
- General-purpose AI applications
Capabilities
Tools
Native function calling, so agents can invoke your endpoints.
System prompt
Honours a dedicated system role, separate from the user turn.
Supported parameters
temperatureControls randomness and creativity of responses.
top_pControls token probability sampling diversity.
max_tokensMax Tokens LimitMaximum number of tokens generated in the response.
toolsDefines external tools available to the model.
tool_choiceDetermines whether the model can use tools.
response_typeSpecifies the output response format.
parallel_tool_callsAllows multiple tools to execute simultaneously.
streamStreams the response in real-time.
Use this model
OpenAI-SDK compatible, with gateway fallback and routing across providers.
Other Neev Cloud models
| Model | Context | Max output | Input | Output | Capabilities |
|---|---|---|---|---|---|
| minimax-m2.7 Chat | — | 66K | $0.2852 | $1.1409 | Tools, System prompt |
| minimax-m2.7-highspeed Chat | — | 66K | $57.12 | $228.49 | Tools, System prompt |
| llama-3.1-8b-instant Chat | — | 131K | $4.76 | $7.62 | Tools, System prompt |
| gpt-oss-120b Chat | — | 131K | $9.52 | $47.60 | Tools, System prompt |
| deepseek-v3-2 Chat | — | 131K | $26.66 | $39.99 | Tools, System prompt |
| glm-4-7 Chat | — | 200K | $57.12 | $209.45 | Tools, System prompt |
| llama-3.3-70b-versatile Chat | — | 131K | $56.17 | $75.21 | Tools, System prompt |
Pricing & provider details
Neev Cloud
neev_cloud/gpt-oss-20b
- Input · per 1M
- $7.14
- Output · per 1M
- $28.56
- Cached input · per 1M
- —
- Context window
- —
- Max output
- 131,072 tokens
- Knowledge cutoff
- Unknown
- Auto-router
- Not supported
Start building today
Route gpt-oss-20b — and every other model in the catalogue — through one endpoint, with failover built in.