Context
—
Max output
131K
Input / 1M
$4.76
Output / 1M
$7.62
Modality
Chat
Cutoff
Dec 2023
About this model
Low-latency model suitable for real-time conversational interfaces, content filtering, and general analysis.
Best suited for
- Real-time chat applications
- Low-latency AI assistants
- Content moderation and filtering
- General-purpose text analysis
- Lightweight agent workflows
Capabilities
Tools
Native function calling, so agents can invoke your endpoints.
System prompt
Honours a dedicated system role, separate from the user turn.
Supported parameters
temperatureControls randomness and creativity of responses.
top_pControls diversity by limiting token probability sampling.
max_tokensMax Tokens LimitMaximum number of tokens generated in the response.
toolsDefines external tools available to the model.
tool_choiceDetermines whether the model can use tools.
response_typeSpecifies the output response format.
parallel_tool_callsAllows multiple tools to execute simultaneously.
streamStreams the response in real-time.
Use this model
OpenAI-SDK compatible, with gateway fallback and routing across providers.
Other Neev Cloud models
| Model | Context | Max output | Input | Output | Capabilities |
|---|---|---|---|---|---|
| minimax-m2.7 Chat | — | 66K | $0.2852 | $1.1409 | Tools, System prompt |
| minimax-m2.7-highspeed Chat | — | 66K | $57.12 | $228.49 | Tools, System prompt |
| gpt-oss-20b Chat | — | 131K | $7.14 | $28.56 | Tools, System prompt |
| gpt-oss-120b Chat | — | 131K | $9.52 | $47.60 | Tools, System prompt |
| deepseek-v3-2 Chat | — | 131K | $26.66 | $39.99 | Tools, System prompt |
| glm-4-7 Chat | — | 200K | $57.12 | $209.45 | Tools, System prompt |
| llama-3.3-70b-versatile Chat | — | 131K | $56.17 | $75.21 | Tools, System prompt |
Pricing & provider details
Neev Cloud
neev_cloud/llama-3.1-8b-instant
- Input · per 1M
- $4.76
- Output · per 1M
- $7.62
- Cached input · per 1M
- —
- Context window
- —
- Max output
- 131,072 tokens
- Knowledge cutoff
- Dec 2023
- Auto-router
- Not supported
Start building today
Route llama-3.1-8b-instant — and every other model in the catalogue — through one endpoint, with failover built in.