Context
131K
Max output
8K
Input / 1M
$0.05
Output / 1M
$0.08
Modality
Chat
Cutoff
Dec 2023
About this model
Llama 3.1 8B on Groq provides low-latency, high-quality responses suitable for real-time conversational interfaces, content filtering systems, and data analysis applications. This model offers a balance of speed and performance with significant cost savings compared to larger models. Technical capabilities include native function calling support, JSON mode for structured output generation, and a 128K token context window for handling large documents.
Best suited for
- Low-latency chat interfaces, edge or device deployments, mobile applications, simple content generation, classification, and high-throughput, cost-efficient workloads.
Built-in tools
Hosted by the gateway — enable them per request without wiring your own endpoint.
Capabilities
Tools
Native function calling, so agents can invoke your endpoints.
System prompt
Honours a dedicated system role, separate from the user turn.
Supported parameters
creativity_levelControls the creativity of responses. Higher values (e.g., 0.7) increase creativity; lower values (e.g., 0.2) make responses more predictable.
max_tokensMax Tokens LimitSpecifies the maximum number of text units (tokens) allowed in a response, limiting its length.
probability_cutoffProbability Cutoff (Top P)Focuses on the most likely words based on a percentage of probability.
log_probabilityIf true, returns the log probabilities of each output token returned in the content of message.
repetition_penaltyThe `frequency_penalty` controls how often the model repeats itself, with higher positive values reducing repetition and negative values encouraging it.
novelty_penaltyDiscourages responses that are too similar to previous ones.
stopThis parameter tells the model to stop generating text when it reaches any of the specified sequences (like a word or punctuation)
toolsLists tool definitions or capabilities available to the model.
tool_choiceDecides whether to use tools or just the model for generating responses.
response_typeDefines the format or type of the generated response.
streamSends the response in real-time as it's being generated.
Use this model
OpenAI-SDK compatible, with gateway fallback and routing across providers.
Other Groq models
| Model | Context | Max output | Input | Output | Capabilities |
|---|---|---|---|---|---|
| llama-3.3-70b-versatile Chat | 33K | 33K | $0.59 | $0.79 | Tools, System prompt |
| meta-llama/llama-4-scout-17b-16e-instruct Chat | — | 8K | $0.11 | $0.34 | Vision, Tools, System prompt |
| openai/gpt-oss-120b Chat | 131K | 65K | — | — | Tools, System prompt |
| openai/gpt-oss-20b Chat | 131K | 65K | — | — | Tools, System prompt |
Pricing & provider details
Groq
groq/llama-3.1-8b-instant
- Input · per 1M
- $0.05
- Output · per 1M
- $0.08
- Cached input · per 1M
- —
- Context window
- 131,072 tokens
- Max output
- 8,000 tokens
- Knowledge cutoff
- Dec 2023
- Auto-router
- Not supported
Start building today
Route llama-3.1-8b-instant — and every other model in the catalogue — through one endpoint, with failover built in.