Rate-limited free · API key · code · last checked 30 Aug 2026
Create a key at https://console.groq.com/docs/rate-limits and call it from your harness.
LPU hardware, fastest time-to-first-token. Limits per API key per model; some models get half allowance. No credit card. Limits visible in x-ratelimit-* response headers.
Source: https://console.groq.com/docs/rate-limits
| Model | Cost | Context | RPM | TPM | RPD | Day tokens | Verified |
|---|---|---|---|---|---|---|---|
| openai/gpt-oss-120b | zsh | 128K | 30 | 8K | 1,000 | 200K tokens/day | 2026-08-30 |
| openai/gpt-oss-20b | zsh | 128K | 30 | 8K | 1,000 | 200K tokens/day | 2026-08-30 |
| qwen/qwen3.6-27b | zsh | 128K | 30 | 8K | 1,000 | 200K tokens/day | 2026-08-30 |
| qwen/qwen3.8-27b | zsh | 128K | 30 | 8K | 1,000 | 2M tokens/day | 2026-08-30 |
| meta-llama/llama-prompt-guard-2-22m | zsh | 128K | 30 | 15K | 14,400 | 500K tokens/day | 2026-08-30 |
| groq/compound | zsh | 128K | 30 | 70K | 250 | Not published | 2026-08-30 |
Machine-readable: /data/groq.json · Markdown