Free
What actually costs nothing across the 34 providers in this registry (99 models, 366 offerings) — and what to pay for when the free options run out. Every fact below is drawn from the catalog and carries a source URL on the linked page; a price of 0 only appears there with a source proving it.
Genuinely free (no payment)
Hetzner Inference (Experiments)
Free while experimental — an "as is" program launched ~2026-07, EU-based, no SLA, with email notice before changes. Currently serving Qwen/Qwen3.8-27B (1 model tracked here, all flagged free; the live GET /api/v1/models lineup is definitive). A 262,144-token context window. Limits per key: 10 requests per 60 seconds, 4M input and 100k output tokens per 60 seconds.
NVIDIA NIM API
No pay-as-you-go: hosted access is free-tier rate limited (~40 RPM soft). Production requires deploying NIM yourself or NVIDIA AI Enterprise. Keys (nvapi- prefix) are generated at build.nvidia.com — free with the free NVIDIA Developer Program. The hosted lineup spans a broad open-model catalog; the models tracked here include Nemotron, DeepSeek, and Llama families.
Cohere
command-a-plus and command-a-reasoning are free until rate limits on the platform — no per-token price is cataloged for either. Both are open weights (command-a-plus is Apache-2.0 licensed), so you can also serve them yourself when the limit bites.
Google Gemini API
gemma-4-31b-it is free-tier only — no paid tier exists for it. The Gemini API also has a general free tier for Gemini models with per-model rate limits — see the rate limits page.
OpenCode Zen
Free tiers sit alongside the paid catalog: several models are served on dedicated free wire IDs — muse-spark-1.3-contributor-free, nemotron-3.5-lightning-free, and nemotron-3-ultra-free — next to ~90+ curated pay-as-you-go models billed at cost.
poolside
Inference at inference.poolside.ai is free for a limited time. Model ids are prefixed poolside/ — currently poolside/laguna-s-2.1 and poolside/laguna-xs-2.1.
Cheapest subscriptions
Every provider in the catalog with a plan block — subscription surfaces have no per-token prices; the quota and price live in the plan. Sorted cheapest first.
| Provider | Price | Quota | Notes |
|---|---|---|---|
| Alibaba Token Plan | $6 / monthly | Personal: Lite $6 (700 credits/5h, 2500/7d), Standard $20 (3000/10k), Pro $70 (12k/40k). Team: Standard $30/seat (25k credits), Pro $100/seat (100k), Max $200/seat (250k), shared packs $700/625k. | Sliding 5-hour + 7-day windows, no monthly reset. sk-sp- keys are NOT interchangeable with Coding Plan keys (shared prefix, different product). Interactive-tool use only (Claude Code, Cursor, Qwen Code, OpenClaw); scripts/backends banned on Personal. |
| OpenCode Go | $10 / monthly | $12 per 5 hours, $30/week, $60/month | $5 first month then $10/mo, on top of the Zen console (same API key mechanism); open coding models only — Grok 4.5, GPT 5.6 Luna, GLM-5.x, Kimi K3/K2.x, MiniMax M3/M2.7, Qwen3.x, DeepSeek V4, MiMo, Hy3; optional Zen-balance fallback. Routes per protocol: /responses for Grok/GPT, /chat/completions for GLM/Kimi/DeepSeek, /messages (Anthropic) for MiniMax/Qwen. |
| Z.ai GLM Coding Plan | $18 / monthly | Lite ~2000 credits/5h, 10000/week; Pro ~12000/5h; Max ~28000/5h | Credits-based since 2026-07-30. Includes GLM-5.3, GLM-5-Turbo, GLM-4.7; GLM-5.2/5.1 auto-route to GLM-5.3. |
| Kimi for Coding | $19 / monthly | Kimi Code has its own rolling 5-hour + weekly limits plus shared agent credits (60/150/360/720 by tier); concurrency-limited | Included with Kimi membership — tiers Moderato $19, Allegretto $39, Allegro $99, Vivace $199 (annual discounts); K3 access from Moderato up. Coding keys are distinct from Open Platform keys and never interchangeable (top 401 cause). |
| MiniMax Token Plan | $20 / monthly | Plus ~1.7B tokens/mo (3-4 concurrent agents); Max ~5.1B; higher tiers to ~$120 | Subscription Key is a separate credential type from pay-as-you-go API keys, issued under Billing > Token Plan. Sources conflict on Plus pricing ($20 vs $40) — verify at the subscribe page. For Claude Code: ANTHROPIC_BASE_URL=https://api.minimax.io/anthropic + ANTHROPIC_AUTH_TOKEN=<Subscription Key>. |
| Ollama Cloud | $20 / monthly | Free: 1 concurrent model; Pro $20: 3; Max $100: 10. Sessions reset every 5 hours plus weekly limits; usage weighted by model cost level; extra balance purchasable on Pro/Max. | Zero data retention; no training on prompts; US-primary hosting. |
| Synthetic | $30 / monthly | $1/day or $30/month per pack, stackable; 500 price-weighted requests / 5h + $24/week credits + 1 concurrent per model | Subscription includes all always-on models, no per-token billing; embeddings free and exempt. |
| Qwen Coding Plan | $50 / monthly | Pro $50/mo: 6,000 requests / 5 rolling hours, 45,000/week, 90,000/month; slots restock daily 00:00 UTC+8 | Strict model allowlist: qwen3.7-plus, qwen3.6-plus, qwen3.5-plus, qwen3-coder-next, qwen3-coder-plus, glm-5, glm-4.7, kimi-k2.5, MiniMax-M2.5. Interactive coding-tool use only. Lite discontinued: new subscriptions ended 2026-03-20, renewals ended 2026-04-13. No Max or Team tier exists on the Coding Plan (Token Plan Team is a separate product). |
Standing deals to know
- Alibaba Token Plan night pricing — qwen3.8-max-preview gets a 10x credit discount during night pricing, 22:00–08:00 UTC+8.
- OpenCode Go first month — $5 for the first month, then $10/month, for open coding models on top of the Zen console.
- Gemini 3.6 Flash intro pricing — 50% off through 2026-12-31 (3.7 Flash carries the same list prices but no published end date).
- MiniMax M3 — the ≤512k-input 50% discount was made permanent; prompts above 512k double.
Free things change without notice — each claim above is source-linked on its provider page, with a verification date.