inference-providers

Free

What actually costs nothing across the 34 providers in this registry (99 models, 366 offerings) — and what to pay for when the free options run out. Every fact below is drawn from the catalog and carries a source URL on the linked page; a price of 0 only appears there with a source proving it.

Genuinely free (no payment)

Hetzner Inference (Experiments)

Free while experimental — an "as is" program launched ~2026-07, EU-based, no SLA, with email notice before changes. Currently serving Qwen/Qwen3.8-27B (1 model tracked here, all flagged free; the live GET /api/v1/models lineup is definitive). A 262,144-token context window. Limits per key: 10 requests per 60 seconds, 4M input and 100k output tokens per 60 seconds.

NVIDIA NIM API

No pay-as-you-go: hosted access is free-tier rate limited (~40 RPM soft). Production requires deploying NIM yourself or NVIDIA AI Enterprise. Keys (nvapi- prefix) are generated at build.nvidia.com — free with the free NVIDIA Developer Program. The hosted lineup spans a broad open-model catalog; the models tracked here include Nemotron, DeepSeek, and Llama families.

Cohere

command-a-plus and command-a-reasoning are free until rate limits on the platform — no per-token price is cataloged for either. Both are open weights (command-a-plus is Apache-2.0 licensed), so you can also serve them yourself when the limit bites.

Google Gemini API

gemma-4-31b-it is free-tier only — no paid tier exists for it. The Gemini API also has a general free tier for Gemini models with per-model rate limits — see the rate limits page.

OpenCode Zen

Free tiers sit alongside the paid catalog: several models are served on dedicated free wire IDs — muse-spark-1.3-contributor-free, nemotron-3.5-lightning-free, and nemotron-3-ultra-free — next to ~90+ curated pay-as-you-go models billed at cost.

poolside

Inference at inference.poolside.ai is free for a limited time. Model ids are prefixed poolside/ — currently poolside/laguna-s-2.1 and poolside/laguna-xs-2.1.

Cheapest subscriptions

Every provider in the catalog with a plan block — subscription surfaces have no per-token prices; the quota and price live in the plan. Sorted cheapest first.

ProviderPriceQuotaNotes
Alibaba Token Plan$6 / monthlyPersonal: Lite $6 (700 credits/5h, 2500/7d), Standard $20 (3000/10k), Pro $70 (12k/40k). Team: Standard $30/seat (25k credits), Pro $100/seat (100k), Max $200/seat (250k), shared packs $700/625k.Sliding 5-hour + 7-day windows, no monthly reset. sk-sp- keys are NOT interchangeable with Coding Plan keys (shared prefix, different product). Interactive-tool use only (Claude Code, Cursor, Qwen Code, OpenClaw); scripts/backends banned on Personal.
OpenCode Go$10 / monthly$12 per 5 hours, $30/week, $60/month$5 first month then $10/mo, on top of the Zen console (same API key mechanism); open coding models only — Grok 4.5, GPT 5.6 Luna, GLM-5.x, Kimi K3/K2.x, MiniMax M3/M2.7, Qwen3.x, DeepSeek V4, MiMo, Hy3; optional Zen-balance fallback. Routes per protocol: /responses for Grok/GPT, /chat/completions for GLM/Kimi/DeepSeek, /messages (Anthropic) for MiniMax/Qwen.
Z.ai GLM Coding Plan$18 / monthlyLite ~2000 credits/5h, 10000/week; Pro ~12000/5h; Max ~28000/5hCredits-based since 2026-07-30. Includes GLM-5.3, GLM-5-Turbo, GLM-4.7; GLM-5.2/5.1 auto-route to GLM-5.3.
Kimi for Coding$19 / monthlyKimi Code has its own rolling 5-hour + weekly limits plus shared agent credits (60/150/360/720 by tier); concurrency-limitedIncluded with Kimi membership — tiers Moderato $19, Allegretto $39, Allegro $99, Vivace $199 (annual discounts); K3 access from Moderato up. Coding keys are distinct from Open Platform keys and never interchangeable (top 401 cause).
MiniMax Token Plan$20 / monthlyPlus ~1.7B tokens/mo (3-4 concurrent agents); Max ~5.1B; higher tiers to ~$120Subscription Key is a separate credential type from pay-as-you-go API keys, issued under Billing > Token Plan. Sources conflict on Plus pricing ($20 vs $40) — verify at the subscribe page. For Claude Code: ANTHROPIC_BASE_URL=https://api.minimax.io/anthropic + ANTHROPIC_AUTH_TOKEN=<Subscription Key>.
Ollama Cloud$20 / monthlyFree: 1 concurrent model; Pro $20: 3; Max $100: 10. Sessions reset every 5 hours plus weekly limits; usage weighted by model cost level; extra balance purchasable on Pro/Max.Zero data retention; no training on prompts; US-primary hosting.
Synthetic$30 / monthly$1/day or $30/month per pack, stackable; 500 price-weighted requests / 5h + $24/week credits + 1 concurrent per modelSubscription includes all always-on models, no per-token billing; embeddings free and exempt.
Qwen Coding Plan$50 / monthlyPro $50/mo: 6,000 requests / 5 rolling hours, 45,000/week, 90,000/month; slots restock daily 00:00 UTC+8Strict model allowlist: qwen3.7-plus, qwen3.6-plus, qwen3.5-plus, qwen3-coder-next, qwen3-coder-plus, glm-5, glm-4.7, kimi-k2.5, MiniMax-M2.5. Interactive coding-tool use only. Lite discontinued: new subscriptions ended 2026-03-20, renewals ended 2026-04-13. No Max or Team tier exists on the Coding Plan (Token Plan Team is a separate product).

Standing deals to know

Free things change without notice — each claim above is source-linked on its provider page, with a verification date.