An OpenAI-compatible API exposes the same /chat/completions shape, so switching is usually a one-line change: swap base_url and the key. That makes these the lowest-friction way to move side projects off paid OpenAI usage, or to add a free fallback behind the same SDK.
Each provider's exact base URL is on its page, and the free models you can pass as model are in the model index.
Top pick — Cloudflare Workers AI
Highest free daily volume for a real side project Details →
30 verified providers
| Provider | Free tier | Rate limits | Gotchas |
|---|---|---|---|
| Google Gemini API (AI Studio) | Gemini 2.5 Flash, 2.5 Flash-Lite, 2.5 Pro (limited), embeddings, TTS models | Varies by model, roughly 5-30 req/min and 20-500 req/day depending on model | no cardno phonecommercial OKOpenAI-compatible |
| Groq | Open-weight models (Llama, Qwen, GPT-OSS) plus Whisper, no credit card required | e.g. llama-3.1-8b-instant: 30 RPM/14.4K RPD/6K TPM/500K TPD; llama-3.3-70b-versatile: 30 RPM/1K RPD/12K TPM/100K TPD; qwen3-32b: 60 RPM/1K RPD/6K TPM/500K TPD; similar for GPT-OSS and Whisper models | no cardphone requiredcommercial OKOpenAI-compatible |
| OpenRouter | 20+ models with a :free suffix, single API across many providers | 20 req/min; 50 req/day under 10 credits purchased lifetime, 1000 req/day once 10+ credits purchased (one-time, not a subscription) | no cardno phonecommercial OKOpenAI-compatible |
| Cloudflare Workers AI | 10,000 Neurons/day, all account plans | 30+ models: LLMs (Llama, Mistral, DeepSeek, Qwen...), embeddings, image, audio | no cardno phonecommercial OKOpenAI-compatible |
| GitHub Models | Included with any GitHub account via the Copilot tier | Copilot Free: ~15 RPM / 150 RPD on "low" tier models, 8K input / 4K output tokens per request. Higher Copilot tiers raise the ceiling | no cardno phoneeval onlyOpenAI-compatible |
| Cohere | Trial (evaluation) API keys covering chat, embed and rerank | 1,000 API calls/month total; 20 req/min chat; 2,000 inputs/min embed; 10 req/min rerank | no cardno phoneeval onlyOpenAI-compatible |
| Cerebras | Access to all Cerebras-hosted models | Officially published per-model: 5 RPM / 30,000 TPM / 1,000,000 TPH / 1,000,000 TPD (e.g. gpt-oss-120b, zai-glm-4.7, gemma-4-31b); limits vary by model | OpenAI-compatible |
| HuggingFace | Free CPU Basic + ZeroGPU for Spaces; Inference Providers has a monthly credit ($0.10/mo on Free plan, $2.00/mo on PRO/Team/Enterprise) | No RPM/TPM published, only credit amounts | no cardOpenAI-compatible |
| SiliconFlow | Several models permanently free (e.g. Qwen2.5-7B-Instruct and others) at $0 cost, plus a $1 welcome credit for paid models | Fixed per-model limits for free models; generic docs cite ranges of 1,000-10,000 RPM and 50,000-5,000,000 TPM depending on model tier — exact limits shown in-account | phone requiredOpenAI-compatible |
| Z.ai (Zhipu AI / GLM) | GLM-4.5-Flash, GLM-4.7-Flash (text), and GLM-4.6V-Flash (vision) are officially listed as $0 cost (input, cached input, and output) on a permanent basis | Not specified with concrete RPM/TPM figures in public docs | no phonecommercial OKOpenAI-compatible |
| IBM watsonx.ai (Lite plan) | Lite plan: 300,000 tokens/month for foundation model inference, 20 CUH/month for ML tooling, 100 pages/month of document text extraction | 2 inference requests per second (explicitly documented for the Lite plan) | card requiredOpenAI-compatible |
| OVHcloud AI Endpoints | Several open-weight models in the catalog (e.g. Qwen3Guard) listed as $0 per token, permanently, via two access modes: anonymous (no account) and authenticated (API key tied to a Public Cloud project) | Anonymous access: 2 requests/min per IP per model. Authenticated (API key): 400 requests/min per project per model. Exceeding either returns HTTP 429 | card requiredOpenAI-compatible |
| Fireworks AI | $1 trial credit | Various open-weight models | no cardOpenAI-compatible |
| Baseten | $30 trial credit | Any supported model, compute-based pricing | OpenAI-compatible |
| Nebius AI Studio | $1 trial credit, valid for 30 days | Various open-weight models | card requiredno phonecommercial OKOpenAI-compatible |
| Novita AI | $1 free credit on signup | Various open-weight models | OpenAI-compatible |
| Alibaba Cloud (Model Studio) | 1,000,000 tokens (example figure, varies by model), international/Singapore region only | Qwen open & proprietary models | no cardOpenAI-compatible |
| SambaNova Cloud | Rate-limited free tier (applies when no payment method is linked) across all models | Free Tier: 20 RPM / 20 RPD / 200,000 TPD across all models; Developer Tier (card required): 60-240 RPM depending on model | no cardcommercial OKOpenAI-compatible |
| Scaleway Generative APIs | 1,000,000 tokens free + 60 min Whisper transcription; billing starts at token 1,000,001 | Gemma, Llama, Mistral, Qwen | OpenAI-compatible |
| NVIDIA NIM | Trial credit, phone verification required | Some models have reduced context windows on the free trial | phone requiredeval onlyOpenAI-compatible |
| Vercel AI Gateway | Free tier with a monthly free credit covering a subset of models at lower rate limits | Free tier is rate-limited per model (HTTP 429 on exceed), lower than paid; routes to many providers rather than hosting models itself | no cardOpenAI-compatible |
| Jina AI | 10M free tokens (one-time) across all models — embeddings (v3/v4), rerankers, classifier; plus a keyless Reader (r.jina.ai) for basic use | Free key: 100 RPM / 100k TPM for embeddings & reranker (2 concurrent); keyless Reader 20 RPM | no cardno phonecommercial OKOpenAI-compatible |
| AssemblyAI | $50 free credit on signup (no card) — pre-recorded & streaming speech-to-text, Speech Understanding, and an OpenAI-compatible LLM Gateway (25+ models) | Free tier: 5 parallel transcriptions; streaming 5 new streams/min; global 20k requests / 5 min | no cardno phonecommercial OKOpenAI-compatible |
| Clarifai | One-time $5 credit across serverless models (GPT-OSS-120B, Claude, Llama), vision, embeddings and image generation | 15 requests/second global default (CONN_THROTTLED on exceed) | no cardphone requiredOpenAI-compatible |
| Arli AI | Free plan ($0): access to all text LLMs (Gemma, Qwen, etc.), capped at ~5 requests per 2-day window, 12K context, 1 request at a time | 1 request at a time; ~5 requests per 2 days across all models; max 12K context; delayed responses | OpenAI-compatible |
| Ollama Cloud | $0 Free plan: access to cloud-hosted open models (Qwen, GPT-OSS, DeepSeek, etc.) via API | Session limits reset every 5 hours and weekly limits every 7 days; 1 concurrent cloud model on the free plan (exact token caps not published) | no cardno phonecommercial OKOpenAI-compatible |
| ModelScope (API-Inference) | ~2,000 free API calls/day across open-weight models (Qwen3, DeepSeek, GLM, Llama, etc.) via API-Inference | ~2,000 calls/day; concurrency/QPS caps applied and dynamically adjusted | no cardeval onlyOpenAI-compatible |
| Moondream Cloud | $5/month usage credits in every workspace (Free plan) for the Moondream vision model — caption, query (VQA), detect, point | Bounded by the $5/month credit | OpenAI-compatible |
| Sarvam AI | ₹100 in free credits on signup, usable across all APIs including the Sarvam-M chat/LLM API and speech (STT/TTS) | Not published; bounded by the ₹100 credit | OpenAI-compatible |
| Tencent Hunyuan | 1,000,000 free tokens for Hunyuan text LLMs (hunyuan-a13b, turbos, translation & vision models), plus a separate 1,000,000-token allotment for hunyuan-embedding | Not published; bounded by the token package | OpenAI-compatible |
FAQ
What does OpenAI-compatible mean?
The API accepts the same request and response format as OpenAI’s /chat/completions, so the official OpenAI SDKs work by changing only base_url and the API key.
How do I switch my code to a free OpenAI-compatible API?
Set base_url to the provider’s endpoint (listed on each provider page), use its free API key, and pass one of its free model IDs as model.
Do all free LLM APIs support the OpenAI format?
No — this list is only the providers confirmed to expose an OpenAI-compatible endpoint against their own documentation.