Free OpenAI-compatible APIs (drop-in base_url)

Point the OpenAI SDK at a different base_url and keep all of your existing code.

An OpenAI-compatible API exposes the same /chat/completions shape, so switching is usually a one-line change: swap base_url and the key. That makes these the lowest-friction way to move side projects off paid OpenAI usage, or to add a free fallback behind the same SDK.

Each provider's exact base URL is on its page, and the free models you can pass as model are in the model index.

Top pick — Cloudflare Workers AI

Highest free daily volume for a real side project Details →

30 verified providers

ProviderFree tierRate limitsGotchas
Google Gemini API (AI Studio)Gemini 2.5 Flash, 2.5 Flash-Lite, 2.5 Pro (limited), embeddings, TTS modelsVaries by model, roughly 5-30 req/min and 20-500 req/day depending on modelno cardno phonecommercial OKOpenAI-compatible
GroqOpen-weight models (Llama, Qwen, GPT-OSS) plus Whisper, no credit card requirede.g. llama-3.1-8b-instant: 30 RPM/14.4K RPD/6K TPM/500K TPD; llama-3.3-70b-versatile: 30 RPM/1K RPD/12K TPM/100K TPD; qwen3-32b: 60 RPM/1K RPD/6K TPM/500K TPD; similar for GPT-OSS and Whisper modelsno cardphone requiredcommercial OKOpenAI-compatible
OpenRouter20+ models with a :free suffix, single API across many providers20 req/min; 50 req/day under 10 credits purchased lifetime, 1000 req/day once 10+ credits purchased (one-time, not a subscription)no cardno phonecommercial OKOpenAI-compatible
Cloudflare Workers AI10,000 Neurons/day, all account plans30+ models: LLMs (Llama, Mistral, DeepSeek, Qwen...), embeddings, image, audiono cardno phonecommercial OKOpenAI-compatible
GitHub ModelsIncluded with any GitHub account via the Copilot tierCopilot Free: ~15 RPM / 150 RPD on "low" tier models, 8K input / 4K output tokens per request. Higher Copilot tiers raise the ceilingno cardno phoneeval onlyOpenAI-compatible
CohereTrial (evaluation) API keys covering chat, embed and rerank1,000 API calls/month total; 20 req/min chat; 2,000 inputs/min embed; 10 req/min rerankno cardno phoneeval onlyOpenAI-compatible
CerebrasAccess to all Cerebras-hosted modelsOfficially published per-model: 5 RPM / 30,000 TPM / 1,000,000 TPH / 1,000,000 TPD (e.g. gpt-oss-120b, zai-glm-4.7, gemma-4-31b); limits vary by modelOpenAI-compatible
HuggingFaceFree CPU Basic + ZeroGPU for Spaces; Inference Providers has a monthly credit ($0.10/mo on Free plan, $2.00/mo on PRO/Team/Enterprise)No RPM/TPM published, only credit amountsno cardOpenAI-compatible
SiliconFlowSeveral models permanently free (e.g. Qwen2.5-7B-Instruct and others) at $0 cost, plus a $1 welcome credit for paid modelsFixed per-model limits for free models; generic docs cite ranges of 1,000-10,000 RPM and 50,000-5,000,000 TPM depending on model tier — exact limits shown in-accountphone requiredOpenAI-compatible
Z.ai (Zhipu AI / GLM)GLM-4.5-Flash, GLM-4.7-Flash (text), and GLM-4.6V-Flash (vision) are officially listed as $0 cost (input, cached input, and output) on a permanent basisNot specified with concrete RPM/TPM figures in public docsno phonecommercial OKOpenAI-compatible
IBM watsonx.ai (Lite plan)Lite plan: 300,000 tokens/month for foundation model inference, 20 CUH/month for ML tooling, 100 pages/month of document text extraction2 inference requests per second (explicitly documented for the Lite plan)card requiredOpenAI-compatible
OVHcloud AI EndpointsSeveral open-weight models in the catalog (e.g. Qwen3Guard) listed as $0 per token, permanently, via two access modes: anonymous (no account) and authenticated (API key tied to a Public Cloud project)Anonymous access: 2 requests/min per IP per model. Authenticated (API key): 400 requests/min per project per model. Exceeding either returns HTTP 429card requiredOpenAI-compatible
Fireworks AI$1 trial creditVarious open-weight modelsno cardOpenAI-compatible
Baseten$30 trial creditAny supported model, compute-based pricingOpenAI-compatible
Nebius AI Studio$1 trial credit, valid for 30 daysVarious open-weight modelscard requiredno phonecommercial OKOpenAI-compatible
Novita AI$1 free credit on signupVarious open-weight modelsOpenAI-compatible
Alibaba Cloud (Model Studio)1,000,000 tokens (example figure, varies by model), international/Singapore region onlyQwen open & proprietary modelsno cardOpenAI-compatible
SambaNova CloudRate-limited free tier (applies when no payment method is linked) across all modelsFree Tier: 20 RPM / 20 RPD / 200,000 TPD across all models; Developer Tier (card required): 60-240 RPM depending on modelno cardcommercial OKOpenAI-compatible
Scaleway Generative APIs1,000,000 tokens free + 60 min Whisper transcription; billing starts at token 1,000,001Gemma, Llama, Mistral, QwenOpenAI-compatible
NVIDIA NIMTrial credit, phone verification requiredSome models have reduced context windows on the free trialphone requiredeval onlyOpenAI-compatible
Vercel AI GatewayFree tier with a monthly free credit covering a subset of models at lower rate limitsFree tier is rate-limited per model (HTTP 429 on exceed), lower than paid; routes to many providers rather than hosting models itselfno cardOpenAI-compatible
Jina AI10M free tokens (one-time) across all models — embeddings (v3/v4), rerankers, classifier; plus a keyless Reader (r.jina.ai) for basic useFree key: 100 RPM / 100k TPM for embeddings & reranker (2 concurrent); keyless Reader 20 RPMno cardno phonecommercial OKOpenAI-compatible
AssemblyAI$50 free credit on signup (no card) — pre-recorded & streaming speech-to-text, Speech Understanding, and an OpenAI-compatible LLM Gateway (25+ models)Free tier: 5 parallel transcriptions; streaming 5 new streams/min; global 20k requests / 5 minno cardno phonecommercial OKOpenAI-compatible
ClarifaiOne-time $5 credit across serverless models (GPT-OSS-120B, Claude, Llama), vision, embeddings and image generation15 requests/second global default (CONN_THROTTLED on exceed)no cardphone requiredOpenAI-compatible
Arli AIFree plan ($0): access to all text LLMs (Gemma, Qwen, etc.), capped at ~5 requests per 2-day window, 12K context, 1 request at a time1 request at a time; ~5 requests per 2 days across all models; max 12K context; delayed responsesOpenAI-compatible
Ollama Cloud$0 Free plan: access to cloud-hosted open models (Qwen, GPT-OSS, DeepSeek, etc.) via APISession limits reset every 5 hours and weekly limits every 7 days; 1 concurrent cloud model on the free plan (exact token caps not published)no cardno phonecommercial OKOpenAI-compatible
ModelScope (API-Inference)~2,000 free API calls/day across open-weight models (Qwen3, DeepSeek, GLM, Llama, etc.) via API-Inference~2,000 calls/day; concurrency/QPS caps applied and dynamically adjustedno cardeval onlyOpenAI-compatible
Moondream Cloud$5/month usage credits in every workspace (Free plan) for the Moondream vision model — caption, query (VQA), detect, pointBounded by the $5/month creditOpenAI-compatible
Sarvam AI₹100 in free credits on signup, usable across all APIs including the Sarvam-M chat/LLM API and speech (STT/TTS)Not published; bounded by the ₹100 creditOpenAI-compatible
Tencent Hunyuan1,000,000 free tokens for Hunyuan text LLMs (hunyuan-a13b, turbos, translation & vision models), plus a separate 1,000,000-token allotment for hunyuan-embeddingNot published; bounded by the token packageOpenAI-compatible

FAQ

What does OpenAI-compatible mean?

The API accepts the same request and response format as OpenAI’s /chat/completions, so the official OpenAI SDKs work by changing only base_url and the API key.

How do I switch my code to a free OpenAI-compatible API?

Set base_url to the provider’s endpoint (listed on each provider page), use its free API key, and pass one of its free model IDs as model.

Do all free LLM APIs support the OpenAI format?

No — this list is only the providers confirmed to expose an OpenAI-compatible endpoint against their own documentation.

More guides

← All guides