Skip to content
free-ai-agent-stack

Free LLM APIs

70 entries · verified 2026-09-24 · schema llm-api.schema.json

Free LLM APIs: 70 verified entries, 58 of which need no credit card, last checked 2026-09-24.

A free LLM API is the starting point for almost every agent project, and it is the category where lists rot fastest: Cerebras removed its no-card free tier in July 2026, Google moved its Pro models behind billing in April, and Half the providers quoted in older guides now require a payment method before the first request.

Every entry below states the free allowance in the vendor's own units, the rate limit, whether a credit card is required, and whether the free tier trains on your prompts. Providers that quietly stopped being free are still listed — marked as no longer free, with the replacement named — because knowing where not to go is half the value.

If you want to start in the next five minutes: Gemini gives the most daily requests without a card, Groq is the fastest, and Ollama is the only option with genuinely no limit at all.

resources catalogued
384
resources catalogued
need no credit card
222
need no credit card
verified within 30 days
100%
verified within 30 days
Filter resources

showing 70 of 70

Alibaba Cloud Model Studio

alibaba-model-studio

Qwen family access with a free per-model token grant — the cleanest route to Qwen3 Coder for free.

Free limit
~1M free tokens per model, valid 90 days
Rate limit
Low RPM on free quota
Context
1M

✅ no credit cardDoes not train on prompts

Grants expire 90 days after activation and do not renew — good for a build sprint, not a standing tier.

Anthropic Claude API

anthropic-claude-apino longer free

No longer free — kept so you do not waste a signup.

No public free API tier; the OSS program grants Claude Max to qualifying maintainers instead.

Free limit
None — paid API only. OSS maintainer program gives a Claude Max seat
Rate limit
Not applicable

💳 card requiredDoes not train on prompts

Free chat on claude.ai is separate from API access — there is no free API key.

Bhashini (ULCA)

bhashini

India's national language platform — free STT, TTS, translation and OCR APIs for 22 scheduled languages.

Free limit
Free for developers via the Bhashini API program
Rate limit
Fair-use, per-API quotas

✅ no credit cardDoes not train on prompts

Requires registration approval and a pipeline configuration — the signup is heavier than a typical vendor, but the API is genuinely free.

Cerebras Cloud

cerebras-clouddegraded

Wafer-scale inference — still the fastest tokens/sec available — but the permanent free tier ended in July 2026.

Free limit
$5 trial credit valid 30 days (card verification required since 2026-07-16)
Rate limit
5 RPM, 30K TPM on trial
Context
8K

💳 card requiredDoes not train on prompts

Downgraded from a no-card permanent free tier to a card-gated 30-day trial. Context is capped at 8K on trial.

Chutes.ai

chutesdegraded

Decentralised open-model inference with a community tier that costs nothing for low-volume agent calls.

Free limit
Community tier: quota-limited free requests across 100+ open models
Rate limit
Community queue priority

✅ no credit cardDoes not train on prompts

Community capacity is best-effort; expect queueing at peak. Fine for batch, risky for user-facing agents.

Cloudflare AI Gateway

cloudflare-ai-gateway

Free caching, rate-limiting and logging proxy placed in front of any provider, including your own free keys.

Free limit
Free on all plans; caching and analytics included
Rate limit
Not applicable (proxy layer)

✅ no credit cardDoes not train on prompts

The single highest-leverage free tool in this list: response caching can stretch a 1,000 req/day quota several-fold.

Cloudflare Workers AI

cloudflare-workers-ai

Serverless inference that runs inside a Worker — no egress cost, no cold-start management, free daily budget.

Free limit
10,000 Neurons/day (roughly 3-8K short LLM calls/day)
Rate limit
Account-level; unchanged by plan

✅ no credit cardDoes not train on prompts

Budget is denominated in Neurons, not tokens — per-model conversion varies, so budget by testing one call.

Dahl Inference

dahl-inference

Anonymous open-model inference on a decentralised GPU network — 100M tokens per key with no account, no card and no signup, on an OpenAI-compatible endpoint.

Free limit
100M tokens per API key
Rate limit
Not published; the 100M-token balance is the binding limit

✅ no credit cardDoes not train on prompts

The 100M allowance is per key, not per day — issuance is unlimited, so you can mint another key when one is spent. Anonymous keys come from POST /tokens. Direct curl calls frequently trip a Cloudflare challenge (403); the browser chat and the documented endpoints work. Runs on the Gonka compute network.

Deepgram

deepgram

Low-latency STT and TTS built for voice agents, with a large one-time credit and no card on signup.

Free limit
$200 one-time credit (roughly 100 hours of Nova-3 transcription)
Rate limit
Standard free limits

✅ no credit cardDoes not train on prompts

Credit is one-time; voice agents burn it fast. Pair with a local Whisper fallback.

DeepSeek API

deepseek-apino longer free

No longer free — kept so you do not waste a signup.

Extremely cheap per-token pricing, but no standing free tier and no signup grant since 2026.

Free limit
None — top-up required (cheapest paid tokens on the market)
Rate limit
Not applicable
Context
128K

💳 card requiredDoes not train on prompts

DeepSeek weights are open — run V4-Flash locally for free rather than paying the API.

Google Cloud Speech & TTS

google-cloud-speech-tts

Production speech APIs with a standing free monthly allowance for STT and WaveNet/Neural2 voices.

Free limit
60 min STT + 1M characters TTS per month
Rate limit
Standard free-tier quotas

💳 card requiredDoes not train on prompts

The STT/TTS free allowance does not expire, unlike the $300 signup credit — this is one of the best long-term free voice stacks.

Google Cloud Vertex AI Free Credits

google-cloud-vertex-free-credits

Vertex AI access to Gemini and Model Garden on new-account credits — the enterprise-labelled alternative to AI Studio.

Free limit
$300 in new-customer credits valid 90 days
Rate limit
Paid-tier quotas from day one

💳 card requiredDoes not train on prompts

Requires a billing account with a card, so it fails the strictest no-card filter — but the $300 covers a serious prototype.

Google Gemini API

google-gemini

Frontier multimodal models on a permanent free tier — Flash and Flash-Lite families only since April 2026.

Free limit
1,500 requests/day on Flash, 1,000/day on Flash-Lite; 250K TPM
Rate limit
10-15 RPM, 250K TPM
Context
1M

✅ no credit cardTrains on your prompts

Free-tier prompts are used to improve Google products outside the EU/UK/EEA. Pro models left the free tier in April 2026.

Groq Cloud

groq

Open-weight models at 300+ tokens/sec on LPU hardware, with the most generous request-per-day allowance of any free tier.

Free limit
1,000 req/day and 200K tokens/day per model
Rate limit
30 RPM, 8K TPM
Context
131K

✅ no credit cardDoes not train on prompts

The 200K tokens/day ceiling usually binds before the request count does. Llama models were removed from the free plan in 2026.

Hack Club AI

hack-club-ai

Free OpenAI-compatible access to Gemini, GPT-5.2, Kimi K2 and 30+ other models for Hack Clubbers — the most generous free tier aimed at teenage developers.

Free limit
Free for Hack Club members
Rate limit
Not published

✅ no credit card

Gated on a Hack Club account, which is aimed at students and teenage hackers. Not a general-purpose public tier — do not list it as one.

LLM7.io

llm7

Gateway you can call with the literal placeholder key "unused" — no account, no email, no card, just a base URL and an OpenAI SDK.

Free limit
10 req/min anonymous; 40 req/min with a free token
Rate limit
10 RPM anonymous, 40 RPM with token
Context
1M

✅ no credit card

The widely-repeated "2 req/s, 20 RPM, 100 req/hr" figures do not match the vendor documentation: anonymous access is 10 req/min and a free token from token.llm7.io raises it to 40 req/min. Anonymous traffic only reaches the turbo model set.

Mistral La Plateforme (Experiment)

mistral-la-plateforme

European-hosted Mistral family with a free Experiment mode worth roughly $10/month of credits, no card required.

Free limit
~$10/month of credits in Experiment mode (about 1B tokens/month on small models)
Rate limit
2 RPM on free mode

✅ no credit cardTrains on your prompts

Experiment mode trains on your prompts by default. The 2 RPM limit makes it a batch/offline worker, not an interactive backend.

Ollama

ollama

The default way to run open models locally — one command to pull, serve and expose an OpenAI-compatible endpoint.

Free limit
Unlimited; limited only by your GPU/RAM
Rate limit
None

✅ no credit cardDoes not train on prompts

No rate limit and no data leaving your machine — the correct answer to 'what is the truly free LLM API'.

OpenAI API

openai-apino longer free

No longer free — kept so you do not waste a signup.

No standing free tier. Data-sharing evaluation tokens are the only free path.

Free limit
None — no signup credits; $5 prepay minimum for API access
Rate limit
Not applicable

💳 card requiredDoes not train on prompts

Unless your org is eligible for the data-sharing token grant, treat OpenAI as paid-only.

OVHcloud AI Endpoints

ovh-ai-endpoints

EU-hosted, GDPR-friendly inference over 20+ open-weight models at $0, usable without an API key on the anonymous tier.

Free limit
Free — $0 per request on the anonymous tier
Rate limit
~2 requests/min per model anonymous

✅ no credit card

The API host is oai.endpoints.kepler.ai.cloud.ovh.net — posting to the marketing domain returns 405. Anonymous calls are accepted and then answered with "API rate limit exceeded", which is how the ~2 req/min cap presents itself; no key is required to reach that point. The published catalogue returns pricing of 0 for prompt, completion and request.

Perplexity API (Sonar)

perplexity-apino longer free

No longer free — kept so you do not waste a signup.

Search-grounded Sonar models are paid API only; the free consumer app is not API access.

Free limit
None — Pro subscribers get $5/month of API credit
Rate limit
Not applicable

💳 card requiredDoes not train on prompts

Tavily, Exa and Brave all have free search tiers — see the MCP list for agent-ready alternatives.

Puter.js

puter-js

Client-side JavaScript SDK that gives a web app AI, storage, database and auth with no backend — the developer pays nothing because each signed-in user spends credits from their own Puter account.

Free limit
Free for developers; end users spend their own Puter credits
Rate limit
Per-user credit balance

✅ no credit card

Usually marketed as "free and unlimited", which is only half true. It is genuinely free for the developer, but not unlimited: once a user exhausts their Puter allocation they pay Puter directly, so an app with heavy anonymous traffic will hit a wall. The open-source developer tools that wrap it as a drop-in OpenAI server are unmaintained third-party bridges and are not listed.

SiliconFlow

siliconflow-free-models

Chinese inference host that keeps a rotating set of open models free to call, including embedding and rerank models.

Free limit
Rotating set of free models (Qwen, GLM, DeepSeek small variants)
Rate limit
3-10 RPM on free models

✅ no credit cardDoes not train on prompts

Free model list changes without notice; the reranker and embedding models are the durable part.

Space Bunny Alpha

space-bunny-alpha

Anonymous stealth model on OpenRouter with a 1M-token context window, free at $0 in and $0 out while the provider stays cloaked.

Free limit
Free ($0/$0 per million tokens) while the preview runs
Rate limit
OpenRouter free-tier limits
Context
1M

✅ no credit card

Confirmed live against OpenRouter's model API on 2026-09-24. Stealth previews are temporary by design: Ox Alpha, the previous drop, was unmasked as Z.AI's GLM-5.3-Flash and removed from the free catalogue within a week — which is exactly why it is not listed here. Needs an OpenRouter account, and prompts are handled by an anonymous provider, so keep secrets out of it.

Voyage AI

voyage-ai

High-quality embeddings and rerankers with an unusually large free token grant for RAG prototypes.

Free limit
200M free tokens across embedding and rerank models
Rate limit
3 RPM, 10K TPM on free

✅ no credit cardDoes not train on prompts

The 3 RPM limit is the real constraint — fine for indexing a corpus offline, painful for live retrieval.

ZenMux

zenmux-free

Unified gateway over 100+ models that keeps a small set of model ids permanently free, including DeepSeek V4 Flash Vision and GLM-5.3.

Free limit
Free model ids, including deepseek/deepseek-v4-flash-vision-exp-free and z-ai/glm-5.3-free
Rate limit
Not published

✅ no credit card

The gateway as a whole is commercial — only the -free suffixed model ids cost nothing. Paid tiers advertise a 5% service-fee discount on top-up, so expect upsell copy on the pricing page.

Free LLM APIs — questions

How many free LLM APIs are there in 2026?

70 free LLM API providers are listed in this catalogue as of 2026-09-24, and 58 of them can be started without a credit card. Each entry records the free limit in the provider's own units, the rate limit, and the date a human last confirmed it.

Which free LLM APIs need no credit card?

58 of the 70 providers here can be used without entering card details. The catalogue flags every provider that requires a card and can filter them out in one click, which matters because a card-required free tier is a trial with extra steps.

Are free LLM APIs actually free?

They are free up to a published quota, not unmetered. Every provider listed states a limit — requests per minute, tokens per day, or a credit balance — and this catalogue records that limit instead of repeating a marketing claim. No legitimate unmetered frontier API exists, so none is listed.

What is the catch with free LLM APIs?

Three catches recur: your prompts may be used for training, commercial use may be restricted, and a free tier can be withdrawn with little notice. All three are recorded per provider here, which is why the entries carry a training flag, a commercial-use flag and a last-verified date.

Keep going