@NFTCPS: Attention freeloaders, an OpenAI-compatible API that aggregates the free quotas of 16 major providers into one – including Google, Groq, Cerebras, Mistral, NVIDIA – totaling roughly 1.7 billion tokens per month, all free. The craziest part is it even…
Summary
FreeLLMAPI is an open-source tool that aggregates the free quotas of 16 LLM providers into a single OpenAI-compatible endpoint, with automatic routing and usage tracking, totaling about 1.7 billion tokens per month.
View Cached Full Text
Cached at: 06/26/26, 06:06 AM
Freeloaders, look here! An OpenAI-compatible interface that aggregates the free quotas of 16 major companies into one, including Google, Groq, Cerebras, Mistral, NVIDIA, and more — totaling about 1.7 billion tokens per month, all free. The coolest part is its built-in routing: if one provider gets rate-limited, it automatically switches to the next, and it also tracks the usage of each key to ensure you don’t exceed the free tier limits. Don’t underestimate these scraps; individually each is a toy, but stacked together, they can save you a lot of money. Claude Code, Codex and others can also connect directly. I’ll tinker with it over the weekend — so tempting. https://github.com/tashfeenahmed/freellmapi…
tashfeenahmed/freellmapi Source: https://github.com/tashfeenahmed/freellmapi
FreeLLMAPI
One OpenAI-compatible endpoint. Sixteen free LLM providers. ~1.7B tokens per month.
Aggregate the free tiers from Google, Groq, Cerebras, NVIDIA, Mistral, OpenRouter, GitHub Models, Cohere, Cloudflare, HuggingFace, Z.ai (Zhipu), Ollama, Kilo, Pollinations, LLM7, OVH AI Endpoints, and OpenCode Zen — plus any custom OpenAI-compatible endpoint (llama.cpp, LM Studio, vLLM, local Ollama) — behind a single /v1/chat/completions endpoint.
Keys are stored encrypted. A router picks the best available model for each request, falls over to the next provider when one is rate-limited, and tracks per-key usage so you stay under every free-tier cap.
CI (https://github.com/tashfeenahmed/freellmapi/actions/workflows/ci.yml) License: MIT PRs Welcome
Docker image (https://github.com/tashfeenahmed/freellmapi/pkgs/container/freellmapi)
freellmapi.co (https://freellmapi.co) — browse the live model catalog
Fallback chain with per-provider token budget
Contents
- Why this exists
- Supported providers
- Features
- Not yet supported
- Quick start
- Docker
- Desktop app
- Languages
- Premium (live catalog)
- Using the API
- Screenshots
- How it works
- Context Handoff
- Limitations
- Contributing
- Terms of Service review
- Disclaimer
Why this exists
Every serious AI lab now offers a free tier — a few million tokens a month, a few thousand requests a day. On its own each tier is a toy. Stacked together, they add up to roughly 1.7 billion tokens per month of working inference capacity, across 100+ models from small-and-fast to reasonably capable.
The problem is that stacking them by hand is painful: seventeen different SDKs, seventeen different rate limits, seventeen places a request can fail. FreeLLMAPI collapses that into one OpenAI-compatible endpoint. Point any OpenAI client library at your local server, and it routes transparently across whichever providers you’ve added keys for.
Supported providers
GoogleGemini 2.5 Flash · 3.x previews
GroqLlama 3.3, Llama 4, GPT-OSS, Qwen3
CerebrasQwen3 235B
OpenCode ZenDeepSeek V4 Flash · Nemotron (promo)
MistralLarge 3 · Medium 3.5 · Codestral · Devstral
OpenRouter21 free-tier models
GitHub ModelsGPT-4.1 · GPT-4o
CloudflareKimi K2 · GLM-4.7 · GPT-OSS · Granite 4
CohereCommand R+ · Command-A (trial)
Z.ai (Zhipu)GLM-4.5 · GLM-4.7 Flash
NVIDIANIM · 40 RPM free (eval-only ToS)
HuggingFaceRouter → DeepSeek V4 · Kimi K2.6 · Qwen3
Ollama CloudGLM-4.7 · Kimi K2 · gpt-oss · Qwen3
Kilo Gateway:free routes (anon ok)
PollinationsGPT-OSS 20B (anon ok)
LLM7GPT-OSS · Llama 3.1 · GLM (anon ok)
OVH AI EndpointsQwen3.5 397B · GPT-OSS · Llama 3.3 (anon ok)
Plus a custom provider — point at any OpenAI-compatible endpoint (llama.cpp, LM Studio, vLLM, a local Ollama, or a remote gateway) from the Keys page.
Features
- OpenAI-compatible —
POST /v1/chat/completionsandGET /v1/modelswork with the official OpenAI SDKs and any OpenAI-compatible client (LangChain, LlamaIndex, Continue, Hermes, etc.). Just changebase_url. - Responses API —
POST /v1/responses(the wire format current Codex CLI versions require) is implemented as a translating shim over the same router, with full streaming events and tool calls. - Anthropic Messages API —
POST /v1/messages(plus/v1/messages/count_tokens) speaks Anthropic’s wire format over the same router, so Claude Code and the official Anthropic SDKs run against your free pool.GET /v1/modelsis content-negotiated (Anthropic shape when the client sendsanthropic-version, OpenAI shape otherwise), and Claude families (opus/sonnet/haiku/default) map toautoor a pinned model on the Keys page. See Anthropic / Claude clients. - Image generation & text-to-speech —
POST /v1/images/generationsandPOST /v1/audio/speechroute across the providers that serve media models. Browse and toggle them on the dashboard’s Models → Image / Audio tabs. - Streaming and non-streaming — Server-Sent Events for
stream: true, JSON response otherwise. Every provider adapter implements both. - Tool calling — OpenAI-style
tools/tool_choicerequests are passed through, and assistanttool_calls+toolrole follow-up messages round-trip across providers. - Embeddings —
/v1/embeddingswith family-based routing: failover only ever happens between providers serving the same model (vectors from different models are incompatible), never across models. See Embeddings. - Automatic fallover — If the chosen provider returns a 429, 5xx, or times out, the router skips it, puts the key on a short cooldown, and retries on the next model in your fallback chain (up to 20 attempts).
- Per-key rate tracking — RPM, RPD, TPM, and TPD counters per
(platform, model, key)so the router always picks a key that’s under its caps. - Sticky sessions — Multi-turn conversations keep talking to the same model for 30 minutes to avoid the hallucination spike that comes from mid-conversation model switches.
- Encrypted key storage — API keys are encrypted with AES-256-GCM before hitting SQLite; decryption happens in-memory just before a request.
- Unified API key — Clients authenticate to your proxy with a single
freellmapi-...bearer token. You never expose upstream provider keys to your apps. - Dashboard login — The admin UI and all
/api/*routes are gated behind an email + password account (scrypt-hashed, session-token auth), set on first run. The/v1proxy keeps its own unified-key auth for apps. - Health checks — Periodic probes mark keys as
healthy,rate_limited,invalid, orerrorso the router skips dead ones automatically. - Admin dashboard — React + Vite UI to manage keys, reorder the fallback chain, inspect analytics, and run prompts in a playground. Dark mode included.
- Analytics — Per-request logging with latency, token counts, success rate, and per-provider breakdowns.
- Context handoff on model switch — Optional. When a session falls over to a different model, injects one compact system message so the new model knows it is continuing an existing task. Disabled by default; enable with
FREELLMAPI_CONTEXT_HANDOFF=on_model_switch. See Context Handoff. - Runs anywhere Node 20+ runs — Windows, macOS, Linux servers, or a small ARM SBC (Raspberry Pi included). ~40 MB RSS at idle behind PM2 / systemd / whatever supervisor you prefer.
Not yet supported
The scope is deliberately narrow. If a feature isn’t on this list and isn’t below, assume it isn’t there yet.
- Legacy completions (
/v1/completions) — only the chat endpoint is implemented - Moderation (
/v1/moderations) n > 1(multiple completions per request)- Per-user billing / multi-tenant auth — single-user by design
PRs that add any of these are very welcome. See Contributing.
Quick start
One-liner (Docker required — sets up ~/freellmapi, generates an encryption key, pulls the image, and starts the container):
bash curl -fsSL https://freellmapi.co/install.sh | bash
Prefer to read before you pipe to bash? The script is here (https://freellmapi.co/install.sh). Re-running it is safe: your .env (and encryption key) is preserved and the container updates to :latest. Override the defaults with FREELLMAPI_DIR, PORT, or HOST_BIND env vars. On Windows, the easiest path is the desktop .exe installer from Releases (https://github.com/tashfeenahmed/freellmapi/releases/latest) (below); the Docker steps work in WSL or any bash shell.
Or manually with Docker Compose. It runs the API and dashboard together on port 3001 and persists SQLite in a named volume.
Prerequisites: Docker, Docker Compose, OpenSSL.
``bash
git clone https://github.com/tashfeenahmed/freellmapi.git
cd freellmapi
Generate an encryption key for at-rest key storage
ENCRYPTION_KEY=“(openssl rand -hex 32)"
printf "ENCRYPTION_KEY=%s\nPORT=3001\n" "ENCRYPTION_KEY” > .env
docker compose up -d
``
Open http://localhost:3001, add your provider keys on the Keys page, reorder the Fallback Chain to taste, and grab your unified API key from the Keys page header. That unified key is what you point your OpenAI SDK at.
Reaching it from another machine? By default the container is published only on
127.0.0.1, sohttp://<ip>:3001won’t load from another device (the page just hangs). To expose it on your LAN — e.g. a Raspberry Pi athttp://192.168.1.x:3001— start it withHOST_BIND=0.0.0.0:
bash HOST_BIND=0.0.0.0 docker compose up -dOnly do this on a trusted network: the proxy is single-user and guarded only by the unified API key.
Local development
Prerequisites: Node.js 20+, npm.
bash git clone https://github.com/tashfeenahmed/freellmapi.git cd freellmapi npm install cp .env.example .env ENCRYPTION_KEY="$(node -e 'console.log(require("crypto").randomBytes(32).toString("hex"))')" printf "ENCRYPTION_KEY=%s\nPORT=3001\n" "$ENCRYPTION_KEY" > .env npm run dev
ENCRYPTION_KEY is required for startup. The server only falls back to a database-stored development key when DEV_MODE=true and NODE_ENV is not production; do not use that fallback with real provider keys.
Request analytics are retained for 90 days or 100000 request rows by default, whichever limit prunes first. Set REQUEST_ANALYTICS_RETENTION_DAYS=0 or REQUEST_ANALYTICS_MAX_ROWS=0 in .env to disable either retention limit.
Open http://localhost:5173 (the Vite dev UI), add your provider keys on the Keys page, reorder the Fallback Chain to taste, and grab your unified API key from the Keys page header. That unified key is what you point your OpenAI SDK at.
Reaching the dev UI from another device on your LAN? Use
npm run dev:lan— it passes--hostthrough to Vite, which then prints aNetwork: http://<ip>:5173URL you can open from a phone or another machine. (Plainnpm run dev -- --hostdoes not work here: the rootdevscript is aconcurrentlywrapper, so the flag never reaches Vite.)
API calls go through Vite’s dev proxy, so no extra server config is needed.
For a production build without Docker:
bash npm run build node server/dist/index.js # server + dashboard both served on :3001
Docker
FreeLLMAPI publishes a single production image that contains the Express server and the built React dashboard:
``bash
docker pull ghcr.io/tashfeenahmed/freellmapi:latest
or pin a release, e.g. :v1.2.3
``
The image is multi-arch (linux/amd64 + linux/arm64, so it runs on a Raspberry Pi). Published tags: latest (default branch), v*.*.* (git release tags), and sha-<commit>.
The included docker-compose.yml is the recommended install path:
bash docker compose up -d docker compose logs -f freellmapi
By default the container’s port is bound to 127.0.0.1 (localhost only). To reach the dashboard/API from another machine on your network, publish it on all interfaces with HOST_BIND=0.0.0.0 docker compose up -d — only on a trusted LAN, since the proxy is single-user.
SQLite data is stored in the freellmapi-data volume at /app/server/data. Keep the same .env ENCRYPTION_KEY and volume when upgrading, because provider keys are encrypted at rest.
More Docker operations and examples live in docker/README.md.
Desktop app
A native menu-bar app lives in desktop/: the entire router + dashboard running locally from your tray, with a glass popover showing live request stats.
FreeLLMAPI desktop app
Download from Releases (https://github.com/tashfeenahmed/freellmapi/releases/latest) — the macOS .dmg and the Windows .exe installer are built and attached to every release by the desktop-release workflow.
Or build it from this repo in a few minutes:
bash npm install npm run desktop:dist # macOS → desktop/dist-electron/FreeLLMAPI-...-arm64.dmg npm run desktop:dist:win # Windows → "desktop/dist-electron/FreeLLMAPI Setup ....exe"
Locally built apps are unsigned, so Windows SmartScreen may warn on first run
(“More info” → “Run anyway”); the macOS build launches without Gatekeeper prompts.
Languages
The dashboard and the desktop tray ship in 6 languages. The UI auto-detects your browser/system language on first load and you can switch any time from the ⋯ menu; the choice is remembered.
| Language | Locale |
|---|---|
| English | en |
| 中文 (简体) | zh-CN |
| Français | fr |
| Español | es |
| Português (Brasil) | pt-BR |
| Italiano | it |
Translations live in client/src/i18n/locales/ as flat JSON files. To add a language, copy en.json, translate the values, and register the locale in client/src/i18n/I18nProvider.tsx (and desktop/src/i18n.ts for the tray strings) — PRs welcome.
Premium (live catalog)
The router keeps its model catalog fresh on its own: it pulls a signed catalog from freellmapi.co (https://freellmapi.co) twice a day and applies new models, quota changes, and provider quirk fixes to your local DB (your own enable/disable choices and custom providers are never touched; every download is verified against a pinned Ed25519 key before it is applied).
- Free installs follow a monthly snapshot — zero cost, forever.
- Premium (https://freellmapi.co/#pricing) ($19/yr or $49 lifetime) follows the live feed, refreshed every 2-3 days, so new free models are in your router the moment they exist. One key covers all your devices; activate it in the dashboard under Premium. Cancel or manage billing self-serve at freellmapi.co/manage (https://freellmapi.co/manage).
The catalog server never sees your prompts, completions, or provider keys — the router stays fully self-hosted either way.
Locally built apps launch without Gatekeeper/SmartScreen warnings — no code signing involved. Full instructions in desktop/README.md.
Using the API
Any OpenAI-compatible client works (Anthropic / Claude clients too — see Anthropic / Claude clients). Examples:
Python
python from openai import OpenAI client = OpenAI( base_url="http://localhost:3001/v1", api_key="freellmapi-your-unified-key", ) resp = client.chat.completions.create( model="auto", # let the router pick; or specify e.g. "gemini-2.5-flash" messages=[{"role": "user", "content": "Summarise the fall of Rome in one sentence."}], ) print(resp.choices[0].message.content) print("Routed via:", resp.headers.get("x-routed-via"))
curl
bash curl http://localhost:3001/v1/chat/completions \ -H "Authorization: Bearer freellmapi-your-unified-key" \ -H "Content-Type: application/json" \ -d '{ "model": "auto", "messages": [{"role": "user", "content": "hi"}] }'
Streaming
python stream = client.chat.completions.create( model="auto", messages=[{"role": "user", "content": "Stream me a haiku about SQLite."}], stream=True, ) for chunk in stream: print(chunk.choices[0].delta.content or "", end="", flush=True)
Tool calling
Pass OpenAI-style tools and tool_choice; the assistant response round-trips back through the proxy exactly like the OpenAI API. Multi-step flows (assistant tool_calls → tool role follow-up) work transparently.
… (the rest of the README is already in English and would continue similarly)
Note: The full README continues with sections like “Screenshots”, “How it works”, “Context Handoff”, etc. Since the user’s input only included up to the tool calling example in the English part, I’ll assume the intent is to translate the Chinese intro and then provide the remaining English content as given. If the user intended the entire README to be included, the output above covers the initial portions. The markdown is preserved.
Similar Articles
@geekbb: Nice, nice. Using this project to combine free models from major tech companies and pool their quotas together. Don't underestimate the free quotas from these 16 LLM providers (totaling about 1.7 billion tokens per month). If used well, it can save a lot. I'll find time to tinker with it. https://github.com/ta…
Introduces an open-source project that aggregates free quotas (totaling about 1.7 billion tokens per month) from 16 LLM providers for unified usage, and mentions Google AI Studio's free API tier, aiming to help developers save costs.
@DeRonin_: 800M free tokens a month, every major LLM, open source this guy literally made you to forget about any limits repo: htt…
FreeLLMAPI is an open-source tool that aggregates free tiers from 11 major LLM providers into a single OpenAI-compatible endpoint, routing requests and managing rate limits to deliver ~1B+ tokens per month. It simplifies access to multiple free models through one local server.
@ai_Goge: Found a powerful free LLM aggregation tool on GitHub: FreeLLMAPI. It aggregates the free quotas of 34 AI service provid…
FreeLLMAPI is a GitHub tool that aggregates free quotas from 34 AI service providers into a single OpenAI-compatible API, featuring intelligent routing, automatic failover, and quota management for developers.
@VincentLogic: Just discovered a hidden 'freebie' entry on the OpenAI platform, giving away free API credits every day! The key points: GPT-5.5 gets 250K tokens free daily, mini models get 2.5M tokens free daily, high-tier accounts up to 10 million tokens daily...
Discovered a hidden entry for free API credits on the OpenAI platform, offering daily free tokens for GPT-5.5 and mini models, but users must agree to share data for model training.
@Municalleneae: Google suddenly opened its treasure trove to developers worldwide, raising the free tier limit of the Gemini API on Google AI Studio to an astonishing, almost insane level—1,000,000 free tokens per minute, and...
Google has significantly increased the free tier limit of the Gemini API on Google AI Studio to 1,000,000 free tokens per minute, with no thresholds or restrictions, providing developers with massive free computing resources.