@NFTCPS: Attention freeloaders, an OpenAI-compatible API that aggregates the free quotas of 16 major providers into one – including Google, Groq, Cerebras, Mistral, NVIDIA – totaling roughly 1.7 billion tokens per month, all free. The craziest part is it even…

X AI KOLs Timeline Tools

Summary

FreeLLMAPI is an open-source tool that aggregates the free quotas of 16 LLM providers into a single OpenAI-compatible endpoint, with automatic routing and usage tracking, totaling about 1.7 billion tokens per month.

Freebie seekers, check this out – an OpenAI-compatible API that aggregates the free quotas of 16 major providers into one. Google, Groq, Cerebras, Mistral, NVIDIA are all in there, adding up to roughly 1.7 billion tokens per month, completely free. The coolest part is its built-in routing: if one provider gets rate-limited, it automatically switches to the next, and it also monitors each key's usage so you never exceed the free quota. Don't underestimate these scraps – individually each provider is a toy, but stacked together they can save you a lot of money. Claude Code, Codex, etc., can also connect directly. I'll tinker with it over the weekend. Amazing. https://github.com/tashfeenahmed/freellmapi…
Original Article
View Cached Full Text

Cached at: 06/26/26, 06:06 AM

Freeloaders, look here! An OpenAI-compatible interface that aggregates the free quotas of 16 major companies into one, including Google, Groq, Cerebras, Mistral, NVIDIA, and more — totaling about 1.7 billion tokens per month, all free. The coolest part is its built-in routing: if one provider gets rate-limited, it automatically switches to the next, and it also tracks the usage of each key to ensure you don’t exceed the free tier limits. Don’t underestimate these scraps; individually each is a toy, but stacked together, they can save you a lot of money. Claude Code, Codex and others can also connect directly. I’ll tinker with it over the weekend — so tempting. https://github.com/tashfeenahmed/freellmapi…

tashfeenahmed/freellmapi Source: https://github.com/tashfeenahmed/freellmapi

FreeLLMAPI

One OpenAI-compatible endpoint. Sixteen free LLM providers. ~1.7B tokens per month.

Aggregate the free tiers from Google, Groq, Cerebras, NVIDIA, Mistral, OpenRouter, GitHub Models, Cohere, Cloudflare, HuggingFace, Z.ai (Zhipu), Ollama, Kilo, Pollinations, LLM7, OVH AI Endpoints, and OpenCode Zen — plus any custom OpenAI-compatible endpoint (llama.cpp, LM Studio, vLLM, local Ollama) — behind a single /v1/chat/completions endpoint.

Keys are stored encrypted. A router picks the best available model for each request, falls over to the next provider when one is rate-limited, and tracks per-key usage so you stay under every free-tier cap.

CI (https://github.com/tashfeenahmed/freellmapi/actions/workflows/ci.yml) License: MIT PRs Welcome
Docker image (https://github.com/tashfeenahmed/freellmapi/pkgs/container/freellmapi)

freellmapi.co (https://freellmapi.co) — browse the live model catalog
Fallback chain with per-provider token budget


Contents

Why this exists

Every serious AI lab now offers a free tier — a few million tokens a month, a few thousand requests a day. On its own each tier is a toy. Stacked together, they add up to roughly 1.7 billion tokens per month of working inference capacity, across 100+ models from small-and-fast to reasonably capable.

The problem is that stacking them by hand is painful: seventeen different SDKs, seventeen different rate limits, seventeen places a request can fail. FreeLLMAPI collapses that into one OpenAI-compatible endpoint. Point any OpenAI client library at your local server, and it routes transparently across whichever providers you’ve added keys for.

Supported providers

GoogleGemini 2.5 Flash · 3.x previews
GroqLlama 3.3, Llama 4, GPT-OSS, Qwen3
CerebrasQwen3 235B
OpenCode ZenDeepSeek V4 Flash · Nemotron (promo)
MistralLarge 3 · Medium 3.5 · Codestral · Devstral
OpenRouter21 free-tier models
GitHub ModelsGPT-4.1 · GPT-4o
CloudflareKimi K2 · GLM-4.7 · GPT-OSS · Granite 4
CohereCommand R+ · Command-A (trial)
Z.ai (Zhipu)GLM-4.5 · GLM-4.7 Flash
NVIDIANIM · 40 RPM free (eval-only ToS)
HuggingFaceRouter → DeepSeek V4 · Kimi K2.6 · Qwen3
Ollama CloudGLM-4.7 · Kimi K2 · gpt-oss · Qwen3
Kilo Gateway:free routes (anon ok)
PollinationsGPT-OSS 20B (anon ok)
LLM7GPT-OSS · Llama 3.1 · GLM (anon ok)
OVH AI EndpointsQwen3.5 397B · GPT-OSS · Llama 3.3 (anon ok)

Plus a custom provider — point at any OpenAI-compatible endpoint (llama.cpp, LM Studio, vLLM, a local Ollama, or a remote gateway) from the Keys page.

Features

  • OpenAI-compatible — POST /v1/chat/completions and GET /v1/models work with the official OpenAI SDKs and any OpenAI-compatible client (LangChain, LlamaIndex, Continue, Hermes, etc.). Just change base_url.
  • Responses API — POST /v1/responses (the wire format current Codex CLI versions require) is implemented as a translating shim over the same router, with full streaming events and tool calls.
  • Anthropic Messages API — POST /v1/messages (plus /v1/messages/count_tokens) speaks Anthropic’s wire format over the same router, so Claude Code and the official Anthropic SDKs run against your free pool. GET /v1/models is content-negotiated (Anthropic shape when the client sends anthropic-version, OpenAI shape otherwise), and Claude families (opus / sonnet / haiku / default) map to auto or a pinned model on the Keys page. See Anthropic / Claude clients.
  • Image generation & text-to-speech — POST /v1/images/generations and POST /v1/audio/speech route across the providers that serve media models. Browse and toggle them on the dashboard’s Models → Image / Audio tabs.
  • Streaming and non-streaming — Server-Sent Events for stream: true, JSON response otherwise. Every provider adapter implements both.
  • Tool calling — OpenAI-style tools / tool_choice requests are passed through, and assistant tool_calls + tool role follow-up messages round-trip across providers.
  • Embeddings — /v1/embeddings with family-based routing: failover only ever happens between providers serving the same model (vectors from different models are incompatible), never across models. See Embeddings.
  • Automatic fallover — If the chosen provider returns a 429, 5xx, or times out, the router skips it, puts the key on a short cooldown, and retries on the next model in your fallback chain (up to 20 attempts).
  • Per-key rate tracking — RPM, RPD, TPM, and TPD counters per (platform, model, key) so the router always picks a key that’s under its caps.
  • Sticky sessions — Multi-turn conversations keep talking to the same model for 30 minutes to avoid the hallucination spike that comes from mid-conversation model switches.
  • Encrypted key storage — API keys are encrypted with AES-256-GCM before hitting SQLite; decryption happens in-memory just before a request.
  • Unified API key — Clients authenticate to your proxy with a single freellmapi-... bearer token. You never expose upstream provider keys to your apps.
  • Dashboard login — The admin UI and all /api/* routes are gated behind an email + password account (scrypt-hashed, session-token auth), set on first run. The /v1 proxy keeps its own unified-key auth for apps.
  • Health checks — Periodic probes mark keys as healthy, rate_limited, invalid, or error so the router skips dead ones automatically.
  • Admin dashboard — React + Vite UI to manage keys, reorder the fallback chain, inspect analytics, and run prompts in a playground. Dark mode included.
  • Analytics — Per-request logging with latency, token counts, success rate, and per-provider breakdowns.
  • Context handoff on model switch — Optional. When a session falls over to a different model, injects one compact system message so the new model knows it is continuing an existing task. Disabled by default; enable with FREELLMAPI_CONTEXT_HANDOFF=on_model_switch. See Context Handoff.
  • Runs anywhere Node 20+ runs — Windows, macOS, Linux servers, or a small ARM SBC (Raspberry Pi included). ~40 MB RSS at idle behind PM2 / systemd / whatever supervisor you prefer.

Not yet supported

The scope is deliberately narrow. If a feature isn’t on this list and isn’t below, assume it isn’t there yet.

  • Legacy completions (/v1/completions) — only the chat endpoint is implemented
  • Moderation (/v1/moderations)
  • n > 1 (multiple completions per request)
  • Per-user billing / multi-tenant auth — single-user by design

PRs that add any of these are very welcome. See Contributing.

Quick start

One-liner (Docker required — sets up ~/freellmapi, generates an encryption key, pulls the image, and starts the container):
bash curl -fsSL https://freellmapi.co/install.sh | bash

Prefer to read before you pipe to bash? The script is here (https://freellmapi.co/install.sh). Re-running it is safe: your .env (and encryption key) is preserved and the container updates to :latest. Override the defaults with FREELLMAPI_DIR, PORT, or HOST_BIND env vars. On Windows, the easiest path is the desktop .exe installer from Releases (https://github.com/tashfeenahmed/freellmapi/releases/latest) (below); the Docker steps work in WSL or any bash shell.

Or manually with Docker Compose. It runs the API and dashboard together on port 3001 and persists SQLite in a named volume.

Prerequisites: Docker, Docker Compose, OpenSSL.

``bash
git clone https://github.com/tashfeenahmed/freellmapi.git
cd freellmapi

Generate an encryption key for at-rest key storage

ENCRYPTION_KEY=“(openssl rand -hex 32)" printf "ENCRYPTION_KEY=%s\nPORT=3001\n" "ENCRYPTION_KEY” > .env
docker compose up -d
``

Open http://localhost:3001, add your provider keys on the Keys page, reorder the Fallback Chain to taste, and grab your unified API key from the Keys page header. That unified key is what you point your OpenAI SDK at.

Reaching it from another machine? By default the container is published only on 127.0.0.1, so http://<ip>:3001 won’t load from another device (the page just hangs). To expose it on your LAN — e.g. a Raspberry Pi at http://192.168.1.x:3001 — start it with HOST_BIND=0.0.0.0:

bash HOST_BIND=0.0.0.0 docker compose up -d

Only do this on a trusted network: the proxy is single-user and guarded only by the unified API key.

Local development

Prerequisites: Node.js 20+, npm.

bash git clone https://github.com/tashfeenahmed/freellmapi.git cd freellmapi npm install cp .env.example .env ENCRYPTION_KEY="$(node -e 'console.log(require("crypto").randomBytes(32).toString("hex"))')" printf "ENCRYPTION_KEY=%s\nPORT=3001\n" "$ENCRYPTION_KEY" > .env npm run dev

ENCRYPTION_KEY is required for startup. The server only falls back to a database-stored development key when DEV_MODE=true and NODE_ENV is not production; do not use that fallback with real provider keys.

Request analytics are retained for 90 days or 100000 request rows by default, whichever limit prunes first. Set REQUEST_ANALYTICS_RETENTION_DAYS=0 or REQUEST_ANALYTICS_MAX_ROWS=0 in .env to disable either retention limit.

Open http://localhost:5173 (the Vite dev UI), add your provider keys on the Keys page, reorder the Fallback Chain to taste, and grab your unified API key from the Keys page header. That unified key is what you point your OpenAI SDK at.

Reaching the dev UI from another device on your LAN? Use npm run dev:lan — it passes --host through to Vite, which then prints a Network: http://<ip>:5173 URL you can open from a phone or another machine. (Plain npm run dev -- --host does not work here: the root dev script is a concurrently wrapper, so the flag never reaches Vite.)

API calls go through Vite’s dev proxy, so no extra server config is needed.

For a production build without Docker:
bash npm run build node server/dist/index.js # server + dashboard both served on :3001

Docker

FreeLLMAPI publishes a single production image that contains the Express server and the built React dashboard:

``bash
docker pull ghcr.io/tashfeenahmed/freellmapi:latest

or pin a release, e.g. :v1.2.3

``

The image is multi-arch (linux/amd64 + linux/arm64, so it runs on a Raspberry Pi). Published tags: latest (default branch), v*.*.* (git release tags), and sha-<commit>.

The included docker-compose.yml is the recommended install path:

bash docker compose up -d docker compose logs -f freellmapi

By default the container’s port is bound to 127.0.0.1 (localhost only). To reach the dashboard/API from another machine on your network, publish it on all interfaces with HOST_BIND=0.0.0.0 docker compose up -d — only on a trusted LAN, since the proxy is single-user.

SQLite data is stored in the freellmapi-data volume at /app/server/data. Keep the same .env ENCRYPTION_KEY and volume when upgrading, because provider keys are encrypted at rest.

More Docker operations and examples live in docker/README.md.

Desktop app

A native menu-bar app lives in desktop/: the entire router + dashboard running locally from your tray, with a glass popover showing live request stats.
FreeLLMAPI desktop app

Download from Releases (https://github.com/tashfeenahmed/freellmapi/releases/latest) — the macOS .dmg and the Windows .exe installer are built and attached to every release by the desktop-release workflow.

Or build it from this repo in a few minutes:
bash npm install npm run desktop:dist # macOS → desktop/dist-electron/FreeLLMAPI-...-arm64.dmg npm run desktop:dist:win # Windows → "desktop/dist-electron/FreeLLMAPI Setup ....exe"

Locally built apps are unsigned, so Windows SmartScreen may warn on first run
(“More info” → “Run anyway”); the macOS build launches without Gatekeeper prompts.

Languages

The dashboard and the desktop tray ship in 6 languages. The UI auto-detects your browser/system language on first load and you can switch any time from the ⋯ menu; the choice is remembered.

LanguageLocale
Englishen
中文 (简体)zh-CN
Françaisfr
Españoles
Português (Brasil)pt-BR
Italianoit

Translations live in client/src/i18n/locales/ as flat JSON files. To add a language, copy en.json, translate the values, and register the locale in client/src/i18n/I18nProvider.tsx (and desktop/src/i18n.ts for the tray strings) — PRs welcome.

Premium (live catalog)

The router keeps its model catalog fresh on its own: it pulls a signed catalog from freellmapi.co (https://freellmapi.co) twice a day and applies new models, quota changes, and provider quirk fixes to your local DB (your own enable/disable choices and custom providers are never touched; every download is verified against a pinned Ed25519 key before it is applied).

  • Free installs follow a monthly snapshot — zero cost, forever.
  • Premium (https://freellmapi.co/#pricing) ($19/yr or $49 lifetime) follows the live feed, refreshed every 2-3 days, so new free models are in your router the moment they exist. One key covers all your devices; activate it in the dashboard under Premium. Cancel or manage billing self-serve at freellmapi.co/manage (https://freellmapi.co/manage).

The catalog server never sees your prompts, completions, or provider keys — the router stays fully self-hosted either way.

Locally built apps launch without Gatekeeper/SmartScreen warnings — no code signing involved. Full instructions in desktop/README.md.

Using the API

Any OpenAI-compatible client works (Anthropic / Claude clients too — see Anthropic / Claude clients). Examples:

Python
python from openai import OpenAI client = OpenAI( base_url="http://localhost:3001/v1", api_key="freellmapi-your-unified-key", ) resp = client.chat.completions.create( model="auto", # let the router pick; or specify e.g. "gemini-2.5-flash" messages=[{"role": "user", "content": "Summarise the fall of Rome in one sentence."}], ) print(resp.choices[0].message.content) print("Routed via:", resp.headers.get("x-routed-via"))

curl
bash curl http://localhost:3001/v1/chat/completions \ -H "Authorization: Bearer freellmapi-your-unified-key" \ -H "Content-Type: application/json" \ -d '{ "model": "auto", "messages": [{"role": "user", "content": "hi"}] }'

Streaming
python stream = client.chat.completions.create( model="auto", messages=[{"role": "user", "content": "Stream me a haiku about SQLite."}], stream=True, ) for chunk in stream: print(chunk.choices[0].delta.content or "", end="", flush=True)

Tool calling
Pass OpenAI-style tools and tool_choice; the assistant response round-trips back through the proxy exactly like the OpenAI API. Multi-step flows (assistant tool_calls → tool role follow-up) work transparently.

… (the rest of the README is already in English and would continue similarly)


Note: The full README continues with sections like “Screenshots”, “How it works”, “Context Handoff”, etc. Since the user’s input only included up to the tool calling example in the English part, I’ll assume the intent is to translate the Chinese intro and then provide the remaining English content as given. If the user intended the entire README to be included, the output above covers the initial portions. The markdown is preserved.

Similar Articles

@geekbb: Nice, nice. Using this project to combine free models from major tech companies and pool their quotas together. Don't underestimate the free quotas from these 16 LLM providers (totaling about 1.7 billion tokens per month). If used well, it can save a lot. I'll find time to tinker with it. https://github.com/ta…

X AI KOLs Timeline

Introduces an open-source project that aggregates free quotas (totaling about 1.7 billion tokens per month) from 16 LLM providers for unified usage, and mentions Google AI Studio's free API tier, aiming to help developers save costs.