Tag
MorrowCache is a local OpenAI-compatible proxy that caches AI model responses based on intent similarity to reduce latency and token costs, using an adjudicator to check for cached answers.
2BA.AI provides EU-hosted AI infrastructure with a flat €20/month fee, offering 4,500 requests per 5-hour window and integration with tools like Cursor and VS Code while ensuring GDPR compliance.
Lorivo is a serverless platform that allows sharing GPU servers for hosting multiple LoRA adapters, simplifying deployment and reducing costs for fine-tuned AI models.
EmperoAI celebrates reaching 2000 followers by launching a free community API endpoint for the Qwen3.8-27B-FP8 AI model, allowing developers to access it with any API key.
TeamoRouter offers free or low-cost access to AI models like DeepSeek V4 Pro, compatible with development tools such as Claude Code and Codex, providing convenient integration through a unified API.
Eric Zakariasson announces that Grok 4.6 is available via an API compatible with OpenAI's format, enabling easy integration as a drop-in replacement.
Hetzner offers free access to the Qwen 3.8 27B AI model via an OpenAI-compatible API, claiming it performs better than GPT-5.6 Luna in tests.
The article introduces Gonka, a decentralized inference network that enables access to the V4-Flash AI model with full 1M context without requiring local GPU ownership, using an OpenAI-compatible interface.
CORS Chat is a browser-based tool for chatting with OpenAI Responses-compatible API endpoints, featuring custom headers, local conversation saving, and progressive SVG rendering.
A free public endpoint for the Qwen3.8-27B AI model has been deployed, offering an OpenAI-compatible API with vision support, tool calls, and a large context window, powered by Hugging Face Inference Endpoints for at least 72 hours.
DeepSeek launches V4-Pro and V4-Flash with flexible reasoning effort, native OpenAI Responses API support, and optimized agent workflows for Codex, available via API and app/web.
FAAAH is a dependency-free, OpenAI-compatible proxy that converts API requests into text files and uses AI coding agents like Claude Code as the backend, enabling reuse of existing subscriptions instead of paying for cloud LLM APIs.
Google Cloud API Gateway now offers model routing in Public Preview, providing a serverless ingress layer that accepts OpenAI-compatible requests and dynamically routes them to Gemini, Claude, or OpenAI models.
Google Cloud API Gateway announces public preview of model routing, letting developers access Gemini, Claude, and OSS models via a single endpoint using OpenAPI specs.
Celeris-1 is a new language model using diffusion-based inference architecture, achieving near-GPT-5 level intelligence with 15x faster response times and high throughput.
Hetzner has launched an experimental LLM inference API service, offering an OpenAI-compatible endpoint with the Qwen3.6-35B-A3B-FP8 model. The service is free during the experiment period, has no SLA, and is intended to gather user feedback.
Ramp is open-sourcing its internal LLM router that automatically selects the best model for each request to optimize cost and performance.
CLIProxyAPI 是一个为 CLI 提供兼容 OpenAI/Gemini/Claude/Codex/Grok API 接口的代理服务器,现已支持 OpenAI Codex 和 Claude Code 的 OAuth 访问,方便开发者使用本地或多账户进行 CLI 访问。
Mesh LLM is a distributed AI computing platform that pools idle GPUs across multiple machines to run large language models, exposing a single OpenAI-compatible API. It leverages iroh's peer-to-peer networking to enable private, decentralized inference without a central server.
Another open-source tool on GitHub, Shimmy, is a single 5MB file written in Rust that provides fast and stable local inference with a full OpenAI-compatible API, targeting Ollama's pain points. It starts in under 100ms and uses about 50MB of memory.