Tag
VDN-Minimax-H3 is an open-source hybrid-attention model that speeds up video generation with near-lossless quality, featuring fast inference and plug-and-play adapters powered by MiniMax H3.
Viggle-Animate is an AI model that replaces characters in videos using a single repainted frame, achieving efficient and accurate motion transfer without complex intermediate representations.
Fast Inference is an LLM API service that offers fast and affordable access to top open-source and closed-source models for developers, with integrations for coding agents and tools like Claude Code and Codex.
IBM releases two new compact AI models, Granite Speech 5.0 Turbo CTC, for extremely fast and accurate English speech transcription, achieving over 12,600 RTFx on NVIDIA H200 GPU.
The user praises the Ling 3.0 Tiny AI model for being fast and efficient on low-end PCs, comparing it favorably to models like Qwen 3.5 9b and Gemma 12.
Molei Tao introduces FLARE, a diffusion language model that achieves near GPT5 performance with significantly faster inference speed.
Philipp Schmid praises an unnamed AI model for being cost-effective and blazing fast.
Liquid AI released two new encoder models, LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, that are fast, easy to train, and strongly multilingual, with speed benchmarks showing over 3.7x improvement on CPU compared to ModernBERT-base.
A single viral post from investor Matt Shumer dramatically boosted Groq's business by effectively demonstrating the value of fast inference, a point Groq founder Jonathan Ross had struggled to convey for years.
Google DeepMind released Nano Banana 2 Lite (also known as Gemini 3.1 Flash Lite Image), positioned as the fastest and cheapest Gemini image model, optimized for speed and scale.
Mimo 2.5 demonstrates fast performance with large context windows using dual RTX Pro 6000 GPUs.
Google introduces DiffusionGemma, an experimental 26B MoE open model that achieves up to 4x faster text generation on GPUs using text diffusion, targeting speed-critical interactive local workflows.
Mosaic is a probabilistic weather model that matches state-of-the-art skill while generating a 24-member, 10-day global forecast in under 12 seconds on a single H100.
Santiago (@svpino) highlights MiniMax-M2.7, a 230B open-weight model that rivals top proprietary models like Opus 4.6 and GPT-5.4, achieving 440+ tokens/s inference on SambaNova at low cost.
Baidu releases ERNIE-Image-Turbo, a distilled text-to-image generation model that achieves fast generation in 8 inference steps while maintaining strong text rendering, instruction following, and structured image generation capabilities.
OpenAI has launched the GPT Image 2.5 series of image generation models, including the high-quality flagship Sunburst and the ultra-fast Flare, now available on the API, ChatGPT, and Codex platforms.
P-Image is Pruna's text-to-image generation model that produces state-of-the-art images in less than a second, offering a combination of speed, affordability, and quality.
Pruna's p-image-edit is a premium AI model on Replicate offering fast state-of-the-art image editing under one second, combining speed, affordability, and high visual quality with precise prompt adherence and text rendering capabilities.