Tag
GLM-5.3-Flash model is scheduled for release on the 1x Spark platform this Thursday.
The article highlights the GLM-5.3-Flash EXL3-3.0bpw AI model with an inference speed of 193.8 tokens per second, attributed to multiple contributors.
Release of GLM-5.3-Flash, an AI language model optimized for fast inference and performance updates.
DeepSeek-V4-Flash-Vision-Exp is an experimental or updated AI model from DeepSeek focusing on vision capabilities.
Gemini 3.7 Flash takes first place on the new AA-AnalystAgent benchmark, excelling in accuracy (60% pass^5), speed (1.32s per task), and cost efficiency across 80 real-world quantitative analysis tasks in multiple domains.
Gemini 3.7 Flash is promoted as a highly intelligent AI model, with a 50% discount exclusively available on OpenRouter until August 27, and is noted for its competitiveness in multimodal and agentic workloads.
Google has released Gemini 3.7 Flash, a new AI model variant that arrives with characteristics different from what was anticipated.
Tweet reports that Ling 3.0 Flash on AMD Strix Halo is significantly faster than Qwen-122b using ROCm-optimized formats, but notes tool calls are broken in certain harnesses.
Google has deprecated temperature, top_p, and top_k parameters starting with Gemini 3.6 Flash and 3.5 Flash-Lite models, as detailed in their developer guide.
This tweet benchmarks Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding, finding that while the Flash series initially excelled at visual understanding, recent versions have plateaued or regressed due to posttraining for coding and reasoning.
Google's Gemini 3.6 Flash family introduces three new models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber.
Google DeepMind announces Gemini 3.6 Flash, a new version with improved production-ready code generation and multimodal capabilities for analyzing charts and documents.
Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, focusing on efficiency, coding, and cybersecurity, but the anticipated Gemini 3.5 Pro was not included due to internal delays.
OpenWiki now supports Google's Gemini AI Studio and Vertex AI, including the newly released Gemini 3.6 Flash and 3.5 Flash Lite models, thanks to community contributions.
Google launches Gemini 3.5 Flash Cyber, a cost-efficient AI security model for vulnerability detection, alongside Gemini 3.6 Flash and 3.5 Flash-Lite, positioning it as a cheaper alternative to Anthropic's Mythos.
Google silently released Gemini 3.6 Flash, an updated version of its efficient language model.
DeepSeek-V4-Flash-DSpark achieves 328 tok/s single inference and 1.7k tok/s batch throughput on 4x RTX PRO 6000 GPUs.
Promotes Gemini 3.5 Flash as a faster, cheaper, and more accurate model for OCR and VQA tasks.
DeepSeek V4 Flash on dual RTX PRO 6000 GPUs completes real coding tasks faster than Anthropic's Sonnet and Opus models while achieving similar quality to Sonnet.
Google may be testing an upgraded Gemini Flash model on LM Arena, showing incremental improvements over the current version, with possible naming as Gemini 3.6 Flash or Gemini 4 Flash.