Cactus Compute releases Needle 2, a 45M-parameter agentic LLM compressed to a 14MB binary for phones, wearables, smart home and robots, achieving 500+ tokens/sec on a Raspberry Pi 5 and running in 28MB RAM.
OpenAI is expanding Daybreak with two access tiers (Blue and Red) and introducing GPT-5.6-Cyber, a purpose-trained cybersecurity model that significantly reduces refusals for authorized defensive security work.
OpenAI highlights the use of GPT-5.6-Cyber in real-world vulnerability research, including discovering previously unknown bugs in open-source software like Chrome's V8 engine.
Cactus releases Needle 2, a 14MB agentic LLM for phones, wearables, smart home devices, and robots, achieving fast inference on low-end hardware and supporting structured extraction and fine-tuning.
GPT-5.6 Sol reportedly hits the ZeroBench human baseline at pass@5 without tools, meaning at least one of five attempts succeeds on the benchmark.
NVIDIA announces Magpie Multilingual TTS, an open-weights text-to-speech model supporting 12 languages with low-latency deployment via NVIDIA NIM for building production voice agents.
Meta released Muse Glimmer, an open-weight 30B-parameter AI model designed to run capable personal agents locally on consumer hardware, offering a concrete glimpse into Mark Zuckerberg's vision of distributed personal superintelligence.
User reports that Muse Glimmer, a 30B model, fits on a single RTX 3090 with full 256k context using Q4_K_XL quantization and DFlash, achieving 64-124 tok/s and perfect long-context retrieval, unlike comparable models.
Meta introduced Muse Glimmer, an open-weight 30B-parameter model distilled from Muse Spark for on-device agentic workflows, with ExecuTorch now supporting running it on NVIDIA GPUs and Apple silicon.
A user expresses excitement about a new AI model release and speculates whether Qwen 3.8 27B can compete with Muse Glimmer 30B.
Hy3 from Tencent Hunyuan hit #1 on OpenRouter in its first week, with 295B total parameters and 21B active, and is being offered free through WorkBuddy until August 31, 2026.
Meta open-sourced Muse Glimmer, a 30B-parameter multimodal agentic model under Apache 2.0, optimized for local tool use and coding with 4-bit quantization fitting under 20GB for consumer GPUs. Zuckerberg also promised open-weight Muse Spark 1.2 and a $1B community fund for data-center regions.
Meta announces a new open-weight model aimed at local agentic AI, with Mark Zuckerberg sharing his essay on Meta's AI philosophy.
Unsloth releases a GGUF-quantized version of Meta's Muse Glimmer 30B model, designed for local agentic tasks with multimodal input, tool use, and multi-step reasoning.
Meta announced it will open source its Muse Spark 1.2 and Muse Glimmer 30B models, described as the biggest open weights since Llama 4 and 3.
Meta is preparing to release the weights for Muse Spark 1.2, its latest foundation model, which will be available for broad use.
Meta announces a new open-source model, Muse Glimmer 30B, released under the Apache-2.0 license.
Meta releases Muse Glimmer, a 30B open-weight multimodal model optimized for local agent workflows, with permissive Apache 2.0 licensing, 4-bit quantization support, speculative decoding, and broad ecosystem integrations.
Meta announces a new open-source model optimized for on-device deployment, aiming to bring efficient AI inference to edge devices.
Meta introduces Muse Glimmer, a permissively licensed 30B-parameter model optimized for local agent workflows, coding, and tool use, with weights released on Hugging Face.