Tag
This article provides an update on the Engram model, detailing its 2.6b parameter architecture with a large Engram table and initial training progress at 100m tokens, showing improved completions with context-aware data offloading.
Addy Osmani praises Opus 5.5 as a cheaper and faster AI model, declaring it his new daily driver for most work tasks.
Qwen image 2.1, an AI model using FP8 for speed, generates premium quality images with a size under 10GB. It is available on Unsloth Studio, which has been updated to support this release.
Anthropic is set to release Sonnet 5.5 and Haiku 5.5, which are upcoming updates to their AI language models.
Miles Brundage observes that Grok 4.7 has slight improvements in safety features.
Anthropic is secretly testing a comprehensive update to its AI model lineup, suggesting potential forthcoming changes to their products.
An unreleased AI model from the Astra family has undergone persona modifications during reinforcement learning training.
Qwen3.8 Max has been upgraded and now leads China's AI leaderboard with a score of 45 on the Artificial Analysis Intelligence Index, surpassing GLM-5.3 and Kimi K3 after a 30-day improvement to the 2.4T MoE model.
OpenAI is retiring GPT-5.5 on October 14 and recommending users switch to GPT-5.6 Sol or the new GPT-6 Astra model.
Elon Musk tweets that Grok 5 may be better than any existing AI model, comparing Grok versions to other models like Opus 5.0 and highlighting improvements in upcoming versions.
DeepSeek has released V4.1 Flash, a 552B MoE model with efficient active parameters, achieving performance close to GPT-5.6 Sol at a much lower cost.
The AI model has doubled its parameters, marking a substantial increase in its computational scale and potential capabilities.
The article describes fine-tuning Qwen3-TTS to add emotion control tags, overcoming training challenges like codec prefix inconsistencies and generation concurrency issues, and discovering that emotion can be manipulated via affine transformations in speaker embeddings.
Yana Welinder highlights that ChatGPT Images 2.5 represents a major improvement for fashion applications, enhancing design realism over the previous version.
Google's AI weather model WeatherNext 3 has been updated to use more raw satellite data and incorporate physical information, improving forecast accuracy by up to 30% for surface temperature, and is now the source for forecasts in Google services.
Box tested Fable 5.1, which delivered a 7 percentage point improvement over Fable 5 in complex enterprise tasks, with notable gains in financial services, technology, and public sector, and it will be available in Box AI Studio alongside new models Claude Fable 5.1 and Claude Mythos 5.1.
A Twitter user discusses owning two NVIDIA DGX Sparks systems, while another user reports improved token throughput with a model update to DFlash2-7, achieving 67 tokens per second.
The tweet introduces the Ox Alpha model, also known as GLM 5.3 Flash, noting rapid updates and outstanding capability iterations. It shares test results, including successfully replicating Factorio in a 4-hour task.
Ling-3.0-tiny uses a Mixture-of-Experts architecture with 7.9B total parameters but only activates 1.3B per run, reducing computational overhead to enable smoother local execution on Mac.
Harvey has post-trained the Kimi K3 model to create Harvey Tenet for long-horizon legal work, achieving improved performance on legal benchmarks and cost-efficiency.