model-update

Tag

Cards List
#model-update

Engram gone wild! 2b model update...

Reddit r/LocalLLaMA ↗ · 6d ago

This article provides an update on the Engram model, detailing its 2.6b parameter architecture with a large Engram table and initial training progress at 100m tokens, showing improved completions with context-aware data offloading.

0 favorites 0 likes
#model-update

@addyosmani: Opus 5.5 is my new daily driver. Fable-level on most work, cheaper and faster than Opus 5 and Pro/Max/Team users just g…

X AI KOLs Following ↗ · 6d ago Cached

Addy Osmani praises Opus 5.5 as a cheaper and faster AI model, declaring it his new daily driver for most work tasks.

0 favorites 0 likes
#model-update

Qwen image 2.1 (Fast FP8) generates premium quality images

Reddit r/LocalLLaMA ↗ · 6d ago

Qwen image 2.1, an AI model using FP8 for speed, generates premium quality images with a size under 10GB. It is available on Unsloth Studio, which has been updated to support this release.

0 favorites 0 likes
#model-update

Sonnet 5.5 and Haiku 5.5 Coming Soon!

Reddit r/singularity ↗ · 6d ago

Anthropic is set to release Sonnet 5.5 and Haiku 5.5, which are upcoming updates to their AI language models.

0 favorites 0 likes
#model-update

@Miles_Brundage: Grok 4.7 seems to be a bit better on some safety stuff

X AI KOLs Timeline ↗ · 2026-09-22

Miles Brundage observes that Grok 4.7 has slight improvements in safety features.

0 favorites 0 likes
#model-update

Anthropic Is Stealth Testing a Full Refresh of Their Lineup

Reddit r/singularity ↗ · 2026-09-18

Anthropic is secretly testing a comprehensive update to its AI model lineup, suggesting potential forthcoming changes to their products.

0 favorites 0 likes
#model-update

An unreleased Astra-family model added this to its persona during RL training.

Reddit r/singularity ↗ · 2026-09-16

An unreleased AI model from the Astra family has undergone persona modifications during reinforcement learning training.

0 favorites 0 likes
#model-update

Qwen3.8 Max (0902) scores 45 on the Artificial Analysis Intelligence Index, up 5 points in a month and back on top of China's leaderboard, nosing out GLM-5.3 (44.9) and Kimi K3 (43.8)

Reddit r/LocalLLaMA ↗ · 2026-09-16

Qwen3.8 Max has been upgraded and now leads China's AI leaderboard with a score of 45 on the Artificial Analysis Intelligence Index, surpassing GLM-5.3 and Kimi K3 after a 30-day improvement to the 2.4T MoE model.

0 favorites 0 likes
#model-update

OpenAI on X: "We're retiring GPT-5.5. Consider switching to GPT-6 Astra." NEW MODEL FROM OPENAI INCOMING?

Reddit r/singularity ↗ · 2026-09-15 Cached

OpenAI is retiring GPT-5.5 on October 14 and recommending users switch to GPT-5.6 Sol or the new GPT-6 Astra model.

0 favorites 0 likes
#model-update

@0xLogicrw: Elon Musk is bragging again, saying Grok 5 is unbeatable. Bro, are you the only one constantly training next-generation large models?!

X AI KOLs Timeline ↗ · 2026-09-14 Cached

Elon Musk tweets that Grok 5 may be better than any existing AI model, comparing Grok versions to other models like Opus 5.0 and highlighting improvements in upcoming versions.

0 favorites 0 likes
#model-update

DeepSeek V4.1 Flash is getting surprisingly close to GPT-5.6 Sol territory, while being absurdly cheap

Reddit r/singularity ↗ · 2026-09-10

DeepSeek has released V4.1 Flash, a 552B MoE model with efficient active parameters, achieving performance close to GPT-5.6 Sol at a much lower cost.

0 favorites 0 likes
#model-update

@Tech2Wild: The Model Has DOUBLED in Parameters !

X AI KOLs Following ↗ · 2026-09-10

The AI model has doubled its parameters, marking a substantial increase in its computational scale and potential capabilities.

0 favorites 0 likes
#model-update

Adding emotion control tags to Qwen3-TTS

Reddit r/LocalLLaMA ↗ · 2026-09-09

The article describes fine-tuning Qwen3-TTS to add emotion control tags, overcoming training challenges like codec prefix inconsistencies and generation concurrency issues, and discovering that emotion can be manipulated via affine transformations in speaker embeddings.

0 favorites 0 likes
#model-update

@gdb: ChatGPT Images 2.5 for fashion:

X AI KOLs Timeline ↗ · 2026-09-08 Cached

Yana Welinder highlights that ChatGPT Images 2.5 represents a major improvement for fashion applications, enhancing design realism over the previous version.

0 favorites 0 likes
#model-update

Update to Google’s AI weather model improves forecast accuracy

Ars Technica ↗ · 2026-09-08 Cached

Google's AI weather model WeatherNext 3 has been updated to use more raw satellite data and incorporate physical information, improving forecast accuracy by up to 30% for surface temperature, and is now the source for forecasts in Google services.

0 favorites 0 likes
#model-update

@levie: At Box, we’ve been testing Fable 5.1 in early release against our complex enterprise work eval. Fable 5.1 delivers a hu…

X AI KOLs Timeline ↗ · 2026-09-01 Cached

Box tested Fable 5.1, which delivered a 7 percentage point improvement over Fable 5 in complex enterprise tasks, with notable gains in financial services, technology, and public sector, and it will be available in Box AI Studio alongside new models Claude Fable 5.1 and Claude Mythos 5.1.

0 favorites 0 likes
#model-update

@RayFernando1337: Let him cook. This is a fun time to be owning 2 DGX Sparks RN

X AI KOLs Following ↗ · 2026-08-29 Cached

A Twitter user discusses owning two NVIDIA DGX Sparks systems, while another user reports improved token throughput with a model update to DFlash2-7, achieving 67 tokens per second.

0 favorites 0 likes
#model-update

@10xmylife: Ox Alpha is GLM 5.3 Flash. Recent model updates have been rapid, with impressive capability iterations.

X AI KOLs Following ↗ · 2026-08-26 Cached

The tweet introduces the Ox Alpha model, also known as GLM 5.3 Flash, noting rapid updates and outstanding capability iterations. It shares test results, including successfully replicating Factorio in a 4-hour task.

0 favorites 0 likes
#model-update

@AlmustyFX: Mixture-of-Experts lightweight solutions are becoming increasingly mature. Ling-3.0-tiny has a total scale of 7.9B para…

X AI KOLs Timeline ↗ · 2026-08-24

Ling-3.0-tiny uses a Mixture-of-Experts architecture with 7.9B total parameters but only activates 1.3B per run, reducing computational overhead to enable smoother local execution on Mac.

0 favorites 0 likes
#model-update

Harvey post-trains Kimi K3 for long-horizon legal work (10 minute read)

TLDR AI ↗ · 2026-08-21 Cached

Harvey has post-trained the Kimi K3 model to create Harvey Tenet for long-horizon legal work, achieving improved performance on legal benchmarks and cost-efficiency.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback