model-update

Tag

Cards List
#model-update

@_mohansolo: We’ve rolled out a new version of Gemini 3.5 Flash in Antigravity that boasts much less and has higher endurance on har…

X AI KOLs Following ↗ · 2026-06-02 Cached

A new version of Gemini 3.5 Flash has been rolled out on Antigravity, with improvements for harder tasks and reset rate limits for all users.

0 favorites 0 likes
#model-update

Claude Opus 4.8: The System Card (40 minute read)

TLDR AI ↗ · 2026-06-01 Cached

Deep analysis of Anthropic's Claude Opus 4.8 system card, detailing incremental improvements in capability, safety evaluations, and alignment risks over Opus 4.7.

0 favorites 0 likes
#model-update

The new Claude scored 0% on "confidently reporting wrong answers" in testing. Here's a prompt that takes advantage of it on anything important.

Reddit r/ArtificialInteligence ↗ · 2026-05-31

Anthropic's Claude Opus 4.8 update dramatically reduces confident but incorrect answers, scoring 0% on reporting flawed results, and a prompt is provided to leverage this improvement for critical self-critique.

0 favorites 0 likes
#model-update

@datacurve: Opus 4.8 is now on DeepSWE. On the default high thinking effort, it scores 6% higher than Opus 4.7 xhigh, while also lo…

X AI KOLs Following ↗ · 2026-05-30 Cached

Opus 4.8 is now available on DeepSWE, scoring 6% higher than Opus 4.7 with reduced average cost per task.

0 favorites 0 likes
#model-update

@KarelDoostrlnck: Do you use ChatGPT in a language other than English? How does the new update feel in your language? We would love your …

X AI KOLs Timeline ↗ · 2026-05-29 Cached

OpenAI shipped a new version of GPT-5.5 Instant with improvements in sycophancy, factuality, and multilingual performance, and is seeking user feedback.

0 favorites 0 likes
#model-update

@FinanceYF5: Claude Opus 4.8 is here: better coding, stronger agent capabilities, dynamic workflows supporting hundreds of parallel sub-agents, same price. Also: A Mythos-level model is expected in weeks. Anthropic has more than one surprise today.

X AI KOLs Timeline ↗ · 2026-05-29 Cached

Claude Opus 4.8 is now available: improved coding and agent capabilities, supports dynamic workflows with hundreds of parallel sub-agents, price unchanged. Anthropic also teases a Mythos-level model within weeks.

0 favorites 0 likes
#model-update

llm-anthropic 0.25.1

Simon Willison's Blog ↗ · 2026-05-28 Cached

llm-anthropic 0.25.1 adds support for Claude Opus 4.8, a fast mode option, and updates default max_tokens behavior.

0 favorites 0 likes
#model-update

@julien_c: now tell me What's the % of weights changed between Opus 4.7 and Opus 4.8 <1%?

X AI KOLs Timeline ↗ · 2026-05-28 Cached

Asking about the percentage of weight changes between Opus 4.7 and Opus 4.8.

0 favorites 0 likes
#model-update

Can liveness detection models generalise to synthetic media generation techniques they were never trained on? [D]

Reddit r/MachineLearning ↗ · 2026-05-21

This discussion examines whether liveness detection models trained on historical deepfake samples can generalize to new synthetic media generation techniques, questioning the update cycle for vendors claiming deepfake detection capabilities.

0 favorites 0 likes
#model-update

MTP PR Merged!!!

Reddit r/LocalLLaMA ↗ · 2026-05-16

A pull request for MTP (likely a model training pipeline or similar) related to LLaMA models has been merged, marking a milestone.

0 favorites 0 likes
#model-update

@OpenAI: Another reason to switch to Codex.

X AI KOLs ↗ · 2026-05-13

OpenAI promotes switching to Codex, highlighting another reason to adopt their AI code generation model.

0 favorites 0 likes
#model-update

@thinkymachines: With the model's simultaneous speech capability, Horace has gotten a lot easier to work with recently.

X AI KOLs Following ↗ · 2026-05-11 Cached

The Horace AI model has become easier to work with following improvements to its simultaneous speech capability. This update was highlighted by thinkymachines.

0 favorites 0 likes
#model-update

@FinanceYF5: 2/ More Concise Responses A key focus of this update is to make responses more concise. OpenAI states that this is a priority area for improvement based on user feedback.

X AI KOLs Following ↗ · 2026-05-10 Cached

OpenAI announces an update focused on making model responses more concise, based on user feedback.

0 favorites 0 likes
#model-update

@KyleHessling1: Guys, I am absolutely astounded. The Qwen 3.6 27b is like a jump to Qwen 4 from Qwen 27B 3.5. I just did a full suite o…

X AI KOLs Following ↗ · 2026-04-22

Early user reports that Qwen 3.6 27B shows dramatic performance gains over 3.5, excelling in front-end design and agentic benchmarks.

0 favorites 0 likes
#model-update

Changes in the system prompt between Claude Opus 4.6 and 4.7

Simon Willison's Blog ↗ · 2026-04-18 Cached

Anthropic released Claude Opus 4.7 with notable system prompt changes including expanded child safety instructions, new tool integrations (Claude in PowerPoint, Chrome, Excel), and behavioral adjustments to reduce verbosity and improve task completion without unnecessary clarification.

0 favorites 0 likes
#model-update

GPT-5.3 Instant: Smoother, more useful everyday conversations

OpenAI Blog ↗ · 2026-03-03 Cached

OpenAI releases GPT-5.3 Instant, an update to ChatGPT's most-used model that improves conversational flow, reduces unnecessary refusals, and decreases hallucinations by up to 26.8% in high-stakes domains. The update focuses on tone, relevance, and practical usability based on user feedback.

0 favorites 0 likes
#model-update

Update to GPT-5 System Card: GPT-5.2

OpenAI Blog ↗ · 2025-12-11 Cached

OpenAI releases GPT-5.2, the latest model in the GPT-5 series, with an updated system card documenting safety mitigations and introducing GPT-5.2 Instant and GPT-5.2 Thinking variants.

0 favorites 0 likes
#model-update

Addendum to GPT-5 System Card: Sensitive conversations

OpenAI Blog ↗ · 2025-10-27 Cached

OpenAI released an update to GPT-5 on October 3 to improve handling of sensitive conversations around mental and emotional distress, reducing inadequate responses by 65-80% through collaboration with 170+ mental health experts. The company published a system card addendum and safety evaluations comparing the new model to the previous August 15 version.

0 favorites 0 likes
#model-update

Tasks that require whole-body control with Gemini Robotics 2

YouTube AI Channels ↗ · 2026-07-30 Cached

Google DeepMind's Gemini Robotics 2 expands physical AI to whole-body control, enabling humanoid robots to perform tasks requiring reaching, bending, and balancing in cluttered spaces.

0 favorites 0 likes
#model-update

Introducing GPT-5.5 with Databricks

YouTube AI Channels ↗ · 2026-05-08 Cached

OpenAI partners with Databricks to release the GPT-5.5 model, achieving a 46% reduction in error rate in agent frameworks, becoming the only model to exceed 50% on benchmarks, with significant improvements in parsing quality and function calling capabilities.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback