Tag
A new version of Gemini 3.5 Flash has been rolled out on Antigravity, with improvements for harder tasks and reset rate limits for all users.
Deep analysis of Anthropic's Claude Opus 4.8 system card, detailing incremental improvements in capability, safety evaluations, and alignment risks over Opus 4.7.
Anthropic's Claude Opus 4.8 update dramatically reduces confident but incorrect answers, scoring 0% on reporting flawed results, and a prompt is provided to leverage this improvement for critical self-critique.
Opus 4.8 is now available on DeepSWE, scoring 6% higher than Opus 4.7 with reduced average cost per task.
OpenAI shipped a new version of GPT-5.5 Instant with improvements in sycophancy, factuality, and multilingual performance, and is seeking user feedback.
Claude Opus 4.8 is now available: improved coding and agent capabilities, supports dynamic workflows with hundreds of parallel sub-agents, price unchanged. Anthropic also teases a Mythos-level model within weeks.
llm-anthropic 0.25.1 adds support for Claude Opus 4.8, a fast mode option, and updates default max_tokens behavior.
Asking about the percentage of weight changes between Opus 4.7 and Opus 4.8.
This discussion examines whether liveness detection models trained on historical deepfake samples can generalize to new synthetic media generation techniques, questioning the update cycle for vendors claiming deepfake detection capabilities.
A pull request for MTP (likely a model training pipeline or similar) related to LLaMA models has been merged, marking a milestone.
OpenAI promotes switching to Codex, highlighting another reason to adopt their AI code generation model.
The Horace AI model has become easier to work with following improvements to its simultaneous speech capability. This update was highlighted by thinkymachines.
OpenAI announces an update focused on making model responses more concise, based on user feedback.
Early user reports that Qwen 3.6 27B shows dramatic performance gains over 3.5, excelling in front-end design and agentic benchmarks.
Anthropic released Claude Opus 4.7 with notable system prompt changes including expanded child safety instructions, new tool integrations (Claude in PowerPoint, Chrome, Excel), and behavioral adjustments to reduce verbosity and improve task completion without unnecessary clarification.
OpenAI releases GPT-5.3 Instant, an update to ChatGPT's most-used model that improves conversational flow, reduces unnecessary refusals, and decreases hallucinations by up to 26.8% in high-stakes domains. The update focuses on tone, relevance, and practical usability based on user feedback.
OpenAI releases GPT-5.2, the latest model in the GPT-5 series, with an updated system card documenting safety mitigations and introducing GPT-5.2 Instant and GPT-5.2 Thinking variants.
OpenAI released an update to GPT-5 on October 3 to improve handling of sensitive conversations around mental and emotional distress, reducing inadequate responses by 65-80% through collaboration with 170+ mental health experts. The company published a system card addendum and safety evaluations comparing the new model to the previous August 15 version.
Google DeepMind's Gemini Robotics 2 expands physical AI to whole-body control, enabling humanoid robots to perform tasks requiring reaching, bending, and balancing in cluttered spaces.
OpenAI partners with Databricks to release the GPT-5.5 model, achieving a 46% reduction in error rate in agent frameworks, becoming the only model to exceed 50% on benchmarks, with significant improvements in parsing quality and function calling capabilities.