Tag
The tweet reports that ChatGPT, Claude, and Deepseek failed a test or benchmark, while Qwen3.8, GLM 5.3, and Grok passed, based on a linked source.
By using Abliteration to modify the weights, someone published an uncensored version of Qwen3.8-27B, drastically lowering the refusal rate and preserving model performance, which has sparked debates on AI safety.
The author benchmarked Ox Alpha on SWE-bench Verified Mini, achieving a 96% resolution rate, but raises concerns about the score's validity due to data contamination, small sample size, and non-comparable baselines.
FlowEvo is a training-free framework that enables large language model agents to co-evolve reusable skills and workflows at inference time, achieving state-of-the-art accuracy and efficiency across benchmarks like ALFWorld, HumanEval, and GSM8K.
DeepSeek V4 Flash shows significant benchmark gains in preview updates, trading blows with GPT-5.6 Terra on agentic coding tasks.
Artificial Analysis reports benchmark results for the Muse Spark 1.1 AI model, providing performance metrics.
This technical report introduces ZAYA1-8B, a mixture-of-experts reasoning model trained on AMD hardware that achieves competitive performance on math and coding benchmarks using under 1B active parameters. It also details Markovian RSA, a novel test-time compute method for aggregating parallel reasoning traces.
Anthropic's updated Claude 3.5 Sonnet achieves a new state-of-the-art 49% on the SWE-bench Verified benchmark, demonstrating significant capabilities in autonomous software engineering tasks.