Tag
Xiaomi introduces MiMo-V2.6, an omnimodal AI model with audio/voice features, claiming performance on par with leading models like Claude Opus 5 and GPT-5.6 Sol.
Grok 4.7 shows significant improvements in legal work and terminal tasks, outperforming competitors on benchmarks like Harvey Legal Agent Benchmark and Terminal-Bench 4.0, while keeping token prices unchanged.
Benzi is an AI coding agent that outperforms competitors on benchmarks by reading less source code and using deterministic tool calls, achieving high efficiency in code intelligence tasks.
Abhinav Anand, a 20-year-old independent builder from Bihar, released Arcle V1, an open-weight 5.84B-parameter unified omni AI model that outperforms Apple's AFM 3B model on multiple benchmarks including MATH-500.
H3 Max is an AI model that generates 5 seconds of 768p video with native stereo audio in under 3 seconds, faster than real-time, and ranks #1 on Design Arena and Artificial Analysis for image-to-video. A demonstration shows its use in continuous video streaming on Twitch.
This paper introduces Meta^n, a method for recursive self-improvement in LLM agents by applying a fixed meta-operation to expand reasoning depth, outperforming prior approaches on benchmarks like ARC-AGI-2.
AutoSaddler is an automatic harness optimization framework that improves LLM agent performance on long-horizon tasks by iteratively updating harnesses using failure signals, achieving substantial gains on benchmarks like GAIA2 and SWE-Bench.
Ornith-1.5 releases a family of open-source LLMs from 9B to 397B parameters, achieving state-of-the-art performance among comparable models and offering multiple deployment-friendly formats.
This paper demonstrates that the structure retention in embedding spaces, measured via nearest-neighbor overlap and ICA differences, strongly correlates with benchmark performance across multiple tasks, offering a predictive metric for model effectiveness.
AntAngelMed is a newly open-sourced 100B-parameter medical language model developed by Zhejiang Health Information Center, Ant Healthcare, and Anzhen'er Medical AI. It achieves top rankings on HealthBench and MedAIBench, utilizing efficient MoE architecture for high-performance inference.
Perceptron Inc. released its flagship video analysis model Mk1, claiming 80-90% lower cost than competitors while achieving strong performance on spatial and video reasoning benchmarks.
Interfaze introduces a hybrid AI model architecture combining CNN/DNN specialization with transformer capabilities, achieving superior accuracy on deterministic tasks like OCR and translation while maintaining cost efficiency at scale.