Tag
OpenBMB releases MiniCPM-RobotTrack, a compact vision-language-action model built on MiniCPM4-0.5B for natural-language-conditioned target tracking, achieving 5+ FPS on unitree Go2, with quality-driven self-evolving data and efficient onboard inference.
Thinking Machines' Inkling model has achieved the top ranking among US open weight models.
InvincibleHunter claims GPT-5.6 Sol is the most impressive model ever, with huge improvements in mathematics and vision capabilities, surpassing other models.
A model's two favorite picks have correctly predicted the World Cup champion for the past 10 tournaments.
Own the agent learning loop and rent the model; private evals, traces, and memory compound over time.
Share the most usage-saving Codex configuration: Model 5.6 Sol, Effort set to Extra High or Max, Speed set to Standard. Ultra speed is ineffective, Fast consumes too many tokens.
Alexandr Wang stated that Muse Spark 1.1 is an industry-competitive agentic and coding model, capable of competing with GPT-5.5 and Opus 4.8 in multiple benchmarks. It is now available via the Meta Model API and Meta AI.
Prism ML releases Ternary-Bonsai-27B-mlx-2bit, a ternary-quantized 27B-parameter language model that achieves ~95% of FP16 performance while fitting in ~7.2 GB, enabling full reasoning on laptops.
Mention of Qwen 3.6 27b model in context of Dspark.
This paper introduces Goku, a million-scale dataset and benchmark for instruction-based video editing, supporting multi-task and structural manipulations. The accompanying model, Goku-Edit, achieves up to +8% improvement on instruction following over open-source models.
DataClaw0 proposes an agentic data tailoring paradigm that uses learnable data processing to structure high-entropy multimodal streams, achieving robust alignment via SFT and GRPO on a novel benchmark.
The article questions the validity of vendor benchmarks for Alpie Core 32B, a 4-bit reasoning coding model optimized for low VRAM and agent workflows, noting a lack of independent benchmark replication.
Maxime Labonne shares that their model is trending on Hugging Face and is surprisingly capable at agentic tasks despite having only 1B active parameters.
A user expresses concern that the Qwen 3.7 27b model may not be shipping.
The thread benchmarks Qwen3.6-27B's thinking mode vs non-thinking on 300+ problems, revealing surprising results for the popular local model and its derivative ecosystem on Hugging Face.
An OpenAI researcher claims that their model's solution to an Erdős problem in discrete geometry is the biggest AI achievement to date, but predicts it will be overshadowed by end of year.
A new breakthrough in video generation models is being reported, potentially signaling significant advancements in AI video capabilities.
GPT-5.5-Cyber is now in limited preview for defenders, offering a capable model for securing critical infrastructure.