Tag
Agnost AI launched its first model, agnost-*******-0.1, trained on production traces from a customer, achieving a 22.9% increase in task success and a 90.2% reduction in some metric.
A Stanford and MIT research paper shows that optimizing the Python harness around LLMs can yield up to a 6x performance gap without changing model weights, with systems like Meta-Harness automating context evolution.
opencode2 has achieved up to 30x faster file reading speeds, attributed to contributions from mathematicians.
A blog post detailing how the author used Codex to optimize a kernel in a GPU Mode contest, achieving a 232x speedup in QR decomposition and sharing learnings on auto-research.
This article describes a drop-in chat template fix for Qwen AI models that improves accuracy, reduces token usage, and speeds up responses for knowledge work and coding tasks.
Fable used Claude AI to one-shot a Rust rewrite of the TerminalTextEffects Python library, reducing startup time from 87ms to 2ms and increasing rendering speed by 9.6x, with zero dependencies and a single 3MB executable.
OpenLoco v26.07 is released, featuring extra zoom levels, cargo statistics, UI improvements, bug fixes, and performance optimizations.
Linux 7.2 kernel merges a performance optimization for anonymous/unnamed pipes, improving throughput by 6-48% and reducing latency by 17-33% by pre-allocating pages outside of mutex lock to avoid contention.
Kimi released and open-sourced Kimi 2.7 Code, a coding model with improved performance, reduced reasoning tokens, and long-horizon coding abilities.
The 'Gentle Coding' technique is empirically validated across 1,500+ tests, showing significant improvements (zero regression) for multiple models including Kimi K2.6, GLM-5.1, GPT 5.4/5.5, and Claude Sonnet 3.5/Opus 4.6 by reducing looping and hallucinations.
GPT-Realtime-2 demonstrates a 15 percentage point improvement over version 1.5 on the Big Bench Audio benchmark, approaching saturation levels.
OpenAI released two new embedding models: text-embedding-3-small (5x cheaper than ada-002 with 40%+ MIRACL improvement) and text-embedding-3-large (best performance with up to 3072 dimensions). Both models show significant performance gains on standard benchmarks while reducing costs.
GPT-5.6 achieves a 20% token reduction and improved performance on the Base44 platform compared to GPT-5.5, accelerating application development and enhancing resource efficiency.