Tag
The author has open-sourced Jev Decisions v1, a dataset of 12 million examples for training AI models on agentic decisions like tool selection and routing, to address data gaps in agent decision-making.
NVIDIA releases Isaac ROS 5.0 with new agentic workflows and support for ROS Lyrical to advance open source robotics development using GPU acceleration.
A new quantized AI model, Sharp-Spark-X2.5-4B-GGUF, is released with improvements for agentic coding on small GPUs and limited RAM, enhancing local coding capabilities for less privileged users.
This paper introduces PlaceReasoner-Beta, a verifier-guided multi-agent framework that reformulates macro placement as a reasoning problem for VLSI physical design, achieving significant improvements in timing and wirelength. It also presents PlaceReasoner-Bench, an open benchmark for evaluating methods using routed PPA and DRC.
Claude Code version 2.1.277 adds support for AGENTS.md as an alternative to CLAUDE.md, allowing customizable project instructions through mods.
Agentic Batch Changes by Sourcegraph is now generally available, enhancing automation and code search capabilities for developers.
Cognition's SWE-2 model, post-trained from Kimi K3, achieves a score of 92.8 on Terminal-Bench 2.1, offering competitive performance with lower cost compared to frontier models like Fable 5.1 and GPT-6 Astra.
The author shares tactics for selling SaaS products like Postiz by leveraging AI tools and viral social media strategies, achieving $2.2m ARR through methods involving ChatGPT Astra and content hacking.
Google DeepMind introduces agentic video understanding for Gemini models, reducing token consumption by up to 88% and improving accuracy in video analysis.
The benchmark reveals that Qwen3.8-Flash-Next-NVFP4 achieves nearly the highest scores among local models, with superior efficiency in request handling and token generation, and notable speed despite being undertrained.
Ornith-1.5-35B-A3B is a mixture-of-experts AI model that activates only 3B parameters per token and outperforms similar-sized models like Qwen and Gemma in coding and agentic benchmarks.
This paper proposes Agentic ESOpt, a method using evolution strategies to enable scalable full-parameter fine-tuning of long-horizon LLM agents with minimal GPU memory requirements.
Bitdrift.ai is introduced as the world's first agentic platform for mobile observability, designed to help monitor and debug mobile applications.
Qwen releases Qwen3.8-27B, an open-weights 27B dense vision-language model with major gains in coding, professional work, and long-horizon agentic tasks, available in FP8 with flexible thinking control.
DeepSeek released DeepSeek-V4-Pro-0813, an open-weights MoE model (1.6T total / 49B active) with MIT license, 1M context, and enhanced agentic capabilities, outperforming its preview and competing with top proprietary models on benchmarks like Terminal-Bench 2.1.
A community member speculates about the upcoming Qwen3.8-27B release, highlighting teased features like VLM, agentic improvements, and a think mode, jokingly suggesting a 'monk' reasoning effort for long-horizon cognitive detachment.
Qwen3.8-27B, a compact 27B dense vision-language model with flexible thinking control and long-context support, is released as the most capable Qwen open model to date, available soon via Hugging Face and Qwen Cloud.
TRACE Bench is a task-driven agentic checklist evaluation framework for roleplay, decomposing role profiles into checklists, using a user agent for natural conversation, and tracing scores back to checklist items and dialogue evidence. It achieves 99.91% coverage, outperforming the MiniMax Role-play Benchmark's 73.74%, and supports closed-loop benchmark evolution across 26 models.
SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier alongside GPT-5.6 Sol and Claude models, with strong agentic performance at lower cost.
A case study describing how an AI coding agent dismantled a core architectural invariant across 189 files in a 717k-line codebase using a specification-first convergence methodology, with no test oracle and no human code review. The agent iteratively refined a 55-page specification and performed correction loops, fixing 201 self-identified mistakes across 31 audit loops.