agentic

Tag

Cards List
#agentic

I open sourced 12M agentic decision examples for training Jev style models

Reddit r/artificial ↗ · 2d ago

The author has open-sourced Jev Decisions v1, a dataset of 12 million examples for training AI models on agentic decisions like tool selection and routing, to address data gaps in agent decision-making.

0 favorites 0 likes
#agentic

NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development

NVIDIA Blog ↗ · 3d ago Cached

NVIDIA releases Isaac ROS 5.0 with new agentic workflows and support for ROS Lyrical to advance open source robotics development using GPU acceleration.

0 favorites 0 likes
#agentic

A better coder for the small-GPU/small-RAM crowd!

Reddit r/LocalLLaMA ↗ · 4d ago

A new quantized AI model, Sharp-Spark-X2.5-4B-GGUF, is released with improvements for agentic coding on small GPUs and limited RAM, enhancing local coding capabilities for less privileged users.

0 favorites 0 likes
#agentic

PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking

arXiv cs.AI ↗ · 4d ago Cached

This paper introduces PlaceReasoner-Beta, a verifier-guided multi-agent framework that reformulates macro placement as a reasoning problem for VLSI physical design, achieving significant improvements in timing and wirelength. It also presents PlaceReasoner-Bench, an open benchmark for evaluating methods using routed PPA and DRC.

0 favorites 0 likes
#agentic

Quoting Thariq Shihipar

Simon Willison's Blog ↗ · 2026-09-18 Cached

Claude Code version 2.1.277 adds support for AGENTS.md as an alternative to CLAUDE.md, allowing customizable project instructions through mods.

0 favorites 0 likes
#agentic

@jdorfman: Agentic Batch Changes by @Sourcegraph is now GA.

X AI KOLs Timeline ↗ · 2026-09-14 Cached

Agentic Batch Changes by Sourcegraph is now generally available, enhancing automation and code search capabilities for developers.

0 favorites 0 likes
#agentic

Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

Hacker News Top ↗ · 2026-09-10 Cached

Cognition's SWE-2 model, post-trained from Kimi K3, achieves a score of 92.8 on Terminal-Bench 2.1, offering competitive performance with lower cost compared to frontier models like Fable 5.1 and GPT-6 Astra.

0 favorites 0 likes
#agentic

@wickedguro: https://x.com/wickedguro/status/2096974398167392500

X AI KOLs Following ↗ · 2026-09-07 Cached

The author shares tactics for selling SaaS products like Postiz by leveraging AI tools and viral social media strategies, achieving $2.2m ARR through methods involving ChatGPT Astra and content hacking.

0 favorites 0 likes
#agentic

Introducing agentic video understanding with Gemini

Google DeepMind Blog ↗ · 2026-09-01 Cached

Google DeepMind introduces agentic video understanding for Gemini models, reducing token consumption by up to 88% and improving accuracy in video analysis.

0 favorites 0 likes
#agentic

Local agentic coding Benchmark : Qwen3.8-Flash-Next NVFP4 vs 27B (and the others...)

Reddit r/LocalLLaMA ↗ · 2026-08-28

The benchmark reveals that Qwen3.8-Flash-Next-NVFP4 achieves nearly the highest scores among local models, with superior efficiency in request handling and token generation, and notable speed despite being undertrained.

0 favorites 0 likes
#agentic

ornith-ai/Ornith-1.5-35B-A3B

Hugging Face Models Trending ↗ · 2026-08-18 Cached

Ornith-1.5-35B-A3B is a mixture-of-experts AI model that activates only 3B parameters per token and outperforms similar-sized models like Qwen and Gemma in coding and agentic benchmarks.

0 favorites 0 likes
#agentic

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Hugging Face Daily Papers ↗ · 2026-08-18 Cached

This paper proposes Agentic ESOpt, a method using evolution strategies to enable scalable full-parameter fine-tuning of long-horizon LLM agents with minimal GPU memory requirements.

0 favorites 0 likes
#agentic

bitdrift.ai

Product Hunt ↗ · 2026-08-17

Bitdrift.ai is introduced as the world's first agentic platform for mobile observability, designed to help monitor and debug mobile applications.

0 favorites 0 likes
#agentic

Qwen 3.8 27B is out: open weights, best local dense model yet

Hacker News Top ↗ · 2026-08-14 Cached

Qwen releases Qwen3.8-27B, an open-weights 27B dense vision-language model with major gains in coding, professional work, and long-horizon agentic tasks, available in FP8 with flexible thinking control.

0 favorites 0 likes
#agentic

@AdinaYakup: DeepSeek V4 Pro's weights are out https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813… - 1.6T / 49B MoE, MIT licens…

X AI KOLs Following ↗ · 2026-08-13 Cached

DeepSeek released DeepSeek-V4-Pro-0813, an open-weights MoE model (1.6T total / 49B active) with MIT license, 1M context, and enhanced agentic capabilities, outperforming its preview and competing with top proprietary models on benchmarks like Terminal-Bench 2.1.

0 favorites 0 likes
#agentic

While waiting for the release of Qwen3.8-27B, let's try to guess what will happen

Reddit r/LocalLLaMA ↗ · 2026-08-13

A community member speculates about the upcoming Qwen3.8-27B release, highlighting teased features like VLM, agentic improvements, and a think mode, jokingly suggesting a 'monk' reasoning effort for long-horizon cognitive detachment.

0 favorites 0 likes
#agentic

@AdinaYakup: Open source summer party is not over yet https://huggingface.co/Qwen/Qwen3.8-27B…

X AI KOLs Timeline ↗ · 2026-08-13 Cached

Qwen3.8-27B, a compact 27B dense vision-language model with flexible thinking control and long-context support, is released as the most capable Qwen open model to date, available soon via Hugging Face and Qwen Cloud.

0 favorites 0 likes
#agentic

TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

arXiv cs.CL ↗ · 2026-08-13 Cached

TRACE Bench is a task-driven agentic checklist evaluation framework for roleplay, decomposing role profiles into checklists, using a user agent for natural conversation, and tracing scores back to checklist items and dialogue evidence. It achieves 99.91% coverage, outperforming the MiniMax Role-play Benchmark's 73.74%, and supports closed-loop benchmark evolution across 26 models.

0 favorites 0 likes
#agentic

SpaceXAI's Grok 4.6 Scores 61 on the Artificial Analysis Intelligence Index

Hacker News Top ↗ · 2026-08-12 Cached

SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier alongside GPT-5.6 Sol and Claude models, with strong agentic performance at lower cost.

0 favorites 0 likes
#agentic

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review

Hugging Face Daily Papers ↗ · 2026-08-12 Cached

A case study describing how an AI coding agent dismantled a core architectural invariant across 189 files in a 717k-line codebase using a specification-first convergence methodology, with no test oracle and no human code review. The agent iteratively refined a 55-page specification and performed correction loops, fixing 201 self-identified mistakes across 31 audit loops.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback