agentic

Tag

Cards List
#agentic

BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking

arXiv cs.AI ↗ · 2026-07-14 Cached

This paper introduces BatteryLake, a governed data lakehouse that uses LLM agents for evidence-grounded metadata extraction and schema mapping, with human-in-the-loop verification, to curate heterogeneous battery aging datasets and release an open benchmark.

0 favorites 0 likes
#agentic

@samsja19: We are also releasing prime-rl 0.7.0 which has full support for verifiers v1 and bring your own harness for training. W…

X AI KOLs Following ↗ · 2026-07-14 Cached

Prime Intellect released verifiers v1 and prime-rl 0.7.0, an RL training tool with full support for verifiers, multiple algorithms like GRPO and OPD, and performance improvements.

0 favorites 0 likes
#agentic

OpenProver: Agentic and Interactive Theorem Proving with Lean 4

arXiv cs.AI ↗ · 2026-07-13 Cached

OpenProver is an open-source system for LLM-driven automated theorem proving using Lean 4, featuring a Planner-Worker-Verifier architecture and both autonomous and interactive modes. It enables reproducible evaluation and human-AI synergy in mathematical proof search.

0 favorites 0 likes
#agentic

@sheriyuo: A 35B-parameter MoE agentic model with only 3B active that claims to match or surpass 100B-class models through post-tr…

X AI KOLs Timeline ↗ · 2026-07-13 Cached

A 35B-parameter MoE model with only 3B active parameters matches or surpasses 100B-class models using post-training RL, achieving significant efficiency gains.

0 favorites 0 likes
#agentic

@qingke_ai: https://x.com/qingke_ai/status/2076354316848550126

X AI KOLs Timeline ↗ · 2026-07-12 Cached

MAD-OPD utilizes a multi-teacher debate mechanism to break through the single-teacher distillation ceiling, enabling small models to surpass large teacher models in tool invocation and code generation tasks.

0 favorites 0 likes
#agentic

@FinanceYF5: Alexandr Wang stated that Muse Spark 1.1 is an industry-competitive agentic and coding model. In multiple agentic benchmarks, it can compete with GPT-5.5 and Opus 4.8. It is now available via the new Meta…

X AI KOLs Following ↗ · 2026-07-11 Cached

Alexandr Wang stated that Muse Spark 1.1 is an industry-competitive agentic and coding model, capable of competing with GPT-5.5 and Opus 4.8 in multiple benchmarks. It is now available via the Meta Model API and Meta AI.

0 favorites 0 likes
#agentic

Agentic Alexa with Long Term Memory and connection to 1000+ apps.

Reddit r/ArtificialInteligence ↗ · 2026-07-10

Amazon has announced an agentic version of Alexa with long-term memory and integration with over 1,000 apps, enhancing its capabilities as a personal AI assistant.

0 favorites 0 likes
#agentic

@_jasonwei: In addition to agents and coding, Muse Spark 1.1 is also really strong at answering health questions, a steadily growin…

X AI KOLs Timeline ↗ · 2026-07-09 Cached

Muse Spark 1.1, a new agentic and coding model from Meta, achieves +5% improvement on HealthBench-Pro, outperforming all competitors except Fable and Mythos.

0 favorites 0 likes
#agentic

Muse Spark 1.1 by Meta AI

Product Hunt ↗ · 2026-07-09

Meta AI released Muse Spark 1.1, a multimodal reasoning model designed for agentic tasks.

0 favorites 0 likes
#agentic

Tess-4-27B by Migel Tissera

Reddit r/LocalLLaMA ↗ · 2026-07-08 Cached

Tess-4-27B is a 27B reasoning model built on Qwen3.6-27B, post-trained on 64K-token long-context agentic traces with weight-scaled reasoning. It is designed for efficient, honest, and agentic task execution, available in open-source formats.

0 favorites 0 likes
#agentic

@LiorOnAI: Muse Image isn't just another image generator. I think it's Meta's first real attempt at making image generation agenti…

X AI KOLs Following ↗ · 2026-07-07 Cached

Meta released Muse Image, an agentic image generation model that plans, searches the web, writes code, and edits before rendering.

0 favorites 0 likes
#agentic

@VraserX: Meta just introduced Muse Image and previewed Muse Video. The interesting part is not just “better images.” It’s image …

X AI KOLs Following ↗ · 2026-07-07 Cached

Meta introduced Muse Image and previewed Muse Video, an agentic image and video generation system that enables precise edits, multiple references, and integration with Instagram context, turning media generation into a full creative operating system.

0 favorites 0 likes
#agentic

@ShunyuYao12: Hy2 -> Hy3 preview -> Hy3 Another massive leap forward, under half a year. Not just a leap of reasoning or agentic capa…

X AI KOLs Following ↗ · 2026-07-06 Cached

Tencent releases Hy3, a 295B MoE AI model that rivals trillion-scale flagships, open-sourced under Apache 2.0 with a free API trial.

0 favorites 0 likes
#agentic

Ran a classic(medival europe) fantasy RP/agentic benchmark across 8 local models Qwen3.6-27B held up better than its size suggests

Reddit r/LocalLLaMA ↗ · 2026-07-04

A benchmark comparing 8 local models on a classic medieval European fantasy role-playing and agentic task found that Qwen3.6-27B performed better than its size would suggest.

0 favorites 0 likes
#agentic

Anthropic launches Claude Sonnet 5, most agentic Sonnet yet, priced at $2/$10 per million tokens through August

Reddit r/artificial ↗ · 2026-07-01

Anthropic released Claude Sonnet 5, its most capable Sonnet model yet, featuring improved reasoning, coding, and tool use, with reduced hallucination. Priced at $2/$10 per million tokens through August.

0 favorites 0 likes
#agentic

Sonnet 5 - its updated tokenizer maps the same text to more tokens (roughly 1.0–1.35× depending on content), so cost per task can be higher.

Reddit r/ArtificialInteligence ↗ · 2026-06-30

Anthropic released Claude Sonnet 5 with improved reasoning, tool use, and coding, but its updated tokenizer maps text to more tokens (up to 1.35×), increasing effective cost per task despite the same listed price; introductory pricing applies until August 31, 2026.

0 favorites 0 likes
#agentic

@dair_ai: Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as…

X AI KOLs Following ↗ · 2026-06-30 Cached

NVIDIA proposes HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution, achieving 100% benchmark completion across several hardware design suites.

0 favorites 0 likes
#agentic

Quoting Jon Udell

Simon Willison's Blog ↗ · 2026-06-28 Cached

Jon Udell argues for reframing 'human in the loop' as 'human agent in the loop,' where humans invite AI agents into collaborative processes rather than being subordinated to machine-driven loops.

0 favorites 0 likes
#agentic

@mylifcc: Major Experiment: Using LLM as an 'Optimization Agent' for Automatic Loop Scheduling! Just read this paper accepted at PACT 2025: 'Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Opt…'

X AI KOLs Timeline ↗ · 2026-06-28 Cached

Introduces a paper accepted at PACT 2025, proposing the ComPilot framework, which uses off-the-shelf LLMs as optimization agents to automatically optimize complex loop nests without fine-tuning, achieving a geometric mean speedup of 3.54x, surpassing the SOTA Pluto.

0 favorites 0 likes
#agentic

@HuggingModels: Meet Gemma 4 12B Agentic Fable5: a locally run GGUF model that thinks, reasons, and uses tools like a pro. It's built f…

X AI KOLs Timeline ↗ · 2026-06-28 Cached

Meet Gemma 4 12B Agentic Fable5, a locally-run GGUF model designed for coding, terminal tasks, and agentic workflows, with 206k downloads.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback