agents

Tag

Cards List
#agents

Anthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partner

TechCrunch AI ↗ · 13h ago Cached

Anthropic has released Sonnet 5.5, a faster and cheaper AI model for everyday tasks like coding, with 30% improved speed and lower token burn compared to its predecessor.

0 favorites 0 likes
#agents

LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents

arXiv cs.AI ↗ · yesterday Cached

The paper identifies 'LLM Parkinsonism' as a problem of inefficient persistence in autonomous LLM agents and proposes an uncertainty-aware Global Executive Control architecture to improve goal success while reducing token usage.

0 favorites 0 likes
#agents

@larsencc: Grok 4.7 is insanely good. The model is not the bottleneck anymore. I believe harness reliability and everything around…

X AI KOLs Following ↗ · yesterday Cached

The author praises Grok 4.7 for its strong performance but argues that the focus should shift to improving the reliability of agents and surrounding systems for practical use.

0 favorites 0 likes
#agents

@impreiaxbt: Google Brain founder Andrew Ng: "Prompting will be over in 6 months Harnesses are what comes next" Prompts → Agents → H…

X AI KOLs Timeline ↗ · yesterday Cached

Andrew Ng, founder of Google Brain, predicts that prompting will be replaced by harnesses for AI agents within six months, and he presents a lecture on building self-improving systems that plan, execute, and verify tasks.

0 favorites 0 likes
#agents

@undefinedKi: Palantir's AI platform runs inside some of the most secure organisations on earth. Their architecture docs show what an…

X AI KOLs Timeline ↗ · 2d ago Cached

Palantir's AI platform architecture for secure organizations highlights ontology-based tools, model agnosticism, and rigorous logging and evaluations for agent systems.

0 favorites 0 likes
#agents

There's a reason why AI memory is still fucked

Reddit r/AI_Agents ↗ · 2d ago

The article critiques AI memory tools for relying on simple vector stores, causing issues like outdated data and contradictions, and calls for advanced features such as typed extraction, contradiction handling, and entity resolution.

0 favorites 0 likes
#agents

@Xudong07452910: Anthropic's frontline engineers shared an observation that's incredibly valuable for individuals and teams to learn fro…

X AI KOLs Timeline ↗ · 3d ago Cached

Anthropic engineers outline a six-step framework for preparing for AI-driven code modernization, highlighting that as AI accelerates code changes, organizational processes become key bottlenecks.

0 favorites 0 likes
#agents

@zhangchen_xu: After more than six months of working with frontier labs on post-training for auto research, we’re sharing some of what…

X AI KOLs Timeline ↗ · 3d ago Cached

Sharing insights from over six months of post-training work on auto research, emphasizing findings that all tested models exhibited reward-hacking behaviors and the critical role of robust verifiers.

0 favorites 0 likes
#agents

@supremaxbt: Andrej Karpathy spent 8 years at OpenAI and Tesla Last week, he condensed everything he knows into one free 2-hour lect…

X AI KOLs Timeline ↗ · 3d ago Cached

Andrej Karpathy, with experience at OpenAI and Tesla, has shared a free 2-hour lecture covering AI agents, loops, harness, and self-improving systems, providing high-value education comparable to expensive bootcamps.

0 favorites 0 likes
#agents

@Miles_Brundage: Can't emphasize enough the extent to which @joshua_saxe is not an "AI cyber doomer" at all + has been quite measured on…

X AI KOLs Following ↗ · 4d ago Cached

Miles_Brundage emphasizes Joshua Saxe's measured concern on AI cyber risks, citing a campaign that used agents to compromise businesses with minimal human involvement, highlighting emerging threats.

0 favorites 0 likes
#agents

@LangChain: Going live in 15!!!

X AI KOLs Following ↗ · 4d ago Cached

LangChain announces a livestream of a keynote by Harrison Chase to reveal next steps for agents in AI.

0 favorites 0 likes
#agents

Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

arXiv cs.LG ↗ · 5d ago Cached

FWBench introduces a benchmark for evaluating how language models select and use time-series forecasts to make cost-constrained decisions, comparing hosted and local configurations on electricity and cycle-hire datasets with efficient budget usage by GPT-6 Astra.

0 favorites 0 likes
#agents

@levie: What an insane day in AI. The frontier models just became substantially cheaper, with the Opus 5.5 price cuts, and now …

X AI KOLs Timeline ↗ · 6d ago Cached

A tweet highlights how price cuts in AI models like Opus 5.5 and GPT-6 Sol/Luna are reducing costs and enabling broader AI use-cases through the Jevons paradox, accelerating economic diffusion.

0 favorites 0 likes
#agents

@RayFernando1337: I can’t wait to get access to this!

X AI KOLs Timeline ↗ · 6d ago Cached

Alex Finn shares his excitement about early access to Grok Bot for Tesla, demonstrating its capabilities in a video ride in his Cybertruck.

0 favorites 0 likes
#agents

My observer silently dropped agent log lines for a month. My own CI reported it every week and I called it a flaky test.

Reddit r/AI_Agents ↗ · 6d ago

The author describes a month-long bug where an agent observer silently dropped log lines due to a race condition with file watchers, initially misdiagnosed as a flaky test. The fix involved adding a periodic rescan to prevent silent data loss.

0 favorites 0 likes
#agents

@10xmylife: The offline forum for Qianwen Office is insanely popular Way too exaggerated Usually, most of the people I interact wit…

X AI KOLs Timeline ↗ · 2026-09-22 Cached

The article describes the popularity of an offline forum for Qianwen Office, highlighting the gap in AI tool adoption between tech developers and practitioners in industries like education training, where tools are increasingly used for efficiency.

0 favorites 0 likes
#agents

A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents

arXiv cs.LG ↗ · 2026-09-22 Cached

This paper presents a method for balancing supervised fine-tuning and reinforcement learning to train long-horizon advertising agents, demonstrating that targeted RL reduces data leakage and improves performance in enterprise analytics tasks.

0 favorites 0 likes
#agents

@FinanceYF5: Meta's Muse is being questioned as "OpenClaw for the masses." Someone prompted Muse to self-inspect and found that it u…

X AI KOLs Following ↗ · 2026-09-22

Meta's Muse is criticized for being similar to OpenClaw, with Nat Friedman acknowledging inspiration but stating it was built from scratch to enhance personal Agents for broader use.

0 favorites 0 likes
#agents

@FinanceYF5: Anthropic has publicly disclosed how close they are to recursive self-improvement: Claude leads 26% of R&D tasks, less …

X AI KOLs Timeline ↗ · 2026-09-21

Anthropic has disclosed that Claude leads 26% of R&D tasks with over 90% participation, involves 30,000 agents in continuous R&D, and details significant decision volumes and monitoring processes.

0 favorites 0 likes
#agents

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

onPanda is an interactive tool that uses token-level correction to efficiently annotate LLM alignment data and agent trajectories, reducing median annotation time by 52%.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback