agentic-workflows

Tag

Cards List
#agentic-workflows

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

arXiv cs.AI · 2026-07-10 Cached

Introduces CausalDS, a benchmark for evaluating causal reasoning in LLM-based data science agents, using synthetic structural causal models and natural language stories to test associational, interventional, and counterfactual reasoning along with tool use and abstention.

0 favorites 0 likes
#agentic-workflows

PatchOptic for Shared-State LLM Workflows with Projected Views and Verified Structured Updates

arXiv cs.LG · 2026-07-08 Cached

Introduces PatchOptic, an interface for shared-state LLM workflows that uses projected reads and verified structured patches to ensure valid updates. Evaluated with PatchBench across 46 cases, showing reduced token cost and leakage while maintaining quality.

0 favorites 0 likes
#agentic-workflows

@PrajwalTomar_: You don't understand how BIG this is. Anthropic just published the 4 ways to make Claude Code work without you. Everyon…

X AI KOLs Timeline · 2026-07-07 Cached

Anthropic published four types of loops for Claude Code to operate autonomously: turn-based, goal-based, time-based, and proactive, allowing different levels of task handoff.

0 favorites 0 likes
#agentic-workflows

Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming

arXiv cs.AI · 2026-07-07 Cached

Forethought is a neurosymbolic reasoning system that treats reasoning as an explicit, verifiable program composed from symbolic and neural primitives. It improves base-model accuracy by about 30% relative and enables small models to match frontier models while being model-agnostic and auditable.

0 favorites 0 likes
#agentic-workflows

@ivanfioravanti: DGX Spark Context Benchmark on Qwen3.6-35B-A3B-UD-Q8_K_XL llamacpp script released by Mia. It's fast! Time to test qual…

X AI KOLs Timeline · 2026-07-06 Cached

Benchmark results for Qwen3.6-35B-A3B-UD-Q8_K_XL on DGX Spark using llama.cpp script by Mia, showing fast token generation times across various context lengths.

0 favorites 0 likes
#agentic-workflows

How do you Mapout AI workflows when one suddenly costs 2× more than usual?

Reddit r/AI_Agents · 2026-07-06

The article discusses common causes of cost spikes in AI workflows, such as retries, repeated tool calls, long-running workflows, and growing context, and asks how teams investigate such issues.

0 favorites 0 likes
#agentic-workflows

@omarsar0: $200/week is not bad. It would cover (5-fold) all my engineering & research work. And I do a ton. It's a good direction…

X AI KOLs Following · 2026-07-03 Cached

The tweet agrees that $200/week is sufficient for engineering and research work, criticizing wasteful spending on expensive models and bloated agentic workflows.

0 favorites 0 likes
#agentic-workflows

@FinanceYF5: Four open-weight models have entered a stage where they can support real agent workflows. OpenRouter published a new article on the Insights blog discussing why the company chose these models in June:

X AI KOLs Following · 2026-06-29 Cached

OpenRouter posted on the Insights blog, pointing out that four open-weight models have reached a stage capable of supporting real agent workflows, and explained why the company chose these models in June.

0 favorites 0 likes
#agentic-workflows

@ehsanik: We're hiring on the computer use team at @AnthropicAI Building inside Anthropic has been a crazy amazing intense but be…

X AI KOLs Following · 2026-06-26 Cached

Anthropic's computer use team is hiring product engineers and researchers, seeking candidates who are passionate about agentic workflows and comfortable with ambiguity.

0 favorites 0 likes
#agentic-workflows

[R] Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost

Reddit r/MachineLearning · 2026-06-25 Cached

This paper demonstrates that compiling agentic workflow procedures into the weights of a small fine-tuned model achieves near-frontier quality at 128–462× cost reduction compared to in-context baselines, addressing perceived barriers of quality, cost, and flexibility.

0 favorites 0 likes
#agentic-workflows

@johnlindquist: Posted my talk from yesterday (with references). "Agentic Power User's Playbook" Watch it here (and grab a ticket to my…

X AI KOLs Following · 2026-06-24 Cached

John Lindquist shares his talk 'Agentic Power User's Playbook,' covering practical workflows, shortcuts, and habits for efficiently managing AI agents.

0 favorites 0 likes
#agentic-workflows

What's the worst thing your AI agent did in production without asking first?

Reddit r/AI_Agents · 2026-06-24

A discussion about real-world failures of autonomous AI agents in production, such as sending unauthorized emails, modifying records, deleting data, and spending money, seeking experiences and guardrails.

0 favorites 0 likes
#agentic-workflows

@posthog: https://x.com/posthog/status/2069472232712389112

X AI KOLs Timeline · 2026-06-23 Cached

PostHog explains why 'loops'—self-prompting agent workflows—are gaining traction, driven by improved model capabilities and real-world results from companies like Stripe and Lovable. The thread details what's needed to engineer a loop and showcases examples like PR babysitting and bug fixing.

0 favorites 0 likes
#agentic-workflows

@rohanpaul_ai: Agents can now have their own email! @atomic_mail just launched something to fix a missing piece in agentic workflows: …

X AI KOLs Following · 2026-06-18 Cached

Atomic Mail launches an API-first email service that gives AI agents their own inboxes, supporting integration with popular agents like Claude Desktop and Cursor via MCP or Agent Skill, and includes Proof-of-Work + reputation to combat spam.

0 favorites 0 likes
#agentic-workflows

Building independent LLM drift detection - sharing the methodology, looking for feedback on the approach

Reddit r/artificial · 2026-06-18

The author shares a methodology for building an external LLM drift detection system that continuously probes model behavior (schema adherence, instruction-following, refusal rates, etc.) to catch silent degradations in API performance, and invites feedback on the approach, pricing, and use cases.

0 favorites 0 likes
#agentic-workflows

@ConsciousRide: 90% of AI Engineering interviews in 2026 come down to these 7 points. 1. LLM Fundamentals: tokenization, transformers &…

X AI KOLs Timeline · 2026-06-17 Cached

A Twitter thread outlines the seven key areas that will dominate AI engineering interviews in 2026, including LLM fundamentals, RAG systems, agentic workflows, inference optimization, evaluation, MLOps, and production realities.

0 favorites 0 likes
#agentic-workflows

The founder's playbook: Building an AI-native startup

Hacker News Top · 2026-06-17 Cached

A practical playbook for building AI-native startups, covering stages from idea to scale with AI-powered exercises and frameworks using Claude.

1 favorites 1 likes
#agentic-workflows

@vboykis: new post: how I develop recently using local models. the tooling is now good enough to do agentic workflows and everyon…

X AI KOLs Following · 2026-06-15 Cached

Vicki Boykis shares her experience using local AI models for development, noting that recent releases like Gemma 4 have made agentic workflows feasible locally with about 75% accuracy of frontier models.

0 favorites 0 likes
#agentic-workflows

@ArizePhoenix: This week in Phoenix, a big one for agentic workflows: Slash commands & skills in PXI - Phoenix's built-in agent now ha…

X AI KOLs Following · 2026-06-12 Cached

Arize Phoenix's built-in agent PXI now supports slash commands and skills, allowing users to invoke custom workflows directly from chat.

0 favorites 0 likes
#agentic-workflows

FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse

arXiv cs.LG · 2026-06-11 Cached

FlowBank introduces a three-stage framework for optimizing agentic workflows in LLM multi-agent systems by precomputing a diverse set of reusable workflows and adaptively selecting the best one per query, achieving higher scores while maintaining cost competitiveness.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback