llm-agent

Tag

Cards List
#llm-agent

agentic-ger: terminology recovery in long-form speech using global context

arXiv cs.CL ↗ · yesterday Cached

Agentic-GER proposes an LLM-based agent for correcting domain-specific terminology in long-form speech transcripts using global context and selective re-transcription, achieving significant improvements in ASR accuracy for Chinese and English.

0 favorites 0 likes
#llm-agent

@0xCodila: Send this Jev prompt to any LLM or AI agent It sets up Jev → analyzes you → finds where you waste time and money → upgr…

X AI KOLs Timeline ↗ · 3d ago Cached

A tweet promotes a 'Jev' prompt for LLMs and AI agents that analyzes users to optimize their AI setup and save time and money.

0 favorites 0 likes
#llm-agent

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

arXiv cs.AI ↗ · 2026-09-15 Cached

ClinAgent is a ReAct-based agent system using agentic RAG to enable natural language querying of clinical trial information, with evaluation across multiple LLM backends showing complementary strengths.

0 favorites 0 likes
#llm-agent

CityPlanner: A Sandbox Agent for Executable Urban Planning

arXiv cs.AI ↗ · 2026-09-11 Cached

CityPlanner introduces a unified sandbox environment and atomic-task reinforcement learning for executable urban planning, outperforming heuristic, task-specific RL, and general LLM-agent baselines on a real-world benchmark.

0 favorites 0 likes
#llm-agent

Built a read-only analytics agent (route → fetch → narrate → ground). Before I let it write anything, what am I missing?

Reddit r/AI_Agents ↗ · 2026-09-09

A developer describes building a read-only analytics agent for restaurant POS systems using Node.js, TypeScript, and Gemini Flash, seeking advice on tool selection, grounding techniques, and patterns for implementing write actions.

0 favorites 0 likes
#llm-agent

Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops

arXiv cs.CL ↗ · 2026-09-02 Cached

This paper diagnoses algorithmic mode collapse in code-level autonomous research loops and proposes Diversity-Aware Proposal Sampling (DAPS) as a lightweight mitigation to preserve semantic diversity and improve generalization.

0 favorites 0 likes
#llm-agent

MacroAgent: Regularity-Aware Macro Legalization with LLM-Agent-Designed Contour Algorithms

arXiv cs.LG ↗ · 2026-08-27 Cached

MacroAgent introduces a novel framework using LLMs to design contour algorithms for macro legalization in VLSI circuits, achieving significant improvements in layout regularity and performance.

0 favorites 0 likes
#llm-agent

VortexChat: An agentic framework for autonomous multi-objective integrated photonic design

arXiv cs.AI ↗ · 2026-08-24 Cached

VortexChat is an LLM-based agentic framework that automates the inverse design of integrated photonic devices from natural language specifications, demonstrated by autonomously fabricating a terahertz multiplexer with high performance.

0 favorites 0 likes
#llm-agent

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

arXiv cs.CL ↗ · 2026-08-21 Cached

Task-CoEvolve is a novel approach for efficient LLM agent harness optimization that adaptively selects validation tasks to reduce evaluation costs while maintaining performance, achieving an 80% reduction in evaluations on benchmarks.

0 favorites 0 likes
#llm-agent

Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection

arXiv cs.AI ↗ · 2026-08-19 Cached

This paper presents an automated agentic approach using Large Language Models to synthesize interpretable Python feature extractors for algorithm selection in constraint satisfaction problems, outperforming expert-curated methods.

0 favorites 0 likes
#llm-agent

Reflexões do meu Agente - Parte 2

Reddit r/AI_Agents ↗ · 2026-08-18

O artigo descreve as reflexões metacognitivas de um agente de codificação LLM no Devin/Cascade, expondo seu raciocínio e pontuações de confiança durante tarefas de análise de código.

0 favorites 0 likes
#llm-agent

Wiring an ai content generator into an agent loop taught me volume was never the constraint

Reddit r/AI_Agents ↗ · 2026-08-12

A developer reflects on building an AI content generation agent loop, discovering that throughput isn't the real constraint—quality control and relevance are, leading to a human-in-the-loop approach that produces fewer but better pieces.

0 favorites 0 likes
#llm-agent

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

Hugging Face Daily Papers ↗ · 2026-08-12 Cached

This paper studies GPU control gates for LLM-agent services, analyzing concurrent cohort scheduling and on-device routing versus host redispatch to reduce host round trips and improve GPU utilization.

0 favorites 0 likes
#llm-agent

ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration

arXiv cs.AI ↗ · 2026-08-11 Cached

This paper presents ZhuLong, an execution-grounded LLM coding agent for EDA scripting that uses API retrieval, documentation inspection, and sandbox execution via MCP tools, augmented by an offline API self-exploration mechanism to infer undocumented API behaviors. It achieves 78.5% Pass@1 on a benchmark of 158 real-world EDA tasks, significantly outperforming a pure LLM baseline.

0 favorites 0 likes
#llm-agent

Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns

arXiv cs.AI ↗ · 2026-08-11 Cached

Agent-MD is a framework that selectively applies LLM reasoning to long-running molecular simulation campaigns, using a deterministic rule-based agent for routine tasks and event-triggered LLM review for exceptional conditions. Demonstrated in GCMC–MD water-vapor desorption simulations, it shows that auditable, reproducible scientific workflows can avoid placing every operation inside an LLM reasoning loop.

0 favorites 0 likes
#llm-agent

Looking for extreme / impossible tasks to properly stress-test my agent.I can’t trust my own judgment anymore

Reddit r/AI_Agents ↗ · 2026-08-10

A developer who built a fully autonomous custom agent architecture that can run for weeks without intervention is asking the community for extreme, adversarial tasks to properly stress-test it, because they can no longer trust their own judgment.

0 favorites 0 likes
#llm-agent

Letting an agent loose on a real iPhone taught me to build the kill switch first

Reddit r/AI_Agents ↗ · 2026-08-09

The author built sidetap, a Python harness that lets an LLM agent control a real iPhone from Windows over USB, with safety features like a kill switch and guardrails. The post details the technical challenge of sideloading WebDriverAgent with a free Apple ID and introduces the open-source tool.

0 favorites 0 likes
#llm-agent

@jinglian: Whoa, Codex can use your iPhone! Introducing: phone-harness > Automate any iOS app like a human > Native iPhone control → no API, no jailbreak > Connect once, control anytime. Set up with just one prompt. Get started right…

X AI KOLs Timeline ↗ · 2026-08-09 Cached

Introducing phone-harness, an open-source tool that uses macOS's iPhone Mirroring to let LLMs (such as Codex) automate any iOS app like a human, without needing an API or jailbreak. It uses screenshots, Vision OCR, and HID events for vision and control.

0 favorites 0 likes
#llm-agent

SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse

arXiv cs.AI ↗ · 2026-08-07 Cached

Presents SkillTrace, a multi-trace provenance auditing framework for LLM-agent skill reuse that extracts expression, implementation, and operational traces, achieving strong accuracy on a benchmark and enabling large-scale wild audits.

0 favorites 0 likes
#llm-agent

An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil Wells

arXiv cs.LG ↗ · 2026-08-06 Cached

This paper proposes an explainable LLM agent layer placed downstream of an open-world learning pipeline for oil well anomaly detection, using the Qwen3.5-397B-A17B model to provide natural-language justifications and novelty naming on the 3W dataset.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback