Tag
This position paper argues against the claim that natural language can fully replace formal languages such as programming languages, proposing an information-theoretic specificity framework and proving a crossover theorem showing formal languages are better for high-specificity tasks.
ExecuGraph is a multi-agent framework for backend code synthesis that leverages execution-based validation and six specialized agents to improve reliability, showing gains particularly with more capable models like DeepSeek-Coder-V2-Lite.
KForge is a cross-platform framework that uses two collaborating LLM-based agents to automatically generate and optimize high-performance compute kernels for diverse AI accelerators, achieving significant speedups on NVIDIA B200 and Intel Arc B580 hardware.
This paper proposes VFEAgent, a multi-agent system that automates finite element analysis by integrating vision-language models with a verification-first code synthesis framework, enabling end-to-end simulation from images and problem descriptions.
Introduces Atomic Decomposition and Recombination (ADR), a framework that generates novel and challenging verifiable code tasks by decomposing and recombining atomic elements, enabling scalable reinforcement learning with verifiable rewards for large language models.
AutoRPA is a framework that automatically distills the decision logic of ReAct-style LLM agents into robust, token-efficient RPA functions for repetitive GUI tasks, reducing token usage by 82-96%.
OpenAI presents a hazard analysis framework for evaluating safety risks associated with code synthesis LLMs like Codex, examining technical, social, political, and economic impacts through a novel evaluation methodology for code generation capabilities.