Tag
This paper introduces the deployment-fidelity gap in decomposed algorithm selection, demonstrating that partition-level evaluations can differ from end-to-end system performance, with implications for reporting and benchmarking.
The paper proposes TraceSQL, a lightweight and traceable verification model for text-to-SQL systems that uses explicit diagnostic features to estimate answerability without reference queries, achieving improved performance over existing baselines on the BIRD benchmark.
The paper proposes DoctorAgents, an agentic AI framework that uses specialized LLM agents to iteratively generate, validate, and refine end-to-end machine learning pipelines for small, heterogeneous clinical temporal datasets, outperforming established AutoML baselines.
Introduces AutoProteinEngine (AutoPE), an LLM-driven agent framework that enables biologists without deep learning expertise to perform multimodal AutoML for protein engineering via natural language, showing improvements over zero-shot and manual fine-tuning approaches.
This paper systematically evaluates five train-test splitting strategies for AutoML, showing that geometry-based methods are less effective than random/stratified splitting in preserving distributional similarity, and proposes an Optimised-Distribution method that achieves 89% similarity.
This paper presents an LLM-driven pipeline using GPT-5, GPT-4o, and Claude Sonnet 4 to automatically design neural network architectures for cross-lingual handwritten OCR, achieving over 93% accuracy across Arabic, English, and Persian scripts without human intervention.
Auto-FL-Research introduces a constrained coding-agent workflow for automatically searching and evaluating federated learning algorithmic recipes, showing performance gains on multiple healthcare and LEAF tasks while also exposing seed-sensitive and search-selected failure cases.
Introduces SAGE, the first end-to-end LLM-driven multi-agent framework for fraud detection, using a Data Diagnostic Tree and Markov decision process with natural-language gradients to optimize models under class imbalance. Experiments show significant F1 improvements over baselines across five datasets.
Introduces the autoresearch project, which breaks down the AI research process into a verifiable loop (fixed environment, single editable file, fixed metric, Git rollback), enabling AI agents to perform controllable and reproducible experiment iterations; also mentions the 12-factor-agents checklist.
This paper introduces yvsoucom-iterkit, a deterministic, log-driven AutoML framework for reproducible pipeline optimization in healthcare risk prediction, evaluated on diabetes and stroke datasets with over 18,000 pipeline configurations, achieving strong performance and revealing structured search spaces with component redundancy.
Researchers from Google and Meta propose AutoTTS, a framework using AI agents to automatically discover and refine test-time scaling strategies for LLMs without human intervention. The agent successfully identified complex, coordinated reasoning mechanisms that outperformed manual baselines at a low computational cost.