Tag
The article discusses the importance of integrating QA into RL task generation through an iterative process to improve AI pipelines and model evaluation, emphasizing the value of intuition in eval design.
The article discusses methods and challenges for evaluating AI agents in production environments, focusing on quality assurance for real-world conversations beyond pre-defined evaluation sets.
The article demonstrates how a NousResearch Hermes local AI agent on NVIDIA RTX Spark autonomously identifies and fixes website issues and runs QA testing on device.
The article discusses the challenge of AI-generated code 'slop' due to human review bottlenecks and argues that software engineering must evolve to focus on system design rather than code readability.
The paper presents a locally deployed multi-agent AI system for structuring radiology reports and performing quality assurance, with radiologist evaluation showing favorable performance.
The article promotes using Replay QA for automated testing of web applications, especially for AI-generated code, highlighting its ease of use and features like continuous QA and root-cause analysis.
Xiaomi highlights over 1,000 days and 4.28 million kilometers of real-world testing for its vehicles, emphasizing meticulous quality and thorough road validation.
DoorDash shared how they automated evaluation of their Ask DoorDash AI assistant, enabling 2,000 daily graded sessions and cutting test time from six hours to 20 minutes while halving error rates.
A discussion on the next bottleneck for AI coding agents: verifying that AI-generated applications are actually correct, and who should be responsible for checking the output.
A blog post reflecting on the nature of software quality, arguing that quality is about gracefully performing development and leaving the codebase better than found, and exploring how to foster or destroy quality in software products.
Introduces the RE-AD framework that uses LLMs to provide real-time validation of data labeling quality, achieving 82% error acceptance and fix rate in production.
The Sovereign Tech Fund investment funds KDE's work on strengthening its PIM infrastructure, including Akonadi, with improvements in quality, protocol support, and ease of use for enterprise adoption.
This article discusses the insufficient testing of AI coding agents, highlighting a critical gap in ensuring their reliability and safety in software development.
This paper introduces an agentic workflow that uses LLMs and VLMs to iteratively generate and improve high-quality mathematical diagrams for K-12 education, addressing the reliability gap in AI-generated visual aids.
A critique of AI coding agents that claim tasks are complete without verifying functionality in a real browser environment.
Blop is a tool that tests your app and automatically repairs broken tests.
Momentic announces a major platform update with an AI-powered knowledge base and autonomous testing agents to address the growing gap between code velocity and software quality.
The article discusses using LLMs as automated QA engineers to perform manual testing tasks, such as integration and regression testing, potentially raising software quality bar.
Nolan Lawson argues that AI coding assistants can be used to write high-quality code slowly by employing multiple models for thorough code review and bug detection, improving codebase health rather than maximizing output speed.
A practitioner at a company handling ~40k conversations/month describes the bottleneck of manual prompt QA and asks how teams are using automated systems to detect regressions and user frustration in production.