quality-assurance

Tag

Cards List
#quality-assurance

@Vtrivedy10: most of the value in understanding models + doing better RL Task generation comes in the QA step the entire process to …

X AI KOLs Timeline · 3d ago Cached

The article discusses the importance of integrating QA into RL task generation through an iterative process to improve AI pipelines and model evaluation, emphasizing the value of intuition in eval design.

0 favorites 0 likes
#quality-assurance

How are people evaluating AI agents after they go into production?

Reddit r/AI_Agents · 2026-08-25

The article discusses methods and challenges for evaluating AI agents in production environments, focusing on quality assurance for real-world conversations beyond pre-defined evaluation sets.

0 favorites 0 likes
#quality-assurance

@NVIDIARTXSpark: Monday morning: website is broken, Slack is blowing up, and no coffee yet. See how a @NousResearch Hermes local AI agen…

X AI KOLs Timeline · 2026-08-20 Cached

The article demonstrates how a NousResearch Hermes local AI agent on NVIDIA RTX Spark autonomously identifies and fixes website issues and runs QA testing on device.

0 favorites 0 likes
#quality-assurance

@dzhng: https://x.com/dzhng/status/2090252351533973768

X AI KOLs Timeline · 2026-08-20 Cached

The article discusses the challenge of AI-generated code 'slop' due to human review bottlenecks and argues that software engineering must evolve to focus on system design rather than code readability.

0 favorites 0 likes
#quality-assurance

Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation

arXiv cs.CL · 2026-08-19 Cached

The paper presents a locally deployed multi-agent AI system for structuring radiology reports and performing quality assurance, with radiologist evaluation showing favorable performance.

0 favorites 0 likes
#quality-assurance

@svpino: Instead of reading your AI-generated code by hand, spend time on a plan to verify that your system is working as intend…

X AI KOLs Timeline · 2026-08-18 Cached

The article promotes using Replay QA for automated testing of web applications, especially for AI-generated code, highlighting its ease of use and features like continuous QA and root-cause analysis.

0 favorites 0 likes
#quality-assurance

@Xiaomi: Over 1,000 days and nights. 4.28 million kilometers of real-world testing. Some things simply can't be rushed. Some roa…

X AI KOLs Timeline · 2026-08-05 Cached

Xiaomi highlights over 1,000 days and 4.28 million kilometers of real-world testing for its vehicles, emphasizing meticulous quality and thorough road validation.

0 favorites 0 likes
#quality-assurance

@undefinedKi: DoorDash just published the full structure behind Ask DoorDash, their new AI assistant. And it's the clearest picture o…

X AI KOLs Following · 2026-08-04 Cached

DoorDash shared how they automated evaluation of their Ask DoorDash AI assistant, enabling 2,000 daily graded sessions and cutting test time from six hours to 20 minutes while halving error rates.

0 favorites 0 likes
#quality-assurance

AI can build apps now, but who checks if it built the right thing?

Reddit r/AI_Agents · 2026-08-02

A discussion on the next bottleneck for AI coding agents: verifying that AI-generated applications are actually correct, and who should be responsible for checking the output.

0 favorites 0 likes
#quality-assurance

How to not die by a thousand cuts or how to think about software quality (2023)

Hacker News Top · 2026-07-29 Cached

A blog post reflecting on the nature of software quality, arguing that quality is about gracefully performing development and leaving the codebase better than found, and exploring how to foster or destroy quality in software products.

0 favorites 0 likes
#quality-assurance

RE-AD: Real-Time Requirement Adherence for Data Labeling

arXiv cs.CL · 2026-07-24 Cached

Introduces the RE-AD framework that uses LLMs to provide real-time validation of data labeling quality, achieving 82% error acceptance and fix rate in production.

0 favorites 0 likes
#quality-assurance

KDE for Enterprise Needs a Strong PIM Infrastructure

Lobsters Hottest · 2026-07-21 Cached

The Sovereign Tech Fund investment funds KDE's work on strengthening its PIM infrastructure, including Akonadi, with improvements in quality, protocol support, and ease of use for enterprise adoption.

0 favorites 0 likes
#quality-assurance

Nobody's Testing AI Coding Agents Enough

Reddit r/AI_Agents · 2026-07-14

This article discusses the insufficient testing of AI coding agents, highlighting a critical gap in ensuring their reliability and safety in software development.

0 favorites 0 likes
#quality-assurance

Exploring Agentic Workflows for Generating High Quality Math Visual Aids

arXiv cs.AI · 2026-07-14 Cached

This paper introduces an agentic workflow that uses LLMs and VLMs to iteratively generate and improve high-quality mathematical diagrams for K-12 education, addressing the reliability gap in AI-generated visual aids.

0 favorites 0 likes
#quality-assurance

Your coding agent says "done." It never actually checked if the thing works in a browser.

Reddit r/AI_Agents · 2026-06-30

A critique of AI coding agents that claim tasks are complete without verifying functionality in a real browser environment.

0 favorites 0 likes
#quality-assurance

Blop

Product Hunt · 2026-06-24

Blop is a tool that tests your app and automatically repairs broken tests.

0 favorites 0 likes
#quality-assurance

A New Era of Software Quality Starts Today (5 minute read)

TLDR AI · 2026-06-24 Cached

Momentic announces a major platform update with an AI-powered knowledge base and autonomous testing agents to address the growing gap between code velocity and software quality.

0 favorites 0 likes
#quality-assurance

A new era for software testing

Hacker News Top · 2026-06-07 Cached

The article discusses using LLMs as automated QA engineers to perform manual testing tasks, such as integration and regression testing, potentially raising software quality bar.

0 favorites 0 likes
#quality-assurance

Using AI to write better code more slowly

Lobsters Hottest · 2026-05-25 Cached

Nolan Lawson argues that AI coding assistants can be used to write high-quality code slowly by employing multiple models for thorough code review and bug detection, improving codebase health rather than maximizing output speed.

0 favorites 0 likes
#quality-assurance

How are teams handling prompt QA at scale?

Reddit r/AI_Agents · 2026-05-20

A practitioner at a company handling ~40k conversations/month describes the bottleneck of manual prompt QA and asks how teams are using automated systems to detect regressions and user frustration in production.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback