data-quality

Tag

Cards List
#data-quality

@bojie_li: Just finished watching Teacher He Jiyan from Zhongguancun College lead 7 PhD students to train a 7B model from scratch …

X AI KOLs Timeline ↗ · 5d ago Cached

Teacher He Jiyan and 7 PhD students from Zhongguancun College trained a 7B model from scratch in 3 months, achieving state-of-the-art performance in the 7B model category, with insights on efficient training and data quality.

0 favorites 0 likes
#data-quality

The data layer becomes harder to ignore when an agent starts acting on it

Reddit r/AI_Agents ↗ · 2026-09-23

James Luan, CTO of Zilliz, argues that data infrastructure becomes more critical as AI agents act on enterprise data, emphasizing the need for data quality, freshness, and proper permissions.

0 favorites 0 likes
#data-quality

We tested the client's data before building the RAG system. 4 in 10 answers were right; here's what got it to 9.

Reddit r/AI_Agents ↗ · 2026-09-22

The article details a pilot project where a team tested client data for a RAG system, initially achieving only 40% accuracy due to issues like conflicting documents and poor retrieval. Through routing, data cleaning, and acronym mapping, accuracy improved to about 90%, highlighting the critical role of data readiness in successful AI transformation.

0 favorites 0 likes
#data-quality

No Easy Fix for Bogus Respondents in Online Opt-In Polls

Hacker News Top ↗ · 2026-09-22 Cached

A Pew Research Center study finds that methods to remove bogus respondents from online opt-in polls, such as trap questions and automated prescreening, can reduce error but do not provide a surefire solution, with voter file matching potentially increasing error.

0 favorites 0 likes
#data-quality

why most enterprise AI projects die before reaching production

Reddit r/artificial ↗ · 2026-09-20

The article discusses why enterprise AI projects often fail to reach production, emphasizing challenges with data quality, legacy systems, and security in deployment.

0 favorites 0 likes
#data-quality

What a life-expectancy chart taught me about reviewing AI-built data pipelines

Reddit r/ArtificialInteligence ↗ · 2026-09-19

The author shares a lesson learned from a life expectancy data project, highlighting the need to preserve metadata like footnotes in AI-built data pipelines to ensure data accuracy and comparability.

0 favorites 0 likes
#data-quality

Data Quality Rule Generation with LLMs

arXiv cs.LG ↗ · 2026-09-10 Cached

This paper introduces LeDQeR, a framework that uses large language models to automatically generate data quality rules for enterprise tools, improving rule coverage and reducing manual maintenance effort.

0 favorites 0 likes
#data-quality

Pretraining progress is mostly coming from data (17 minute read)

TLDR AI ↗ · 2026-09-09 Cached

The article investigates AI pretraining progress from 2019 to 2025, finding that data improvements contribute 3.24 times more to compute efficiency gains than model improvements at a 1e19 FLOPs budget.

0 favorites 0 likes
#data-quality

How Output Format Confounds Data Quality and Capability in Instruction Tuning

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper demonstrates that output formats confound data quality metrics and model capability assessments in instruction tuning, causing significant accuracy shifts and rendering current practices ineffective without interface-aware adjustments.

0 favorites 0 likes
#data-quality

What part of AI do you think we still fundamentally misunderstand?

Reddit r/artificial ↗ · 2026-09-02

The article discusses the gap between impressive AI demos and successful production systems, highlighting overlooked challenges like bad data and poor evaluation, and invites input on fundamental misunderstandings in AI.

0 favorites 0 likes
#data-quality

The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

arXiv cs.AI ↗ · 2026-09-01 Cached

This paper investigates how agentic AI architectures can complete online surveys and pass attention checks, analyzing vulnerabilities from attack and defense perspectives. It evaluates multiple open-source models and offers strategies for data quality control in the age of AI.

0 favorites 0 likes
#data-quality

@Miles_Brundage: Can’t speak to the details here but I will say that 1. one generally does not hear the best things about the data indus…

X AI KOLs Timeline ↗ · 2026-08-26 Cached

A discussion highlighting poor practices in the data industry, specifically regarding RLVR training data for AI models, with widespread sloppiness noted.

0 favorites 0 likes
#data-quality

Is (or will) AI learn backwards? (Since most of its training data is now AI-generated data)

Reddit r/ArtificialInteligence ↗ · 2026-08-21

The article raises concerns about AI models potentially training on data generated by AI itself, which could lead to issues like model collapse, citing examples such as Deezer removing millions of AI-generated songs and the high proportion of AI-created blog articles.

0 favorites 0 likes
#data-quality

Training AI on AI slop

Reddit r/ArtificialInteligence ↗ · 2026-08-21

The article discusses the problem of training AI models on low-quality data generated by other AI systems, known as 'slop,' and its implications for model reliability and performance.

0 favorites 0 likes
#data-quality

When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators

arXiv cs.CL ↗ · 2026-08-20 Cached

This study evaluates LLMs as data quality annotators on e-commerce tasks, finding they outperform baselines when background knowledge is required but offer limited advantages for tasks with strong lexical signals, while demonstrating high consistency across runs.

0 favorites 0 likes
#data-quality

Do enterprise AI projects actually fail because the AI isn't good enough?

Reddit r/artificial ↗ · 2026-08-18

The article questions if AI model quality is the primary reason for enterprise AI project failures, suggesting that data and context issues are often the real culprits, and fixing them can make AI implementation easier.

0 favorites 0 likes
#data-quality

Marketers are Addicted to Bad Data (2020)

Hacker News Top ↗ · 2026-08-13 Cached

The article criticizes marketers for relying on bad data, highlighting issues such as ad blockers, misleading metrics, and fraudulent audiences that lead to ineffective decision-making in marketing.

0 favorites 0 likes
#data-quality

Does pre-generative-AI data become more valuable as the internet fills with synthetic material?

Reddit r/artificial ↗ · 2026-08-12

A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.

0 favorites 0 likes
#data-quality

After a few months running an AI report generator for a client, the writing was never the hard part

Reddit r/AI_Agents ↗ · 2026-08-10

A developer shares lessons from running an AI report generator in production, arguing that data quality and validation matter far more than the model's writing ability, since fluent but incorrect reports are dangerous.

0 favorites 0 likes
#data-quality

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

arXiv cs.CL ↗ · 2026-08-05 Cached

Proposes a scalable subdocument deduplication framework for LLM pretraining that separates duplicate detection from copy retention, using frequency- and length-aware policies. Experiments on FineWeb-Edu and a code web corpus show improved model performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback