data-quality

Tag

Cards List
#data-quality

Does pre-generative-AI data become more valuable as the internet fills with synthetic material?

Reddit r/artificial · 3d ago

A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.

0 favorites 0 likes
#data-quality

After a few months running an AI report generator for a client, the writing was never the hard part

Reddit r/AI_Agents · 5d ago

A developer shares lessons from running an AI report generator in production, arguing that data quality and validation matter far more than the model's writing ability, since fluent but incorrect reports are dangerous.

0 favorites 0 likes
#data-quality

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

arXiv cs.CL · 2026-08-05 Cached

Proposes a scalable subdocument deduplication framework for LLM pretraining that separates duplicate detection from copy retention, using frequency- and length-aware policies. Experiments on FineWeb-Edu and a code web corpus show improved model performance.

0 favorites 0 likes
#data-quality

Wrappers and snakeoil ideas seem to be winning more deals

Reddit r/ArtificialInteligence · 2026-07-29

The article argues that companies making bold AI promises and selling 'wrappers' are winning more deals than those focusing on hard engineering problems like data quality, governance, and enterprise integrations, reflecting a market that rewards hype over substance.

0 favorites 0 likes
#data-quality

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

arXiv cs.LG · 2026-07-28 Cached

Introduces a scalable, inference-only data valuation pipeline that approximates Shapley values to audit LLM alignment datasets, reducing manual audit search space by 99.1% and uncovering hidden label failures in HelpSteer2 and HH-RLHF.

0 favorites 0 likes
#data-quality

Training data needs a real go/no-go gate before training [D]

Reddit r/MachineLearning · 2026-07-27

The author proposes a formal pre-training control layer that audits training data artifacts and provides a verdict (PASS/FAIL) based on explicit criteria, as a missing gate between data preparation and training, and invites discussion on its practicality.

0 favorites 0 likes
#data-quality

Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA

arXiv cs.CL · 2026-07-27 Cached

This paper studies internalizing documents into LoRA adapters for closed-book QA and finds that once adapter capacity is sufficient, training data quality (especially answer conciseness) is the dominant factor, outperforming architectural changes and achieving higher accuracy than BM25-RAG.

0 favorites 0 likes
#data-quality

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

arXiv cs.LG · 2026-07-24 Cached

DataPrep-Bench is a unified benchmark evaluating LLMs' capabilities in training data construction and quality evaluation across six domains, including a skill-guided agent (Data-Construction-Skill) and a distribution-based evaluator (DAS) that achieves strong cross-model correlation.

0 favorites 0 likes
#data-quality

agent memory is less useful if it cannot forget bad examples

Reddit r/AI_Agents · 2026-07-20

The article argues that for effective agent memory, it is crucial to forget bad examples, as retaining them degrades performance.

0 favorites 0 likes
#data-quality

AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets

arXiv cs.AI · 2026-07-20 Cached

AgentFAIR is a multi-agent framework that uses LLM evaluators and a critic to assess FAIR compliance of geospatial datasets, achieving sub-principle agreement of 89% and Fleiss' κ=0.71 in expert studies, at a cost of $0.054 per dataset.

0 favorites 0 likes
#data-quality

Gartner predicts 60% of AI projects will be abandoned through 2026. From what I see doing ECM/content work, the reason isn't the model.

Reddit r/ArtificialInteligence · 2026-07-10

Gartner predicts 60% of AI projects will be abandoned through 2026. The author, working in enterprise content management, argues the real blocker is messy unstructured legacy content and inconsistent metadata, not the AI models themselves.

0 favorites 0 likes
#data-quality

UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing

arXiv cs.CL · 2026-07-10 Cached

UltraX proposes a function-calling refinement framework for large-scale pre-training data that introduces insertion alongside deletion and modification, enabling fine-grained instance-level editing. It builds a reliable program-supervision generation pipeline and demonstrates improved data efficiency and model performance when pretraining 1B models from scratch.

0 favorites 0 likes
#data-quality

The foundational elements of AI architecture that IT leaders need to scale

MIT Technology Review · 2026-07-07 Cached

The article outlines four foundational elements of AI architecture—data quality, context engineering, governance, and human expertise—that IT leaders should prioritize to scale AI systems reliably as models evolve.

0 favorites 0 likes
#data-quality

Show HN: Pulpie – Models for Cleaning the Web

Hacker News Top · 2026-07-06 Cached

Feyn introduces Pulpie, a family of Pareto-optimal models for extracting main content from HTML pages, achieving near state-of-the-art quality at one twentieth the cost. The smallest model, pulpie-orange-small, matches leading extractors in ROUGE-5 F1 while being smaller and much faster, enabling scalable web cleaning for pre-training and inference.

0 favorites 0 likes
#data-quality

Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap

arXiv cs.CL · 2026-07-01 Cached

This paper identifies and corrects label errors and test-train overlap in the RVL-CDIP document classification dataset, finding 12% label errors and 35% duplication. Correction improves classification accuracy and out-of-distribution generalization.

0 favorites 0 likes
#data-quality

Agriculture is ready for AI, but its data isn’t

MIT Technology Review · 2026-06-30 Cached

AI has great potential in agriculture, but its effectiveness depends on clean and complete data foundations; the industry faces unique data challenges from IoT devices, weather feeds, and land-specific variables.

0 favorites 0 likes
#data-quality

@Phoenixyin13: This latest blockbuster paper from Meta FAIR aims to tell the AI industry an important bellwether: "Large model data is ushering in the era of intelligent scientists." In this paper, a 4B small model precisely refined by Autodata not only crushes the same-scale models trained with traditional synthetic data on legal reasoning tasks, but also...

X AI KOLs Timeline · 2026-06-27 Cached

Meta FAIR's latest paper proposes the Autodata method, which uses an intelligent data scientist Agent to autonomously generate and optimize high-quality data, enabling a 4B small model to defeat a 397B large model on legal reasoning tasks. This indicates that data quality can bridge the gap in parameter count, providing new insights for data pipelines and scaling.

0 favorites 0 likes
#data-quality

Training Dynamics of Neural Software Defect Predictors under Coupled Data-Quality Issues

arXiv cs.LG · 2026-06-25 Cached

This paper investigates how training dynamics of neural networks for software defect prediction are affected by coupled data-quality issues such as class imbalance and overlap, proposing an interaction-aware empirical protocol.

0 favorites 0 likes
#data-quality

AI is getting better at analysis. The problem is still the data.

Reddit r/ArtificialInteligence · 2026-06-24

The author argues that AI analysis quality is limited more by data access and reliability than by reasoning, and that structured datasets dramatically improve outputs.

0 favorites 0 likes
#data-quality

Google display wrong flags for world cup 2026

Hacker News Top · 2026-06-20 Cached

Google's World Cup 2026 match schedule widget displays incorrect flags for countries like Norway and England due to likely data mapping or asset mismanagement, highlighting gaps in automated data quality checks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback