Tag
Teacher He Jiyan and 7 PhD students from Zhongguancun College trained a 7B model from scratch in 3 months, achieving state-of-the-art performance in the 7B model category, with insights on efficient training and data quality.
James Luan, CTO of Zilliz, argues that data infrastructure becomes more critical as AI agents act on enterprise data, emphasizing the need for data quality, freshness, and proper permissions.
The article details a pilot project where a team tested client data for a RAG system, initially achieving only 40% accuracy due to issues like conflicting documents and poor retrieval. Through routing, data cleaning, and acronym mapping, accuracy improved to about 90%, highlighting the critical role of data readiness in successful AI transformation.
A Pew Research Center study finds that methods to remove bogus respondents from online opt-in polls, such as trap questions and automated prescreening, can reduce error but do not provide a surefire solution, with voter file matching potentially increasing error.
The article discusses why enterprise AI projects often fail to reach production, emphasizing challenges with data quality, legacy systems, and security in deployment.
The author shares a lesson learned from a life expectancy data project, highlighting the need to preserve metadata like footnotes in AI-built data pipelines to ensure data accuracy and comparability.
This paper introduces LeDQeR, a framework that uses large language models to automatically generate data quality rules for enterprise tools, improving rule coverage and reducing manual maintenance effort.
The article investigates AI pretraining progress from 2019 to 2025, finding that data improvements contribute 3.24 times more to compute efficiency gains than model improvements at a 1e19 FLOPs budget.
This paper demonstrates that output formats confound data quality metrics and model capability assessments in instruction tuning, causing significant accuracy shifts and rendering current practices ineffective without interface-aware adjustments.
The article discusses the gap between impressive AI demos and successful production systems, highlighting overlooked challenges like bad data and poor evaluation, and invites input on fundamental misunderstandings in AI.
This paper investigates how agentic AI architectures can complete online surveys and pass attention checks, analyzing vulnerabilities from attack and defense perspectives. It evaluates multiple open-source models and offers strategies for data quality control in the age of AI.
A discussion highlighting poor practices in the data industry, specifically regarding RLVR training data for AI models, with widespread sloppiness noted.
The article raises concerns about AI models potentially training on data generated by AI itself, which could lead to issues like model collapse, citing examples such as Deezer removing millions of AI-generated songs and the high proportion of AI-created blog articles.
The article discusses the problem of training AI models on low-quality data generated by other AI systems, known as 'slop,' and its implications for model reliability and performance.
This study evaluates LLMs as data quality annotators on e-commerce tasks, finding they outperform baselines when background knowledge is required but offer limited advantages for tasks with strong lexical signals, while demonstrating high consistency across runs.
The article questions if AI model quality is the primary reason for enterprise AI project failures, suggesting that data and context issues are often the real culprits, and fixing them can make AI implementation easier.
The article criticizes marketers for relying on bad data, highlighting issues such as ad blockers, misleading metrics, and fraudulent audiences that lead to ineffective decision-making in marketing.
A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.
A developer shares lessons from running an AI report generator in production, arguing that data quality and validation matter far more than the model's writing ability, since fluent but incorrect reports are dangerous.
Proposes a scalable subdocument deduplication framework for LLM pretraining that separates duplicate detection from copy retention, using frequency- and length-aware policies. Experiments on FineWeb-Edu and a code web corpus show improved model performance.