Tag
Presents Doc2DB-Bench, a benchmark for evaluating LLM-based extraction of relational databases from long documents, with 203 instances across 42 schemas and seven domains.
DataPrep-Bench is a unified benchmark evaluating LLMs' capabilities in training data construction and quality evaluation across six domains, including a skill-guided agent (Data-Construction-Skill) and a distribution-based evaluator (DAS) that achieves strong cross-model correlation.
DataEvolver is a self-evolving multi-agent framework that leverages feedback from rejected samples to iteratively enhance data quality for text-rich image generation, achieving 85.3% OCR-F1 improvement on TextScenesHQ and 35.3% on LongTextBench using PixArt-alpha.
This paper introduces CHILLGuard, a fine-grained Chinese LLM content safety guardrail built on a new 5-macro, 31-micro category risk taxonomy and a scalable multi-stage data construction pipeline. The model achieves state-of-the-art performance, improving F1 score by 15.92% over existing baselines.