Tag
The Era by Eon Benchmark generates a complete fictional enterprise with simulators and databases to evaluate LLM agents on enterprise tools, providing exact ground truth for accurate grading.
A developer shares the challenge of creating a gold standard evaluation dataset for an AI product with no users, considering synthetic data generation and adversarial testing to avoid post-launch restructuring.
This position paper argues that ground truth datasets in machine learning are not objective truths but human constructions shaped by choices, and advocates for articulating these choices to improve reliability, transparency, and accountability.
Introduces a generative framework that uses LLM agents to inject behavioral anomalies into simulated trajectories and applies kinematic and map constraints to produce realistic anomalous mobility data with ground truth.