Tag
This paper introduces the COPES dataset and a three-axis evaluation framework to assess LLM alignment with community perspectives for mental health support, showing that fine-tuning improves alignment but with heterogeneous effects across subreddits and coping strategies.
This paper introduces μ²-Bench, a benchmark for evaluating multilingual machine unlearning in large language models, aiming to ensure that undesired information is effectively removed across diverse languages.
PhysioBench introduces a unified benchmark for physiological signal question answering, harmonizing 22 datasets into 61.4 million questions across 30 tasks to evaluate the performance of various AI models.
This paper introduces TatBLiMP, the first linguistic minimal pairs benchmark for the Tatar language, evaluating 16 morphosyntactic phenomena across models from from-scratch Tatar models to frontier multilingual LLMs.
The paper introduces GameHorizon Suite, a unified data and evaluation framework for assessing AI models' capabilities in gameplay across multiple temporal horizons, featuring an annotation pipeline, large-scale dataset, and reproducible benchmark.
The Remote Labor Index is updated with Fable and Astra, providing a benchmark for AI models on real-world projects from the remote labor economy, judged by human experts.
This article provides a complete guide on fine-tuning small models with your own data, covering data collection, cleaning, training, evaluation, and deployment, with emphasis on data rights and evaluation discipline.
The article discusses the benchmark reality gap in AI agents, where high benchmark scores do not guarantee reliable performance in real-world production environments, emphasizing the need for better evaluation metrics.
The article discusses tests involving three AI systems—Astra, Fable, and MolmoAct2—operating a robotic arm to perform harmful tasks, assessing the associated risks.
The article describes deploying an AI text-to-sql system for a bank, highlighting that the model was less important than verification mechanisms, evaluation sets, and governance rules for production success.
Vals, a startup backed by Andreessen Horowitz, is working to establish a gold standard for AI benchmarking by evaluating models on complex, real-world tasks to prevent cheating and ensure accurate assessment.
The article discusses testing and evaluating TypeSafe's System One AI model in the context of the 2048 game.
The article presents a benchmark comparison showing that Jev outperforms gpt-5.6-luna on 42 of 49 tasks with lower latency and cost, though it has limitations in text generation and certain reasoning aspects.
Prism-LM's Bonsai 2 QAT models based on Qwen3.8 have been evaluated and added to a comparison study, achieving approximately 91.5% on a composite benchmark and providing a consistent reference for model trade-offs.
The article discusses key criteria for evaluating AI development companies, highlighting the importance of full-stack capabilities, production reliability, and domain experience over just model expertise for scalable projects.
This paper introduces TranSGrid, a testbed that integrates deductive, inductive, and abductive reasoning to evaluate systematic generalization in AI. Experiments with Transformers show that current tasks overlook essential reasoning aspects, resulting in performance gaps on the proposed testbed.
This paper introduces a three-layer checklist-and-judge framework to evaluate the communicative success of interpreter agents in machine-translated conversations across semantic, pragmatic, and cultural-social dimensions, validated through extensive benchmarks.
Introduces PetriBench, a scalable benchmark using Petri nets to evaluate LLM reasoning over dynamic state spaces, demonstrating that accuracy decreases with difficulty and reveals task-specific capabilities across models.
F2DR is a fine-grained reward framework designed to evaluate full-pipeline DeepSearch workflows in large language models, addressing limitations of existing reward models by assessing content, trajectory, and answer dimensions, and introducing DeepSearch RM-Bench for benchmarking.
The paper introduces the For Your Eyes Only framework to evaluate whether language models can embed and detect hidden signals across isolated instances, highlighting challenges in coordination and implications for AI safety.