Tag
The article reports on a performance comparison between glm-4-flash and TypeSafe Jev on 299 real user intents, showing glm-4-flash's higher accuracy but TypeSafe Jev's faster speed and lower cost.
LLM Ass Bench is a benchmark or tool for evaluating Large Language Models, with a focus on prompts.
OpenAI has outlined four priority areas and principles for supporting independent third-party assessments of AI safety, emphasizing rigor, security, and accountability in frontier AI development.
QontoFAQ introduces a new benchmark and metric for information retrieval, focusing on improving relevance measurement for answering product questions, with associated code and dataset released.
JevBench v1.3.0 is a reproducible benchmark for Jev-class decision models, evaluating and ranking 52 systems based on intelligence, calibration, speed, and cost.
The paper proposes that intransitive preferences in language models arise from multiple latent orderings and introduces a mixture Bradley–Terry model to analyze them, showing better explanatory power across various models and tasks.
This exploratory pilot study evaluates personal information output from conversational interactions in generative AI systems, finding limited impact from model design differences and suggesting inferred profiles are constructed from contextual information.
The paper introduces Whiteboard, the first benchmark for evaluating imagination in large language models by cross-referencing it with hallucination, and reveals a counterintuitive negative correlation between the two across 79 state-of-the-art LLMs.
Kev is a family of small decision models built on Qwen3.5, offering pretrained weights and training code for yes/no, multiple-choice, and rating questions. It includes a web playground and is compatible with TypeSafe's System One API.
A user compares capy to other AI harnesses and reports it performed better with half the cost and time, while Garry Tan endorses it as a top tool for agentic coding.
This article likely discusses the benchmark results for the Grok 4.7 AI model, comparing its performance across various tasks.
Kev is a family of small decision models built on Qwen3.5, offering open-source training code and pretrained weights for local deployment with support for various question types.
OpenMAS-GCom introduces a diagnostic benchmark for evaluating graph-enhanced multi-agent systems by using controlled interventions to attribute performance differences to specific organizational components.
This paper introduces a behavior-aware role-playing framework called SIBPersona to enhance the fidelity of impersonating social media influencers by integrating situation-dependent behavioral strategies and an evaluation protocol for obscure individuals.
This paper introduces the COPES dataset and a three-axis evaluation framework to assess LLM alignment with community perspectives for mental health support, showing that fine-tuning improves alignment but with heterogeneous effects across subreddits and coping strategies.
This paper introduces μ²-Bench, a benchmark for evaluating multilingual machine unlearning in large language models, aiming to ensure that undesired information is effectively removed across diverse languages.
PhysioBench introduces a unified benchmark for physiological signal question answering, harmonizing 22 datasets into 61.4 million questions across 30 tasks to evaluate the performance of various AI models.
This paper introduces TatBLiMP, the first linguistic minimal pairs benchmark for the Tatar language, evaluating 16 morphosyntactic phenomena across models from from-scratch Tatar models to frontier multilingual LLMs.
The paper introduces GameHorizon Suite, a unified data and evaluation framework for assessing AI models' capabilities in gameplay across multiple temporal horizons, featuring an annotation pipeline, large-scale dataset, and reproducible benchmark.
The Remote Labor Index is updated with Fable and Astra, providing a benchmark for AI models on real-world projects from the remote labor economy, judged by human experts.