evaluation

Tag

Cards List
#evaluation

299 real user intents tested Jev against production base line. Here is the result.

Reddit r/AI_Agents · 4h ago

The article reports on a performance comparison between glm-4-flash and TypeSafe Jev on 299 real user intents, showing glm-4-flash's higher accuracy but TypeSafe Jev's faster speed and lower cost.

0 favorites 0 likes
#evaluation

LLM Ass Bench

Hacker News Top · 6h ago Cached

LLM Ass Bench is a benchmark or tool for evaluating Large Language Models, with a focus on prompts.

0 favorites 0 likes
#evaluation

@OpenAI: As part of our efforts to pace the frontier, we’re committed to supporting independent assessments with deep levels of …

X AI KOLs · 9h ago Cached

OpenAI has outlined four priority areas and principles for supporting independent third-party assessments of AI safety, emphasizing rigor, security, and accountability in frontier AI development.

0 favorites 0 likes
#evaluation

QontoFAQ: A better Information Retrieval Benchmark [R]

Reddit r/MachineLearning · 13h ago

QontoFAQ introduces a new benchmark and metric for information retrieval, focusing on improving relevance measurement for answering product questions, with associated code and dataset released.

0 favorites 0 likes
#evaluation

Show HN: JevBench, a reproducible benchmark for typed decision models

Hacker News Top · 13h ago Cached

JevBench v1.3.0 is a reproducible benchmark for Jev-class decision models, evaluating and ranking 52 systems based on intelligence, calibration, speed, and cost.

0 favorites 0 likes
#evaluation

Multiple latent orderings better predict language model preferences

arXiv cs.LG · 22h ago Cached

The paper proposes that intransitive preferences in language models arise from multiple latent orderings and introduces a mixture Bradley–Terry model to analyze them, showing better explanatory power across various models and tasks.

0 favorites 0 likes
#evaluation

Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems

arXiv cs.CL · 22h ago Cached

This exploratory pilot study evaluates personal information output from conversational interactions in generative AI systems, finding limited impact from model design differences and suggesting inferred profiles are constructed from contextual information.

0 favorites 0 likes
#evaluation

Is Imagination Derived from Hallucination? A Cross-Taxonomy Evaluation of Imagination and Hallucination in Large Language Models

arXiv cs.CL · 22h ago Cached

The paper introduces Whiteboard, the first benchmark for evaluating imagination in large language models by cross-referencing it with hallucination, and reveals a counterintuitive negative correlation between the two across 79 state-of-the-art LLMs.

0 favorites 0 likes
#evaluation

Kev (GitHub Repo)

TLDR AI · yesterday Cached

Kev is a family of small decision models built on Qwen3.5, offering pretrained weights and training code for yes/no, multiple-choice, and rating questions. It includes a web playground and is compatible with TypeSafe's System One API.

0 favorites 0 likes
#evaluation

@0xluffy: just did an eval on capy vs one of the big harnesses out there capy did better with half the cost and half the time

X AI KOLs Timeline · yesterday Cached

A user compares capy to other AI harnesses and reports it performed better with half the cost and time, while Garry Tan endorses it as a top tool for agentic coding.

0 favorites 0 likes
#evaluation

Grok 4.7 benchmarks

Reddit r/singularity · yesterday

This article likely discusses the benchmark results for the Grok 4.7 AI model, comparing its performance across various tasks.

0 favorites 0 likes
#evaluation

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

Hacker News Top · yesterday Cached

Kev is a family of small decision models built on Qwen3.5, offering open-source training code and pretrained weights for local deployment with support for various question types.

0 favorites 0 likes
#evaluation

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

arXiv cs.LG · yesterday Cached

OpenMAS-GCom introduces a diagnostic benchmark for evaluating graph-enhanced multi-agent systems by using controlled interventions to attribute performance differences to specific organizational components.

0 favorites 0 likes
#evaluation

From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers

arXiv cs.CL · yesterday Cached

This paper introduces a behavior-aware role-playing framework called SIBPersona to enhance the fidelity of impersonating social media influencers by integrating situation-dependent behavioral strategies and an evaluation protocol for obscure individuals.

0 favorites 0 likes
#evaluation

Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

arXiv cs.CL · yesterday Cached

This paper introduces the COPES dataset and a three-axis evaluation framework to assess LLM alignment with community perspectives for mental health support, showing that fine-tuning improves alignment but with heterogeneous effects across subreddits and coping strategies.

0 favorites 0 likes
#evaluation

$\mu^2$-Bench: A Multilingual Machine Unlearning Benchmark

arXiv cs.CL · yesterday Cached

This paper introduces μ²-Bench, a benchmark for evaluating multilingual machine unlearning in large language models, aiming to ensure that undesired information is effectively removed across diverse languages.

0 favorites 0 likes
#evaluation

PhysioBench: A Unified Benchmark for Physiological Signal Question Answering

arXiv cs.CL · yesterday Cached

PhysioBench introduces a unified benchmark for physiological signal question answering, harmonizing 22 datasets into 61.4 million questions across 30 tasks to evaluate the performance of various AI models.

0 favorites 0 likes
#evaluation

TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar

arXiv cs.CL · yesterday Cached

This paper introduces TatBLiMP, the first linguistic minimal pairs benchmark for the Tatar language, evaluating 16 morphosyntactic phenomena across models from from-scratch Tatar models to frontier multilingual LLMs.

0 favorites 0 likes
#evaluation

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Hugging Face Daily Papers · 2d ago Cached

The paper introduces GameHorizon Suite, a unified data and evaluation framework for assessing AI models' capabilities in gameplay across multiple temporal horizons, featuring an annotation pipeline, large-scale dataset, and reproducible benchmark.

0 favorites 0 likes
#evaluation

Remote Labor Index updated with Fable and Astra

Reddit r/singularity · 2d ago

The Remote Labor Index is updated with Fable and Astra, providing a benchmark for AI models on real-world projects from the remote labor economy, judged by human experts.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback