evaluation

Tag

Cards List
#evaluation

Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

arXiv cs.CL · 2d ago Cached

This paper introduces the COPES dataset and a three-axis evaluation framework to assess LLM alignment with community perspectives for mental health support, showing that fine-tuning improves alignment but with heterogeneous effects across subreddits and coping strategies.

0 favorites 0 likes
#evaluation

$\mu^2$-Bench: A Multilingual Machine Unlearning Benchmark

arXiv cs.CL · 2d ago Cached

This paper introduces μ²-Bench, a benchmark for evaluating multilingual machine unlearning in large language models, aiming to ensure that undesired information is effectively removed across diverse languages.

0 favorites 0 likes
#evaluation

PhysioBench: A Unified Benchmark for Physiological Signal Question Answering

arXiv cs.CL · 2d ago Cached

PhysioBench introduces a unified benchmark for physiological signal question answering, harmonizing 22 datasets into 61.4 million questions across 30 tasks to evaluate the performance of various AI models.

0 favorites 0 likes
#evaluation

TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar

arXiv cs.CL · 2d ago Cached

This paper introduces TatBLiMP, the first linguistic minimal pairs benchmark for the Tatar language, evaluating 16 morphosyntactic phenomena across models from from-scratch Tatar models to frontier multilingual LLMs.

0 favorites 0 likes
#evaluation

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Hugging Face Daily Papers · 2d ago Cached

The paper introduces GameHorizon Suite, a unified data and evaluation framework for assessing AI models' capabilities in gameplay across multiple temporal horizons, featuring an annotation pipeline, large-scale dataset, and reproducible benchmark.

0 favorites 0 likes
#evaluation

Remote Labor Index updated with Fable and Astra

Reddit r/singularity · 2d ago

The Remote Labor Index is updated with Fable and Astra, providing a benchmark for AI models on real-world projects from the remote labor economy, judged by human experts.

0 favorites 0 likes
#evaluation

@yibie: https://x.com/yibie/status/2101502455544451502

X AI KOLs Timeline · 3d ago Cached

This article provides a complete guide on fine-tuning small models with your own data, covering data collection, cleaning, training, evaluation, and deployment, with emphasis on data rights and evaluation discipline.

0 favorites 0 likes
#evaluation

Your AI Agent Scores Well on Benchmarks. So Why Does It Still Fail in Production?

Reddit r/AI_Agents · 3d ago

The article discusses the benchmark reality gap in AI agents, where high benchmark scores do not guarantee reliable performance in real-world production environments, emphasizing the need for better evaluation metrics.

0 favorites 0 likes
#evaluation

Astra, Fable, and MolmoAct2 were put to the test by tasking them with 4 harmful operations through a robotic arm, to see just how risky things can get

Reddit r/singularity · 3d ago

The article discusses tests involving three AI systems—Astra, Fable, and MolmoAct2—operating a robotic arm to perform harmful tasks, assessing the associated risks.

0 favorites 0 likes
#evaluation

We put an AI agent in front of a bank's data warehouse. The part that mattered was not the model.

Reddit r/AI_Agents · 3d ago

The article describes deploying an AI text-to-sql system for a bank, highlighting that the model was less important than verification mechanisms, evaluation sets, and governance rules for production success.

0 favorites 0 likes
#evaluation

Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking

TechCrunch AI · 3d ago Cached

Vals, a startup backed by Andreessen Horowitz, is working to establish a gold standard for AI benchmarking by evaluating models on complex, real-world tasks to prevent cheating and ensure accurate assessment.

0 favorites 0 likes
#evaluation

kicking the tires on jev (TypeSafe's System One model) with 2048

Lobsters Hottest · 3d ago

The article discusses testing and evaluating TypeSafe's System One AI model in the context of the 2048 game.

0 favorites 0 likes
#evaluation

I benchmarked Jev against gpt-5.6-luna!

Reddit r/ArtificialInteligence · 4d ago

The article presents a benchmark comparison showing that Jev outperforms gpt-5.6-luna on 42 of 49 tasks with lower latency and cost, though it has limitations in text generation and certain reasoning aspects.

0 favorites 0 likes
#evaluation

Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison

Reddit r/LocalLLaMA · 4d ago

Prism-LM's Bonsai 2 QAT models based on Qwen3.8 have been evaluated and added to a comparison study, achieving approximately 91.5% on a composite benchmark and providing a consistent reference for model trade-offs.

0 favorites 0 likes
#evaluation

What do you actually look for when evaluating AI development companies?

Reddit r/artificial · 5d ago

The article discusses key criteria for evaluating AI development companies, highlighting the importance of full-stack capabilities, production reliability, and domain experience over just model expertise for scalable projects.

0 favorites 0 likes
#evaluation

What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis

arXiv cs.AI · 5d ago Cached

This paper introduces TranSGrid, a testbed that integrates deductive, inductive, and abductive reasoning to evaluate systematic generalization in AI. Experiments with Transformers show that current tasks overlook essential reasoning aspects, resulting in performance gaps on the proposed testbed.

0 favorites 0 likes
#evaluation

Evaluating Communicative Success in Machine-Translated Conversation

arXiv cs.CL · 5d ago Cached

This paper introduces a three-layer checklist-and-judge framework to evaluate the communicative success of interpreter agents in machine-translated conversations across semantic, pragmatic, and cultural-social dimensions, validated through extensive benchmarks.

0 favorites 0 likes
#evaluation

PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

arXiv cs.CL · 5d ago Cached

Introduces PetriBench, a scalable benchmark using Petri nets to evaluate LLM reasoning over dynamic state spaces, demonstrating that accuracy decreases with difficulty and reveals task-specific capabilities across models.

0 favorites 0 likes
#evaluation

F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

arXiv cs.CL · 5d ago Cached

F2DR is a fine-grained reward framework designed to evaluate full-pipeline DeepSearch workflows in large language models, addressing limitations of existing reward models by assessing content, trajectory, and answer dimensions, and introducing DeepSearch RM-Bench for benchmarking.

0 favorites 0 likes
#evaluation

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

arXiv cs.CL · 5d ago Cached

The paper introduces the For Your Eyes Only framework to evaluate whether language models can embed and detect hidden signals across isolated instances, highlighting challenges in coordination and implications for AI safety.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback