LLM Ass Bench
Summary
LLM Ass Bench is a benchmark or tool for evaluating Large Language Models, with a focus on prompts.
View Cached Full Text
Cached at: 09/22/26, 09:53 PM
Similar Articles
Benchmarking LLMs
A study or report on benchmarking large language models, likely comparing performance across various tasks.
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
OmniAssistBench is a benchmark for evaluating omni-modal large language models as real-time video assistants, revealing that current models struggle with visual prompts, context retention, and timely responses.
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
DLawBench is a new benchmark for evaluating large language models in multi-turn legal consultation, covering Chinese and US law with four client types. Experiments show significant room for improvement, with the best model achieving only 0.562 on legal reasoning.
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
This article introduces Magis-Bench, a benchmark for evaluating large language models on magistrate-level legal tasks such as judicial reasoning and sentence drafting, using data from Brazilian judicial exams.
BenchMIRT: What are LLM benchmarks actually measuring?
BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.