legal-benchmark

Tag

Cards List
#legal-benchmark

Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks

arXiv cs.CL · 2026-08-05 Cached

This paper proposes automatically auditing answer correctness and legal authority grounding jointly in legal benchmarks, showing that LLMs often produce correct answers while citing wrong governing provisions under ordinary prompts.

0 favorites 0 likes
#legal-benchmark

@akshay_pachaar: Don't train the model, evolve the harness. I read a brilliant blog post from Hugging Face where they took a frozen open…

X AI KOLs Following · 2026-07-03 Cached

The article discusses a Hugging Face experiment where an automated loop rewrites only the code (harness) around a frozen model, raising its benchmark score from 0% to near Sonnet 4.6 at lower cost, demonstrating that many benchmark failures stem from the harness, not the model itself.

0 favorites 0 likes
#legal-benchmark

@gabepereyra: Harvey partnered with @appliedcompute to train a legal agent. We optimized each part of the agent stack, including the …

X AI KOLs Following · 2026-06-22 Cached

Harvey partnered with Applied Compute to train a legal agent, optimizing the agent stack and post-training the GLM-5.1 model using reward signals from their Legal Agent Benchmark.

0 favorites 0 likes
#legal-benchmark

TW-LegalBench: Measuring Taiwanese Legal Understanding

arXiv cs.CL · 2026-06-18 Cached

TW-LegalBench is a benchmark for evaluating large language models on Taiwanese legal understanding, including over 16,000 multiple-choice questions, 117 essay questions, and 14,000 legal judgment prediction instances. Results show top models exceed the passing threshold for lawyers but fall short for judges, highlighting challenges in reliable legal text generation.

0 favorites 0 likes
#legal-benchmark

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification

arXiv cs.CL · 2026-04-22 Cached

Researchers release LegalBench-BR, the first public benchmark for evaluating LLMs on Brazilian legal text classification, showing LoRA-fine-tuned BERTimbau dramatically outperforms GPT-4o mini and Claude 3.5 Haiku.

0 favorites 0 likes
#legal-benchmark

harveyai/harvey-labs

GitHub Trending (daily) · 2026-08-09 Cached

Harvey AI released Harvey LAB, an open-source benchmark for evaluating LLM agents on realistic legal work, featuring a dataset of 1,671 tasks across 24+ practice areas and an execution harness.

1 favorites 1 likes
← Back to home

Submit Feedback