factual-consistency

Tag

Cards List
#factual-consistency

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

arXiv cs.CL · 2026-07-09 Cached

This paper proposes a multi-factor scoring system for evaluating LLM responses, integrating accuracy, conciseness, factual consistency, readability, and coherence. Applied to the TruthfulQA dataset, it reveals strengths and limitations of mainstream models, offering a transparent evaluation framework.

0 favorites 0 likes
#factual-consistency

@omarsar0: If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use L…

X AI KOLs Following · 2026-06-27 Cached

BinEval is a new framework that decomposes LLM evaluation criteria into atomic binary questions, improving interpretability and enabling targeted prompt optimization, achieving strong results on factual consistency benchmarks.

0 favorites 0 likes
#factual-consistency

Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics

arXiv cs.CL · 2026-05-27 Cached

This paper introduces a method to improve factual consistency in text summarization by aggregating scores from multiple weak metrics via preference learning, achieving consistent factuality gains across various language models.

0 favorites 0 likes
← Back to home

Submit Feedback