fine-grained

Tag

Cards List
#fine-grained

Quantifying the Affective Gap: A Zero-Shot Evaluation of LLMs on Fine-Grained Emotion Taxonomies

arXiv cs.CL · 2026-07-02 Cached

This paper presents a zero-shot evaluation of three LLMs (Claude, GPT-5.4, Gemini) on a 13-class emotion classification task, finding no model exceeds 39.9% accuracy and revealing systematic failures on specific emotions such as love, confusion, and shame.

0 favorites 0 likes
#fine-grained

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

Hugging Face Daily Papers · 2026-06-24 Cached

Physics Question Scene Graph (PQSG) is a hierarchical question-based pipeline using VLMs to evaluate video generation models' physical plausibility with fine-grained violation detection. It introduces the FinePhyEval dataset and shows higher correlation with human judgments than prior work.

0 favorites 0 likes
#fine-grained

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv cs.CL · 2026-06-16 Cached

This paper introduces CHILLGuard, a fine-grained Chinese LLM content safety guardrail built on a new 5-macro, 31-micro category risk taxonomy and a scalable multi-stage data construction pipeline. The model achieves state-of-the-art performance, improving F1 score by 15.92% over existing baselines.

0 favorites 0 likes
#fine-grained

GENIE: A Fine-Grained Measure for Novelty

arXiv cs.CL · 2026-06-12 Cached

GENIE is a fine-grained evaluation metric that measures the novelty of LLM responses along task-specific features, providing more insight than holistic metrics.

0 favorites 0 likes
#fine-grained

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

Hugging Face Daily Papers · 2026-05-30 Cached

FineVerify is a self-verification framework for agentic search that decomposes questions into sub-questions, verifies sampled candidates, and selects the best one, achieving substantial accuracy improvements over baselines on multiple benchmarks, including enabling GPT-5-mini to surpass GPT-5 on BrowseComp-Plus.

0 favorites 0 likes
#fine-grained

@AdinaYakup: Qwen @Alibaba_Qwen just dropped a new Text to Image benchmark + a judge model https://huggingface.co/collections/Qwen/q…

X AI KOLs Following · 2026-05-28 Cached

Qwen released a new Text-to-Image benchmark with 56 fine-grained evaluation facets, measuring creativity beyond prompt alignment, and includes a human-aligned judge model.

0 favorites 0 likes
#fine-grained

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

arXiv cs.LG · 2026-05-21 Cached

Introduces ClaimDiff-RL, a reinforcement learning framework for long-form image captioning that uses typed, verifiable claim differences as reward units to separately measure and balance hallucination and missing facts, improving faithfulness and coverage.

0 favorites 0 likes
#fine-grained

Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models

arXiv cs.LG · 2026-05-20

A new framework for automated benchmark generation enables fine-grained, comprehensive evaluation of foundation models with lower error rates and richer metadata, as demonstrated on ML, Corporate Finance, and Personal Finance benchmarks.

0 favorites 0 likes
#fine-grained

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

arXiv cs.CL · 2026-05-20 Cached

This paper introduces REFLECT, a meta-evaluation benchmark for assessing the reliability of LLM judges in evaluating deep research agents. Experiments show current LLM judges remain unreliable, with overall accuracies below 55% across reasoning, tool-use, and report-quality failures.

0 favorites 0 likes
#fine-grained

Towards Fine-Grained and Verifiable Concept Bottleneck Models

arXiv cs.LG · 2026-05-15 Cached

This paper proposes a fine-grained concept bottleneck model framework that grounds each concept in localized visual evidence, enabling direct verification of concept correctness and improving transparency in medical imaging tasks.

0 favorites 0 likes
#fine-grained

DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

arXiv cs.AI · 2026-05-15 Cached

This paper introduces DVMap, a framework for fine-grained pluralistic value alignment in LLMs that uses high-consensus demographic-value mapping instead of coarse national labels, achieving strong generalization across demographics, countries, and values.

0 favorites 0 likes
#fine-grained

CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark

arXiv cs.CL · 2026-04-21 Cached

Researchers from Peking University introduce CFMS, the first fine-grained Chinese multimodal sarcasm detection benchmark with 2,796 image-text pairs and a triple-level annotation framework (sarcasm identification, target recognition, explanation generation), along with a novel RL-augmented in-context learning method (PGDS) that significantly outperforms existing baselines.

0 favorites 0 likes
← Back to home

Submit Feedback