open-ended-generation

Tag

Cards List
#open-ended-generation

Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper investigates the effects of inference-time interventions and weight consolidation on open-ended AI generation, finding that training on value-filtered candidates improves mean quality but does not exceed classic heuristic performance.

0 favorites 0 likes
#open-ended-generation

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper proposes a human-grounded framework to measure distributional breadth (cultural reach) of LLM-generated content in open-ended tasks, introducing metrics LLM Coverage and In-Boundary Rate. Experiments show current LLMs produce plausible but narrow content concentrated near the center of human response space.

0 favorites 0 likes
#open-ended-generation

Rubrics as Privileged Information for Open-Ended Generation

arXiv cs.LG ↗ · 2026-08-05 Cached

This paper introduces Rubrics as Privileged Information (RuPI), extending on-policy self-distillation to open-ended generation by conditioning the teacher on rubrics as soft privileged information. The method outperforms rubric-as-reward RL and reference-completion distillation across multiple LLMs and benchmarks like HealthBench.

0 favorites 0 likes
#open-ended-generation

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

arXiv cs.CL ↗ · 2026-08-03 Cached

CalibratedRubric is a task-adaptive framework for building compact, measurable rubric banks for open-ended LLM evaluation, using Bayesian measurability filtering and IRT-based selection to improve human-gold agreement and rank fidelity across financial, healthcare, general, and legal benchmarks.

0 favorites 0 likes
#open-ended-generation

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

arXiv cs.AI ↗ · 2026-08-03 Cached

EarlyDx is a new large-scale benchmark for evaluating LLMs on open-ended, evidence-supported diagnosis generation at emergency department admission, built from 154,834 MIMIC-IV encounters. It reveals that even frontier and medical-specialized models struggle to synthesize admission-time evidence, with post-training only partially improving inference-dependent recall.

0 favorites 0 likes
#open-ended-generation

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

arXiv cs.CL ↗ · 2026-07-30 Cached

SERPO introduces a self-evolving rubric policy optimization framework for test-time reinforcement learning in open-ended generation, replacing answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters, achieving significant improvements on health and research benchmarks.

0 favorites 0 likes
#open-ended-generation

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Hugging Face Daily Papers ↗ · 2026-07-21 Cached

This paper introduces a two-level meta-rubric framework for evaluating factual completeness in open-ended generation, instantiated as the GAMUT benchmark. It features 1,813 questions across 10 domains and finds the benchmark challenging and discriminative, with top models scoring 58.7%.

0 favorites 0 likes
#open-ended-generation

Where You Inject Diversity Matters: A Unified Framework for Diverse Generation

arXiv cs.CL ↗ · 2026-06-10 Cached

This paper introduces a unified framework for test-time diverse generation in large language models, categorizing methods by where diversity is injected (surface-level vs. specification-level). It proposes specification-level methods that generate diverse intermediate specifications, achieving better output diversity across five open-ended tasks and four backbone models while maintaining quality.

0 favorites 0 likes
#open-ended-generation

G-Zero: Self-Play for Open-Ended Generation from Zero Data

Hugging Face Daily Papers ↗ · 2026-05-11 Cached

This paper introduces G-Zero, a verifier-free framework that enables autonomous large language model self-improvement through co-evolutionary training using intrinsic rewards and hint-based guidance. It aims to overcome the limitations of proxy LLM judges in open-ended tasks by deriving supervision from internal distributional dynamics.

0 favorites 0 likes
← Back to home

Submit Feedback