Tag
This paper empirically investigates the relationship between salience and answerability of questions in naturalistic dialogue, finding a robust but low positive correlation that is weaker than in monologic text, indicating that conversational structure is less predictable.
This paper introduces Coupled Usage–Sense Processes (CUSP), a hierarchical model for analyzing lexical semantic change by quantifying timing, mechanisms, and attributions in word usage over time.
This paper proposes a knowledge graph-based evaluation framework for assessing the contextual understanding of large language models in question answering, introducing a new similarity measure called S3KG that achieves superior performance over existing baselines.
This paper introduces ℓp-LoRA, a principled method for automatic rank allocation in low-rank adaptation using ℓp regularization, demonstrating competitive performance on NLP tasks.
The study introduces SHAP-RTL, a rendering layer that corrects the visualization of SHAP and LIME explanations for right-to-left languages, addressing issues like token sequence and script shaping while preserving original attribution values.
This paper replicates a distributional-semantics extractive summarization method for Hindi and evaluates it on standard corpora, finding that sentence position is the only contributing feature and current Hindi benchmarks fail to incentivize advanced content selection.
The article likely presents research on contrastive language models, exploring the use of contrastive learning techniques in language model development.
The paper proposes retrieval strategies to filter irrelevant HTML content for LLM-based autonomous web agents, improving performance on benchmarks like WebArena.
This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.
This paper introduces UGTPhon, a benchmark for grapheme-to-phoneme conversion in user-generated text, and presents a compositional approach that improves performance by leveraging canonical forms.
This paper presents design principles and the 'AGIMUD' software for enabling socio-affective interactions between humans and multiple AI agents in simulated dynamic worlds, leveraging generative AI and distributed processing.
This paper introduces Arafa, a large-scale Arabic fact-checking dataset generated using LLMs, aimed at addressing the scarcity of resources for automatic fact-checking in Arabic.
The paper proposes a Context-to-Answer-Aligned Memory Compression (CMC) framework that compresses long input contexts into compact memory embeddings to reduce LLM inference costs without modifying decoder weights, achieving significant performance and efficiency gains.
Peerify is a pipeline for automatically verifying peer-review claims against manuscript evidence, using a benchmark of 800 claims from NeurIPS 2024 and ICLR 2024, demonstrating that retrieval-centered verification outperforms entailment baselines.
OpenAI has released GPT-6 Sol and GPT-6 Luna, which are faster and more affordable models based on GPT-6 Astra's advances, aimed at scalable use.
This paper presents a unified framework for evaluating multimodal synthetic data using semantic quantization and cross-modal metrics, emphasizing the need for explicit evaluation with permutation baselines and coverage reporting.
Hemmingway-1 is a finetuned version of the Qwen3.8-27B AI model, designed to generate text in a more human-like writing style.
The user describes challenges and partial solutions for making an AI agent reliably convert raw meeting notes into structured action items with owners and due dates, highlighting issues like hallucination and missed context.
The paper introduces CaLR, a framework that reformulates reasoning as constrained latent optimization using causal topology to enhance diffusion language models, achieving state-of-the-art performance on complex benchmarks.
This paper analyzes reasoning traces in large reasoning models for machine translation, finding that reasoning benefits are conditional and introducing Hierarchical Meta-Summarization to understand trace patterns.