Tag
This paper introduces a benchmark dataset for Bangla idioms and evaluates recent large language models on idiom-related tasks, revealing substantial variability in performance across models.
FrameBench is a new benchmark for evaluating whether large language models can distinguish context-dependent frame-semantic interpretations of verbs, constructed for English and Japanese using FrameNet resources and released with code.
MineTRACE is a web-based evidence-grounded interactive reasoning system that integrates geochemical, geophysical, and geological data to provide transparent mineral prospectivity scores and natural language interaction, supporting efficient and verifiable mineral exploration.
This paper explores how persona prompting with different attribute selection methods affects the alignment of large language models with human responses in social surveys. It finds that effectiveness depends on human response variation and the choice of attributes.
MultiGhostBench is a multilingual benchmark for long-form LLM-generated text attribution under distribution shifts, featuring 928 books in six languages and highlighting performance degrades and no single method consistently best across settings.
This study proposes a two-stage framework for sentence-level depression symptom recognition using candidate generation and definition-guided verification, achieving best accuracy and F1 scores among evaluated methods.
The article explores Pangram, an AI startup that detects AI-generated text, its role in publishing scandals, and questions about the trustworthiness of its detection model.
This article provides a detailed explanation of the Transformer architecture, covering attention mechanisms, QKV, and residual connections, while tracing its historical development from N-gram to LSTM.
This paper investigates the curse of multilinguality in lexical normalization, finding that training a single model on multiple languages leads to decreased per-language accuracy, with optimal performance when languages are trained in small groups.
CUDA-Harness is a framework that uses agentic techniques to generate and optimize CUDA kernels from natural language descriptions, addressing challenges in Text2CUDA by connecting high-level semantics with low-level implementation and verification.
AI Historian is an AI agent system that helps historians organize and verify person-centred temporal clues from dispersed historical narratives, reducing the cost of historical research while achieving high accuracy in temporal localization.
The article demonstrates a Python/Flask example using Telnyx Call Control and AI inference to create a natural language IVR system, allowing callers to verbally state their needs instead of navigating fixed menus.
This paper presents the first application of data science to evaluate the UK Honours system using natural language processing, introducing a novel sentiment analysis algorithm called Minos to assess public opinion on honours recipients.
This paper presents a production-grade framework that uses large language models to convert natural-language pricing policies into executable decisions for tourism pricing, achieving significant efficiency gains and auditability in real-world deployment.
This paper investigates the interpretability of DAPF-based models for dementia detection, revealing that while DAPF achieves strong performance, its token-level explanations lack faithfulness.
This paper introduces a causal graph-based attention mechanism to enhance retrieval precision in Retrieval-Augmented Generation (RAG) systems, showing improvements in keyword-stuffing regimes of proprietary knowledge bases.
The paper introduces Peony, the first benchmark for evaluating large language models on comprehending poetic logic in modern Chinese poetry, and evaluates six mainstream LLMs, revealing their limitations in this specialized task.
The paper introduces an ontology-driven framework to quantify and enforce structural consistency in document-level relation extraction datasets, reducing logical contradictions and improving model generalization when using distant supervision data.
This paper introduces ImmigrationReason, a large-scale structured dataset of U.S. immigration appeals for legal reasoning research, addressing the gap in administrative adjudication data for NLP studies.
This study develops an ambiguity taxonomy to evaluate large language model performance on clinical registry abstraction from unprocessed EMR data, finding that LLM accuracy is significantly lower than human abstractors and declines as task ambiguity increases.