Tag
该论文从心理测量学视角评估 LLM-as-a-judge,使用 Many-Facet Rasch 模型将评分分解为潜在质量、评分者严苛度与残差难度,发现人类与 LLM 在汇总对齐之外仍存在明显的残差难度结构错配。
IQ Tester is an AI agent that estimates cognitive profiles with 30 adaptive questions, delivering detailed assessment reports based on psychometric frameworks across multiple reasoning domains.
This research investigates whether near-tied rankings of large language models remain robust to changes in benchmark item composition using psychometric methods, finding that while overall rankings are stable, individual close orderings can reverse based on item selection.
BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.
Introduces an analytically exact framework for controlled behavioral evaluation of LLMs, using fully crossed factorial experiments and exact token-level probability mass functions to isolate causal biases that aggregate benchmarks obscure.
This paper describes using a fine-tuned multimodal LLM based on Qwen3.5 to simulate student response probabilities and estimate item difficulty parameters for multiple-choice assessments, approximating 3PL and MCM curves.
This paper introduces NLP Psychometrics, a framework that treats psychological prediction from text as a psychometric problem. Using LLM personas, emotional profiles, and syntactic-semantic networks with random forest regressors, it explains up to 76% of variance in mental health scores and shows promise and limits of synthetic data for psychometric prediction.
This paper applies Item Response Theory to eight safety benchmarks across 192 language models, identifying three latent factors, enabling 97-99% cost reduction via adaptive testing, and supporting sandbagging detection and model auditing.
This paper introduces LLM-NRM, an option-level psychometric framework for multiple-choice benchmarks that models the full distribution over answer choices rather than binary correctness, showing that incorrect responses carry useful measurement information and improving ability estimation and benchmarking efficiency.
CogArena introduces a procedurally generated 13-paradigm benchmark to evaluate whether LLMs exhibit separable cognitive abilities or a single general competence, finding only weak support for stable five-dimensional profiles across 55 models.
This paper validates the Spanish version of the Large Language Models Dependency Scale (LLM-D12-SP), confirming its two-factor structure and good psychometric properties.
This paper proposes a fine-tuned multi-agent framework for detecting OCEAN personality traits from life narratives, using LLM sub-agents conditioned to adopt high, low, or neutral perspectives and a judge LLM that aggregates outputs to mitigate biases and improve interpretability.
This paper proposes an evaluation framework for predicting item parameters from text embeddings using regularized regression and reliability/design ceilings. Results show that difficulty is substantially predictable from text, while discrimination and pseudo-guessing are limited by reliability, not text signal.
This paper introduces a psychometric battery to separate framing artifacts from genuine moral judgment in LLMs, finding that frontier models have a coherent internal moral scale but display a yes/no bias that is purely a surface-level artifact of answer order and wording, not a real disposition to reject.
This paper presents a psychometric audit of large language models as synthetic survey respondents, finding that they fail to match human joint distributions and reliability, making them unsuitable replacements.
This paper introduces EduArt, an educational-level benchmark with 871 human-authored questions for evaluating art history knowledge and visual reasoning in multimodal LLMs, revealing that single-format benchmarks overestimate model capabilities.
Proposes M-QCDNet, a structure-aware deep learning architecture that embeds multilayer Q-matrices for enhanced psychometric interpretability in cognitive diagnosis.
This paper evaluates 42 large language models on their ability to measure item discrimination in reading comprehension assessments, finding weak alignment with human-calibrated measures and highlighting it as an open challenge for psychometric evaluation.
This paper introduces a psychometric datasheet protocol for evaluating LLM judges as measurement instruments, measuring dark current, positional false preference, stable cross-sensitivity, and target sensitivity. A case study on three open-weight models reveals significant differences in judge quality and behavior.
This paper examines when and why self-reported psychometric measures predict the actual behavior of large language models, finding that fine-grained, behavior-specific instruments (Theory of Planned Behavior) achieve human-level coherence within a shared conversation, while broad traits like Big 5 do not.