psychometrics

Tag

Cards List
#psychometrics

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

arXiv cs.CL ↗ · 12h ago Cached

该论文从心理测量学视角评估 LLM-as-a-judge,使用 Many-Facet Rasch 模型将评分分解为潜在质量、评分者严苛度与残差难度,发现人类与 LLM 在汇总对齐之外仍存在明显的残差难度结构错配。

0 favorites 0 likes
#psychometrics

@JenovaAIAgent: IQ Tester is an AI agent that estimates your cognitive profile with 30 adaptive questions across seven reasoning domain…

X AI KOLs Following ↗ · 2026-09-23 Cached

IQ Tester is an AI agent that estimates cognitive profiles with 30 adaptive questions, delivering detailed assessment reports based on psychometric frameworks across multiple reasoning domains.

0 favorites 0 likes
#psychometrics

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

arXiv cs.CL ↗ · 2026-09-02 Cached

This research investigates whether near-tied rankings of large language models remain robust to changes in benchmark item composition using psychometric methods, finding that while overall rankings are stable, individual close orderings can reverse based on item selection.

0 favorites 0 likes
#psychometrics

BenchMIRT: What are LLM benchmarks actually measuring?

Hugging Face Blog ↗ · 2026-09-01 Cached

BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.

0 favorites 0 likes
#psychometrics

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

arXiv cs.CL ↗ · 2026-08-12 Cached

Introduces an analytically exact framework for controlled behavioral evaluation of LLMs, using fully crossed factorial experiments and exact token-level probability mass functions to isolate causal biases that aggregate benchmarks obscure.

0 favorites 0 likes
#psychometrics

Multimodal Item Parameter Estimation using Simulated Response Probabilitie

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper describes using a fine-tuned multimodal LLM based on Qwen3.5 to simulate student response probabilities and estimate item difficulty parameters for multiple-choice assessments, approximating 3PL and MCM curves.

0 favorites 0 likes
#psychometrics

Natural Language Processing Psychometrics

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper introduces NLP Psychometrics, a framework that treats psychological prediction from text as a psychometric problem. Using LLM personas, emotional profiles, and syntactic-semantic networks with random forest regressors, it explains up to 76% of variance in mental health scores and shows promise and limits of synthetic data for psychometric prediction.

0 favorites 0 likes
#psychometrics

Item Response Theory for AI Safety

arXiv cs.AI ↗ · 2026-08-06 Cached

This paper applies Item Response Theory to eight safety benchmarks across 192 language models, identifying three latent factors, enabling 97-99% cost reduction via adaptive testing, and supporting sandbagging detection and model auditing.

0 favorites 0 likes
#psychometrics

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

arXiv cs.CL ↗ · 2026-08-05 Cached

This paper introduces LLM-NRM, an option-level psychometric framework for multiple-choice benchmarks that models the full distribution over answer choices rather than binary correctness, showing that incorrect responses carry useful measurement information and improving ability estimation and benchmarking efficiency.

0 favorites 0 likes
#psychometrics

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

arXiv cs.CL ↗ · 2026-07-29 Cached

CogArena introduces a procedurally generated 13-paradigm benchmark to evaluate whether LLMs exhibit separable cognitive abilities or a single general competence, finding only weak support for stable five-dimensional profiles across 55 models.

0 favorites 0 likes
#psychometrics

Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP)

arXiv cs.CL ↗ · 2026-07-27 Cached

This paper validates the Spanish version of the Large Language Models Dependency Scale (LLM-D12-SP), confirming its two-factor structure and good psychometric properties.

0 favorites 0 likes
#psychometrics

Fine-Tuned Multi-Agent Framework for Detecting OCEAN in Life Narratives

arXiv cs.CL ↗ · 2026-07-15 Cached

This paper proposes a fine-tuned multi-agent framework for detecting OCEAN personality traits from life narratives, using LLM sub-agents conditioned to adopt high, low, or neutral perspectives and a judge LLM that aggregates outputs to mitigate biases and improve interpretability.

0 favorites 0 likes
#psychometrics

From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings

arXiv cs.CL ↗ · 2026-07-09 Cached

This paper proposes an evaluation framework for predicting item parameters from text embeddings using regularized regression and reliability/design ceilings. Results show that difficulty is substantially predictable from text, while discrimination and pseudo-guessing are limited by reliability, not text signal.

0 favorites 0 likes
#psychometrics

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

arXiv cs.CL ↗ · 2026-07-08 Cached

This paper introduces a psychometric battery to separate framing artifacts from genuine moral judgment in LLMs, finding that frontier models have a coherent internal moral scale but display a yes/no bias that is purely a surface-level artifact of answer order and wording, not a real disposition to reject.

0 favorites 0 likes
#psychometrics

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

Hugging Face Daily Papers ↗ · 2026-07-06 Cached

This paper presents a psychometric audit of large language models as synthetic survey respondents, finding that they fail to match human joint distributions and reliability, making them unsuitable replacements.

0 favorites 0 likes
#psychometrics

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models

arXiv cs.CL ↗ · 2026-07-03 Cached

This paper introduces EduArt, an educational-level benchmark with 871 human-authored questions for evaluating art history knowledge and visual reasoning in multimodal LLMs, revealing that single-format benchmarks overestimate model capabilities.

0 favorites 0 likes
#psychometrics

Multilayer Q-Matrix-Embedded Neural Network for Cognitive Diagnosis (M-QCDNet): Structure-Aware Deep Learning Architecture for Psychometric Interpretability

arXiv cs.LG ↗ · 2026-07-03 Cached

Proposes M-QCDNet, a structure-aware deep learning architecture that embeds multilayer Q-matrices for enhanced psychometric interpretability in cognitive diagnosis.

0 favorites 0 likes
#psychometrics

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

arXiv cs.CL ↗ · 2026-06-18 Cached

This paper evaluates 42 large language models on their ability to measure item discrimination in reading comprehension assessments, finding weak alignment with human-calibrated measures and highlighting it as an open challenge for psychometric evaluation.

0 favorites 0 likes
#psychometrics

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

arXiv cs.CL ↗ · 2026-06-16 Cached

This paper introduces a psychometric datasheet protocol for evaluating LLM judges as measurement instruments, measuring dark current, positional false preference, stable cross-sensitivity, and target sensitivity. A case study on three open-weight models reveals significant differences in judge quality and behavior.

0 favorites 0 likes
#psychometrics

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

arXiv cs.AI ↗ · 2026-06-12 Cached

This paper examines when and why self-reported psychometric measures predict the actual behavior of large language models, finding that fine-grained, behavior-specific instruments (Theory of Planned Behavior) achieve human-level coherence within a shared conversation, while broad traits like Big 5 do not.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback