psychometrics

Tag

Cards List
#psychometrics

When it comes to predicting people’s preferences, it pays to consider “the power of three”

MIT News — Artificial Intelligence ↗ · 2026-06-11 Cached

MIT researchers present a paper showing that using three-way comparisons instead of pairwise comparisons can significantly improve the accuracy of random utility models for predicting human preferences.

0 favorites 0 likes
#psychometrics

Annotator Positionality as Signal: Psychometric Weighting for Anti-Autistic Ableism Detection

arXiv cs.CL ↗ · 2026-05-27 Cached

This paper introduces a bias-aware evaluation framework for detecting anti-autistic ableist language in LLMs, using psychometrically-weighted ground truth based on annotator positionality. It finds that LLMs frequently misclassify community-reclaimed language as ableist and rely on surface-level keyword matching rather than context.

0 favorites 0 likes
#psychometrics

Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment

arXiv cs.AI ↗ · 2026-05-20 Cached

Introduces Generative-Evaluative Agreement (GEA), a validity criterion for LLM-enabled adaptive assessments, and measures it on a two-stage adaptive test, finding that the model recovers about half the intended variance with systematic bias.

0 favorites 0 likes
#psychometrics

Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning

arXiv cs.CL ↗ · 2026-05-19 Cached

The paper proposes fine-tuning transformer encoders end-to-end for response-free item difficulty modelling of multiple-choice reading comprehension items, with component-wise and multi-task variants, showing that multi-task learning improves in small-sample regimes.

0 favorites 0 likes
#psychometrics

Can We Trust AI-Inferred User States. A Psychometric Framework for Validating the Reliability of Users States Classification by LLMs in Operational Environments

arXiv cs.AI ↗ · 2026-05-18 Cached

This paper empirically tests the psychometric reliability of LLM-based user state classification, finding that only 31 of 213 metrics met reliability criteria, questioning trust in real-time adaptive systems.

0 favorites 0 likes
#psychometrics

Uneven Evolution of Cognition Across Generations of Generative AI Models

Reddit r/singularity ↗ · 2026-05-11 Cached

This paper introduces a psychometric framework and the AIQ Benchmark to evaluate the cognitive profiles of generative AI models, revealing uneven evolution with strong verbal skills but stagnant perceptual reasoning.

0 favorites 0 likes
#psychometrics

We gave 45 psychological questionnaires to 50 LLMs. What we found was not “personality.”

Reddit r/artificial ↗ · 2026-05-07

Researchers analyzed 50 LLMs across 45 psychometric questionnaires, identifying a 'Pinocchio Dimension' that measures how models endorse inner experiences rather than reflecting true personality traits.

0 favorites 0 likes
#psychometrics

The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring

arXiv cs.CL ↗ · 2026-04-20 Cached

A new cross-domain benchmark (Metacognitive Monitoring Battery) with 524 items evaluates LLM self-monitoring capabilities across six cognitive domains using human psychometric methodology. Applied to 20 frontier LLMs, it reveals three distinct metacognitive profiles and shows that accuracy rank and metacognitive sensitivity rank are largely inverted.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback