Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP)
Summary
This paper validates the Spanish version of the Large Language Models Dependency Scale (LLM-D12-SP), confirming its two-factor structure and good psychometric properties.
View Cached Full Text
Cached at: 07/27/26, 07:40 AM
# Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP) Source: [https://arxiv.org/abs/2607.22041](https://arxiv.org/abs/2607.22041) [View PDF](https://arxiv.org/pdf/2607.22041) > Abstract:There is a growing need for reliable and culturally validated instruments to assess psychological dependency on large language models \(LLMs\), particularly as LLMs are increasingly used for task execution, decision\-making, and communication in organizational and work\-related settings\. This need is especially relevant for Spanish\-speaking populations, where LLM adoption is rapidly expanding, yet validated psychometric tools remain scarce\. The present study reports the first validation of the Spanish version of the Large Language Model Dependency Scale \(LLM\-D12\-SP\), extending prior validations conducted in English\- and Arabic\-speaking samples\. The LLM\-D12 is a two\-dimensional instrument assessing Instrumental Dependency \(reliance on LLMs for performing tasks and supporting decisions\) and Relationship Dependency \(psychological reliance on LLMs for companionship and social interaction\)\. A total of 386 Spanish\-speaking participants \(M = 28\.0 years, SD = 6\.1; 55% male\) completed the LLM\-D12\-SP\. Confirmatory factor analysis supported the original two\-factor structure\. The scale demonstrated good internal consistency \(Cronbach's alpha = 0\.89 total; 0\.86 Instrumental; 0\.85 Relationship\)\. Discriminant validity analyses indicated that the two subscales represent related but distinct constructs\. External validation showed that both dependency dimensions were positively associated with internet addiction and perceived trustworthiness of LLMs, while showing weak or no association with need for cognition\. Together with prior English and Arabic validations, these findings establish cross\-linguistic support for the scale's structure and provide a psychometrically sound tool for investigating psychological aspects of LLM use in organizational contexts\. ## Submission history From: Mo El\-Haj \[[view email](https://arxiv.org/show-email/aad691d8/2607.22041)\] **\[v1\]**Fri, 24 Jul 2026 07:09:56 UTC \(2,657 KB\)
Similar Articles
Evaluating Developmental Cognition Capabilities of LLMs
This paper introduces the Developmental Sentence Completion Test (DSCT) to evaluate Large Language Models' ability to recognize developmental cognitive stages in text, finding that models perform better on synthetic personas than on real human responses.
When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening
This paper introduces a SCID-anchored benchmark of 555 interviews to evaluate five LLMs for psychiatric screening, finding that while models show potential, they tend to discount symptom evidence in the presence of preserved functioning or protective context, requiring careful validation.
Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System
This paper proposes a multi-factor scoring system for evaluating LLM responses, integrating accuracy, conciseness, factual consistency, readability, and coherence. Applied to the TruthfulQA dataset, it reveals strengths and limitations of mainstream models, offering a transparent evaluation framework.
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses
This paper presents a five-stage framework integrating large language models into survey research, addressing declining response rates, sample bias, and fraudulent completions. Using 2024 Hurricane Milton survey data, the authors propose a theory-informed LLM (A-TLM) that outperforms classical imputation methods in missing-data scenarios and demonstrates manageable hallucination risk through grounded refusal.
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
This paper examines when and why self-reported psychometric measures predict the actual behavior of large language models, finding that fine-grained, behavior-specific instruments (Theory of Planned Behavior) achieve human-level coherence within a shared conversation, while broad traits like Big 5 do not.