Interpreting Learning Under Competing Models: Joint and Stepwise Approaches for Dynamic Cognitive Diagnosis

arXiv cs.LG Papers

Summary

This paper compares joint and stepwise approaches for estimating learning under competing cognitive diagnostic models, using data from reading games. It shows that the choice of approach can change conclusions about learner progress, and joint analysis is more reliable when item-skill structure is uncertain.

arXiv:2606.06804v1 Announce Type: new Abstract: Digital learning environments record learners' responses to individual items, making it possible to study the development of specific skills rather than overall scores. Drawing conclusions about learning from these data requires a model that links responses to latent skills and tracks how mastery changes over time. When the skills measured by each item are unknown, the analyst must decide whether to estimate this structure, the Q-matrix, jointly with the learning process, or to establish it first and study learning afterwards. We show that this decision can change substantive conclusions about how learners develop. Using dynamic cognitive diagnostic models, we analyse data from two reading games measuring vocabulary and comprehension from Grade 2 to Grade 3, with item-text embeddings providing prior information for the unknown Q-matrix. A joint analysis and a bias-corrected stepwise analysis agree that most learners move toward mastering both skills, but disagree about how many remain only partially proficient at Grade 3, changing how reading progress would be reported. A simulation study identifies when the two analyses diverge and shows that joint analysis is more reliable when the item-skill structure is uncertain and the item pool changes between grades. We provide R code for both analyses.
Original Article
View Cached Full Text

Cached at: 06/08/26, 09:18 AM

# Interpreting Learning Under Competing Models: Joint and Stepwise Approaches for Dynamic Cognitive Diagnosis
Source: [https://arxiv.org/html/2606.06804](https://arxiv.org/html/2606.06804)
Yawen MaSchool of Mathematical Sciences, Lancaster University, Lancaster, LA1 4YF, Lancashire, United KingdomKate CainDepartment of Psychology, Lancaster University, Lancaster, LA1 4YF, Lancashire, United KingdomGabriel WallinCorresponding author\. School of Mathematical Sciences, Lancaster University, Lancaster, LA1 4YF, United Kingdom\.g\.wallin@lancaster\.ac\.ukSchool of Mathematical Sciences, Lancaster University, Lancaster, LA1 4YF, Lancashire, United Kingdom

###### Abstract

Digital learning environments record learners’ responses to individual items, making it possible to study the development of specific skills rather than overall scores\. Drawing conclusions about learning from these data requires a model that links responses to latent skills and tracks how mastery changes over time\. When the skills measured by each item are unknown, the analyst must decide whether to estimate this structure, the Q\-matrix, jointly with the learning process, or to establish it first and study learning afterwards\. We show that this decision can change substantive conclusions about how learners develop\. Using dynamic cognitive diagnostic models, we analyse data from two reading games measuring vocabulary and comprehension from Grade 2 to Grade 3, with item\-text embeddings providing prior information for the unknown Q\-matrix\. A joint analysis and a bias\-corrected stepwise analysis agree that most learners move toward mastering both skills, but disagree about how many remain only partially proficient at Grade 3, changing how reading progress would be reported\. A simulation study identifies when the two analyses diverge and shows that joint analysis is more reliable when the item\-skill structure is uncertain and the item pool changes between grades\. We provide R code for both analyses\.

Keywords:Cognitive Diagnostic Models; Dynamic Cognitive Diagnosis;QQ\-matrix Estimation; Bayesian Inference; Learning Transitions; Natural Language Processing\.

## 1Introduction

A single test score says little about which skills a learner has and has not mastered\. Learning involves changes in how component skills are acquired and coordinated over time, and a total score cannot reveal them\. Reading is a good example\. Skilled reading depends on more than accurate and fluent word recognition; it also requires listening comprehension\(larrc2015learning;gough1986decoding\), which itself is based on vocabulary, sentence processing, and the identification of ideas within prose\(larrc2017pressure;oakhill2012precursors\)\. Digital learning environments make it possible to observe these component skills at the level of individual items\. As learners interact with an educational tool, the tool produces log files that record not only whether each item was answered correctly, but also timestamps, actions, response times, repeated attempts, and early exit behaviour\. These records allow researchers to move beyond aggregated performance and to study how a learner’s knowledge develops through repeated interaction with the measurement instrument\.

This richness creates an interpretive problem: how should these records be translated into statements about latent skills and their development? Because learners are observed repeatedly as they progress, the data describe a mastery profile that changes over time rather than a fixed one, and how that change is described depends on the statistical model\. The model does more than determine estimation accuracy; it shapes the substantive account of what was learned, by whom, and when\. Cognitive diagnostic models \(CDMs\) are well suited to this task, because they classify learners by mastery or non\-mastery of several latent attributes rather than placing them on a single scale\(haertel1984application;junker2001cognitive;templin2010diagnostic\)\. Dynamic extensions of CDMs allow mastery to change across occasions, which makes them appropriate for longitudinal digital data\(wang2018tracking;wang2020development;zhan2018cognitive;zhan2019using;liang2023latent\)\. Applying them to real data, however, forces several decisions: how to use the response data, how to specify or estimate the item\-attribute structure \(theQQ\-matrix\), how to incorporate external information, and whether learning transitions are estimated jointly with the measurement model or in separate steps\.

This article treats the last of these decisions as a substantive question rather than a technical one\. Using data from two digital reading games, we ask whether estimating the measurement model and the learning process jointly, or in separate steps, leads to different conclusions about learners’ mastery, their transitions between mastery states, and the variables associated with learning\. A joint analysis estimates the measurement model, theQQ\-matrix, the mastery trajectories, and the transition model together\. A bias\-corrected stepwise analysis estimates the measurement model first and then fits the transition model from the resulting classifications, with a correction for classification error\. We find that the two analyses agree on the broad direction of learning but differ in how confidently learners are assigned to full rather than partial mastery, and that this difference is largest when the measurement structure is uncertain, as in our data\.

Because the games were not designed around a predefined item\-attribute structure, the attributes each item measures are not known in advance, and theQQ\-matrix must be estimated\. We use item text to inform this estimation: sentence embeddings of the item content provide prior information about the item\-attribute structure, which the response data can then revise\. Text\-informed priors of this kind are a recent addition to cognitive diagnosis rather than a standard component\.ma2026nlpintroduced them within a jointly estimated CDM\. Implementing them within a bias\-corrected stepwise procedure, as we do here, has not to our knowledge been done before, and it is what allows the comparison to be fair: the joint and stepwise analyses then draw on the same item text, response data, and covariates, so any difference between them reflects the estimation strategy and not the information available to each model\.

The remainder of the article is organised as follows\. We first review the use of response data for modelling learning, including dynamic extensions and the role of theQQ\-matrix, and describe how item text can inform the measurement structure when theQQ\-matrix is unknown\. We then discuss how log\-derived covariates and learner characteristics support interpretation, and set out the joint and stepwise strategies\. We apply the framework to the reading\-game data and examine how the modelling choice affects conclusions about learning, and we use simulations to establish when the two strategies recover theQQ\-matrix, the mastery profiles, and the transition parameters\. The article closes with the implications and limitations of the approach\.

### 1\.1Using Response Data to Model Learning Processes

Response data have long been the primary source of evidence in educational and psychological measurement\. In traditional applications, item responses are often summarized by total scores or modeled using item response theory to estimate a general proficiency level\. Such approaches are useful when the goal is to locate students on a common ability scale, but they provide limited information about the specific skills that students have or have not mastered\. This limitation is important in learning environments where the goal is not only to evaluate performance but also to provide diagnostic information that can guide feedback, instruction, and intervention\.

CDMs address this problem by linking observed responses to a vector of latent attribute mastery indicators\(haertel1984application;junker2001cognitive;templin2010diagnostic\)\. Instead of representing students using a single continuous trait, CDMs classify students into mastery profiles that describe their status on multiple skills\. This makes CDMs especially useful for educational applications where interpretable information about component skills is needed\. For reading data, such a framework is appealing because reading development involves multiple related but distinct skills, and students may show different patterns of mastery across decoding, fluency, vocabulary, and other comprehension\-related attributes\.

When data are collected repeatedly over time, static CDMs are insufficient because they do not directly model changes in mastery\. Dynamic CDMs and related longitudinal variants extend the CDM framework by allowing latent attribute profiles to evolve across time\(wang2018tracking;wang2020development;zhan2018cognitive;zhan2019using;zhan2020partial;liang2023latent\)\. These models provide a way to study learning transitions rather than only cross\-sectional mastery status\. In digital learning environments, this dynamic perspective is crucial because students interact with tasks repeatedly and may acquire skills during the period of observation\. However, the interpretation of such learning transitions depends on the measurement structure, the available external information, and the estimation strategy used to connect observed responses to latent mastery states\.

### 1\.2Using Item Text to Inform the Measurement

A central component of CDMs is theQQ\-matrix, which specifies which latent attributes are required by each item\. TheQQ\-matrix determines the meaning of the latent attributes and directly affects item parameter estimation, mastery profiles, and conclusions about student learning\. Misspecification of theQQ\-matrix can therefore lead to biased inferences about both items and learners\(rupp2008effects;chen2015statistical\)\. This issue becomes even more consequential in dynamic CDMs, because uncertainty in the measurement structure can influence the interpretation of learning transitions\.

In many real\-world digital learning environments, theQQ\-matrix is not fully known, and different domain experts may provide inconsistent specifications\. As a result, the relation between item content and latent diagnostic attributes may remain uncertain\. A growing body of research has therefore treated theQQ\-matrix as an object of inference rather than as a fixed input\. Bayesian and data\-driven approaches have been proposed to estimate, validate, or revise theQQ\-matrix using response data\(chen2018bayesian;culpepper2016revisiting;gu2021sufficient;fang2019identifiability\)\. These developments are important because they recognize that measurement structure is itself uncertain in many applied settings\.

At the same time, response data alone may not always identify the item\-attribute structure with sufficient certainty\. This problem is especially likely in short assessments, sparse adaptive trajectories, or settings in which items are heterogeneous in content\. Digital learning environments often provide additional information that can help reduce this uncertainty\. In particular, item text and response options contain semantic information about what an item asks students to do\. Recent developments in natural language processing make it possible to represent such text using embedding\-based methods that capture semantic relationships among words, sentences, or item components\(vaswani2017attention;devlin2019bert;reimers2019sentencebert\)\. For CDMs, text\-derived information need not determine theQQ\-matrix directly\. Instead, text\-derived information serves as structured prior information, allowing item content to inform the estimation of plausible item\-attribute relations\.

Among embedding\-based approaches, Sentence\-BERT \(SBERT\) is particularly useful for representing assessment text because it produces sentence\-level embeddings that can be efficiently compared using similarity measures such as cosine similarity or Euclidean distance\(reimers2019sentencebert\)\. SBERT extends the BERT network to generate semantically meaningful vector representations, allowing semantically similar texts to be located close to one another in a shared embedding space\. Compared with standard BERT representations, SBERT substantially improves computational efficiency for large\-scale semantic similarity and clustering tasks while maintaining strong performance on semantic textual similarity benchmarks\(reimers2019sentencebert\)\. In the context of educational assessment, these properties make SBERT especially suitable for quantifying semantic relationships among items and response options\. Such semantic similarities may provide useful information about the complexity and potential attribute requirements of items, which can then be incorporated into the prior structure of theQQ\-matrix estimation process\. Recent work by\(ma2026nlp\)used SBERT\-derived information to construct informative priors forQQ\-matrix estimation and demonstrated its effectiveness on a different dataset\.

### 1\.3Incorporating External Information for Interpreting Learning Processes

Digital learning data contain more than item responses\. They may also include log\-based summaries, response times, number of attempts, success frequencies, early exit behaviors, learner demographics, and prior achievement\. These sources of information are useful because they describe aspects of the learning process that are not fully captured by binary correctness\. For example, two students may produce the same response pattern but differ substantially in time spent, persistence, number of attempts, or prior literacy ability\. Such differences may affect how learning trajectories should be interpreted\.

The present study analysed log files to understand student learning on two games in a supplementary digital reading support\. In addition, an out\-of\-game assessment, the Dynamic Indicators of Basic Early Literacy Skills\(universityoforegon2018dibels\), was administered to measure each student’s initial literacy ability\. Students were classified into four performance levels: well below benchmark, below benchmark, at benchmark, and above benchmark\. These categories were used to place students at an appropriate starting level in the app\. Additional categorical covariates included race, special educational needs \(SEN\), English language learner status \(ELL\), and gender\. Continuous log\-based covariates included the average number of attempts, the number of correctly answered questions, and average response time\. Descriptive statistics for the continuous and categorical covariates are summarized in Tables[2](https://arxiv.org/html/2606.06804#S3.T2)and[1](https://arxiv.org/html/2606.06804#S3.T1), respectively\.

Including such information can improve the substantive interpretation of learning trajectories by connecting latent changes to observable features of students and their learning environments\. However, incorporating external information also creates methodological challenges\. If covariates are measured with error, unevenly observed, or strongly related to the adaptive item selection process, then their inclusion may affect both estimation and interpretation\. For this reason, external information should be treated not merely as additional predictors, but as part of an interpretive framework\. In applied educational settings, researchers often want to know not only whether students learned, but under what conditions learning occurred and whether different groups of students followed different developmental patterns\. Dynamic CDMs provide a structured way to address these questions, but the conclusions depend on how response data, item information, and external covariates are combined\.

### 1\.4Joint and Stepwise Strategies for Dynamic CDMs

A final modeling decision concerns whether the measurement and structural components of a dynamic CDM should be estimated jointly or in separate steps\. Stepwise approaches have a long history in latent class, latent transition, latent profile, growth mixture, and related finite mixture models\. In these approaches, researchers first estimate a measurement model and then relate estimated latent classes or profiles to covariates, transitions, or distal outcomes\(Bakk2013;ClarkMuthen2009;Asparouhov2014;Bakk2018;Bolck2004;Vermunt2010\)\. The main advantage of stepwise methods is flexibility\. They allow researchers to establish a measurement model before adding structural components and to modify covariate or transition models without refitting the full model\.

However, simple stepwise procedures may treat assigned latent classes as if they were observed without error\. This can bias structural parameter estimates when classification uncertainty is ignored\(Croon2002\)\. Bias\-adjusted three\-step methods address this problem by incorporating information about classification error into subsequent structural analyses\(Asparouhov2014;Bakk2013;Bakk2018;Bolck2004;Vermunt2010\)\. These developments are highly relevant for dynamic CDMs because learning trajectories are usually inferred from uncertain latent mastery classifications rather than directly observed\.

Joint approaches represent a different strategy\. They estimate measurement parameters, latent mastery profiles, transition parameters, and covariate effects within a single model\. By accounting for uncertainty across model components, joint estimation may provide more stable inference when tests are short, sample sizes are modest, latent states are difficult to distinguish, or theQQ\-matrix is uncertain\(ma2026dynamic\)\. This feature is especially relevant for transition parameters, which are often difficult to estimate accurately when transition models are fitted after latent profile classification\(liang2023latent\)\. Related one\-step approaches have been used in latent variable modeling and cognitive diagnosis to estimate measurement and structural components simultaneously\(DelaTorre2004;Ayers2013;Park2014;wang2019joint\)\. However, joint estimation can also be less convenient when researchers wish to modify parts of the structural model or compare alternative measurement specifications\.

The distinction between joint and stepwise strategies is therefore not only computational, but also interpretive\. A stepwise analysis asks how learning conclusions change after a measurement model has first been established and then linked to transitions and external variables\. A joint analysis asks which learning trajectories are most coherent with the measurement model, transition model, covariates, and observed responses simultaneously\. These two inferential pathways can agree when the measurement model is strong and classification uncertainty is limited, but they may diverge when theQQ\-matrix is unknown, item information is sparse, or latent states are difficult to distinguish\. This article studies that difference as an applied measurement problem\.

## 2Methods

### 2\.1Cognitive Diagnostic Measurement Model

Cognitive diagnostic models \(CDMs\) are widely used in educational, psychological, and behavioral measurement to analyze item response data and provide fine\-grained classifications of individuals’ mastery or non\-mastery of discrete latent attributes\(huff2007diagnostic;zenisky2012developing;roussos2007cognitive\)\. In this paper, an attribute refers to a specific skill or ability required for answering items correctly\. CDMs provide a framework for identifying which attributes are likely to have been mastered by each individual and examining which attributes are required by each item\. In this framework, the relationship between items and attributes is commonly represented by aQQ\-matrix, whose entries indicate whether a given item requires a given attribute\(tatsuoka1983rule;de2009dina\)\. Different CDMs make different assumptions about how required attributes combine to produce a correct response\. Non\-compensatory models assume that all required attributes must be mastered, whereas compensatory models allow mastery of some required attributes to compensate for non\-mastery of others\. More general models, such as the generalized DINA model, relax the assumptions of the deterministic inputs, noisy “and” gate \(DINA\) and deterministic inputs, noisy “or” gate \(DINO\) models by allowing the effects of required attributes to vary across items\(de2011generalized\)\.

In this paper, the response of studentiito itemjjat timettis denoted byYi​j​t∈\{0,1\}Y\_\{ijt\}\\in\\\{0,1\\\}, wherei=1,…,Ni=1,\\ldots,Ndenotes students,t=1,…,Tt=1,\\ldots,Tdenotes time points,j=1,…,Jtj=1,\\ldots,J\_\{t\}denotes items administered at timett, andk=1,…,Kk=1,\\ldots,Kdenotes latent attributes\. The latent mastery profile of studentiiat timettis

𝜶i​t=\(αi​1​t,…,αi​K​t\)⊤,αi​k​t∈\{0,1\},\\bm\{\\alpha\}\_\{it\}=\(\\alpha\_\{i1t\},\\ldots,\\alpha\_\{iKt\}\)^\{\\top\},\\qquad\\alpha\_\{ikt\}\\in\\\{0,1\\\},
whereαi​k​t=1\\alpha\_\{ikt\}=1indicates that studentiihas mastered attributekkat timett\. The item\-attribute relationship at timettis represented by aJt×KJ\_\{t\}\\times KQQ\-matrix,

Qt=\(qj​k​t\),qj​k​t∈\{0,1\}\.Q\_\{t\}=\(q\_\{jkt\}\),\\qquad q\_\{jkt\}\\in\\\{0,1\\\}\.
The entryqj​k​t=1q\_\{jkt\}=1indicates that itemjjat timettrequires attributekk\.

A commonly used non\-compensatory CDM is the DINA model\(junker2001cognitive;de2009dina\)\. Under the DINA model, the ideal response indicator is

ηi​j​t\\displaystyle\\eta\_\{ijt\}=∏k=1Kαi​k​tqj​k​t\.\\displaystyle=\\prod\_\{k=1\}^\{K\}\\alpha\_\{ikt\}^\{q\_\{jkt\}\}\.\(2\.1\)
Thus,ηi​j​t=1\\eta\_\{ijt\}=1only when studentiihas mastered all attributes required by itemjjat timett\. Letgj​tg\_\{jt\}denote the guessing parameter andsj​ts\_\{jt\}denote the slipping parameter\. The guessing parameter represents the probability of a correct response when the ideal response is incorrect, and the slipping parameter represents the probability of an incorrect response when the ideal response is correct\. The item response probability is

P​\(Yi​j​t=1∣𝜶i​t,Qt,gj​t,sj​t\)\\displaystyle P\(Y\_\{ijt\}=1\\mid\\bm\{\\alpha\}\_\{it\},Q\_\{t\},g\_\{jt\},s\_\{jt\}\)=\(1−sj​t\)ηi​j​t​gj​t1−ηi​j​t\.\\displaystyle=\(1\-s\_\{jt\}\)^\{\\eta\_\{ijt\}\}g\_\{jt\}^\{1\-\\eta\_\{ijt\}\}\.\(2\.2\)
Although the proposed Bayesian framework can be implemented with different CDM measurement components, the empirical and simulation analyses in this paper use the DINA model\. This choice is made for three reasons\. First, the DINA model has a simple non\-compensatory structure, which is straightforward to interpret when an item is assumed to require mastery of all specified attributes\. Second, its guessing and slipping parameters provide a computationally efficient measurement component for joint estimation with an unknownQQ\-matrix and dynamic latent mastery trajectories\. Third, using the DINA model maintains consistency with previous applications of this modeling framework\(ma2026dynamic\)\. For simplicity, however, without loss of generality, we present the proposed approach under the DINA measurement model111An R package has been developed to allow switching between different measurement models and will be made available upon acceptance of the manuscript\.\.

Under the DINA measurement component, the likelihood contribution at timettis

p​\(Yt∣𝜶t,Qt,𝐠t,𝐬t\)\\displaystyle p\(Y\_\{t\}\\mid\\bm\{\\alpha\}\_\{t\},Q\_\{t\},\\mathbf\{g\}\_\{t\},\\mathbf\{s\}\_\{t\}\)=∏i=1N∏j=1Jt\{\(1−sj​t\)ηi​j​t​gj​t1−ηi​j​t\}Yi​j​t​\{sj​tηi​j​t​\(1−gj​t\)1−ηi​j​t\}1−Yi​j​t\.\\displaystyle=\\prod\_\{i=1\}^\{N\}\\prod\_\{j=1\}^\{J\_\{t\}\}\\left\\\{\(1\-s\_\{jt\}\)^\{\\eta\_\{ijt\}\}g\_\{jt\}^\{1\-\\eta\_\{ijt\}\}\\right\\\}^\{Y\_\{ijt\}\}\\left\\\{s\_\{jt\}^\{\\eta\_\{ijt\}\}\(1\-g\_\{jt\}\)^\{1\-\\eta\_\{ijt\}\}\\right\\\}^\{1\-Y\_\{ijt\}\}\.\(2\.3\)

### 2\.2UnknownQQ\-matrix and Text\-informed Prior

In practice, such as in educational technology applications, theQQ\-matrix is not known with certainty\. An expert\-specified item\-attribute structure may not always be available, and different experts may disagree about which attributes are required for a given item\. To address this uncertainty, we propose estimatingQtQ\_\{t\}from the response data while incorporating item\-level text information as prior information\.

Letτj​t\\tau\_\{jt\}denote a standardized text\-derived signal for itemjjat timett\. Theτj​t\\tau\_\{jt\}is standardized before entering the prior forQtQ\_\{t\}\. Letπj​k​t∈\(0,1\)\\pi\_\{jkt\}\\in\(0,1\)denote the prior probability that itemjjat timettrequires attributekk, that is, the prior inclusion probability thatqj​k​t=1q\_\{jkt\}=1\. We model this probability as:

logit​\(πj​k​t\)\\displaystyle\\mathrm\{logit\}\(\\pi\_\{jkt\}\)=logit​\(θ\)−λ​τj​t,\\displaystyle=\\mathrm\{logit\}\(\\theta\)\-\\lambda\\tau\_\{jt\},\(2\.4\)qj​k​t∣θ,λ,τj​t\\displaystyle q\_\{jkt\}\\mid\\theta,\\lambda,\\tau\_\{jt\}∼Bernoulli​\(πj​k​t\)\.\\displaystyle\\sim\\mathrm\{Bernoulli\}\(\\pi\_\{jkt\}\)\.\(2\.5\)
Here,θ∈\(0,1\)\\theta\\in\(0,1\)controls the overall sparsity of theQQ\-matrix, andλ∈ℝ\\lambda\\in\\mathbb\{R\}controls how strongly the text\-derived signal modifies this prior inclusion probability\. Whenλ=0\\lambda=0, the prior reduces to a Bernoulli prior in which eachqj​k​tq\_\{jkt\}equals one with probabilityθ\\theta\. The negative sign reflects the working assumption that items with stronger text\-based specificity tend to require fewer attributes, although the posterior distribution can override this prior when the response data provide contrary evidence\.

In this application, the text\-derived signal is used as auxiliary information about the likely complexity of each item\. Because the available text information is summarised at the item level, the prior modifies the overall probability that an item requires an attribute but does not, by itself, determine which attribute is required\. Attribute\-specificQQ\-matrix structure is therefore informed by the combination of response data, admissibility constraints, and the dynamic model\. This choice reflects the information available in the empirical application and is intended to provide a text\-informed prior rather than a standalone NLP\-basedQQ\-matrix estimation procedure\.

To ensure identifiability, theQQ\-matrices are restricted to satisfy three conditions\. First, each item must require at least one attribute, so zero rows in theQQ\-matrix are not allowed\. Second, each attribute must be measured by at least three items at each time point\. Third, each attribute must have at least one single\-attribute item, meaning that for every attribute there is at least one item that requires that attribute and no other attributes\(gu2021sufficient\)\.

The prior distributions are specified using weakly informative ranges that reflect the scale of each parameter\. The baseline inclusion probability satisfiesθ∈\(0,1\)\\theta\\in\(0,1\)and is assigned a beta prior,

θ\\displaystyle\\theta∼Beta​\(aθ,bθ\)\.\\displaystyle\\sim\\mathrm\{Beta\}\(a\_\{\\theta\},b\_\{\\theta\}\)\.\(2\.6\)We specified aBeta​\(6,4\)\\mathrm\{Beta\}\(6,4\)prior forθ\\theta, which has a mean of0\.60\.6and a variance of approximately0\.02180\.0218\.

The text effectλ\\lambdais a real\-valued coefficient and is assigned a normal prior,

λ\\displaystyle\\lambda∼N​\(0,σλ2\)\.\\displaystyle\\sim N\(0,\\sigma\_\{\\lambda\}^\{2\}\)\.\(2\.7\)In the empirical analyses, we setσλ=0\.5\\sigma\_\{\\lambda\}=0\.5, so thatλ∼N​\(0,0\.52\)\\lambda\\sim N\(0,0\.5^\{2\}\)\.

The DINA item parameters satisfygj​t∈\(0,1\)g\_\{jt\}\\in\(0,1\)andsj​t∈\(0,1\)s\_\{jt\}\\in\(0,1\)\. They are assigned beta priors, in line with previous works\(ma2026dynamic\),

gj​t\\displaystyle g\_\{jt\}∼Beta​\(1,1\),\\displaystyle\\sim\\mathrm\{Beta\}\(1,1\),\(2\.8\)sj​t\\displaystyle s\_\{jt\}∼Beta​\(1,1\)\.\\displaystyle\\sim\\mathrm\{Beta\}\(1,1\)\.\(2\.9\)

### 2\.3Dynamic Structural Model

The structural component describes students’ initial mastery and subsequent learning transitions over time\. LetZi​0Z\_\{i0\}denote covariates measured before the first assessment occasion, and letZi,t−1Z\_\{i,t\-1\}denote covariates available before the transition from timet−1t\-1to timett\. These covariates may include student\-level background variables, initial language ability measures, or time\-varying learning process variables\.

For the initial time point, mastery of attributekkis modeled using a logistic regression:

logit​\{P​\(αi​k​1=1∣Zi​0\)\}\\displaystyle\\mathrm\{logit\}\\\{P\(\\alpha\_\{ik1\}=1\\mid Z\_\{i0\}\)\\\}=β0​k\+Zi​0⊤​𝜷k\.\\displaystyle=\\beta\_\{0k\}\+Z\_\{i0\}^\{\\top\}\\bm\{\\beta\}\_\{k\}\.\(2\.10\)Here,β0​k\\beta\_\{0k\}denotes the intercept for initial mastery of attributekk, and𝜷k\\bm\{\\beta\}\_\{k\}denotes the corresponding regression coefficients\.

For transitions between time points, the main parameter of interest is the acquisition probability from non\-mastery to mastery\. Fort=2,…,Tt=2,\\ldots,T, this probability is modeled as

logit\{P\(αi​k​t=1∣αi​k,t−1=0,Zi,t−1\)\}\\displaystyle\\mathrm\{logit\}\\\{P\(\\alpha\_\{ikt\}=1\\mid\\alpha\_\{ik,t\-1\}=0,Z\_\{i,t\-1\}\)\\\}=γ01,k,0\+Zi,t−1⊤​𝜸01,k\.\\displaystyle=\\gamma\_\{01,k,0\}\+Z\_\{i,t\-1\}^\{\\top\}\\bm\{\\gamma\}\_\{01,k\}\.\(2\.11\)
Here,γ01,k,0\\gamma\_\{01,k,0\}denotes the intercept for acquiring attributekk, and𝜸01,k\\bm\{\\gamma\}\_\{01,k\}denotes the effect of covariates on the acquisition probability\. When loss of mastery is allowed, the transition from mastery to non\-mastery is modeled as

logit\{P\(αi​k​t=0∣αi​k,t−1=1,Zi,t−1\)\}\\displaystyle\\mathrm\{logit\}\\\{P\(\\alpha\_\{ikt\}=0\\mid\\alpha\_\{ik,t\-1\}=1,Z\_\{i,t\-1\}\)\\\}=γ10,k,0\+Zi,t−1⊤​𝜸10,k\.\\displaystyle=\\gamma\_\{10,k,0\}\+Z\_\{i,t\-1\}^\{\\top\}\\bm\{\\gamma\}\_\{10,k\}\.\(2\.12\)
The empirical and simulation analyses focus primarily on𝜸01,k\\bm\{\\gamma\}\_\{01,k\}because these parameters directly describe learning, which describes the probability of acquiring an attribute among students who had not previously mastered it\. The loss parameters𝜸10,k\\bm\{\\gamma\}\_\{10,k\}are included when the application allows mastery status to decline over time\.

The acquisition and loss probabilities for one attribute and two time points can be written as

P\(αi​k​2=1∣αi​k​1=0,Zi​1\)\\displaystyle P\(\\alpha\_\{ik2\}=1\\mid\\alpha\_\{ik1\}=0,Z\_\{i1\}\)=exp⁡\(γ01,k,0\+Zi​1⊤​𝜸01,k\)1\+exp⁡\(γ01,k,0\+Zi​1⊤​𝜸01,k\),\\displaystyle=\\frac\{\\exp\(\\gamma\_\{01,k,0\}\+Z\_\{i1\}^\{\\top\}\\bm\{\\gamma\}\_\{01,k\}\)\}\{1\+\\exp\(\\gamma\_\{01,k,0\}\+Z\_\{i1\}^\{\\top\}\\bm\{\\gamma\}\_\{01,k\}\)\},\(2\.13\)P\(αi​k​2=0∣αi​k​1=1,Zi​1\)\\displaystyle P\(\\alpha\_\{ik2\}=0\\mid\\alpha\_\{ik1\}=1,Z\_\{i1\}\)=exp⁡\(γ10,k,0\+Zi​1⊤​𝜸10,k\)1\+exp⁡\(γ10,k,0\+Zi​1⊤​𝜸10,k\)\.\\displaystyle=\\frac\{\\exp\(\\gamma\_\{10,k,0\}\+Z\_\{i1\}^\{\\top\}\\bm\{\\gamma\}\_\{10,k\}\)\}\{1\+\\exp\(\\gamma\_\{10,k,0\}\+Z\_\{i1\}^\{\\top\}\\bm\{\\gamma\}\_\{10,k\}\)\}\.\(2\.14\)
The regression parameters are assigned normal priors\. Continuous covariates were standardized before model fitting\. In the joint empirical analysis, the initial mastery parameters and transition parameters were assigned as:

β0​k,𝜷k\\displaystyle\\beta\_\{0k\},\\bm\{\\beta\}\_\{k\}∼N​\(0,1\),\\displaystyle\\sim N\(0,1\),\(2\.15\)γ01,k,0,𝜸01,k,γ10,k,0,𝜸10,k\\displaystyle\\gamma\_\{01,k,0\},\\bm\{\\gamma\}\_\{01,k\},\\gamma\_\{10,k,0\},\\bm\{\\gamma\}\_\{10,k\}∼N​\(0,1\)\.\\displaystyle\\sim N\(0,1\)\.\(2\.16\)For the loss transition, the intercept was assignedγ10,k,0∼N​\(−2,1\)\\gamma\_\{10,k,0\}\\sim N\(\-2,1\), while the covariate coefficients𝜸10,k\\bm\{\\gamma\}\_\{10,k\}were assignedN​\(0,1\)N\(0,1\)\.

### 2\.4Joint Bayesian Estimation

The joint Bayesian procedure estimates the measurement model, the unknownQQ\-matrix, the latent mastery trajectories, item parameters, and structural regression parameters simultaneously\. LetY=\{Y1,…,YT\}Y=\\\{Y\_\{1\},\\ldots,Y\_\{T\}\\\}denote the observed item responses across all time points, and letZZdenote the collection of covariates used in the initial mastery and transition models\. Let𝝉=\{𝝉1,…,𝝉T\}\\bm\{\\tau\}=\\\{\\bm\{\\tau\}\_\{1\},\\ldots,\\bm\{\\tau\}\_\{T\}\\\}denote the item\-level text\-derived signals, where𝝉t=\(τ1​t,…,τJt​t\)⊤\\bm\{\\tau\}\_\{t\}=\(\\tau\_\{1t\},\\ldots,\\tau\_\{J\_\{t\}t\}\)^\{\\top\}at timett\. The joint posterior distribution is proportional to

p​\(Q1:T,𝜶1:T,𝐠1:T,𝐬1:T,𝜷,𝜸,θ,λ∣Y,Z,𝝉\)\\displaystyle p\(Q\_\{1:T\},\\bm\{\\alpha\}\_\{1:T\},\\mathbf\{g\}\_\{1:T\},\\mathbf\{s\}\_\{1:T\},\\bm\{\\beta\},\\bm\{\\gamma\},\\theta,\\lambda\\mid Y,Z,\\bm\{\\tau\}\)∝∏t=1Tp​\(Yt∣𝜶t,Qt,𝐠t,𝐬t\)​p​\(Qt∣θ,λ,𝝉t\)\\displaystyle\\qquad\\propto\\prod\_\{t=1\}^\{T\}p\(Y\_\{t\}\\mid\\bm\{\\alpha\}\_\{t\},Q\_\{t\},\\mathbf\{g\}\_\{t\},\\mathbf\{s\}\_\{t\}\)p\(Q\_\{t\}\\mid\\theta,\\lambda,\\bm\{\\tau\}\_\{t\}\)×∏i=1Np\(𝜶i​1∣Zi​0,𝜷\)∏t=2Tp\(𝜶i​t∣𝜶i,t−1,Zi,t−1,𝜸\)\\displaystyle\\hskip 18\.49988pt\\times\\prod\_\{i=1\}^\{N\}p\(\\bm\{\\alpha\}\_\{i1\}\\mid Z\_\{i0\},\\bm\{\\beta\}\)\\prod\_\{t=2\}^\{T\}p\(\\bm\{\\alpha\}\_\{it\}\\mid\\bm\{\\alpha\}\_\{i,t\-1\},Z\_\{i,t\-1\},\\bm\{\\gamma\}\)×p​\(𝐠1:T\)​p​\(𝐬1:T\)​p​\(𝜷\)​p​\(𝜸\)​p​\(θ\)​p​\(λ\)\.\\displaystyle\\hskip 18\.49988pt\\times p\(\\mathbf\{g\}\_\{1:T\}\)p\(\\mathbf\{s\}\_\{1:T\}\)p\(\\bm\{\\beta\}\)p\(\\bm\{\\gamma\}\)p\(\\theta\)p\(\\lambda\)\.\(2\.17\)
The first product contains the measurement likelihood and the text\-informed prior for the unknownQQ\-matrix at each time point\. The measurement likelihood is defined by the DINA response model, whereasp​\(Qt∣θ,λ,𝝉t\)p\(Q\_\{t\}\\mid\\theta,\\lambda,\\bm\{\\tau\}\_\{t\}\)denotes the constrained text\-informed prior described above\. The second line describes the structural model for initial mastery and subsequent transitions\. Specifically,p​\(𝜶i​1∣Zi​0,𝜷\)p\(\\bm\{\\alpha\}\_\{i1\}\\mid Z\_\{i0\},\\bm\{\\beta\}\)gives the initial mastery distribution, andp​\(𝜶i​t∣𝜶i,t−1,Zi,t−1,𝜸\)p\(\\bm\{\\alpha\}\_\{it\}\\mid\\bm\{\\alpha\}\_\{i,t\-1\},Z\_\{i,t\-1\},\\bm\{\\gamma\}\)gives the transition distribution from timet−1t\-1to timett\.

Posterior inference is conducted using Markov chain Monte Carlo sampling\. At each iteration, the algorithm updates the latent mastery profiles, item parameters, structural regression parameters, text\-informed prior parameters, and admissibleQQ\-matrices\. The admissibility restrictions onQtQ\_\{t\}are enforced throughout sampling, so posterior draws of theQQ\-matrix always satisfy the structural conditions described above\. Posterior summaries, including posterior means, credible intervals, and classification probabilities, are used to evaluate item\-attribute relationships, student mastery trajectories, and covariate effects on learning transitions\.

### 2\.5Bias\-Corrected Stepwise Estimation

The stepwise strategy separates measurement estimation from structural estimation\. The first step estimates the measurement model separately at each time point\. The second step converts posterior mastery probabilities into hard classifications and estimates classification error probabilities\. The third step fits the dynamic structural model using those hard classifications while correcting for classification error\.

#### 2\.5\.1Step 1: Measurement Model Estimation

At each time point, the following measurement posterior is estimated independently:

p​\(𝜶t,Qt,𝐠t,𝐬t,θ,λ∣Yt,𝐓t\)\\displaystyle p\(\\bm\{\\alpha\}\_\{t\},Q\_\{t\},\\mathbf\{g\}\_\{t\},\\mathbf\{s\}\_\{t\},\\theta,\\lambda\\mid Y\_\{t\},\\mathbf\{T\}\_\{t\}\)∝p​\(Yt∣𝜶t,Qt,𝐠t,𝐬t\)​p​\(Qt∣θ,λ,𝐓t\)\\displaystyle\\propto p\(Y\_\{t\}\\mid\\bm\{\\alpha\}\_\{t\},Q\_\{t\},\\mathbf\{g\}\_\{t\},\\mathbf\{s\}\_\{t\}\)p\(Q\_\{t\}\\mid\\theta,\\lambda,\\mathbf\{T\}\_\{t\}\)×p​\(𝜶t\)​p​\(𝐠t\)​p​\(𝐬t\)​p​\(θ\)​p​\(λ\)\.\\displaystyle\\qquad\\times p\(\\bm\{\\alpha\}\_\{t\}\)p\(\\mathbf\{g\}\_\{t\}\)p\(\\mathbf\{s\}\_\{t\}\)p\(\\theta\)p\(\\lambda\)\.\(2\.18\)
The first\-step measurement model is estimated independently of the longitudinal structural model\. This step can be understood as fitting a cross\-sectional cognitive diagnostic model\. A cognitive diagnostic model is a restricted latent class model in which each latent class corresponds to an attribute mastery profile\. For example, withKKbinary attributes, studentii’s mastery profile𝜶i​t\\bm\{\\alpha\}\_\{it\}belongs to one of2K2^\{K\}possible latent classes\. In a standard latent class formulation, the measurement model requires a population distribution over these latent classes, such as

P​\(𝜶i​t=c\)=πc​t,c=1,…,2K\.\\displaystyle P\(\\bm\{\\alpha\}\_\{it\}=c\)=\\pi\_\{ct\},\\qquad c=1,\\ldots,2^\{K\}\.
This latent class distribution describes how likely each mastery profile is in the student population at timett\. In the present implementation, we use an attribute\-wise version of this idea\. Instead of assigning a probability to each full mastery profile, we assign a marginal mastery probability to each attribute:

αi​k​t∣ρk​t\\displaystyle\\alpha\_\{ikt\}\\mid\\rho\_\{kt\}∼Bernoulli​\(ρk​t\),\\displaystyle\\sim\\mathrm\{Bernoulli\}\(\\rho\_\{kt\}\),\(2\.19\)ρk​t\\displaystyle\\rho\_\{kt\}∼Beta​\(aρ,bρ\)\.\\displaystyle\\sim\\mathrm\{Beta\}\(a\_\{\\rho\},b\_\{\\rho\}\)\.\(2\.20\)In the empirical stepwise analysis, we setaρ=1a\_\{\\rho\}=1andbρ=1b\_\{\\rho\}=1\. The parameterρk​t∈\(0,1\)\\rho\_\{kt\}\\in\(0,1\)represents the marginal probability that a student has mastered attributekkat timettin the first\-step measurement model\. This parameter is not interpreted as a learning, acquisition, or transition parameter\. The learning process is modeled later through the corrected structural model\.

#### 2\.5\.2Step 2: Hard Classification and Classification Error Probabilities

Let

p^i​k​t\\displaystyle\\widehat\{p\}\_\{ikt\}=P​\(αi​k​t=1∣Yt\)\\displaystyle=P\(\\alpha\_\{ikt\}=1\\mid Y\_\{t\}\)\(2\.21\)denote the posterior probability of mastery from the first\-step measurement model\. The hard assigned class is

Wi​k​t\\displaystyle W\_\{ikt\}=I​\(p^i​k​t≥0\.5\)\.\\displaystyle=I\(\\widehat\{p\}\_\{ikt\}\\geq 0\.5\)\.\(2\.22\)The assigned classWi​k​tW\_\{ikt\}is treated as an observed but error\-prone indicator of the unobserved true mastery state, denotedLi​k​tL\_\{ikt\}\.

For each attributekkand timett, the classification error probability matrix is

Mk​t​\(w,l\)\\displaystyle M\_\{kt\}\(w,l\)=P​\(Wi​k​t=w∣Li​k​t=l\),w,l∈\{0,1\}\.\\displaystyle=P\(W\_\{ikt\}=w\\mid L\_\{ikt\}=l\),\\hskip 18\.49988ptw,l\\in\\\{0,1\\\}\.\(2\.23\)Its entries are estimated from the posterior probabilities obtained in Step 1:

M^k​t​\(w,l\)\\displaystyle\\widehat\{M\}\_\{kt\}\(w,l\)=∑i=1NP​\(Li​k​t=l∣Yt\)​I​\(Wi​k​t=w\)∑i=1NP​\(Li​k​t=l∣Yt\)\.\\displaystyle=\\frac\{\\sum\_\{i=1\}^\{N\}P\(L\_\{ikt\}=l\\mid Y\_\{t\}\)I\(W\_\{ikt\}=w\)\}\{\\sum\_\{i=1\}^\{N\}P\(L\_\{ikt\}=l\\mid Y\_\{t\}\)\}\.\(2\.24\)Thus, if a student is assignedWi​k​t=0W\_\{ikt\}=0, the structural model evaluates how compatible that assignment is with both possible true states throughM^k​t​\(0,0\)\\widehat\{M\}\_\{kt\}\(0,0\)andM^k​t​\(0,1\)\\widehat\{M\}\_\{kt\}\(0,1\)\. If a student is assignedWi​k​t=1W\_\{ikt\}=1, the corresponding correction probabilities areM^k​t​\(1,0\)\\widehat\{M\}\_\{kt\}\(1,0\)andM^k​t​\(1,1\)\\widehat\{M\}\_\{kt\}\(1,1\)\.

#### 2\.5\.3Step 3: Corrected Structural Model

In the third step, the structural model is fitted to the assigned classificationsWW, but the likelihood marginalizes over the possible true latent statesLL\. For two occasions and one attribute, the corrected likelihood contribution is

P​\(Wi​k​1=w1,Wi​k​2=w2∣Zi\)\\displaystyle P\(W\_\{ik1\}=w\_\{1\},W\_\{ik2\}=w\_\{2\}\\mid Z\_\{i\}\)=∑l1=01∑l2=01P​\(Li​k​1=l1∣Zi​0,𝜷k\)\\displaystyle=\\sum\_\{l\_\{1\}=0\}^\{1\}\\sum\_\{l\_\{2\}=0\}^\{1\}P\(L\_\{ik1\}=l\_\{1\}\\mid Z\_\{i0\},\\bm\{\\beta\}\_\{k\}\)×P\(Li​k​2=l2∣Li​k​1=l1,Zi​1,𝜸k\)M^k​1\(w1,l1\)M^k​2\(w2,l2\)\.\\displaystyle\\qquad\\times P\(L\_\{ik2\}=l\_\{2\}\\mid L\_\{ik1\}=l\_\{1\},Z\_\{i1\},\\bm\{\\gamma\}\_\{k\}\)\\widehat\{M\}\_\{k1\}\(w\_\{1\},l\_\{1\}\)\\widehat\{M\}\_\{k2\}\(w\_\{2\},l\_\{2\}\)\.\(2\.25\)For example, when the observed assigned trajectory is\(Wi​k​1,Wi​k​2\)=\(1,1\)\(W\_\{ik1\},W\_\{ik2\}\)=\(1,1\), the likelihood is the sum of four possible true trajectories:

P​\(Wi​k​1=1,Wi​k​2=1∣Zi\)\\displaystyle P\(W\_\{ik1\}=1,W\_\{ik2\}=1\\mid Z\_\{i\}\)=P\(Li​k​1=0∣Zi​0\)P\(Li​k​2=0∣Li​k​1=0,Zi​1\)M^k​1\(1,0\)M^k​2\(1,0\)\\displaystyle=P\(L\_\{ik1\}=0\\mid Z\_\{i0\}\)P\(L\_\{ik2\}=0\\mid L\_\{ik1\}=0,Z\_\{i1\}\)\\widehat\{M\}\_\{k1\}\(1,0\)\\widehat\{M\}\_\{k2\}\(1,0\)\+P\(Li​k​1=0∣Zi​0\)P\(Li​k​2=1∣Li​k​1=0,Zi​1\)M^k​1\(1,0\)M^k​2\(1,1\)\\displaystyle\\qquad\+P\(L\_\{ik1\}=0\\mid Z\_\{i0\}\)P\(L\_\{ik2\}=1\\mid L\_\{ik1\}=0,Z\_\{i1\}\)\\widehat\{M\}\_\{k1\}\(1,0\)\\widehat\{M\}\_\{k2\}\(1,1\)\+P\(Li​k​1=1∣Zi​0\)P\(Li​k​2=0∣Li​k​1=1,Zi​1\)M^k​1\(1,1\)M^k​2\(1,0\)\\displaystyle\\qquad\+P\(L\_\{ik1\}=1\\mid Z\_\{i0\}\)P\(L\_\{ik2\}=0\\mid L\_\{ik1\}=1,Z\_\{i1\}\)\\widehat\{M\}\_\{k1\}\(1,1\)\\widehat\{M\}\_\{k2\}\(1,0\)\+P\(Li​k​1=1∣Zi​0\)P\(Li​k​2=1∣Li​k​1=1,Zi​1\)M^k​1\(1,1\)M^k​2\(1,1\)\.\\displaystyle\\qquad\+P\(L\_\{ik1\}=1\\mid Z\_\{i0\}\)P\(L\_\{ik2\}=1\\mid L\_\{ik1\}=1,Z\_\{i1\}\)\\widehat\{M\}\_\{k1\}\(1,1\)\\widehat\{M\}\_\{k2\}\(1,1\)\.\(2\.26\)For multiple time points, this correction generalizes to

P​\(Wi​k​1,…,Wi​k​T∣Zi\)\\displaystyle P\(W\_\{ik1\},\\ldots,W\_\{ikT\}\\mid Z\_\{i\}\)=∑l1=01⋯​∑lT=01P​\(Li​k​1=l1∣Zi​0,𝜷k\)\\displaystyle=\\sum\_\{l\_\{1\}=0\}^\{1\}\\cdots\\sum\_\{l\_\{T\}=0\}^\{1\}P\(L\_\{ik1\}=l\_\{1\}\\mid Z\_\{i0\},\\bm\{\\beta\}\_\{k\}\)×∏t=2TP\(Li​k​t=lt∣Li​k,t−1=lt−1,Zi,t−1,𝜸k\)∏t=1TM^k​t\(Wi​k​t,lt\)\.\\displaystyle\\qquad\\times\\prod\_\{t=2\}^\{T\}P\(L\_\{ikt\}=l\_\{t\}\\mid L\_\{ik,t\-1\}=l\_\{t\-1\},Z\_\{i,t\-1\},\\bm\{\\gamma\}\_\{k\}\)\\prod\_\{t=1\}^\{T\}\\widehat\{M\}\_\{kt\}\(W\_\{ikt\},l\_\{t\}\)\.\(2\.27\)The stepwise posterior for the structural parameters is therefore

p​\(𝜷,𝜸∣W,Z,M^\)\\displaystyle p\(\\bm\{\\beta\},\\bm\{\\gamma\}\\mid W,Z,\\widehat\{M\}\)∝∏i=1N∏k=1KP​\(Wi​k​1,…,Wi​k​T∣Zi,M^k​1,…,M^k​T\)​p​\(𝜷\)​p​\(𝜸\)\.\\displaystyle\\propto\\prod\_\{i=1\}^\{N\}\\prod\_\{k=1\}^\{K\}P\(W\_\{ik1\},\\ldots,W\_\{ikT\}\\mid Z\_\{i\},\\widehat\{M\}\_\{k1\},\\ldots,\\widehat\{M\}\_\{kT\}\)p\(\\bm\{\\beta\}\)p\(\\bm\{\\gamma\}\)\.\(2\.28\)

### 2\.6Comparison of the Two Strategies

The two approaches differ in how they condition on the measurement model\. The joint approach targets

p​\(𝜶1:T,Q1:T,𝐠1:T,𝐬1:T,𝜷,𝜸∣Y,Z,T\),\\displaystyle p\(\\bm\{\\alpha\}\_\{1:T\},Q\_\{1:T\},\\mathbf\{g\}\_\{1:T\},\\mathbf\{s\}\_\{1:T\},\\bm\{\\beta\},\\bm\{\\gamma\}\\mid Y,Z,T\),\(2\.29\)so learning parameters are estimated while accounting for posterior uncertainty in the latent states, item parameters, andQQ\-matrix\. The stepwise approach targets

p​\(𝜷,𝜸∣W,Z,M^\),\\displaystyle p\(\\bm\{\\beta\},\\bm\{\\gamma\}\\mid W,Z,\\widehat\{M\}\),\(2\.30\)whereWWandM^\\widehat\{M\}are functions of the first\-step measurement posterior\. Thus, the stepwise approach preserves the separation between measurement and structural modeling, while the joint approach estimates all components simultaneously\.

This distinction is important for interpretation\. In the joint strategy, an estimated acquisition effect reflects evidence from the full response process, the text\-informedQQ\-matrix prior, item\-level uncertainty, and the dynamic structural model at the same time\. In the stepwise strategy, an estimated acquisition effect reflects the association between covariates and the corrected assigned mastery trajectories after the measurement model has already been established\. When the measurement model is strong and classification error is small, the two approaches should provide similar conclusions\. When theQQ\-matrix is uncertain, tests are short, or classifications are unstable, the two strategies may lead to different interpretations of learning\.

#### 2\.6\.1Implementation

The stepwise procedure was implemented in three stages\. First, the DINA measurement model was estimated separately at each time point usingnimble\. TheQQ\-matrix was treated as unknown and estimated jointly with the item parameters and latent attribute indicators in the first\-step measurement model\. Second, posterior mastery probabilities from the first\-step model were used to obtain hard classifications and to estimate the classification error probability matrices\. Third, the structural transition model was fitted using a Bayesian implementation of the corrected marginal likelihood\.

This implementation preserves the bias\-corrected logic of the original stepwise approach\(liang2023latent\)\. The measurement and structural components remain separated, and classification error is corrected through the estimated classification error probability matrices\. Compared with the earlier work, the present implementation extends the stepwise procedure to the case where theQQ\-matrix is unknown\. The measurement model is estimated innimblewith theQQ\-matrix treated as unknown, and the structural parameters are estimated using Bayesian posterior sampling rather than maximum likelihood optimization\.

The joint procedure was implemented by estimating the measurement model, unknownQQ\-matrix, latent mastery trajectories, item parameters, and structural transition parameters simultaneously within a single Bayesian model\. Unlike the stepwise procedure, the joint procedure does not fix first\-step mastery classifications or rely on a separate classification error correction\. Instead, uncertainty in theQQ\-matrix, item parameters, and latent mastery states is estimated directly into posterior inference for the structural transition parameters\. The code for implementing the proposed framework is available on the Open Science Framework \(OSF\)222Code is available at:[https://osf\.io/tjqbp/overview?view\_only=b3672f24ba8f43ebb99f86f2d0ec8ef4](https://osf.io/tjqbp/overview?view_only=b3672f24ba8f43ebb99f86f2d0ec8ef4)\., and an accompanying R package is under development\.

## 3Empirical Study

The empirical study used log files provided by Amplify from the Boost Reading platform \([https://amplify\.com/programs/boost\-reading/](https://amplify.com/programs/boost-reading/)\)\. We focused on two reading\-related games:PunchlineandField Observer\. These two games were selected because they target different reading skills and provide repeated item responses that can be linked to specific attributes\.Punchlinemainly targets vocabulary knowledge, whereasField Observermainly targets comprehension skills related to key ideas and details\. These games also focus on critical comprehension\-related skills that predict reading comprehension in this age group, when word recognition skills are becoming fluent\(garcia2014decoding;oakhill2012precursors\)\.

### 3\.1Game Description and Item Selection

Punchlineis a vocabulary game based on jokes and word meanings\. The game contains 11 levels\. Each level includes several questions related to words with multiple meanings\. In the original game design, students complete two response steps\. First, they choose between two sentence options and identify the option whose meaning is most consistent with the question\. Second, after selecting the sentence, they identify the target word that carries two different meanings\. The game also includes an instructional component in which students are shown explanations of the two meanings of the target word\. For the present analysis, only the first response step was used\. That is, each analyzed response corresponded to choosing the correct option from two alternatives\. The second response step, in which students selected the word with two meanings from the chosen sentence, was not included in the scored item responses because this learning step was not consistently represented in the scoring structure used for the analysis\. Each level contained six analyzed questions, with every two questions associated with one relevant topic\. Thus, each level covered three target words, and the analytic responses reflected whether students selected the sentence option that matched the intended meaning in context\.

Field Observeris a comprehension game focused on key ideas and details\. The game contains 12 levels\. In this game, students search for hidden animals and answer comprehension questions\. The order in which animals are found is not fixed, so the order of questions may vary across students\. For each question, students are presented with a question and a set of evidence statements\. They select an answer option and then map the answer back to the corresponding evidence\. In the original game design, a correct response requires both selecting the correct answer and identifying the corresponding supporting evidence\. In the present analysis, the item was represented using the question together with the true evidence as the item stem, and the response reflected whether the student selected the correct answer\. Distractor evidence was not used in constructing the item text for the text\-informed prior\. Therefore, forField Observer, each analyzed item was defined by the combination of the question and the true supporting evidence\.

For both games, we analyzed item responses from the same cohort of students at two time points: U\.S\. Grade 2 and Grade 3\. To ensure that the same students had sufficient participation at both time points, we conducted exploratory data analysis to balance sample size and item length \([6](https://arxiv.org/html/2606.06804#S6)\)\. Based on this trade\-off, the final sample contained 1,978 students\. At each time point, the model used 12 selected items: sixPunchlineitems and sixField Observeritems\. Across the two time points, this corresponded to 24 items\. These items were used to estimate the latent mastery profiles and learning transitions in the proposed framework\.

### 3\.2Descriptive Statistics

The descriptive statistics for the categorical and continuous covariates are summarized in Tables[1](https://arxiv.org/html/2606.06804#S3.T1)and[2](https://arxiv.org/html/2606.06804#S3.T2)\. As shown in Table[1](https://arxiv.org/html/2606.06804#S3.T1), a large proportion of students were classified as “Above Benchmark” or “At Benchmark” based on their initial DIBELS benchmark level\. This pattern is partly consistent with the selected item set, which consisted of Grade 2 and Grade 3 content\. Students classified as “Below Benchmark” or “Well Below Benchmark” may have been less likely to encounter these items if the adaptive platform placed them in earlier or below\-grade content\.

The categorical covariates contained substantial not\-specified responses for several demographic variables, particularly race, SEN, and ELL\. The gender distribution included a higher proportion of male students than female students\. The sample also included students from diverse racial and ethnic backgrounds, including White, Asian, Black or African\-American, Multiracial, American Indian or Alaska Native, Native Hawaiian or Other Pacific Islander, Other, and not specified categories\.

Table[2](https://arxiv.org/html/2606.06804#S3.T2)presents summary statistics for the log\-derived behavioral variables\. On average, students used more attempts inPunchlinethan inField Observer\. Performance onPunchlineshowed greater variability, whereas students answered nearly all selectedField Observeritems correctly on average\. Average response time was shorter forPunchlinethan forField Observer, which is consistent with the comprehension\-based structure ofField Observer, where students needed to process both a question and supporting evidence\.

Table 1:Summary of Categorical Variables\.
Note: SEN = special educational needs; ELL = English language learner\. Missing indicates students without matched demographic recordsTable 2:Summary of Continuous Variables\.
Note: Response time is summarized using the transformed response\-time variable used in the analysis
### 3\.3Empirical Comparison of Joint and Stepwise Results

The empirical comparison focuses on the text\-informedQQ\-matrix prior because the main inferential question is how the joint and stepwise strategies support different interpretations of learning when the measurement structure is uncertain\. The goal of the empirical analysis is not to isolate the incremental contribution of the text\-derived prior relative to a non\-text prior\. Rather, the text\-informed prior provides a common measurement framework under which the joint and stepwise strategies can be compared using the same available item\-level information\. Both empirical analyses used three Markov chains with 10,000 burn\-in iterations and 10,000 monitored iterations\. In the joint text\-prior model, the posterior mean ofλ\\lambdawas 0\.075, with posterior standard deviation 0\.264 and a 95% credible interval of\(−0\.440,0\.561\)\(\-0\.440,0\.561\)\. In the stepwise text\-prior measurement models, the estimatedλ\\lambdawas \-0\.164 at time 1 and 0\.170 at time 2\.

Table[3](https://arxiv.org/html/2606.06804#S3.T3)places the text\-prior joint and stepwiseQQ\-matrix and item\-parameter estimates in the same table\. The two approaches agree on several item\-attribute patterns, especially for high\-performing items at Time 2, but they also produce differentQQ\-matrix rows for a number of items\. These differences are important because they show that the interpretation of the latent attributes can depend on whether measurement uncertainty is incorporated jointly through the dynamic model or resolved first in separate measurement steps\.

Table 3:Text\-prior empirical estimates of theQQ\-matrix, guessing \(gg\), and slipping \(ss\) parameters from the joint and stepwise approaches\. Stepwise estimates are from the first\-step measurement models fitted separately at each time point\.Table[4](https://arxiv.org/html/2606.06804#S3.T4)reports the profile transition summaries for the text\-prior joint and stepwise approaches\. Both approaches indicate movement toward full mastery of both attributes at time 2\. The joint model assigns 1,730 students to profile\(1,1\)\(1,1\)at time 2, whereas the stepwise classification summary assigns 1,659 students to this profile\. The stepwise summary also places more students in the time 2 profile\(0,1\)\(0,1\)than the joint model\. Because the stepwise transition table is based on hard classifications from the separate first\-step measurement models, these differences should be interpreted as differences in the empirical classification pathway rather than as direct discrepancies in a single common posterior distribution\. Substantively, both analyses suggest that most students showed evidence of broad mastery by the second time point\. The difference between the two approaches is therefore not a disagreement about whether students generally moved toward mastery, but about how confidently students should be assigned to full mastery rather than partial mastery at Grade 3\. Under the joint model, the interpretation is broad consolidation of both attributes\. Under the stepwise model, a larger subgroup remains classified as having mastered only one of the two attributes\. This distinction matters because it affects whether the empirical learning process is interpreted primarily as general movement toward full mastery or as a more heterogeneous pattern in which some students remain partially mastered at the second occasion\.

Table 4:Text\-prior transition matrices of attribute profiles from Time 1 to Time 2 under the joint and stepwise approaches\. Each cell reports count \(percentage\)\. Stepwise counts are based on first\-step hard classifications before the third\-step CEP correction is applied to structural estimation\. Profile labels: 00 = no mastery, 10 =A1A\_\{1\}only, 01 =A2A\_\{2\}only, and 11 = mastery of both attributes\.For an applied reading programme, this difference is more than statistical\. The two analyses would support different summaries of the same cohort\. Under the joint model, most children appear to have consolidated both vocabulary and comprehension by Grade 3, and a progress report based on it would describe near\-uniform mastery\. Under the stepwise model, a larger group is credited with only one of the two skills, and a report based on it would flag a sizeable minority as still developing one component\. The underlying data are identical; what differs is whether uncertainty in the measurement model is carried into the classification or resolved before it\. A programme deciding who to target for additional instruction would identify a different set of children under each analysis\.

Table[5](https://arxiv.org/html/2606.06804#S3.T5)summarizes statistically credible covariate effects for initial mastery and acquisition transitions under both approaches\. For initial mastery, both methods identify the number of correctly answeredPunchlinequestions as a strong positive predictor and the number of attempts inPunchlineas a negative predictor\. For acquisition transitions, both methods identify the number of correctly answeredField Observerquestions as a positive predictor and “Well Below Benchmark” status as a negative predictor\. The stepwise model additionally identifies negative effects for the number of correctly answeredPunchlinequestions and the number of attempts inField Observeron acquisition of comprehension skill, as well as a positive effect of “Above Benchmark” status for comprehension skill\. These results suggest that the most stable empirical signals are associated with behavioural indicators from the learning environment and initial literacy benchmark status, rather than with demographic variables\. For initial mastery, the positive effect of correctly answeredPunchlinequestions and the negative effect ofPunchlineattempts are consistent with the interpretation that students who answered more vocabulary\-related items correctly and required fewer attempts were more likely to begin the observed period with stronger mastery\. For acquisition transitions, the positive effect of correctly answeredField Observerquestions and the negative effect of “Well Below Benchmark” status indicate that comprehension\-related task performance and baseline literacy status are the clearest markers of subsequent mastery acquisition\.

The additional effects identified only by the stepwise model should be interpreted cautiously\. Because the stepwise approach first converts posterior mastery probabilities into hard classifications and then corrects for classification error, some covariate effects may reflect the specific classification pathway induced by the first\-stage measurement models\. These stepwise\-only effects may indicate that the stepwise model distinguishes more sharply between students who were already high performing and students whose response patterns provide evidence of transition into mastery\. In contrast, the joint model carries uncertainty in theQQ\-matrix, item parameters, and latent states directly into the transition model, leading to a smaller set of acquisition effects that are supported across the full dynamic measurement process\.

Table 5:Significant covariates for initial mastery \(βz\\beta\_\{z\}\) and acquisition transition probability \(γ01\\gamma\_\{01\}\) under the text\-prior joint and stepwise approaches\. Only covariates with 95% credible intervals for odds ratios excluding 1 are shown\.Note\. OR = odds ratio; RT = response time; NLM = number of correctly answered questions; NRA = number of attempts; WB = well below benchmark\.Detailed posterior odds ratios and 95% credible intervals are reported in Tables[6](https://arxiv.org/html/2606.06804#S3.T6)\-[9](https://arxiv.org/html/2606.06804#S3.T9)\. Tables[6](https://arxiv.org/html/2606.06804#S3.T6)and[7](https://arxiv.org/html/2606.06804#S3.T7)report covariate effects on initial mastery, with joint rows followed by stepwise rows\. Tables[8](https://arxiv.org/html/2606.06804#S3.T8)and[9](https://arxiv.org/html/2606.06804#S3.T9)report acquisition transition effects in the same format\. Across these detailed tables, the strongest empirical signals are concentrated in log\-derived behavioral variables and benchmark categories, while demographic effects are generally more uncertain\.

Table 6:Posterior means of odds ratios \(OR\) for initial mastery covariate effectsβz\\beta\_\{z\}under the text\-prior joint and stepwise approaches, with 95% credible intervals\. Statistically significant results are shown inbold\. Part 1 of 2\.
Note: RT = response time; NLM = number of correctly answered questions; NRA = number of attempts\.
Table 7:Posterior means of odds ratios \(OR\) for initial mastery covariate effectsβz\\beta\_\{z\}under the text\-prior joint and stepwise approaches, with 95% credible intervals\. Statistically significant results are shown inbold\. Part 2 of 2\.
Note: SEN = special educational needs; ELL = English language learner; WB = well below benchmark; BB = below benchmark; AB = above benchmark\. Race effects are relative to White students\.
Table 8:Posterior means of odds ratios \(OR\) for acquisition transition effectsγ01\\gamma\_\{01\}under the text\-prior joint and stepwise approaches, with 95% credible intervals\. Statistically significant results are shown inbold\. Part 1 of 2\.
Note: RT = response time; NLM = number of correctly answered questions; NRA = number of attempts\.
Table 9:Posterior means of odds ratios \(OR\) for acquisition transition effectsγ01\\gamma\_\{01\}under the text\-prior joint and stepwise approaches, with 95% credible intervals\. Statistically significant results are shown inbold\. Part 2 of 2\.
Note: SEN = special educational needs; ELL = English language learner; WB = well below benchmark; BB = below benchmark; AB = above benchmark\. Race effects are relative to White students\.

## 4Simulations

### 4\.1Simulation Design

The simulation study was designed to reflect the scale of the empirical application while also evaluating the model under varying levels of measurement information\. For each simulation condition, we conducted 100 independent replications\. We considered a dynamic CDM with two attributes measured across two time points\. Sample size varied across conditions asN∈\{1000,2000,4000\}N\\in\\\{1000,2000,4000\\\}, and the number of items administered at each time pointttvaried asJt∈\{6,12,24\}J\_\{t\}\\in\\\{6,12,24\\\}\. The middle condition,N=2000N=2000andJt=12J\_\{t\}=12at each time point, was chosen to approximate the empirical setting\. The trueQQ\-matrices used under each simulation condition are provided in Table[16](https://arxiv.org/html/2606.06804#S9.T16)in the Appendix\. The prior forθ\\thetawas specified asBeta​\(6,4\)\\mathrm\{Beta\}\(6,4\), with a prior mean of0\.60\.6\. The prior forλ\\lambdawas specified as𝒩​\(0,0\.52\)\\mathcal\{N\}\(0,0\.5^\{2\}\)\.

### 4\.2Simulation Results

The simulation results are summarized in Tables[10](https://arxiv.org/html/2606.06804#S4.T10)\-[13](https://arxiv.org/html/2606.06804#S4.T13)\. Each fit used three independent Markov chains, with 3,000 burn\-in iterations followed by 3,000 monitored iterations\. Standard errors in the tables were computed using 1,000 nonparametric bootstrap resamples over successful replications\.

The average running time per text\-prior fit was approximately 33\.1 to 216\.0 minutes for the joint approach and 21\.5 to 119\.4 minutes for the stepwise approach across different simulation settings\. These runs were conducted on a MacBook Pro \(13\-inch, M1, 2020\) equipped with an Apple M1 chip \(8\-core: 4 performance and 4 efficiency cores\) and 16 GB of unified memory\.

Table[10](https://arxiv.org/html/2606.06804#S4.T10)reports recovery of the time\-specificQQ\-matrices\. The joint model recovered theQQ\-matrix nearly perfectly across the successful conditions, with entry\-level ACC values equal or close to one and very small PIP errors\. The stepwise approach also recoveredQ1Q\_\{1\}well, with ACC ranging from 0\.917 to 1\.000, but recovery ofQ2Q\_\{2\}was less stable\. StepwiseQ2Q\_\{2\}ACC ranged from 0\.593 to 0\.843 across reported conditions, indicating that uncertainty in the second time point was more difficult to recover in the separate first\-stage measurement models\.

Table 10:Recovery of theQQ\-matrix under the joint and stepwise approaches\. Values are reported as mean \(bootstrap SE\)- Note\.PIP RMSE and PIP MAE summarize posterior inclusion probability accuracy\.

Table[11](https://arxiv.org/html/2606.06804#S4.T11)presents recovery of latent attribute profiles\. The joint model achieved high profile agreement rates at both time points, with PAR values ranging from 0\.926 to 0\.989 forα1\\alpha\_\{1\}and from 0\.923 to 0\.954 forα2\\alpha\_\{2\}among reported conditions\. The stepwise approach gave comparable recovery forα1\\alpha\_\{1\}in some conditions, but recovery ofα2\\alpha\_\{2\}was consistently weaker, with PAR ranging from 0\.668 to 0\.779\. These results suggest that classification error from the separate measurement steps can affect the transition stage, whereas the joint model stabilizes the two time points by estimating measurement and transition components simultaneously\.

Table 11:Recovery of attribute profiles under the joint and stepwise approaches\. Values are reported as mean \(bootstrap SE\)- Note\.PAR denotes profile agreement rate; AAR denotes attribute\-wise agreement rate\.

Table[12](https://arxiv.org/html/2606.06804#S4.T12)summarizes guessing and slipping parameter recovery\. Both approaches recovered item parameters accurately when the stepwise model converged\. Item\-parameter RMSE and MAE were small across conditions, and they generally decreased as sample size increased\. Differences between the joint and stepwise approaches were much smaller for item parameters than for transition\-related quantities\. This suggests that the largest practical distinction between the two frameworks is not in estimating the DINA measurement parameters once the measurement model is identifiable, but rather in how uncertainty about theQQ\-matrix and latent classifications is incorporated into later components of the dynamic model\.

Table 12:Estimation accuracy of item parameters under the joint and stepwise approaches\. Values are reported as mean \(bootstrap SE\)- Note\.ggandssdenote guessing and slipping parameters at Time 1 and Time 2\.

Table[13](https://arxiv.org/html/2606.06804#S4.T13)reports estimation accuracy for the initial mastery parametersβ0\\beta\_\{0\}andβZ\\beta\_\{Z\}and the acquisition transition parametersγ01\\gamma\_\{01\}\. The joint model consistently produced smaller errors forγ01\\gamma\_\{01\}than the stepwise approach in the reported conditions\. For example, underN=2000,Jt=12N=2000,J\_\{t\}=12, the jointγ01\\gamma\_\{01\}RMSE was 0\.080, compared with 0\.326 for the stepwise model\. UnderN=2000,Jt=24N=2000,J\_\{t\}=24, the corresponding RMSE values were 0\.099 and 0\.428\. Differences forβ0\\beta\_\{0\}andβZ\\beta\_\{Z\}were smaller, but the transition parameters showed a clear advantage for joint estimation\. This is consistent with the expectation that transition effects are particularly sensitive to classification uncertainty in latent mastery profiles\.

Table 13:Estimation accuracy of regression parameters under the joint and stepwise approaches\. Values are reported as mean \(bootstrap SE\)\.- Note\.β0\\beta\_\{0\}denotes initial mastery intercepts,βZ\\beta\_\{Z\}denotes initial mastery covariate effects, andγ01\\gamma\_\{01\}denotes acquisition transition parameters\.

Overall, the simulation results indicate that joint estimation is more robust when theQQ\-matrix is unknown and measurement information is limited\. The stepwise procedure can give reasonable item\-parameter estimates when the first\-step measurement model is identified by the data, but it performs worse than the joint strategy for the second time point and for transition parameters\. This pattern is consistent with the role of the joint model in carrying uncertainty aboutQQ\-matrix recovery and latent mastery classifications directly into the dynamic transition component\.

## 5Discussion

This study compared joint and stepwise estimation strategies for longitudinal CDMs when theQQ\-matrix is unknown and informed by item text information\. Across the simulation conditions, the joint approach generally produced more stable recovery of latent attribute profiles, transition parameters, and covariate effects, particularly when the number of items was limited or when classification uncertainty increased over time\. By contrast, the stepwise approach often produced reasonable recovery for measurement parameters at the first time point, but the performance is weaker than that of the joint approach for the second time point and for the longitudinal transition components\.

The distinction between the two strategies is therefore not merely computational, but also interpretive\. The joint approach treats the measurement model, latent transitions, covariate effects, and observed responses as components of a single probabilistic system that are estimated simultaneously\. In contrast, the stepwise approach separates measurement from longitudinal modelling by first estimating latent profiles and then using those identified profiles in later stages of analysis\. When measurement information is sufficiently strong, both approaches can lead to similar conclusions\. However, the two strategies may diverge because uncertainty from earlier measurement stages affects subsequent transition estimation differently\.

The results further suggest that the largest practical distinction between the two frameworks is not primarily in estimating DINA item parameters themselves, but rather in how uncertainty surrounding theQQ\-matrix and latent mastery profiles is incorporated into later components of the longitudinal model\. Under the stepwise framework, classification decisions from earlier stages may introduce instability into transition and regression parameters estimation when those classifications are uncertain\. The joint framework instead allows measurement and transition information to inform one another simultaneously, leading to more coherent estimation across time points\.

For applied reading assessment, both strategies agree that children in this cohort generally progress toward mastering vocabulary and comprehension between Grade 2 and Grade 3\. They differ in how many children are reported as having fully consolidated both skills, and that difference is largest exactly where digital reading data are weakest: short item sets, item pools that change between grades, and an unknown attribute structure\. In such settings, a programme that classifies children for targeted support should be aware that a stepwise approach can place more children in partial\-mastery categories than a joint analysis of the same data, not because more children are struggling, but because uncertainty resolved early is read as a definite gap\. Where the measurement model is strong the choice is inconsequential; where it is not, the joint analysis gives the more defensible basis for reporting and for deciding who needs further support\.

These findings also highlight the importance of carefully interpreting theQQ\-matrix in longitudinal settings\. In many educational applications, the assumption that the same attributes are measured identically over time may not fully hold because items, task formats, or instructional emphases can change across grades or school years\. In the present study, the item pools differed across time points, making the recovery of attribute structures particularly important for maintaining interpretability of longitudinal transitions\. This issue becomes increasingly relevant in digital learning environments, where item content evolves dynamically and item pools may be continuously updated\.

In the empirical analysis, item text contributed little to the estimated structure\. It entered as a prior on theQQ\-matrix, and the posterior forλ\\lambdawas close to zero, so the response data effectively determined the estimated structure\. We do not read this as evidence that item text is uninformative\. The responses in these data were strong enough to identify theQQ\-matrix on their own, leaving the prior with little to do\. The value of the text component here is that it lets the joint and stepwise strategies draw on the same item information, so the two can be compared on equal terms\.ma2026nlpshow that the text\-informed prior contributes more when the item response signal is weaker\.

Several limitations should also be noted\. First, the proposed Bayesian estimation procedure relies on Markov chain Monte Carlo sampling and can be computationally intensive, particularly under the joint framework and larger item pools\. Second, the empirical application was based on a selected subset of frequently administered items rather than the full item pool\. Consequently, the findings may depend partly on the characteristics of the selected items and the resulting sample restrictions\. Third, the current study focused on the DINA model with binary attributes\. Future research could extend the framework to higher\-order or more complex CDMs that allow hierarchical or continuous latent structures\. Finally, the log\-derived behavioural covariates used in the empirical study may themselves contain measurement error, which was not explicitly modelled in the current analysis\. Incorporating uncertainty in process\-based covariates may further improve inference in longitudinal diagnostic modelling\.

Overall, the results suggest that joint estimation provides a more robust framework for longitudinal cognitive diagnosis when theQQ\-matrix is uncertain and measurement information is limited\. Although stepwise procedures may remain attractive because of their flexibility and lower computational burden, the joint framework offers a more coherent representation of uncertainty across measurement and learning processes over time\.

## References

## 6Appendix A\. Data Preprocessing

The empirical data were obtained from four game\-grade combinations: Grade 2Punchline, Grade 3Punchline, Grade 2Field Observer, and Grade 3Field Observer\. The total number of students in each dataset was as follows: 27,649 for Grade 2Punchline, 4,593 for Grade 3Punchline, 13,044 for Grade 2Field Observer, and 18,220 for Grade 3Field Observer\. To construct a longitudinal sample, we identified students who appeared in both Grade 2 and Grade 3 datasets\.

We conducted item selection separately for each game and grade based on student participation\. For each game and grade, we computed \(a\) the number of students who attempted each item and \(b\) the proportion of participating students\. Items were then ranked according to their frequency of appearance to identify those with the highest coverage\. For each game\-grade combination, we selected the six most frequently attempted items after balancing the trade\-off between including more items and retaining a larger sample size\. This produced 12 selected items at each grade, with six items fromPunchlineand six items fromField Observer\.

The selected items were:

- •Grade 2Punchline: 1,2,3,7,8,9
- •Grade 3Punchline: 1,2,3,31,32,33
- •Grade 2Field Observer: 1,3,5,7,9,11
- •Grade 3Field Observer: 31,33,35,37,39,41

We then restricted the dataset to students who completed all selected items across both games and both time points \(Grade 2 and Grade 3\)\. This resulted in a final analytic sample of 1,978 students and ensured a balanced longitudinal structure for subsequent modelling\.

To identify the most informative subset of items, we conducted an exploratory data analysis \(EDA\) on question usage frequency\. For each question, the number and proportion of students who attempted the item were computed\. The top 30 most frequently used questions in each game are shown in Figures[1](https://arxiv.org/html/2606.06804#S6.F1)and[2](https://arxiv.org/html/2606.06804#S6.F2)\. Based on these distributions, a subset of high\-frequency questions was selected to maximize sample size while maintaining sufficient coverage across levels\.

The covariates were derived from students’ gameplay data across the two games\. The number of attempts was computed as the average number of attempts per level for each student, and then averaged across the two games\. The number of questions correct was defined as the total number of questions answered correctly across both games\. Response time \(RT\) was calculated as the average response time per level for each student and subsequently averaged across the two games\. These definitions and computation procedures are consistent with those used in the previous study byma2026dynamic\.

![Refer to caption](https://arxiv.org/html/2606.06804v1/x1.png)Figure 1:Top 30 most frequently attempted questions in punchline, separately for Grade 2 and Grade 3\. Counts represent the number of students who attempted each question\. Questions selected for the analysis are among the highest\-frequency items![Refer to caption](https://arxiv.org/html/2606.06804v1/x2.png)Figure 2:Top 30 most frequently attempted questions in field observer, separately for Grade 2 and Grade 3\. Counts represent the number of students who attempted each question\. Questions selected for the analysis are among the highest\-frequency items
## 7Appendix B: Example of Questions

Table 14:Example items from PunchlineNote:Each Punchline item requires students to select the correct answer and identify the relevant meaning of the ambiguous word\.

Table 15:Example items from Field ObserverNote:Each item requires students to select both a correct answer and the supporting evidence sentence\.

## 8Appendix C\. Distribution of Text Information

![Refer to caption](https://arxiv.org/html/2606.06804v1/x3.png)Figure 3:Distribution ofτ\\tauvalues in the item pool![Refer to caption](https://arxiv.org/html/2606.06804v1/x4.png)Figure 4:Distribution of standardisedτ\\tauvalues in the item poolThe distribution of the item pool is shown in Figure[3](https://arxiv.org/html/2606.06804#S8.F3), and the standardized version is shown in Figure[4](https://arxiv.org/html/2606.06804#S8.F4)\. Because both Grade 2 and Grade 3 were drawn from the same underlying item pool, the distribution of item\-levelτ\\tauis identical across grades\. Differences between grades arise only from the item sampling process and subsequent student responses\. The histogram and QQ plot indicate thatτ\\tauis approximately normally distributed, with only mild deviations in the tails\. The density panel shows that the two games differ in shape:Punchlineitems form a single peak near zero, whereasField Observeritems are bimodal, with one mode just below zero and a second above it\. The sorted values reveal a smooth distribution, providing no evidence of extreme outliers\.

## 9Appendix D\. TrueQQ\-matrix in Simulation

The shorter test forms were constructed as nested subsets of the full 24\-item test\. The trueQQ\-matrices for the simulation withJ=24J=24andK=2K=2are shown in Table[16](https://arxiv.org/html/2606.06804#S9.T16)\. TheQQ\-matrices were held fixed across all simulation replicates\.

Table 16:TrueQQ\-matrices forJ=24J=24at Time 1 and Time 2

Similar Articles

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

arXiv cs.CL

Proposes J-Access, an inference-time audit using the Jacobian lens to measure residual knowledge accessibility in unlearned LLMs, finding that accessibility predicts recovery speed but that directly minimizing it fails to promote genuine deletion.

Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States

arXiv cs.CL

This paper investigates whether language models can correctly judge how diagnostic evidence supports or challenges different causal claims, introducing paired prompts that vary only the causal target. Linear readouts from the penultimate transformer hidden state on models like Qwen2.5-7B-Instruct show moderate balanced accuracy (0.654-0.659) and recover 18–21 out of 49 pairs, indicating some linear decodability of causal relevance.

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling

arXiv cs.AI

This paper introduces DyCon, a training-free framework that uses step-level embeddings to model evolving task difficulty and dynamically control reasoning depth in Large Reasoning Models, effectively reducing overthinking and improving efficiency without sacrificing accuracy.

Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication

arXiv cs.LG

This paper addresses the problem of treatment-induced label indeterminacy in clinical prediction, using post-cardiac-arrest neurological prognostication as a case study. The authors propose a framework that incorporates expert annotations of counterfactual outcomes and highlight tradeoffs between accuracy on certain versus uncertain cases.