能力-残差解耦建模用于情感认知诊断

arXiv cs.AI 论文

摘要

本文提出一种能力-残差解耦框架,用于情感认知诊断,通过建模认知残差以减轻情感污染,实验表明该方法提升了性能并改进了情感对齐。

arXiv:2609.21214v1 Announce Type: new Abstract: Cognitive diagnosis infers students' concept mastery from response logs. However, students' responses are not determined by mastery alone: non-cognitive factors such as emotion, engagement, and fatigue can also affect performance. Affective cognitive diagnosis therefore extends conventional cognitive diagnosis by incorporating affective states. Existing methods often assume that the cognitive diagnosis backbone has already explained ability, item, and concept effects, so the remaining errors can be attributed mainly to affect. We argue that this assumption can be insufficient in real educational data: item calibration bias, systematic concept bias, personalized student-concept deviations, and latent student-item matching can form stable cognitive residuals. Without an explicit modeling pathway, these residuals may leak into affective representations, producing affect contamination. To address this problem, we propose an ability-residual decoupled framework for affective cognitive diagnosis. The model first captures unmodeled cognitive residuals through student, item, concept, student-concept, and low-rank student-item components, and then uses an affective module to modulate guess/slip effects. A Q-matrix-constrained concept residual attention mechanism adaptively aggregates only item-relevant concept residuals. Experiments on ASSIST2017, ASSIST2012, ASSIST2009, and Junyi with six cognitive diagnosis backbones show response-prediction gains across the reported comparisons and generally improved affect alignment when affect labels are available. Ablation studies, leakage probes, principal component analysis visualization, long-tail analysis, and case studies further indicate that ability residuals absorb stable cognitive bias, reduce cognitive contamination in the affective branch, and enhance the robustness and predictive accuracy of cognitive diagnosis models.
查看原文
查看缓存全文

缓存时间: 2026/09/21 09:18

# Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis
Source: [https://arxiv.org/html/2609.21214](https://arxiv.org/html/2609.21214)
Boyuan ZhaoEmail:[20243259@snnu\.edu\.cn](mailto:[email protected])Affiliation:Key Laboratory of Modern Teaching Technology, Shaanxi Normal University, 199 Chang’an South Road, Xi’an, 710062, Shaanxi, ChinaMeng YeEmail:[yemeng@snnu\.edu\.cn](mailto:[email protected])Affiliation:Key Laboratory of Modern Teaching Technology, Shaanxi Normal University, 199 Chang’an South Road, Xi’an, 710062, Shaanxi, China

###### Abstract

Cognitive diagnosis infers students’ concept mastery from response logs\. However, students’ responses are not determined by mastery alone: non\-cognitive factors such as emotion, engagement, and fatigue can also affect performance\. Affective cognitive diagnosis therefore extends conventional cognitive diagnosis by incorporating affective states\. Existing methods often assume that the cognitive diagnosis backbone has already explained ability, item, and concept effects, so the remaining errors can be attributed mainly to affect\. We argue that this assumption can be insufficient in real educational data: item calibration bias, systematic concept bias, personalized student\-concept deviations, and latent student\-item matching can form stable cognitive residuals\. Without an explicit modeling pathway, these residuals may leak into affective representations, producing affect contamination\. To address this problem, we propose an ability\-residual decoupled framework for affective cognitive diagnosis\. The model first captures unmodeled cognitive residuals through student, item, concept, student\-concept, and low\-rank student\-item components, and then uses an affective module to modulate guess/slip effects\. A Q\-matrix\-constrained concept residual attention mechanism adaptively aggregates only item\-relevant concept residuals\. Experiments on ASSIST2017, ASSIST2012, ASSIST2009, and Junyi with six cognitive diagnosis backbones show response\-prediction gains across the reported comparisons and generally improved affect alignment when affect labels are available\. Ablation studies, leakage probes, principal component analysis visualization, long\-tail analysis, and case studies further indicate that ability residuals absorb stable cognitive bias, reduce cognitive contamination in the affective branch, and enhance the robustness and predictive accuracy of cognitive diagnosis models\. Our code is available at[https://github\.com/psychosiwa/ARACD](https://github.com/psychosiwa/ARACD)\.

###### keywords

cognitive diagnosis, affective cognitive diagnosis, ability residual, affect contamination, concept attention, intelligent education

## 1Introduction

Cognitive diagnosis \(CD\) infers students’ mastery of fine\-grained knowledge concepts, skills, or cognitive attributes from their responses in exercises, tests, or online learning platforms\[[1](https://arxiv.org/html/2609.21214#bib.bib1)\],\[[2](https://arxiv.org/html/2609.21214#bib.bib2)\]\. Unlike evaluation methods that report an overall score or a holistic ability estimate, CD emphasizes interpretable diagnostic results, i\.e\., which concepts a student has mastered and which remain deficient\. Therefore, CD can support personalized exercise recommendation, learning path planning, formative assessment, and adaptive testing in intelligent tutoring systems\[[2](https://arxiv.org/html/2609.21214#bib.bib2),[3](https://arxiv.org/html/2609.21214#bib.bib3),[4](https://arxiv.org/html/2609.21214#bib.bib4)\]\. As student response logs continue to accumulate in large\-scale online education, CD has emerged as a critical technology connecting educational measurement, student modeling, and intelligent educational decision\-making\[[5](https://arxiv.org/html/2609.21214#bib.bib5),[6](https://arxiv.org/html/2609.21214#bib.bib6),[7](https://arxiv.org/html/2609.21214#bib.bib7)\]\.

Research on cognitive diagnosis models \(CDMs\) has evolved from psychometric models to neural diagnostic models\. Classic approaches include DINA, IRT, and MIRT\. Specifically, DINA uses the Q\-matrix to characterize the concepts required by an item and explains response noise through discrete mastery patterns and item parameters, such as guessing and slipping\[[1](https://arxiv.org/html/2609.21214#bib.bib1)\],\[[8](https://arxiv.org/html/2609.21214#bib.bib8)\]\. Accordingly, IRT estimates response probabilities from the matching relationship between continuous latent ability and item difficulty\[[9](https://arxiv.org/html/2609.21214#bib.bib9)\], while MIRT extends ability to a multidimensional latent space for multi\-skill assessment tasks\[[10](https://arxiv.org/html/2609.21214#bib.bib10)\]\. Deep learning has recently been introduced into CD, affording models that can represent more complex interactions among students, items, and cognitive attributes\. For instance, NCDM strengthens student\-item\-concept interaction modeling with neural networks while retaining interpretability\[[5](https://arxiv.org/html/2609.21214#bib.bib5)\]\. Alternative approaches, such as RCD and SCD, improve diagnosis under sparse\-data conditions by using relational structures and self\-supervised signals, respectively\[[11](https://arxiv.org/html/2609.21214#bib.bib11)\],\[[12](https://arxiv.org/html/2609.21214#bib.bib12)\]\. Higher\-order ability modeling, global\-local cognitive modeling, and attention mechanisms have also been studied to feature complex ability structures, learning dynamics, and item\-related concept relations\[[7](https://arxiv.org/html/2609.21214#bib.bib7)\],\[[13](https://arxiv.org/html/2609.21214#bib.bib13)\]\.

Although CDMs are designed to explain response behavior through cognitive ability or mastery, an observed response may be driven by latent cognitive factors, latent affective factors, or their joint influence\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\. Existing studies report distinct roles for different affective factors, with control\-value theory explaining how achievement emotions shape motivation, learning strategies, and performance\[[14](https://arxiv.org/html/2609.21214#bib.bib14)\]\. Empirical work reveals that affective states evolve dynamically during complex learning\[[15](https://arxiv.org/html/2609.21214#bib.bib15)\], and that boredom and frustration differ in their incidence, persistence, and consequences across learning environments\[[16](https://arxiv.org/html/2609.21214#bib.bib16)\]\. Furthermore, affect\-detection research has introduced methods for estimating learner states\[[17](https://arxiv.org/html/2609.21214#bib.bib17)\], whereas affect\-aware tutoring studies use these estimates to adapt learner\-system interactions and feedback\[[18](https://arxiv.org/html/2609.21214#bib.bib18),[19](https://arxiv.org/html/2609.21214#bib.bib19),[20](https://arxiv.org/html/2609.21214#bib.bib20)\]\. Together, these findings lay the path for incorporating affective factors into cognitive diagnosis\. Consequently, researchers have introduced affective states into cognitive diagnosis frameworks, leading to a new perspective, affective cognitive diagnosis\. These models estimate students’ non\-cognitive states through an affective module and use them to modulate guess/slip effects, thereby explaining response fluctuations among students with comparable knowledge mastery\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\. Notably, when real affect labels are unavailable, these methods can draw on supervised or instance\-level contrastive learning and use response labels to construct weak positive and negative relations for learning response\-derived affective representations\[[22](https://arxiv.org/html/2609.21214#bib.bib22)\],\[[23](https://arxiv.org/html/2609.21214#bib.bib23)\]\.

Existing affective cognitive diagnosis methods implicitly assume that the base cognitive diagnosis backbone has already explained student, item, and concept\-related cognitive structure sufficiently, so that the remaining response variation mainly reflects affective states\[[5](https://arxiv.org/html/2609.21214#bib.bib5)\],\[[11](https://arxiv.org/html/2609.21214#bib.bib11)\],\[[12](https://arxiv.org/html/2609.21214#bib.bib12)\],\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\. Nevertheless, this assumption can be insufficient in real educational data, as items may contain difficulty calibration bias; concepts may exhibit systematic prediction bias; a student’s mastery of a specific concept may deviate from his or her overall ability; and a student\-item pair may involve personalized matching effects due to wording familiarity, problem\-solving strategies, or hidden skill requirements\[[7](https://arxiv.org/html/2609.21214#bib.bib7)\],\[[24](https://arxiv.org/html/2609.21214#bib.bib24)\],\[[25](https://arxiv.org/html/2609.21214#bib.bib25)\]\. When these stable cognitive residuals are not explicitly modeled, the affective module may be forced to absorb them\. Consequently, the learned affective representation may encode cognitive\-residual information associated with items, concepts, and student ability in addition to non\-cognitive states\. This phenomenon is defined as affect contamination, namely cognitive residual contamination in affective representations\. Such contamination weakens interpretability and may force the model to misattribute cognitive residual variation to affective factors, thereby reducing diagnostic reliability\[[17](https://arxiv.org/html/2609.21214#bib.bib17)\],\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\.

Motivated by this problem, this study develops an ability\-residual decoupling framework for affective cognitive diagnosis\. It instantiates it as ARACD or ARCACD depending on whether real affect supervision is available\. Inspired by bias decomposition and low\-rank interaction modeling in recommender systems\[[24](https://arxiv.org/html/2609.21214#bib.bib24),[25](https://arxiv.org/html/2609.21214#bib.bib25)\], the proposed approach uses multiple structured bias terms to express unmodeled cognitive residual variation, including item, student, concept\-shared, and student\-concept personalized biases\. Furthermore, we introduce a low\-rank student\-item interaction term to capture latent personalized matching effects without parameterizing a complete student\-item interaction matrix\. Concept\-level residuals are aggregated by a Q\-matrix\-constrained attention mechanism that assigns student–item\-pair\-conditioned weights to item\-related concept\-shared and student\-concept residuals\. The resulting ability residual first corrects the base prediction on the logit scale, after which the affective module performs guess/slip modulation\. This ordered design places stable cognitive residual correction before affective modulation, assigning cognitive residual and affective effects to separate modeling pathways\[[5](https://arxiv.org/html/2609.21214#bib.bib5),[13](https://arxiv.org/html/2609.21214#bib.bib13)\]\. Figure 1 illustrates the affect contamination and ability\-residual decoupling: stable cognitive residuals not explained by the backbone may leak into the affective branch, whereas a dedicated ability\-residual pathway provides a more appropriate structural destination for this residual information\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/introduction.png)Figure 1:Schematic illustration of affect contamination and ability\-residual decoupling\. Panel \(a\) illustrates that a student’s observed response to an exercise may reflect both cognitive ability and affective state\. Panel \(b\) illustrates the proposed allocation of modeling roles: stable cognitive residual variation is assigned to the ability\-residual pathway, reducing its leakage into the affective branch and allowing the affective representation to focus on non\-cognitive modulation\.The main contributions of this paper are as follows\.

1. 1\.We identify the affect contamination problem in affective cognitive diagnosis based on the observation that affective representations may absorb unmodeled cognitive residuals and become entangled with item, concept, and ability bias\. This work complements existing affect\-aware cognitive diagnosis studies, which focus on prediction improvement while paying less attention to representation contamination\[[17](https://arxiv.org/html/2609.21214#bib.bib17)\],\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\.
2. 2\.We propose an ability\-residual decoupling architecture that separates item, student, concept, student\-concept personalized, and low\-rank student\-item interaction residuals from affective modeling\. This strategy allows the affective module to focus more on non\-cognitive factors\.
3. 3\.We construct a Q\-matrix\-masked concept residual attention mechanism, which adaptively aggregates concept\-shared residuals and student\-concept personalized residuals as interpretable concept values within the item\-related concept set\. This approach improves the accuracy and interpretability of residual modeling\.
4. 4\.We validate the proposed framework on ASSIST2017, ASSIST2012, ASSIST2009, and Junyi using multiple cognitive diagnosis backbones, including DINA, IRT, MIRT, NCDM, RCD, and SCD\. Our method’s effectiveness is systematically evaluated on response prediction, affect prediction, and representation decoupling through ablation experiments, residual group analysis, affect contamination diagnostics, and representation visualization\[[1](https://arxiv.org/html/2609.21214#bib.bib1)\],\[[5](https://arxiv.org/html/2609.21214#bib.bib5)\],\[[9](https://arxiv.org/html/2609.21214#bib.bib9),[10](https://arxiv.org/html/2609.21214#bib.bib10),[11](https://arxiv.org/html/2609.21214#bib.bib11),[12](https://arxiv.org/html/2609.21214#bib.bib12)\],\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\.

## 2Related Work

### 2\.1Cognitive Diagnosis and Affect\-Aware Extensions

CD models estimate students’ mastery of fine\-grained knowledge concepts or cognitive attributes from response logs, supporting personalized instruction, formative assessment, and adaptive testing\[[2](https://arxiv.org/html/2609.21214#bib.bib2),[3](https://arxiv.org/html/2609.21214#bib.bib3),[4](https://arxiv.org/html/2609.21214#bib.bib4)\]\. Early psychometric and cognitive diagnosis models have highly interpretable structures\. Classic CDMs include DINA, IRT, and MIRT\. Specifically, DINA relies on the Q\-matrix to represent discrete item\-concept requirements and uses item parameters such as guessing and slipping to explain response noise beyond mastery states\[[1](https://arxiv.org/html/2609.21214#bib.bib1)\],\[[8](https://arxiv.org/html/2609.21214#bib.bib8)\]\. Besides, IRT models response probability as a function of continuous student ability and item difficulty\[[9](https://arxiv.org/html/2609.21214#bib.bib9)\], while MIRT extends unidimensional ability to a multidimensional latent vector for multi\-skill or multi\-concept assessment\[[10](https://arxiv.org/html/2609.21214#bib.bib10)\]\. These models preserve clear educational\-measurement meanings among students, items, and concepts, although parametric assumptions can limit their expressive power\.

As researchers have accumulated large\-scale response data on online education platforms, deep learning has been incorporated into cognitive diagnosis, giving rise to neural CDMs\. An NCDM combines response probabilities with neural networks while retaining interpretable variables and monotonicity constraints\[[5](https://arxiv.org/html/2609.21214#bib.bib5)\]\. In this case, RCD models structural relations among students, items, and concepts, whereas SCD introduces self\-supervised signals to alleviate sparsity and label deficiency\[[11](https://arxiv.org/html/2609.21214#bib.bib11)\],\[[12](https://arxiv.org/html/2609.21214#bib.bib12)\]\. Higher\-order and global\-local models further use hierarchical structures, attention, or local concept relations\[[7](https://arxiv.org/html/2609.21214#bib.bib7)\],\[[13](https://arxiv.org/html/2609.21214#bib.bib13)\]\. These developments strengthen the representation of cognitive structure through more expressive student\-, item\-, and concept\-level interactions, but their primary modeling target remains cognitive ability or mastery\.

Observed responses are also affected by non\-cognitive factors\. Indeed, achievement emotions, confusion, frustration, and boredom can influence motivation, engagement, and performance\[[14](https://arxiv.org/html/2609.21214#bib.bib14),[15](https://arxiv.org/html/2609.21214#bib.bib15),[16](https://arxiv.org/html/2609.21214#bib.bib16)\]\. At the same time, affect\-detection and affect\-aware tutoring studies provide empirical and computational support for incorporating these factors into learner modeling\[[17](https://arxiv.org/html/2609.21214#bib.bib17),[18](https://arxiv.org/html/2609.21214#bib.bib18),[19](https://arxiv.org/html/2609.21214#bib.bib19),[20](https://arxiv.org/html/2609.21214#bib.bib20)\]\. Affect\-aware cognitive diagnosis extends conventional CDMs by estimating affective states and uses these estimations to modulate guessing and slipping\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\. When external affect labels are unavailable, response\-conditioned contrastive learning can provide weak supervision for learning response\-derived affective representations\[[22](https://arxiv.org/html/2609.21214#bib.bib22)\],\[[23](https://arxiv.org/html/2609.21214#bib.bib23)\]\.

Although conventional and affect\-aware CDMs improve cognitive\-structure modeling and non\-cognitive modulation, respectively, existing affect\-aware extensions assume that the cognitive backbone has already explained stable student\-, item\-, and concept\-related variation sufficiently\[[5](https://arxiv.org/html/2609.21214#bib.bib5)\],\[[11](https://arxiv.org/html/2609.21214#bib.bib11)\],\[[12](https://arxiv.org/html/2609.21214#bib.bib12)\],\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\. In real educational data, item calibration bias, concept\-level systematic bias, personalized student\-concept deviations, and latent student\-item matching may remain\[[7](https://arxiv.org/html/2609.21214#bib.bib7)\],\[[24](https://arxiv.org/html/2609.21214#bib.bib24)\],\[[25](https://arxiv.org/html/2609.21214#bib.bib25)\]\. Notably, if these residuals are not explicitly modeled, the affective branch may absorb them and act as a generic error compensator\. In this case, response improvements alone cannot establish that the learned affective representation is aligned with non\-cognitive states\[[17](https://arxiv.org/html/2609.21214#bib.bib17)\],\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\. Motivated by this gap, this study assigns stable cognitive residual variation to a dedicated ability\-residual pathway before affective modulation\.

### 2\.2Representation Decoupling

Representation learning aims to capture factors that are useful for downstream tasks\[[26](https://arxiv.org/html/2609.21214#bib.bib26)\]\. Interestingly, a representation that improves prediction is not necessarily aligned with the target factor\. Indeed, disentangled representation learning reveals that learned factors depend on task objectives, supervision signals, and structural inductive biases\. Without appropriate inductive biases, general learning objectives cannot guarantee that a model will recover the expected factors\[[27](https://arxiv.org/html/2609.21214#bib.bib27)\]\. This limitation is especially important when multiple latent factors are statistically correlated, because a model may encode shortcut features or spurious correlations that are useful for prediction but do not correspond to the intended factor\[[28](https://arxiv.org/html/2609.21214#bib.bib28)\]\.

Representation decoupling overcomes this limitation by encouraging different latent components to carry different types of information\. Depending on the task, decoupling can be introduced through architectural separation, factor\-specific supervision, contrastive objectives, or adversarial constraints\. Although these mechanisms do not guarantee complete independence, they reduce unnecessary information leakage and improve interpretability\. Typical diagnostic tools are linear probes, which test whether a representation contains information about a factor that the model intends to separate\[[29](https://arxiv.org/html/2609.21214#bib.bib29)\]\. Although these do not enforce decoupling, they provide empirical evidence about whether the intended separation has been achieved\. This connection motivates the leakage probes used in this work to evaluate cognitive information remaining in affective representations\.

### 2\.3Residual Modeling

Residual learning is a general modeling concept widely used to let models learn the remaining function relative to a base mapping\[[31](https://arxiv.org/html/2609.21214#bib.bib31)\]\. Bias correction and low\-rank interaction modeling are widely used in recommender systems and educational prediction tasks to explain stable individual differences and sparse interaction structures\[[24](https://arxiv.org/html/2609.21214#bib.bib24)\],\[[25](https://arxiv.org/html/2609.21214#bib.bib25)\],\[[32](https://arxiv.org/html/2609.21214#bib.bib32),[33](https://arxiv.org/html/2609.21214#bib.bib33),[34](https://arxiv.org/html/2609.21214#bib.bib34)\]\. Matrix factorization recommender models typically model user bias, item bias, and low\-rank user\-item interactions to explain personalized differences that the global mean and the main model cannot capture\[[24](https://arxiv.org/html/2609.21214#bib.bib24)\],\[[25](https://arxiv.org/html/2609.21214#bib.bib25)\],\[[32](https://arxiv.org/html/2609.21214#bib.bib32)\]\. This concept has a natural interpretation when transferred to educational data: student bias corresponds to a student’s overall response tendency or baseline ability offset, item bias corresponds to item difficulty calibration error, and low\-rank student\-item interactions correspond to wording familiarity, problem\-solving strategy matching, hidden skill requirements, or other unobserved matching relations\. Similarly, response prediction in knowledge tracing and online assessment systems involves sparse interactions among students, items, and exercise records\. Therefore, it requires additional structures to characterize individual differences and historical response patterns\[[4](https://arxiv.org/html/2609.21214#bib.bib4)\],\[[6](https://arxiv.org/html/2609.21214#bib.bib6)\],\[[33](https://arxiv.org/html/2609.21214#bib.bib33)\],\[[34](https://arxiv.org/html/2609.21214#bib.bib34)\]\.

From this decoupling perspective, residual modeling offers the structural mechanism for assigning stable cognitive variation to a dedicated pathway rather than to the affective representation\. In affective cognitive diagnosis, ability residuals are not merely numerical corrections to response logits but allocate modeling roles\. Thus, residual patterns that recur with students, items, concepts, or student\-item combinations should be explained mainly by the ability\-residual pathway\. In contrast, the affective module effectively handles fluctuations associated with transient states, engagement, or affective changes\. Accordingly, the ability\-residual module includes student, item, concept\-shared, student\-concept personalized, and low\-rank student\-item components\. Hence, residual modeling supports representation decoupling in addition to increasing model capacity\.

## 3Problem Formulation

Let𝒮\\mathcal\{S\}be the set of students, beℐ\\mathcal\{I\}the set of items, and𝒦\\mathcal\{K\}the set of concepts\. For a student\-item interaction\(s,i\)\(s,i\), wheres∈𝒮s\\in\\mathcal\{S\}andi∈ℐi\\in\\mathcal\{I\}, the concept association vector of itemiiis denoted by𝐪i∈\{0,1\}K\\mathbf\{q\}\_\{i\}\\in\\\{0,1\\\}^\{K\}, and the response label isys​i∈\{0,1\}y\_\{si\}\\in\\\{0,1\\\}\. If affect labels or affect vectors are available in the dataset, they are denoted by𝐚s​i\\mathbf\{a\}\_\{si\}\.

An observed response may reflect a latent cognitive state𝐜s​i\\mathbf\{c\}\_\{si\}, a latent affective state𝐚s​i\\mathbf\{a\}\_\{si\}, or their joint influence\[[21](https://arxiv.org/html/2609.21214#bib.bib21)\]\. Letps​ip\_\{si\}denote the probability of a correct response:

ps​i=P⁡\(ys​i=1∣𝐜s​i,𝐚s​i,i,𝐪i\)p\_\{si\}=P\(y\_\{si\}=1\\mid\\mathbf\{c\}\_\{si\},\\mathbf\{a\}\_\{si\},i,\\mathbf\{q\}\_\{i\}\)\(1\)
A conventional cognitive diagnosis model learns the following response probability:

y^s​i=P⁡\(ys​i=1∣s,i,𝐪i\)\.\\hat\{y\}\_\{si\}=P\(y\_\{si\}=1\\mid s,i,\\mathbf\{q\}\_\{i\}\)\.\(2\)
Affective cognitive diagnosis further introduces an affective representation𝐚s​i\\mathbf\{a\}\_\{si\}, which is usually written as

y^s​i=fθ​\(s,i,𝐪i,𝐚s​i\)\.\\hat\{y\}\_\{si\}=f\_\{\\theta\}\(s,i,\\mathbf\{q\}\_\{i\},\\mathbf\{a\}\_\{si\}\)\.\(3\)
However, when the base cognitive diagnosis backbone cannot sufficiently explain cognitive factors, the learned affective representation may contain the target non\-cognitive state and the incorrectly absorbed cognitive residual, which is formulated as:

𝐚s​i=𝐚s​ia​f​f​e​c​t\+Δs​ic​o​g,\\mathbf\{a\}\_\{si\}=\\mathbf\{a\}^\{affect\}\_\{si\}\+\\Delta^\{cog\}\_\{si\},\(4\)
where𝐚s​ia​f​f​e​c​t\\mathbf\{a\}^\{affect\}\_\{si\}is the non\-cognitive component related to affective state andΔs​ic​o​g\\Delta^\{cog\}\_\{si\}denotes cognitive residual variation that enters the affective space, without being explained by the base backbone, such as variation associated with item calibration bias, concept bias, or student\-concept deviations\.

Throughout this paper,*cognitive residuals*refer to stable cognitive variation not explained by the base cognitive diagnosis backbone\. In contrast, the*ability residual*rs​ia​b​i​l​i​t​yr^\{ability\}\_\{si\}is the structured, model\-estimated correction introduced to capture such variation\.*Ability residual*is used in a broad diagnostic sense and is not limited to a student\-specific ability offset, incorporating item, concept, student\-concept, and low\-rank student\-item components\.

Besides,Δs​ic​o​g\\Delta^\{cog\}\_\{si\}is a conceptual latent variation rather than an observed training target\. By contrast,rs​ia​b​i​l​i​t​yr^\{ability\}\_\{si\}is a structured correction estimated by the model from the available data and learning objectives\. The two quantities are not equivalent, and response labels alone cannot uniquely identify their decomposition or recover either quantity as ground truth\.

This paper aims to explicitly model the ability residualrs​ia​b​i​l​i​t​yr^\{ability\}\_\{si\}so that the final response probability satisfies:

y^s​i=fθ​\(s,i,𝐪i,𝐚s​ia​f​f​e​c​t,rs​ia​b​i​l​i​t​y\),\\hat\{y\}\_\{si\}=f\_\{\\theta\}\(s,i,\\mathbf\{q\}\_\{i\},\\mathbf\{a\}^\{affect\}\_\{si\},r^\{ability\}\_\{si\}\),\(5\)
and to reduce the entanglement between𝐚s​ia​f​f​e​c​t\\mathbf\{a\}^\{affect\}\_\{si\}andΔs​ic​o​g\\Delta^\{cog\}\_\{si\}through a structured allocation of modeling roles\. In other words, the ability residual module explains unmodeled cognitive residual variation, while the affective module explains how non\-cognitive states modulate guessing, slipping, and response tendency\.

## 4Method

### 4\.1Overall Framework

The proposed framework has three core components: a base cognitive diagnosis backbone, an ability residual module, and an affective module\. The base backboneFC​DF\_\{CD\}outputs a base response probability according to student ability, item difficulty, item discrimination, and the Q\-matrix:

ys​ib​a​s​e=FC​D​\(𝐞s,𝐝i,𝐝id​i​s​c,𝐪i\),y^\{base\}\_\{si\}=F\_\{CD\}\(\\mathbf\{e\}\_\{s\},\\mathbf\{d\}\_\{i\},\\mathbf\{d\}^\{disc\}\_\{i\},\\mathbf\{q\}\_\{i\}\),\(6\)
where𝐞s\\mathbf\{e\}\_\{s\}denotes student ability or mastery state,𝐝i\\mathbf\{d\}\_\{i\}is item difficulty or concept difficulty,𝐝id​i​s​c\\mathbf\{d\}^\{disc\}\_\{i\}denotes item discrimination, and𝐪i\\mathbf\{q\}\_\{i\}denotes the concepts involved in the item\. DINA, IRT, MIRT, NCDM, RCD, or SCD can replace this backbone\.

Then, the ability residual module introducesrs​ia​b​i​l​i​t​yr^\{ability\}\_\{si\}to capture cognitive residual variation not explained by the base backbone and applies this correction on the logit scale:

zs​ia​b​i​l​i​t​y=logit⁡\(ys​ib​a​s​e\)\+rs​ia​b​i​l​i​t​y,ys​i∗=σ⁡\(zs​ia​b​i​l​i​t​y\)\.z^\{ability\}\_\{si\}=\\operatorname\{logit\}\(y^\{base\}\_\{si\}\)\+r^\{ability\}\_\{si\},\\qquad y^\{\*\}\_\{si\}=\\sigma\(z^\{ability\}\_\{si\}\)\.\(7\)
Finally, the affective module generates an affective representation and modulates the residual\-corrected response probability through the guess/slip mechanism\. Unlike directly concatenating affective representations into the backbone, this structure explicitly distinguishes ability residuals from affective factors\. Therefore, it is consistent with the modeling assumption on affect contamination presented in this paper\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/ARACD_framework_cropped.png)Figure 2:Overall ARACD algorithmic framework and ability residual module\.
### 4\.2Base Cognitive Diagnosis Backbone

The base cognitive diagnosis backbone is treated as a replaceable module\. Considering NCDM as an example, the student knowledge mastery, item concept difficulty, and item discrimination vectors jointly form an interaction representation:

𝐱s​i=𝐝id​i​s​c⊙\(𝐞s−𝐝ik​c\)⊙𝐪i,\\mathbf\{x\}\_\{si\}=\\mathbf\{d\}^\{disc\}\_\{i\}\\odot\(\\mathbf\{e\}\_\{s\}\-\\mathbf\{d\}^\{kc\}\_\{i\}\)\\odot\\mathbf\{q\}\_\{i\},\(8\)
ys​ib​a​s​e=σ⁡\(M​L​P​\(𝐱s​i\)\)\.y^\{base\}\_\{si\}=\\sigma\(MLP\(\\mathbf\{x\}\_\{si\}\)\)\.\(9\)
where⊙\\odotdenotes element\-wise multiplication,𝐝id​i​s​c\\mathbf\{d\}^\{disc\}\_\{i\}is item discrimination, and𝐝ik​c\\mathbf\{d\}^\{kc\}\_\{i\}is the item’s concept difficulty representation\. For IRT and MIRT, the backbone formulation reduces to a matching function between unidimensional or multidimensional latent ability and item parameters\. Accordingly, for DINA, the backbone estimates response probability based on whether the student has mastered all concepts required by the item and on guess/slip parameters\. The proposed framework is designed as a plug\-in that can be attached to different cognitive diagnosis models\.

### 4\.3Ability Residual Modeling

The ability residual explains cognitive residual variation not captured by the base backbone\. However, to improve modeling stability and interpretability, the ability residual is divided into a student\-item residual layer and a concept residual layer\. The former layer models overall offsets and personalized interactions that do not depend on the concept structure, including student global bias, item global bias, and low\-rank student\-item latent interactions\. The latter layer uses Q\-matrix constraints to aggregate concept\-shared residuals and student\-concept personalized residuals within the concepts related to the item\. This division separates non\-concept\-level response bias from concept\-level knowledge bias, making the residual structure more compact and interpretable\.

This modeling strategy absorbs cognitive residuals from non\-affective sources because each residual term corresponds directly to a common stable error source in educational measurement\. Specifically, item global bias absorbs item difficulty calibration errors; student global bias absorbs a student’s overall response tendency or baseline ability offset; concept\-shared residuals absorb systematic prediction bias on certain concepts; student\-concept personalized residuals absorb deviations in a student’s mastery of specific concepts; and low\-rank student\-item interactions absorb student\-item\-specific errors caused by wording familiarity, strategy matching, or hidden skill requirements\. These factors commonly appear with the structure of students, items, or concepts, rather than being determined instantaneously by the affective state in a single interaction\. Therefore, explicitly placing them in the ability residual pathway is a more appropriate modeling pathway for cognitive bias not explained by the base backbone\.

During training, the ability residual corrects the base backbone output on the logit scale and is connected to the final response loss\. If a certain type of error repeatedly appears for the same student, item, concept, or similar student\-item combinations, gradients update the corresponding residual parameters and assign this stable error to the ability residual pathway\. In contrast, the affective module modulates the final response probability through guess/slip and is more suitable for explaining response fluctuations induced by non\-cognitive states\. This structural division does not assume that the two are completely separable\. Instead, it reduces the pressure on the affective branch to absorb cognitive bias through parameter allocation and educational structural constraints, thereby mitigating affect representation contamination\.

rs​ia​b​i​l​i​t​y=wSI​rs​is​t​u​\-​i​t​e​m\+wC​rs​ic​o​n​c​e​p​t,r^\{ability\}\_\{si\}=w\_\{\\mathrm\{SI\}\}r^\{stu\\text\{\-\}item\}\_\{si\}\+w\_\{\\mathrm\{C\}\}r^\{concept\}\_\{si\},\(10\)
wherers​is​t​u​\-​i​t​e​mr^\{stu\\text\{\-\}item\}\_\{si\}is the student\-item residual layer,rs​ic​o​n​c​e​p​tr^\{concept\}\_\{si\}is the concept residual layer, andwSIw\_\{\\mathrm\{SI\}\}andwCw\_\{\\mathrm\{C\}\}are the scaling coefficients of the two layers\. HerewSIw\_\{\\mathrm\{SI\}\}is an overall weight for the student\-item residual layer rather than a parameter that varies separately with sample\(s,i\)\(s,i\)\. In implementation, it can be set as either a learnable parameter or fixed to 1\.

Allbbterms in the ability residual are internal learnable parameters rather than externally provided labels or manually calculated rules\. During training, the model retrieves these residual terms from the corresponding parameter tables according to student, item, and concept IDs, and updates them end\-to\-end through the response prediction loss and affect representation learning loss\. To avoid ambiguity in notation, the main symbols used in ability residual decomposition and concept attention are summarized below\.

Table 1:Main notation used in the ability residual decomposition and concept residual attention\.SymbolMeaningssThe student in the considered student–item pairiiThe item in the considered student–item pairkkThekk\-th concept𝐪i\\mathbf\{q\}\_\{i\}Concept mask of itemii𝒦i\\mathcal\{K\}\_\{i\}The set of concepts involved in itemii,𝒦i=\{k:qi​k=1\}\\mathcal\{K\}\_\{i\}=\\\{k:q\_\{ik\}=1\\\}bsb\_\{s\}Stable offset indicating that studentssis more or less likely than the overall level to answer correctlybib\_\{i\}Difficulty calibration error or item\-specific systematic offset of itemiibk,𝐛b\_\{k\},\\mathbf\{b\}Systematic residual shared by conceptkkacross all students and items, and the vector of all concept\-shared residualsbs,k,𝐛sc​o​n​c​e​p​tb\_\{s,k\},\\mathbf\{b\}^\{concept\}\_\{s\}Personalized mastery deviation of studentsson conceptkkrelative to his or her overall ability, and the vector of all concept\-specific personalized residuals for studentssvs,k,𝐯sv\_\{s,k\},\\mathbf\{v\}\_\{s\}Composite concept residual value for studentsson conceptkk,vs,k=bk\+bs,kv\_\{s,k\}=b\_\{k\}\+b\_\{s,k\}𝐮s,𝐯i\\mathbf\{u\}\_\{s\},\\mathbf\{v\}\_\{i\}Latent student\-item matching relation, whose inner product approximates the complete interaction residual matrixγS​I\\gamma\_\{SI\}Strength coefficient of the low\-rank student\-item interaction residual𝐞sa​t​t\\mathbf\{e\}^\{att\}\_\{s\}Student\-specific embedding used to generate the student–item\-pair\-conditioned query𝐯ia​t​t\\mathbf\{v\}^\{att\}\_\{i\}Item\-specific embedding used to generate the student–item\-pair\-conditioned query𝐪s​ia​t​t\\mathbf\{q\}^\{att\}\_\{si\}Student–item\-pair\-conditioned query,𝐪s​ia​t​t=L​N​\(tanh⁡\(𝐞sa​t​t\+𝐯ia​t​t\)\)\\mathbf\{q\}^\{att\}\_\{si\}=LN\(\\tanh\(\\mathbf\{e\}^\{att\}\_\{s\}\+\\mathbf\{v\}^\{att\}\_\{i\}\)\)𝐤ka​t​t,𝐊a​t​t\\mathbf\{k\}^\{att\}\_\{k\},\\mathbf\{K\}^\{att\}Attention key of thekk\-th concept and all concept keyses​i​k,𝐞s​ie\_\{sik\},\\mathbf\{e\}\_\{si\}Concept attention scores before Q\-mask applicationmi​k,𝐦im\_\{ik\},\\mathbf\{m\}\_\{i\}Q\-matrix mask that restricts attention to item\-related conceptsαs​i​k,𝜶s​i\\alpha\_\{sik\},\\boldsymbol\{\\alpha\}\_\{si\}Q\-masked concept attention weights,αs​i​k=softmaxk⁡\(es​i​k\+mi​k\)\\alpha\_\{sik\}=\\operatorname\{softmax\}\_\{k\}\(e\_\{sik\}\+m\_\{ik\}\)rs​ia​t​tr^\{att\}\_\{si\}Output of the concept residual layer,rs​ia​t​t=∑kαs​i​k​vs,kr^\{att\}\_\{si\}=\\sum\_\{k\}\\alpha\_\{sik\}v\_\{s,k\}These residual biases are usually initialized to 0 so that the model initially degenerates as much as possible to the base cognitive diagnosis backbone\. Low\-rank interaction vectors and attention vectors are initialized with small random values\. As training proceeds, if a certain type of student, item, or concept exhibits stable errors under the base backbone, gradients push the corresponding residual parameters away from 0, absorbing this part of cognitive bias into the ability residual pathway\.

#### 4\.3\.1Student\-Item Residual Layer

The student\-item residual layer is formulated as:

rs​is​t​u​\-​i​t​e​m=rss​t​u​d​e​n​t\+rii​t​e​m\+rs​ii​n​t​e​r​a​c​t​i​o​n,r^\{stu\\text\{\-\}item\}\_\{si\}=r^\{student\}\_\{s\}\+r^\{item\}\_\{i\}\+r^\{interaction\}\_\{si\},\(11\)
which absorbs overall offsets from students and items and models fine\-grained latent student\-item matching\. Specifically,

rss​t​u​d​e​n​t=bs,rii​t​e​m=bi,rs​ii​n​t​e​r​a​c​t​i​o​n=γS​I​⟨𝐮s,𝐯i⟩d\.r^\{student\}\_\{s\}=b\_\{s\},\\qquad r^\{item\}\_\{i\}=b\_\{i\},\\qquad r^\{interaction\}\_\{si\}=\\gamma\_\{SI\}\\frac\{\\langle\\mathbf\{u\}\_\{s\},\\mathbf\{v\}\_\{i\}\\rangle\}\{\\sqrt\{d\}\}\.\(12\)
wherebsb\_\{s\}andbib\_\{i\}correspond to the student global bias and item global bias in the table above,𝐮s\\mathbf\{u\}\_\{s\}and𝐯i\\mathbf\{v\}\_\{i\}are the low\-rank interaction vectors of the student and item, respectively,ddis the low\-rank dimension, andγS​I\\gamma\_\{SI\}is the strength coefficient of the low\-rank student\-item interaction residual, corresponding to student\_item\_rank\_scale\. WhenγS​I\\gamma\_\{SI\}is small, the low\-rank interaction term has a weak influence on the final residual\. WhenγS​I\\gamma\_\{SI\}is large, the model makes stronger use of personalized student\-item matching to correct predictions from the base backbone\. The main experiments useγS​I=1\.5\\gamma\_\{SI\}=1\.5by default\. To avoid the high parameter cost and overfitting risk of directly modeling the complete student\-item interaction matrix, we use the low\-rank decomposition𝐮s⊤​𝐯i\\mathbf\{u\}\_\{s\}^\{\\top\}\\mathbf\{v\}\_\{i\}to represent personalized residuals at the student\-item level\. This design reduces computational and storage overhead while imposing a low\-rank constraint on the interaction structure, thereby improving generalization\.

#### 4\.3\.2Concept Residual Layer

The concept residual layer is defined as:

rs​ic​o​n​c​e​p​t=rs​ia​t​t,r^\{concept\}\_\{si\}=r^\{att\}\_\{si\},\(13\)
and models systematic concept\-level bias and student\-concept personalized bias\. This layer introduces the Q\-matrix\-constrained concept residual attention mechanism to aggregate item\-related concepts through attention weighting, as detailed in Section 4\.4\. Specifically,

rs​ia​t​t=∑k=1Kαs​i​k​\(bk\+bs,k\),r^\{att\}\_\{si\}=\\sum\_\{k=1\}^\{K\}\\alpha\_\{sik\}\(b\_\{k\}\+b\_\{s,k\}\),\(14\)
wherebkb\_\{k\}is the shared concept residual bias of thekk\-th concept, representing systematic bias of that concept across all students and items;bs,kb\_\{s,k\}is the personalized residual bias of studentsson conceptkk, representing the student’s mastery deviation on that specific concept\. Their sum provides the concept\-layer value, which denotes the composite concept residual of the current student on that concept\.𝒦i=\{k:qi​k=1\}\\mathcal\{K\}\_\{i\}=\\\{k:q\_\{ik\}=1\\\}is the set of concepts involved in itemii, and the Q\-matrix mask zeroes out the attention weights of irrelevant concepts\.

This decomposition has clear interpretability\. Indeed, in the student\-item residual layer,bsb\_\{s\}andbib\_\{i\}absorb overall student and item offsets, respectively, and the low\-rank interaction termrs​ii​n​t​e​r​a​c​t​i​o​nr^\{interaction\}\_\{si\}reflect latent matching relations induced by item wording, problem\-solving strategy, or hidden skill requirements\. In the concept residual layer,bk\+bs,kb\_\{k\}\+b\_\{s,k\}explains systematic concept bias and student\-concept personalized mastery deviations\. This two\-layer architecture facilitates ablation analysis and allows direct statistics on the contributions of different residual sources to the final response probability\.

### 4\.4Q\-Matrix\-Constrained Concept Residual Attention

Next, we introduce a Q\-matrix\-constrained concept residual attention mechanism into the concept residual layerrs​ic​o​n​c​e​p​tr^\{concept\}\_\{si\}\. Among the concepts associated with itemii, the mechanism adaptively determines the relative contribution of each explicit concept residual to the concept\-level correction for studentss’s predicted response to itemii\. Specifically, a student–item\-pair\-conditioned query is constructed by combining a student\-specific and an item\-specific attention embedding\. Its compatibility with each candidate concept key provides the basis for residual allocation: a higher compatibility score assigns a larger relative weight to the corresponding concept residual\. The Q\-matrix further restricts this allocation to concepts required by the item\. It should be noted that attention is used only for concept\-layer aggregation, not for the student\-item residual layer\.

The proposed module is a structured residual aggregation rather than a conventional hidden\-state attention\[[30](https://arxiv.org/html/2609.21214#bib.bib30)\]\. It does not use attention to create a new hidden value\. Still, it leaves each scalarvs,kv\_\{s,k\}untransformed and uses attention only to determine its relative allocation\. The resulting correction remains directly decomposable into shared and student\-specific concept residuals\.

During implementation, we directly learn concept key vectors, whose role is equivalent to𝐤ka​t​t=𝐖k​𝐜k\\mathbf\{k\}^\{att\}\_\{k\}=\\mathbf\{W\}\_\{k\}\\mathbf\{c\}\_\{k\}in the theoretical expression\. Specifically, the model does not explicitly store𝐖k\\mathbf\{W\}\_\{k\}and𝐜k\\mathbf\{c\}\_\{k\}as two separate objects but combines their product into a learnable vector𝐤ka​t​t\\mathbf\{k\}^\{att\}\_\{k\}\. We use the symbols in the table above and describe the computation for a student–item pair\(s,i\)\(s,i\)\.

First, the query for the student–item pair is constructed from the student and the item:

𝐪s​ia​t​t=L​N​\(tanh⁡\(𝐞sa​t​t\+𝐯ia​t​t\)\),𝐪s​ia​t​t∈ℝda​t​t\.\\mathbf\{q\}^\{att\}\_\{si\}=LN\\left\(\\tanh\(\\mathbf\{e\}^\{att\}\_\{s\}\+\\mathbf\{v\}^\{att\}\_\{i\}\)\\right\),\\qquad\\mathbf\{q\}^\{att\}\_\{si\}\\in\\mathbb\{R\}^\{d\_\{att\}\}\.\(15\)
where𝐞sa​t​t\\mathbf\{e\}^\{att\}\_\{s\}and𝐯ia​t​t\\mathbf\{v\}^\{att\}\_\{i\}are obtained from student and item embedding tables, respectively, and both are vectors of lengthda​t​td\_\{att\}\.tanh⁡\(⋅\)\\tanh\(\\cdot\)denotes the hyperbolic tangent activation function, which compresses the combined student\-side and item\-side query information into a bounded nonlinear space\. Moreover,L​N​\(⋅\)LN\(\\cdot\)denotes Layer Normalization, which stabilizes the scale of the query vector and avoids unstable fluctuations in subsequent dot\-product scores caused by excessively large vector norms\. For each conceptkk, the model directly learns a key vector:

𝐤ka​t​t∈ℝda​t​t,𝐊a​t​t=\[𝐤1a​t​t;⋯;𝐤Ka​t​t\]∈ℝK×da​t​t\.\\mathbf\{k\}^\{att\}\_\{k\}\\in\\mathbb\{R\}^\{d\_\{att\}\},\\qquad\\mathbf\{K\}^\{att\}=\[\\mathbf\{k\}^\{att\}\_\{1\};\\cdots;\\mathbf\{k\}^\{att\}\_\{K\}\]\\in\\mathbb\{R\}^\{K\\times d\_\{att\}\}\.\(16\)
Therefore, for a student–item pair\(s,i\)\(s,i\), the unmasked attention score on conceptkkis defined as:

es​i​k=\(𝐪s​ia​t​t\)T​𝐤ka​t​tda​t​t,es​i​k∈ℝ\.e\_\{sik\}=\\frac\{\(\\mathbf\{q\}^\{att\}\_\{si\}\)^\{T\}\\mathbf\{k\}^\{att\}\_\{k\}\}\{\\sqrt\{d\_\{att\}\}\},\\qquad e\_\{sik\}\\in\\mathbb\{R\}\.\(17\)
By concatenating the scores of all concepts, vector𝐞s​i∈ℝK\\mathbf\{e\}\_\{si\}\\in\\mathbb\{R\}^\{K\}is obtained\. To ensure that attention is allocated only over item\-related concepts, the item\-level Q\-matrix mask𝐦i∈ℝK\\mathbf\{m\}\_\{i\}\\in\\mathbb\{R\}^\{K\}is introduced:

mi​k=\{0,qi​k=1,−∞,qi​k=0,m\_\{ik\}=\\begin\{cases\}0,&q\_\{ik\}=1,\\\\ \-\\infty,&q\_\{ik\}=0,\\end\{cases\}\(18\)
whereqi​kq\_\{ik\}is the 0/1 mark of itemiion conceptkkin the Q\-matrix\. The final attention weight is:

αs​i​k=softmaxk⁡\(es​i​k\+mi​k\),𝜶s​i∈ℝK\.\\alpha\_\{sik\}=\\operatorname\{softmax\}\_\{k\}\(e\_\{sik\}\+m\_\{ik\}\),\\qquad\\boldsymbol\{\\alpha\}\_\{si\}\\in\\mathbb\{R\}^\{K\}\.\(19\)
The concept residual value remains an interpretable scalar\. For conceptkk,vs,kv\_\{s,k\}is a scalar and for all concepts it forms the vector𝐯s=\[vs,1,…,vs,K\]∈ℝK\\mathbf\{v\}\_\{s\}=\[v\_\{s,1\},\\ldots,v\_\{s,K\}\]\\in\\mathbb\{R\}^\{K\}\. In the proposed solution, the shared concept residual vector𝐛∈ℝK\\mathbf\{b\}\\in\\mathbb\{R\}^\{K\}is added to the student\-concept residual vector𝐛sc​o​n​c​e​p​t∈ℝK\\mathbf\{b\}^\{concept\}\_\{s\}\\in\\mathbb\{R\}^\{K\}:

vs,k=bk\+bs,k,𝐯s∈ℝK\.v\_\{s,k\}=b\_\{k\}\+b\_\{s,k\},\\qquad\\mathbf\{v\}\_\{s\}\\in\\mathbb\{R\}^\{K\}\.\(20\)
The Q\-matrix defines the candidate set𝒦i\\mathcal\{K\}\_\{i\}, whereas the query\-key compatibility scores determine the relative allocation within that set\. Thus,αs​i​k\\alpha\_\{sik\}reflects student–item\-pair–concept compatibility rather than directly the magnitude ofvs,kv\_\{s,k\}, which is used only after the weights are obtained\. Therefore, it is interpreted as a residual\-allocation coefficient, and not as a mastery probability or causal importance score\.

Concept residual aggregation is a weighted sum along the concept dimension,KK:

rs​ia​t​t=∑k=1Kαs​i​k​vs,k,rs​ia​t​t∈ℝ\.r^\{att\}\_\{si\}=\\sum\_\{k=1\}^\{K\}\\alpha\_\{sik\}v\_\{s,k\},\\qquad r^\{att\}\_\{si\}\\in\\mathbb\{R\}\.\(21\)
For a single\-concept item, the masked softmax givesαs​i​k∗=1\\alpha\_\{sik^\{\*\}\}=1andrs​ia​t​t=vs,k∗r^\{att\}\_\{si\}=v\_\{s,k^\{\*\}\}, and adaptive allocation is activated only for multi\-concept items\.

During implementation, the shared concept term and the student\-concept personalized term can be separately summarized:

rs​is​h​a​r​e​d=∑k=1Kαs​i​k​bk,rs​is​t​u​d​e​n​t​\-​c​o​n​c​e​p​t=∑k=1Kαs​i​k​bs,k\.r^\{shared\}\_\{si\}=\\sum\_\{k=1\}^\{K\}\\alpha\_\{sik\}b\_\{k\},\\qquad r^\{student\\text\{\-\}concept\}\_\{si\}=\\sum\_\{k=1\}^\{K\}\\alpha\_\{sik\}b\_\{s,k\}\.\(22\)
Their sum gives the concept residual layerrs​ic​o​n​c​e​p​t=rs​ia​t​tr^\{concept\}\_\{si\}=r^\{att\}\_\{si\}, which is then combined with the student\-item residual layerrs​is​t​u​\-​i​t​e​mr^\{stu\\text\{\-\}item\}\_\{si\}to form the final ability residual\.

### 4\.5Affective Module

The affective module learns the non\-cognitive state of a student in the current interaction\. Given the student affective representation𝐚s\\mathbf\{a\}\_\{s\}and the item concept difficulty representation𝐝ik​c\\mathbf\{d\}^\{kc\}\_\{i\}, the interaction\-level affective representation is formulated as:

𝐚s​i=gθ​\(𝐚s,𝐝ik​c\),\\mathbf\{a\}\_\{si\}=g\_\{\\theta\}\(\\mathbf\{a\}\_\{s\},\\mathbf\{d\}^\{kc\}\_\{i\}\),\(23\)
where𝐚s\\mathbf\{a\}\_\{s\}is a learnable student\-level affective representation obtained by student ID lookup,𝐝ik​c\\mathbf\{d\}^\{kc\}\_\{i\}is the item concept difficulty representation in the base cognitive diagnosis backbone and provides the knowledge requirement context of the current item, andgθ​\(⋅\)g\_\{\\theta\}\(\\cdot\)is a learnable mapping function used to generate the affective representation of the current student\-item interaction\. These parameters are learned end\-to\-end through the training objective, rather than manually specified by external affective rules\.

Let guess and slip be directly determined by the affective representation:

gs​i=σ⁡\(𝐰gT​𝐚s​i\+bg\),g\_\{si\}=\\sigma\(\\mathbf\{w\}\_\{g\}^\{T\}\\mathbf\{a\}\_\{si\}\+b\_\{g\}\),\(24\)
ss​i=σ⁡\(𝐰s​l​i​pT​𝐚s​i\+bs​l​i​p\)\.s\_\{si\}=\\sigma\(\\mathbf\{w\}\_\{slip\}^\{T\}\\mathbf\{a\}\_\{si\}\+b\_\{slip\}\)\.\(25\)
where𝐰g\\mathbf\{w\}\_\{g\}andbgb\_\{g\}are the linear weight and bias of the guess branch, and𝐰s​l​i​p\\mathbf\{w\}\_\{slip\}andbs​l​i​pb\_\{slip\}are the linear weight and bias of the slip branch\. The termsbgb\_\{g\}andbs​l​i​pb\_\{slip\}belong only to the intercepts of the affective modulation branch\. They are not part of the student, item, or concept residuals in the ability residual module\.

The final response probability is presented in Eq\. \(26\) and indicates that whenys​i∗y^\{\*\}\_\{si\}is low, the guess term may increase the probability of a correct response; whenys​i∗y^\{\*\}\_\{si\}is high, the slip term may decrease the probability of a correct response\. The final probability is therefore jointly determined by the residual\-corrected prediction, the guess term, and the slip term\.

y^s​i=\(1−ss​i\)​ys​i∗\+gs​i​\(1−ys​i∗\)\.\\hat\{y\}\_\{si\}=\(1\-s\_\{si\}\)y^\{\*\}\_\{si\}\+g\_\{si\}\(1\-y^\{\*\}\_\{si\}\)\.\(26\)

### 4\.6Affective Representation Learning With and Without Affect Labels

Depending on whether real affect labels are available in the dataset, ACD follows two affect\-aware training strategies\. When the training data contain external affect annotations, supervised affect prediction is used\. When real affect annotations are unavailable, contrastive affect representation learning is used by constructing positive and negative samples based on response labels\. Letℬ\\mathcal\{B\}be the current training batch,𝐚ig​t∈\[0,1\]da​f​f\\mathbf\{a\}^\{gt\}\_\{i\}\\in\[0,1\]^\{d\_\{aff\}\}be the real affect vector of sampleii, and𝐚^i\\hat\{\\mathbf\{a\}\}\_\{i\}be the model’s affective representation\. Given that the original ACD treats affective states as continuous values, the supervised affect perception module uses mean squared error:

ℒa=1\|ℬ\|​∑i∈ℬ‖𝐚^i−𝐚ig​t‖22\.\\mathcal\{L\}\_\{a\}=\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\sum\_\{i\\in\\mathcal\{B\}\}\\left\\\|\\hat\{\\mathbf\{a\}\}\_\{i\}\-\\mathbf\{a\}^\{gt\}\_\{i\}\\right\\\|\_\{2\}^\{2\}\.\(27\)
This loss corresponds to the affective perception loss in ACD and makes the affective representation induced by student affective traits and item difficulty close to the external affect annotations\. In ARACD,𝐚^i\\hat\{\\mathbf\{a\}\}\_\{i\}still participates in final response prediction through guess/slip modulation\. Besides, the ability residual has no independent supervision label, as its parameters are updated jointly by backpropagation through the response prediction path and the affect prediction path\.

For CACD/ARCACD settings without real affect labels, this study follows the contrastive affect perception idea in the original ACD and replacesℒa\\mathcal\{L\}\_\{a\}with a contrastive loss\. Its basic assumption is that student\-item interactions with the same response outcome are more likely to correspond to similar affective states\. In contrast, interactions with different response outcomes are more likely to correspond to different affective states\. For any sampleiiin a batch, the positive and negative sample sets are constructed as:

𝒫i=\{j∈ℬ:j≠i,yj=yi\},𝒩i=\{k∈ℬ:yk≠yi\}\.\\mathcal\{P\}\_\{i\}=\\\{j\\in\\mathcal\{B\}:j\\neq i,\\ y\_\{j\}=y\_\{i\}\\\},\\qquad\\mathcal\{N\}\_\{i\}=\\\{k\\in\\mathcal\{B\}:y\_\{k\}\\neq y\_\{i\}\\\}\.\(28\)
whereyi∈\{0,1\}y\_\{i\}\\in\\\{0,1\\\}is the response label of sampleii\. Letsim⁡\(⋅,⋅\)\\operatorname\{sim\}\(\\cdot,\\cdot\)denote cosine similarity between two affective vectors andτ\\taube the temperature coefficient\. Following CACD, the contrastive affective loss is defined as:

ℒc​a=−1\|ℐ\|∑i∈ℐ1\|𝒫i\|∑j∈𝒫ilogexp⁡\(sim⁡\(𝐚^i,𝐚^j\)/τ\)∑k∈𝒩iexp⁡\(sim⁡\(𝐚^i,𝐚^k\)/τ\),\\mathcal\{L\}\_\{ca\}=\-\\frac\{1\}\{\|\\mathcal\{I\}\|\}\\sum\_\{i\\in\\mathcal\{I\}\}\\frac\{1\}\{\|\\mathcal\{P\}\_\{i\}\|\}\\sum\_\{j\\in\\mathcal\{P\}\_\{i\}\}\\log\\frac\{\\exp\\left\(\\operatorname\{sim\}\(\\hat\{\\mathbf\{a\}\}\_\{i\},\\hat\{\\mathbf\{a\}\}\_\{j\}\)/\\tau\\right\)\}\{\\sum\_\{k\\in\\mathcal\{N\}\_\{i\}\}\\exp\\left\(\\operatorname\{sim\}\(\\hat\{\\mathbf\{a\}\}\_\{i\},\\hat\{\\mathbf\{a\}\}\_\{k\}\)/\\tau\\right\)\},\(29\)
where

ℐ=\{i∈ℬ:\|𝒫i\|\>0,\|𝒩i\|\>0\}\\mathcal\{I\}=\\\{i\\in\\mathcal\{B\}:\|\\mathcal\{P\}\_\{i\}\|\>0,\\ \|\\mathcal\{N\}\_\{i\}\|\>0\\\}\(30\)
is the set of valid anchors\. This training strategy does not introduce real affect labels for individual samples\. Instead, it uses weak positive and negative relations formed by response labels to constrain the affective representation space, thereby providing latent affective cues for guess/slip modulation\.

### 4\.7Training Objective

The response prediction loss follows the binary cross\-entropy used in ACD\. For a batchℬ\\mathcal\{B\}, lety^i\\hat\{y\}\_\{i\}be the final predicted probability andyiy\_\{i\}the true response label\. The loss is formulated as:

ℒC​D​M=−1\|ℬ\|∑i∈ℬ\[yilogy^i\+\(1−yi\)log\(1−y^i\)\]\.\\mathcal\{L\}\_\{CDM\}=\-\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\sum\_\{i\\in\\mathcal\{B\}\}\\left\[y\_\{i\}\\log\\hat\{y\}\_\{i\}\+\(1\-y\_\{i\}\)\\log\(1\-\\hat\{y\}\_\{i\}\)\\right\]\.\(31\)
In supervised ACD/ARACD settings with real affect labels, response prediction loss and affect prediction loss are optimized jointly:

ℒs​u​p=ℒC​D​M\+λa​ℒa,\\mathcal\{L\}\_\{sup\}=\\mathcal\{L\}\_\{CDM\}\+\\lambda\_\{a\}\\mathcal\{L\}\_\{a\},\(32\)
whereλa\\lambda\_\{a\}is the trade\-off coefficient of the affect prediction loss\.ℒC​D​M\\mathcal\{L\}\_\{CDM\}constrains the final response probability, andℒa\\mathcal\{L\}\_\{a\}constrains the predicted affective vector to align with the external affect annotations\.

In CACD/ARCACD settings without real affect labels,ℒc​a\\mathcal\{L\}\_\{ca\}replaces the supervised affect MSE, and the training objective becomes:

ℒu​n​s​u​p=ℒC​D​M\+λc​ℒc​a,\\mathcal\{L\}\_\{unsup\}=\\mathcal\{L\}\_\{CDM\}\+\\lambda\_\{c\}\\mathcal\{L\}\_\{ca\},\(33\)
whereλc\\lambda\_\{c\}is the trade\-off coefficient of the contrastive affect perception loss\. This setting does not computeℒa\\mathcal\{L\}\_\{a\}\. The affective representation influences response prediction through the guess/slip branch and receives weak supervision from response labels throughℒc​a\\mathcal\{L\}\_\{ca\}\. Apart from the supervised affect loss or contrastive affect loss, this study does not introduce additional decorrelation losses or residual losses\. Necessary control of parameter scale is handled by optimizer weight decay or parameter clipping, but is not reported as an independent training objective\.

## 5Experimental Setup

### 5\.1Research Questions

The study is organized around the following research questions\.

RQ1: Can ability\-residual decoupling stably improve response prediction performance?This work compares Base, ACD/CACD, and ARACD/ARCACD on ASSIST2017, ASSIST2012, ASSIST2009, and Junyi, and examines whether ACC, AUC, RMSE, and MAE have consistent improvements across datasets and backbones\.

RQ2: Does the ability residual help affect alignment?ACD and ARACD are compared using Affect RMSE and Affect MAE to test whether explicitly modeling cognitive residuals reduces the explanatory pressure on the affective branch caused by item, concept, and ability bias\.

RQ3: Which residual components account for the benefit of ability residuals?Hierarchical ablations are conducted among Full, Only stu\-item residual, Only concept residual, and w/o ability residual to evaluate the contributions of the student\-item residual layer, the concept residual layer, and their combination to response prediction and affect prediction\.

RQ4: How do the attention dimension and low\-rank dimension affect response prediction and affect alignment?The concept attention dimensionda​t​td\_\{att\}and the student\-item low\-rank dimensionr​a​n​krankare modified to examine the sensitivity of response metrics and affect metrics\.

RQ5: Do the two residual layers capture response\-related structure rather than random perturbations?The student\-item residual and concept residual are visualized as a two\-dimensional residual space to examine whether they form response\-related patterns\.

RQ6: Can ability residuals mitigate affect contamination and improve model reliability?Synthetic contamination probes, affect representation leakage diagnostics, and PCA visualization are used to analyze whether affective representations mix in cognitive bias information, and whether ability residuals can transfer this bias to the ability\-residual pathway\.

RQ7: Can the proposed framework improve long\-tail student diagnosis?We further reanalyze low\-activity students to examine whether ability residuals improve diagnosis for students with few interactions\.

### 5\.2Datasets

The experiments use the educational datasets: ASSIST2017, ASSIST2012, ASSIST2009, and Junyi\. All contain student response records, item IDs, concept annotations, and response correctness labels\. ASSIST2017 and ASSIST2012 contain real affect labels and are used for supervised ACD/ARACD experiments\. ASSIST2009 and Junyi do not use real affect supervision and are used for CACD/ARCACD experiments without affect labels\. For each dataset, a fixed train/validation/test split ratio of 7:2:1 is used, while the data processing pipeline and evaluation protocol are consistent across all models\.

Table 2

Statistics of the four educational datasets\.

DatasetStudentsItemsConceptsASSIST201717093162102ASSIST20122899847124265ASSIST2009416317746123Junyi706272737722

### 5\.3Baseline Models and Compared Methods

Six representative models are used as base backbones: DINA, IRT, MIRT, NCDM, RCD, and SCD to examine whether the ability residual framework can be attached as a general plug\-in to different cognitive diagnosis backbones\. These six backbones cover discrete cognitive diagnosis, unidimensional and multidimensional IRT, neural cognitive diagnosis, relational structure modeling, and self\-supervised cognitive diagnosis\. Therefore, these models allow us to evaluate the stability of the proposed method under different modeling assumptions and expressive capacities\. In addition to base cognitive diagnosis backbones, ACD and CACD are used as affect\-aware comparison methods to distinguish between ”introducing an affective module” and ”further introducing ability residuals”\.

Table 3

Cognitive diagnosis backbones used in the experiments\.

BackboneDescriptionDINAA discrete cognitive diagnosis model that assumes students need to master all concepts required by an item to answer it correctly with high probability, and models guessing and slipping through guess/slip parameters\.IRTA unidimensional item response theory model that estimates response probability based on the matching relation between student ability and item parameters, serving as a classical psychometric baseline\.MIRTA multidimensional item response theory model that extends student ability into a multidimensional latent vector to characterize ability differences in multi\-skill or multi\-concept assessments\.NCDMA neural cognitive diagnosis model that uses neural networks to fit response probabilities based on interpretable variables such as student mastery, item concept difficulty, and item discrimination\.RCDA relation\-aware cognitive diagnosis model that enhances interaction representations through relation structures among students, items, and concepts, and is used to evaluate the generalization of ability residuals on structured backbones\.SCDA self\-supervised cognitive diagnosis model that introduces self\-supervised signals into diagnostic representations to improve robustness under sparse response scenarios\.
Table 4

Affect\-aware comparison methods\.

MethodDescriptionACDA supervised affective cognitive diagnosis model that adds an affect\-aware module to the base cognitive diagnosis backbone, uses real affect labels to constrain affective representations, and modulates response probability through guess/slip\.CACDA contrastive affective cognitive diagnosis model for settings without real affect labels; it uses response labels to construct positive and negative sample relations for learning response\-derived affective representations and modulates response prediction through these representations\.
On ASSIST2017 and ASSIST2012, where real affect labels are available, we compare three methods for each backbone: Base uses only the base cognitive diagnosis backbone without an affective module or ability residual; ACD adds an affect\-aware module to the same backbone and performs supervised affect prediction with real affect labels, but does not explicitly model ability residuals; ARACD adds the ability residual module and Q\-matrix\-constrained concept residual attention on top of ACD\. This setting tests whether ability residuals can further improve response prediction and affect prediction beyond affect supervision\.

On ASSIST2009 and Junyi, where real affect labels are not used, we compare Base, CACD, and ARCACD\. CACD does not use per\-sample real affect labels but constructs positive and negative sample relations from response labels for contrastive affective representation learning\. ARCACD further introduces the ability residual module into CACD’s contrastive affective representation learning framework\. This setting tests whether ability residuals can still combine with response\-derived affective representations and improve response prediction when external affect annotations are absent\.

### 5\.4Evaluation Metrics and Settings

The main experiments use two types of metrics: response prediction metrics, including ACC, AUC, RMSE, and MAE, and affect alignment metrics, used only in supervised ACD/ARACD settings with real affect labels, including Affect RMSE and Affect MAE\. For the former type, ACC denotes prediction accuracy after thresholding at 0\.5; AUC measures whether the model assigns higher probabilities to correct responses; and RMSE and MAE measure probability error against the binary response label, with lower values being better\.

All metrics measure the error between the predicted affective vectors and real affect labels, and all methods use the same data preprocessing, train/validation/test split, and evaluation scripts to ensure comparability\. ASSIST2017 and ASSIST2012 use supervised affective learning, optimizing response prediction loss and real affect prediction loss\. Additionally, ASSIST2009 and Junyi do not use real affect labels and adopt response\-label\-based contrastive affective representation learning\. Except for hyperparameter sensitivity experiments, the main experiments keep the same training configuration within each dataset and backbone, select model checkpoints according to validation performance, and report final metrics on the test set\. The experimental setup comprises a single NVIDIA GeForce RTX 3070 Ti GPU with 8 GB of memory\. For model structural parameters, the main experiments use a concept attention dimensionda​t​t=16d\_\{att\}=16, a student\-item low\-rank interaction dimensionr​a​n​k=1024rank=1024, and a low\-rank student\-item residual strength coefficientγS​I=1\.5\\gamma\_\{SI\}=1\.5by default\. All models are optimized with Adam using a learning rate of5×10−45\\times 10^\{\-4\}\. The supervised experiments on ASSIST2017 and ASSIST2012 use a batch size of 512, train for at most 35 epochs with early\-stopping patience values of 5 and 2, respectively, and optimizeℒC​D​M\+ℒa\\mathcal\{L\}\_\{CDM\}\+\\mathcal\{L\}\_\{a\}\. The contrastive experiments on ASSIST2009 and Junyi optimizeℒC​D​M\+ℒc​a\\mathcal\{L\}\_\{CDM\}\+\\mathcal\{L\}\_\{ca\}, use a batch size of 256 and patience 5, and train for at most 20 and 30 epochs, respectively; RCD and SCD on Junyi use a batch size of 64 because of memory demand\.

## 6Experimental Results and Analysis

This section reports response prediction, affect alignment, ablation, hyperparameter sensitivity, residual maps, contamination diagnostics, long\-tail results, and a student\-level case study according to the research questions\. Sections 6\.1–6\.7 correspond to RQ1–RQ7, respectively, while Section 6\.8 provides supplementary student\-level evidence\.

### 6\.1Response Prediction Results \(RQ1\)

This section answers RQ1, namely whether ability\-residual decoupling can stably improve response prediction performance under multi\-dataset and multi\-backbone settings\. Table 5 reports the supervised affective learning results on ASSIST2017 and ASSIST2012\. For each backbone, Base, ACD, and ARACD are retained, and four response metrics are reported: ACC, AUC, RMSE, and MAE\. Table 6 lists the results without real affect labels on ASSIST2009 and Junyi, comparing Base, CACD, and ARCACD to test whether ability residuals can combine with response\-label contrastive affective representation learning and bring stable gains\.

Table 5:Response prediction results on ASSIST2017 and ASSIST2012\.BackboneMethodASSIST2017ASSIST2012ACC↑\\uparrowAUC↑\\uparrowRMSE↓\\downarrowMAE↓\\downarrowACC↑\\uparrowAUC↑\\uparrowRMSE↓\\downarrowMAE↓\\downarrowDINABase0\.64710\.69500\.46810\.43830\.71180\.69230\.44010\.3654ACD0\.71030\.77600\.43740\.40080\.73520\.73970\.42340\.3489ARACD0\.73900\.80790\.42020\.34490\.74940\.76090\.41390\.3340IRTBase0\.65820\.72240\.46680\.40860\.72860\.72420\.42890\.3440ACD0\.71520\.78430\.43380\.37320\.74010\.75180\.41940\.3354ARACD0\.73940\.80820\.42010\.34420\.74890\.76020\.41420\.3334MIRTBase0\.68010\.73990\.46610\.43210\.73550\.72320\.44670\.3788ACD0\.72320\.79350\.42890\.38210\.74130\.75380\.41860\.3453ARACD0\.73910\.80780\.42020\.34550\.74900\.76060\.41400\.3341NCDMBase0\.69080\.75200\.45270\.41910\.73970\.74510\.42180\.3344ACD0\.72300\.79380\.42960\.40180\.74210\.75020\.41850\.3336ARACD0\.73940\.80780\.42030\.34610\.74870\.76170\.41380\.3316RCDBase0\.71410\.77950\.43540\.37130\.74010\.74770\.42140\.3373ACD0\.72190\.79150\.42960\.36450\.74450\.75700\.41750\.3381ARACD0\.73640\.80560\.42100\.35760\.74990\.76110\.41310\.3328SCDBase0\.71460\.78020\.43500\.36760\.74150\.75490\.41880\.3371ACD0\.72590\.79300\.42890\.36080\.74600\.76130\.41530\.3305ARACD0\.73570\.80550\.42120\.35980\.74970\.76180\.41320\.3301

Table 5 highlights that ARACD provides consistently improved response prediction across the two supervised datasets\. Compared with Base, ARACD improves ACC and AUC and reduces RMSE and MAE for every reported backbone\-dataset pair\. Compared with ACD, ARACD yields higher ACC and AUC and lower RMSE and MAE in all 12 comparisons\. These patterns suggest that ability\-residual modeling can complement supervised affect modeling by accounting for part of the remaining cognitive residual variation\. However, the magnitude of the gain still depends on the backbone and dataset\.

We further run CACD and ARCACD on ASSIST2009 and Junyi under the same six\-backbone setting to supplement cross\-dataset evidence\. Table 6 reports ACC, AUC, RMSE, and MAE for the two datasets side by side\. For ASSIST2009, Base results for each backbone are included as response prediction controls without affective modules or ability residuals\.

Table 6:Response prediction results on ASSIST2009 and Junyi\.BackboneMethodASSIST2009JunyiACC↑\\uparrowAUC↑\\uparrowRMSE↓\\downarrowMAE↓\\downarrowACC↑\\uparrowAUC↑\\uparrowRMSE↓\\downarrowMAE↓\\downarrowDINABase0\.70150\.72060\.45900\.44730\.72810\.71150\.43340\.3866CACD0\.72400\.75130\.43440\.39350\.74560\.75390\.41000\.3575ARCACD0\.77860\.83960\.38960\.27990\.76580\.75820\.40130\.3298IRTBase0\.71750\.73540\.43910\.39790\.73350\.72070\.41130\.3554CACD0\.72550\.75710\.42890\.35810\.76860\.75830\.40090\.3284ARCACD0\.78190\.84280\.38580\.28870\.76920\.75910\.40010\.3251MIRTBase0\.68940\.69860\.45740\.44140\.73550\.72430\.42150\.3877CACD0\.72420\.74990\.43130\.38100\.75800\.75370\.40500\.3436ARCACD0\.78070\.84150\.38820\.27860\.76830\.75970\.40040\.3271NCDMBase0\.71880\.75130\.43580\.39210\.72620\.71620\.44530\.4001CACD0\.72530\.75730\.43320\.38450\.73960\.72770\.42890\.3705ARCACD0\.78840\.84870\.38230\.27400\.76820\.75960\.40070\.3338RCDBase0\.72490\.76180\.42710\.35180\.75950\.75890\.40520\.3327CACD0\.72840\.75840\.42910\.36700\.77060\.76310\.39860\.3174ARCACD0\.77870\.83840\.39200\.27470\.77150\.76340\.39710\.3168SCDBase0\.72570\.76210\.42630\.35940\.75700\.74990\.40970\.3572CACD0\.72730\.75990\.43200\.33580\.77020\.75700\.40450\.3506ARCACD0\.77820\.83870\.39560\.26320\.77040\.75830\.40430\.3501

Table 6 reveals that on ASSIST2009, ARCACD achieves higher ACC/AUC and lower RMSE/MAE than both Base and CACD on all six backbones\. In particular, NCDM\+ARCACD reaches ACC/AUC of 0\.7884/0\.8487 under the targeted configuration with fixedda​t​t=16,r​a​n​k=1024d\_\{att\}=16,rank=1024\. In supplementary runs that combine ability residuals with the CACD contrastive loss, DINA, IRT, MIRT, RCD, and SCD obtain AUC values of approximately 0\.838–0\.843\. On the Junyi subset, ARCACD also yields consistent improvements on all four response metrics across six backbones\. Overall, in the absence of real affect labels, joint modeling of ability residuals and contrastive affective representations yields consistent response\-prediction gains across the reported datasets and backbones\.

### 6\.2Affect Alignment Analysis \(RQ2\)

This section answers RQ2, namely whether explicit ability residuals can help the affective branch better align with real affect when real affect labels are available\. Table 7 reports affect metrics and presents ASSIST2017 and ASSIST2012 side by side\. Since Base does not contain an affective branch, the table includes only ACD and ARACD trained with real affect label supervision\. The metrics are Affect RMSE and Affect MAE\.

Table 7:Affect alignment results under real affect labels\.BackboneMethodASSIST2017ASSIST2012Affect RMSE↓\\downarrowAffect MAE↓\\downarrowAffect RMSE↓\\downarrowAffect MAE↓\\downarrowDINAACD0\.26140\.19140\.21840\.1502ARACD0\.19900\.14890\.19370\.1239IRTACD0\.24610\.17850\.24530\.1792ARACD0\.19250\.13890\.19190\.1179MIRTACD0\.26400\.19000\.24950\.1814ARACD0\.19220\.13900\.19130\.1203NCDMACD0\.22660\.16400\.24220\.1703ARACD0\.19440\.14090\.19220\.1196RCDACD0\.20550\.14410\.19660\.1245ARACD0\.20380\.15710\.19420\.1213SCDACD0\.20540\.14360\.19960\.1269ARACD0\.20370\.15320\.19450\.1246

Table 7 provides generally supportive but non\-uniform evidence for improved affect alignment after introducing ability residuals\. On ASSIST2012, ARACD obtains lower Affect RMSE and Affect MAE than the corresponding ACD on all six backbones\. On ASSIST2017, ARACD lowers both affect metrics for DINA, IRT, MIRT, and NCDM, and reduces Affect RMSE for RCD and SCD, although Affect MAE increases on these two backbones\. These findings highlight that ability residuals reduce the pressure on the affective branch to absorb cognitive bias in many settings\. Still, affect\-alignment gains should not be interpreted as guaranteed for every backbone and metric\. This result also indicates that similar response metrics do not necessarily imply affective representation alignment; therefore, later contamination diagnostics separately analyze whether affective representations absorb cognitive residuals\.

### 6\.3Ablation Study \(RQ3\)

This section answers RQ3 by comparing the full model with three variants: Only stu\-item residual, Only concept residual, and w/o ability residual\. ASSIST2012/ASSIST2017 use supervised ARACD training, whereas ASSIST2009/Junyi use ARCACD training without affect labels\. All variants are evaluated under the same data split and model\-selection protocol within each dataset\.

Table 8:Hierarchical ablation of supervised ability residual layers on ASSIST2017 and ASSIST2012\.VariantASSIST2017ASSIST2012ACC↑\\uparrowAUC↑\\uparrowRMSE↓\\downarrowMAE↓\\downarrowAffect RMSE↓\\downarrowAffect MAE↓\\downarrowACC↑\\uparrowAUC↑\\uparrowRMSE↓\\downarrowMAE↓\\downarrowAffect RMSE↓\\downarrowAffect MAE↓\\downarrowFull ARACD0\.73940\.80780\.42030\.34610\.19440\.14090\.74870\.76170\.41380\.33160\.19220\.1196Only stu\-item residual0\.73850\.80510\.42190\.34940\.19880\.14560\.74440\.75370\.41700\.33280\.19420\.1210Only concept residual0\.72880\.79940\.42860\.38390\.20820\.14840\.74570\.75260\.41560\.33210\.19730\.1258w/o ability residual0\.72300\.79380\.42960\.40180\.22660\.16400\.74210\.75020\.41850\.33360\.24220\.1703

Table 8 reveals that in the supervised setting, the complete model achieves the best result on every reported response and affect metric across ASSIST2017 and ASSIST2012\. The student\-item\-only and concept\-only variants are consistently weaker than the complete model, indicating that the two residual layers provide complementary information\.

Table 9:Hierarchical ablation of ability residual layers on ASSIST2009 and Junyi\.VariantASSIST2009JunyiACC↑\\uparrowAUC↑\\uparrowRMSE↓\\downarrowMAE↓\\downarrowACC↑\\uparrowAUC↑\\uparrowRMSE↓\\downarrowMAE↓\\downarrowFull ARCACD0\.78840\.84870\.38230\.27400\.76820\.75960\.40070\.3232Only stu\-item residual0\.78040\.83880\.38720\.29090\.75990\.75170\.40450\.3354Only concept residual0\.73330\.78720\.41160\.33880\.76360\.74780\.40320\.3338w/o ability residual0\.72530\.75730\.43320\.38450\.73960\.72770\.42890\.3705

Accordingly, Table 9 reveals a similar pattern in the setting without affect labels\. On both ASSIST2009 and Junyi, the full model achieves the best ACC/AUC and the lowest RMSE/MAE among all variants\. The student\-item layer explains most of the gain, while the concept layer provides complementary concept\-related corrections when combined with it\.

### 6\.4Hyperparameter Analysis \(RQ4\)

This section answers RQ4 by examining two capacity parameters on NCDM: the attention dimensionda​t​td\_\{att\}and the student\-item low\-rank dimensionr​a​n​krank\. ASSIST2017 uses supervised affect prediction, while ASSIST2009 uses contrastive ARCACD without affect labels\. Figures 3 and 4 report the complete sensitivity comparisons: Figure 3 uses fixedr​a​n​k=1024rank=1024and sweepsda​t​t=4,8,16,32d\_\{att\}=4,8,16,32, while Figure 4 uses fixedda​t​t=16d\_\{att\}=16and sweepsr​a​n​krankfrom 4 to 2048\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_fixed_rank1024_response_assist2017_readable.png)

\(a\) ASSIST2017 response

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_fixed_rank1024_response_assist2009_readable.png)

\(b\) ASSIST2009 response

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_fixed_rank1024_affect_assist2017_readable.png)

\(c\) ASSIST2017 affect

Figure 3:Sensitivity to the concept\-attention dimensionda​t​td\_\{att\}with fixedr​a​n​k=1024rank=1024\. Panels \(a\) and \(b\) report ACC/AUC and RMSE/MAE; panel \(c\) reports Affect RMSE/MAE\.Figure 3 demonstrates thatda​t​td\_\{att\}has a limited but consistent effect oncer​a​n​k=1024rank=1024is fixed\. Both datasets obtain their best or near\-best response metrics aroundda​t​t=16d\_\{att\}=16, while increasing it to 32 introduces no stable gain\. Thus, concept residual attention only needs moderate capacity when the student\-item residual layer is sufficiently expressive\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_fixed_att16_response_assist2017_readable.png)

\(a\) ASSIST2017 response

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_fixed_att16_response_assist2009_readable.png)

\(b\) ASSIST2009 response

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_fixed_att16_affect_assist2017_readable.png)

\(c\) ASSIST2017 affect

Figure 4:Sensitivity to the student\-item low\-rank dimensionr​a​n​krankwith fixedda​t​t=16d\_\{att\}=16\. Panels \(a\) and \(b\) report ACC/AUC and RMSE/MAE; panel \(c\) reports Affect RMSE/MAE\.Figure 4 highlights thatr​a​n​krankis the more important capacity source\. Increasingr​a​n​krankfrom 4 to 1024 markedly improves ASSIST2009 response metrics and mildly improves ASSIST2017, whiler​a​n​k=2048rank=2048yields limited additional benefit and can increase some errors\. Affect prediction on ASSIST2017 follows the same pattern, with lower error aroundr​a​n​k=1024rank=1024\. Overall,r​a​n​k=1024rank=1024provides a practical balance between response accuracy, affect alignment, and capacity saturation\.

### 6\.5Residual Map Analysis \(RQ5\)

This section answers RQ5 by visualizing the residual space usingrs​is​t​u​\-​i​t​e​mr^\{stu\\text\{\-\}item\}\_\{si\}andrs​ic​o​n​c​e​p​tr^\{concept\}\_\{si\}as two\-dimensional coordinates\. Figure 5 presents the sampled test interactions from all four datasets, colored by true response labels\. The residual spaces exhibit response\-related directions rather than random perturbations: ASSIST2017 and ASSIST2009 are dominated by the student\-item residual direction, ASSIST2012 shows a smoother but still visible offset, and Junyi uses both residual types\. This provides representation\-level evidence that the two residual layers absorb complementary cognitive bias\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_aracd_arcacd_stuitem_concept_residual_map_4datasets_att4_rank1024_10k_2012more.png)Figure 5:Ability\-residual interaction maps for ASSIST2017, ASSIST2012, ASSIST2009, and Junyi\.
### 6\.6Affect Contamination Diagnostics \(RQ6\)

This section answers RQ6 by testing whether ability residuals reduce cognitive leakage into affective representations\. Since response metrics alone cannot determine whether the affective branch has learned a semantically clean non\-cognitive state, we construct a synthetic contamination probe\. A known bias term is added to the response\-generation process, composed of student, item, and concept effects, and ARACD is trained withda​t​t=16d\_\{att\}=16andr​a​n​k=1024rank=1024without exposing the bias term\. Then we test whether this term is easier to recover from the affective branch or the ability residual branch\. Three diagnostics are reported: mean abs corr\(a​f​f​e​c​t,b​i​a​s\)\(affect,bias\), affect probeR2R^\{2\}, and mean abs corr\(ability residual,bias\)\.

Table 10

Affect contamination leakage diagnostics on ASSIST2017\.

VariantAffect RMSE↓\\downarrowAffect MAE↓\\downarrowMean abs corr\(affect,bias\)↓\\downarrowAffect probeR2R^\{2\}↓\\downarrowMean abs corr\(residual,bias\)↑\\uparrowFull0\.19350\.13830\.06650\.05680\.6890w/o A residual0\.22980\.16640\.58780\.4716n/a

Table 10 and Figure 6 provide evidence consistent with a redistribution of synthetic cognitive bias away from the affective branch\. Without ability residuals, the mean abs corr\(a​f​f​e​c​t,b​i​a​s\)\(affect,bias\)and affect probeR2R^\{2\}reach 0\.5878 and 0\.4716, respectively\. With ability residuals, they drop to 0\.0665 and 0\.0568, while residual\-bias correlation reaches 0\.6890, and Affect RMSE/MAE also decrease\. These diagnostics combined suggest that more of the synthetic contamination signal is represented in the ability\-residual pathway rather than in the affective representation\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_supervised_leakage_probe.png)Figure 6:Contamination leakage probe of the affective branch\.As a complementary view, we project supervised ARACD affect representations from ASSIST2017 into a PCA space and color them by response label and ability residual magnitude\. Ideally, the affective representation should preserve response\-relevant non\-cognitive variation needed for guess/slip modulation, while not being primarily structured by the magnitude of ability residuals\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_supervised_affect_pca_residual_response.png)Figure 7:PCA visualization of supervised ARACD affective representations on ASSIST2017, colored by ability\-residual group \(left\) and response label \(right\)\.Table 11:Absolute correlations between supervised ARACD affect dimensions and residuals/response labels\.Affect dimAbs corr\. withabs residual↓\\downarrowAbs corr\. withresponse↑\\uparrow00\.11670\.211410\.09580\.208120\.07850\.193330\.00520\.0011

Figure 7 and Table 11 support this expectation, as the affect dimensions correlate only weakly with absolute ability residuals \(at most 0\.1167\), while the first three dimensions show clearer response\-label correlations\. This does not imply a fully isolated affect space, since affect still participates in response prediction\. Still, it indicates that ability residuals reduce the tendency of the affect branch to encode stable cognitive bias\.

### 6\.7Long\-Tail Student Results \(RQ7\)

This section answers RQ7 by evaluating low\-activity students, defined by the lower tertile of training interactions\. These students are challenging cases because sparse records provide limited evidence for estimating mastery states, affective representations, and personalized residuals\. In this scenario, we compare ACD/RCD/SCD/ARACD on ASSIST2017 and CACD/RCD/SCD/ARCACD on ASSIST2009, using ACC, AUC, RMSE, ECE, and Brier\. ECE measures the confidence\-accuracy gap over probability bins\[[35](https://arxiv.org/html/2609.21214#bib.bib35)\], while Brier score measures the mean squared error between predicted probability and the binary label\[[36](https://arxiv.org/html/2609.21214#bib.bib36)\]; lower values are better for both\.

Table 12

Results for long\-tail students\.

DatasetMethodLow\-activity ACC↑\\uparrowLow\-activity AUC↑\\uparrowLow\-activity RMSE↓\\downarrowLow\-activity ECE↓\\downarrowLow\-activity Brier↓\\downarrowASSIST2017ACD0\.69150\.74420\.45470\.05280\.2067ASSIST2017RCD0\.69220\.74780\.44790\.02630\.2007ASSIST2017SCD0\.68870\.74660\.44810\.02090\.2008ASSIST2017ARACD0\.70010\.76500\.44050\.01750\.1941ASSIST2009CACD0\.65500\.73650\.45960\.07720\.2112ASSIST2009RCD0\.70720\.74970\.44040\.03030\.1939ASSIST2009SCD0\.70950\.74890\.44130\.03940\.1948ASSIST2009ARCACD0\.70970\.75330\.43980\.03810\.1934

Table 12 reveals that ARACD performs best on all reported low\-activity metrics in ASSIST2017\. On ASSIST2009, ARCACD achieves the highest AUC and the lowest RMSE/Brier, while ACC and ECE remain close to strong RCD/SCD baselines\. These results indicate that ability residuals are especially helpful for ranking and probability\-error reduction in sparse student groups\.

### 6\.8Case Study

We further analyze student S453 from ASSIST2017 to illustrate how ability residuals correct individual response bias\. This student has 388 training interactions and 109 test interactions\. After retraining the independent Base NCDM used in the case analysis, Base NCDM obtains ACC/MAE/AUC of 0\.6606/0\.4371/0\.7361 on this student\. Notably, ACD still leaves several systematic probability errors for this student, with ACC/MAE/AUC of 0\.6330/0\.4729/0\.7031, while ARACD raises ACC to 0\.7523, MAE reduces to 0\.3529, and AUC improves to 0\.7571, correcting 16 ACD errors while introducing only 3 new errors\. This suggests that ability residuals provide directional personalized calibration rather than uniform probability shifts\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_student_level_case_study_s453_legend.png)

Legend

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_student_level_case_study_s453_panel1.png)

\(a\) Probability trajectories

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_student_level_case_study_s453_panel2.png)

\(b\) ARACD error reduction over ACD

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_student_level_case_study_s453_panel3.png)

\(c\) Residual decomposition

Figure 8:Student\-level case study for S453\. Panels \(a\)–\(c\) show probability trajectories, ARACD error reduction over ACD, and residual decomposition\.Table 13

Representative interactions and residual decomposition for student S453\.

AttemptItemConceptsyyBase probACD probARACD probrs​t​u​\-​i​t​e​mr^\{stu\\text\{\-\}item\}rc​o​n​c​e​p​tr^\{concept\}6011803810\.47870\.49960\.91671\.0712\-0\.16318014731510\.72240\.48620\.88080\.57180\.0934142263510\.36780\.49450\.88760\.8574\-0\.1367621354600\.28500\.52480\.2414\-0\.3903\-0\.138029401400\.19930\.43670\.1582\-0\.6087\-0\.2034611713800\.46210\.52440\.3337\-0\.2446\-0\.1631

Table 13 presents representative corrected interactions\. Fory=1y=1, positive student\-item residuals raise ACD probabilities near 0\.49 to 0\.8808–0\.9167; fory=0y=0, negative residuals lower probabilities to 0\.1582–0\.3337\. The opposite residual directions across labels show that ARACD performs bidirectional student\-item calibration\.

![Refer to caption](https://arxiv.org/html/2609.21214v1/fig_student_level_affect_summary_s453.png)Figure 9:Student\-level affect analysis for S453\. Mean values of the four observed affect dimensions are compared with ACD and ARACD predictions\.Table 14

Affect alignment statistics for student S453\.

Affect dimMean trueACD predARACD predACD MAE↓\\downarrowARACD MAE↓\\downarrowA10\.21330\.12930\.24300\.11070\.0966A20\.69410\.74860\.69060\.11380\.1092A30\.05910\.07010\.09240\.10830\.1295A40\.06510\.21010\.12020\.22510\.1507

Figure 9 and Table 14 further demonstrate that ARACD preserves affect alignment for S453: overall affect MAE decreases from 0\.1395 to 0\.1215, with improvements on A1, A2, and A4\. Thus, the response improvement does not come at the expense of affect prediction; ability residuals absorb stable student\-item offsets while reducing pressure on the affective representation\. Overall, this case study shows that ability residuals can correct individualized response bias while keeping affect predictions aligned with external labels\.

## 7Conclusion

This paper studies the affect contamination problem in affective cognitive diagnosis\. Unlike approaches that treat the affective module mainly as a compensator for backbone prediction errors, this study focuses on the base cognitive diagnosis backbone, which may still leave stable cognitive bias unmodeled\. Indeed, if such bias is not explicitly modeled, affective representations may absorb information about items, concepts, student ability, or student\-item matching, making them less specific to non\-cognitive states\. Motivated by this observation, this work introduces an ability\-residual decoupling framework that assigns stable cognitive residuals first to the student\-item residual layer and the concept residual layer, and then to the affective module model guess/slip modulation\. This provides a structured separation between cognitive residuals and non\-cognitive factors\.

Experimental results support the proposed approach from multiple perspectives\. Specifically, the main experiments highlight that the ability\-residual framework can be combined with supervised ACD \(with real affect labels\) or contrastive CACD \(without real affect labels\), yielding response\-prediction gains across all tested backbones in the reported main comparisons\. On datasets with real affect labels, ARACD generally improves affect alignment, although not every backbone and metric benefits\. Hierarchical ablation further suggests that the student\-item residual layer is the main source of personalized response correction, while the concept residual layer complements concept\-related bias\. Furthermore, the experiments reveal that Q\-matrix\-constrained attention can aggregate relevant concept residuals in an interpretable way\. Contamination leakage diagnostics reveal that removing ability residuals substantially increases the correlation and linear decodability between affective representations and synthetic cognitive bias\. In contrast, the complete model transfers more of this stable bias to the ability residual pathway\. PCA visualization, long\-tail student reanalysis, and student\-level cases further indicate that the framework not only improves overall prediction performance but also gives a more explicit separation of modeling roles in representation organization, low\-interaction student diagnosis, and student\-level correction\.

Overall, this paper concludes that affective cognitive diagnosis should not focus only on improving response prediction but should also examine whether affective representations contain cognitive residuals that do not belong to affective states\. Ability\-residual decoupling provides a structured method for this problem, allowing the model to explain stable cognitive residual variation and model non\-cognitive modulation, thereby improving the reliability, interpretability, and response prediction stability of affective diagnosis\. Future work may further incorporate real multimodal affect observations and cross\-platform educational data to evaluate the generalizability and limits of ability\-residual decoupling in more complex learning scenarios\.

## Statements and Declarations

#### Funding

No funding was received for conducting this study\.

#### Competing interests

The authors have no relevant financial or non\-financial interests to disclose\.

#### Ethics approval and consent to participate

Not applicable\. This study used publicly available, de\-identified educational datasets and involved no new data collection from human participants\.

#### Consent for publication

Not applicable\.

#### Data availability

The public datasets analysed in this study \(ASSIST2017, ASSIST2012, ASSIST2009, and Junyi\) are available from their original public repositories\. The source code is available at[https://github\.com/psychosiwa/ARACD](https://github.com/psychosiwa/ARACD)\. Additional processed data and experimental outputs are available from the corresponding author on reasonable request\.

#### Materials availability

Not applicable\.

#### Code availability

#### Author contributions

Boyuan Zhao: methodology, software, validation, formal analysis, investigation, data curation, visualization, and writing–original draft\. Meng Ye: conceptualization, supervision, project administration, and writing–review and editing\. Both authors read and approved the final manuscript\.

## References

- \(1\)B\. W\. Junker and K\. Sijtsma, ”Cognitive assessment models with few assumptions, and connections with nonparametric item response theory,”*Applied Psychological Measurement*, vol\. 25, no\. 3, pp\. 258\-272, 2001\.
- \(2\)J\. P\. Leighton and M\. J\. Gierl,*Cognitive Diagnostic Assessment for Education: Theory and Applications*\. Cambridge, UK: Cambridge University Press, 2007\.
- \(3\)K\. VanLehn, ”The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems,”*Educational Psychologist*, vol\. 46, no\. 4, pp\. 197\-221, 2011\.
- \(4\)M\. Feng, N\. Heffernan, and K\. R\. Koedinger, ”Addressing the assessment challenge with an online system that tutors as it assesses,”*User Modeling and User\-Adapted Interaction*, vol\. 19, no\. 3, pp\. 243\-266, 2009\.
- \(5\)F\. Wang, Q\. Liu, E\. Chen, Z\. Huang, Y\. Chen, Y\. Yin, Z\. Huang, and S\. Wang, ”Neural cognitive diagnosis for intelligent education systems,” in*Proc\. AAAI Conference on Artificial Intelligence*, 2020, pp\. 6153\-6161\.
- \(6\)C\. Piech, J\. Bassen, J\. Huang, S\. Ganguli, M\. Sahami, L\. J\. Guibas, and J\. Sohl\-Dickstein, ”Deep knowledge tracing,” in*Advances in Neural Information Processing Systems*, vol\. 28, 2015\.
- \(7\)Y\. Su, S\. Shen, L\. Zhu, L\. Wu, Z\. Huang, Z\. Cheng, Q\. Liu, and S\. Wang, ”Global and local neural cognitive modeling for student performance prediction,”*Expert Systems with Applications*, vol\. 237, Art\. no\. 121637, 2024\.
- \(8\)J\. de la Torre, ”DINA model and parameter estimation: A didactic,”*Journal of Educational and Behavioral Statistics*, vol\. 34, no\. 1, pp\. 115\-130, 2009\.
- \(9\)F\. M\. Lord,*Applications of Item Response Theory to Practical Testing Problems*\. Hillsdale, NJ, USA: Lawrence Erlbaum Associates, 1980\.
- \(10\)M\. D\. Reckase,*Multidimensional Item Response Theory*\. New York, NY, USA: Springer, 2009\.
- \(11\)W\. Gao, Q\. Liu, Z\. Huang, Y\. Yin, H\. Bi, M\.\-C\. Wang, J\. Ma, S\. Wang, and Y\. Su, ”RCD: Relation map driven cognitive diagnosis for intelligent education systems,” in*Proc\. 44th International ACM SIGIR Conference on Research and Development in Information Retrieval*, 2021, pp\. 501\-510\.
- \(12\)S\. Wang, Z\. Zeng, X\. Yang, and X\. Zhang, ”Self\-supervised graph learning for long\-tailed cognitive diagnosis,” in*Proc\. AAAI Conference on Artificial Intelligence*, vol\. 37, no\. 1, pp\. 110\-118, 2023\.
- \(13\)T\. Huang, Y\. Chen, J\. Geng, H\. Yang, and S\. Hu, ”A higher\-order neural cognitive diagnosis model with hierarchical attention networks,”*Expert Systems with Applications*, vol\. 273, Art\. no\. 126848, 2025\.
- \(14\)R\. Pekrun, ”The control\-value theory of achievement emotions: Assumptions, corollaries, and implications for educational research and practice,”*Educational Psychology Review*, vol\. 18, no\. 4, pp\. 315\-341, 2006\.
- \(15\)S\. D’Mello and A\. Graesser, ”Dynamics of affective states during complex learning,”*Learning and Instruction*, vol\. 22, no\. 2, pp\. 145\-157, 2012\.
- \(16\)R\. S\. J\. d\. Baker, S\. K\. D’Mello, M\. M\. T\. Rodrigo, and A\. C\. Graesser, ”Better to be frustrated than bored: The incidence, persistence, and impact of learners’ cognitive\-affective states during interactions with three different computer\-based learning environments,”*International Journal of Human\-Computer Studies*, vol\. 68, no\. 4, pp\. 223\-241, 2010\.
- \(17\)R\. A\. Calvo and S\. D’Mello, ”Affect detection: An interdisciplinary review of models, methods, and their applications,”*IEEE Transactions on Affective Computing*, vol\. 1, no\. 1, pp\. 18\-37, 2010\.
- \(18\)B\. P\. Woolf, W\. Burleson, I\. Arroyo, T\. Dragon, D\. Cooper, and R\. Picard, ”Affect\-aware tutors: recognising and responding to student affect,”*International Journal of Learning Technology*, vol\. 4, nos\. 3/4, pp\. 129\-164, 2009\.
- \(19\)C\. Conati and H\. Maclaren, ”Empirically building and evaluating a probabilistic model of user affect,”*User Modeling and User\-Adapted Interaction*, vol\. 19, no\. 3, pp\. 267\-303, 2009\.
- \(20\)S\. K\. D’Mello, R\. W\. Picard, and A\. C\. Graesser, ”Toward an affect\-sensitive AutoTutor,”*IEEE Intelligent Systems*, vol\. 22, no\. 4, pp\. 53\-61, 2007\.
- \(21\)S\. Wang, Z\. Zeng, X\. Yang, K\. Xu, and X\. Zhang, ”Boosting neural cognitive diagnosis with student’s affective state modeling,” in*Proc\. AAAI Conference on Artificial Intelligence*, vol\. 38, no\. 1, pp\. 620\-627, 2024\.
- \(22\)P\. Khosla, P\. Teterwak, C\. Wang, A\. Sarna, Y\. Tian, P\. Isola, A\. Maschinot, C\. Liu, and D\. Krishnan, ”Supervised contrastive learning,” in*Advances in Neural Information Processing Systems*, vol\. 33, pp\. 18661\-18673, 2020\.
- \(23\)T\. Chen, S\. Kornblith, M\. Norouzi, and G\. Hinton, ”A simple framework for contrastive learning of visual representations,” in*Proc\. International Conference on Machine Learning*, PMLR, vol\. 119, pp\. 1597\-1607, 2020\.
- \(24\)Y\. Koren, R\. Bell, and C\. Volinsky, ”Matrix factorization techniques for recommender systems,”*Computer*, vol\. 42, no\. 8, pp\. 30\-37, 2009\.
- \(25\)C\. C\. Aggarwal,*Recommender Systems: The Textbook*\. Cham, Switzerland: Springer, 2016\.
- \(26\)Y\. Bengio, A\. Courville, and P\. Vincent, ”Representation learning: A review and new perspectives,”*IEEE Transactions on Pattern Analysis and Machine Intelligence*, vol\. 35, no\. 8, pp\. 1798\-1828, 2013\.
- \(27\)F\. Locatello, S\. Bauer, M\. Lucic, G\. Raetsch, S\. Gelly, B\. Schoelkopf, and O\. Bachem, ”Challenging common assumptions in the unsupervised learning of disentangled representations,” in*Proc\. International Conference on Machine Learning*, PMLR, vol\. 97, pp\. 4114\-4124, 2019\.
- \(28\)R\. Geirhos, J\.\-H\. Jacobsen, C\. Michaelis, R\. Zemel, W\. Brendel, M\. Bethge, and F\. A\. Wichmann, ”Shortcut learning in deep neural networks,”*Nature Machine Intelligence*, vol\. 2, pp\. 665\-673, 2020\.
- \(29\)G\. Alain and Y\. Bengio, ”Understanding intermediate layers using linear classifier probes,” arXiv:1610\.01644, 2016\.
- \(30\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin, ”Attention is all you need,” in*Advances in Neural Information Processing Systems*, vol\. 30, 2017\.
- \(31\)K\. He, X\. Zhang, S\. Ren, and J\. Sun, ”Deep residual learning for image recognition,” in*Proc\. IEEE Conference on Computer Vision and Pattern Recognition*, 2016, pp\. 770\-778\.
- \(32\)S\. Rendle, ”Factorization machines,” in*Proc\. IEEE International Conference on Data Mining*, 2010, pp\. 995\-1000\.
- \(33\)N\. Thai\-Nghe, L\. Drumond, T\. Horvath, and L\. Schmidt\-Thieme, ”Recommender system for predicting student performance,”*Procedia Computer Science*, vol\. 1, no\. 2, pp\. 2811\-2819, 2010\.
- \(34\)J\.\-J\. Vie and H\. Kashima, ”Knowledge tracing machines: Factorization machines for knowledge tracing,” in*Proc\. AAAI Conference on Artificial Intelligence*, vol\. 33, no\. 1, pp\. 750\-757, 2019\.
- \(35\)M\. Pakdaman Naeini, G\. Cooper, and M\. Hauskrecht, ”Obtaining well calibrated probabilities using Bayesian binning,”*Proc\. AAAI Conference on Artificial Intelligence*, vol\. 29, no\. 1, 2015\.
- \(36\)G\. W\. Brier, ”Verification of forecasts expressed in terms of probability,”*Monthly Weather Review*, vol\. 78, no\. 1, pp\. 1\-3, 1950\.

相似文章

基于EEG-fNIRS的跨被试连续情感回归中共享与个体结构建模

arXiv cs.AI

本文研究基于同步EEG-fNIRS数据的零样本跨被试连续效价-唤醒度回归问题,将情感分解为刺激共享成分与个体成分,其中个体成分通过无标签的α频带跨通道同步性进行校准;该方法在性能上超越EEGNet和ASAC-Net等基线模型,并对多种替代架构与特征进行了系统的负面结果搜索。