DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
Summary
This paper proposes DysLexLens, a low-resource LLM framework for analyzing dyslexic learners' experiences with AI tools using online forum data, featuring dictionary-driven filtering, knowledge-graph reasoning, and evaluation metrics.
View Cached Full Text
Cached at: 06/29/26, 05:26 AM
# DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners’ Insights from Online Forums Source: [https://arxiv.org/html/2606.27619](https://arxiv.org/html/2606.27619) andAtie Kia, Phongpadid Nandavong, Dominique Carlon, Jeremy Nguyen, Abhik Banerjee, James Marshall, Anthony McCosker, Yong\-Bin KangSwinburne University of TechnologyMelbourne, VICAustralia[drezazadegan, akia, bnandavong, dcarlon, jdnguyen, abanerjee, jgmarshall, amccosker, ykang@swin\.edu\.au](https://arxiv.org/html/2606.27619v1/mailto:drezazadegan,%20akia,%20bnandavong,%20dcarlon,%20jdnguyen,%20abanerjee,%20jgmarshall,%20amccosker,%[email protected]) ###### Abstract\. Dyslexic learners increasingly use artificial intelligence \(AI\) tools to support reading, writing, organisation, and study\-related tasks\. However, their lived experiences with these tools remain largely underexamined\. This paper proposes DysLexLens, a low\-resource LLM framework, designed to analyse dyslexic learners’ experience with AI through online forum discussions\. DysLexLens is designed as an end\-to\-end, evidence\-traceable architecture which transforms noisy social media posts into a dictionary\-driven corpora, provides knowledge\-graph \(KG\)\-based question reasoning, generates verifiable query responses, and enables response evaluation through quantitative and human\-grounded assessment\. DysLexLens has four key features\. First, it employs a dictionary\-driven filtering method to construct a more focused Reddit corpus on dyslexia and AI, filtering out noisy and weakly related posts to improve the relevance of data collected from low\-resource forum contexts\. Second, it integrates LLM\-assisted semantic analysis with KG\-based query reasoning to uncover meaningful patterns\. Third, it has quantitative evaluation metrics \(RAGAS and Query Robustness\) to measure LLM\-generated response performance\. Fourth, it provides structured qualitative validation guidelines for assessing response quality, with a specific focus on hallucination and evidence alignment\. We demonstrate the effectiveness of DysLexLens using dyslexia\-related Reddit forum data and 30 questions\. The results show its potential generalisability to other low\-resource forum data contexts\. DysLexLens, sample data, questions and evaluation results are available at Github111https://github\.com/SIRI\-HAC\-Program/DysLexLensto support reproducibility\. DysLexLens, Dyslexia, LLM, Dyslexic Reddit Forum, Dyslexic Learners, Inclusive AI ††copyright:none## 1\.Introduction Dyslexia is widely recognised as a neurobiological learning difficulty that affects accurate and fluent word reading, spelling, and decoding, despite adequate intelligence, motivation, and educational opportunity\(Lyonet al\.,[2003](https://arxiv.org/html/2606.27619#bib.bib14)\)\. Its underlying causes are multifactorial, encompassing neurobiological, psychological, and environmental influences\(Cattset al\.,[2024](https://arxiv.org/html/2606.27619#bib.bib15)\)\. Recent prevalence estimates suggest that dyslexia affects 5%\-17% of the population due to reading difficulty, the discrepancy between expected and observed reading performance, and some literacy\-related measures used in assessment\(Wagneret al\.,[2020](https://arxiv.org/html/2606.27619#bib.bib16)\)\. Learners with dyslexia often face important educational challenges, such as difficulties in reading fluency, written communication, comprehension, organisation, and classroom participation\. These challenges create a strong need to better understand the forms of support, coping strategies, and practical solutions to improve their learning outcomes\. The rapid rise of AI tools has introduced new possibilities for personalised and accessible learning support for dyslexic learners\. To help dyslexic learners, existing works have primarily focused on two directions\. The first is technology\-centred, including reviews of assistive technologies for dyslexic learners\(Smith and Hattingh,[2020](https://arxiv.org/html/2606.27619#bib.bib17); Lergaet al\.,[2021](https://arxiv.org/html/2606.27619#bib.bib18)\)\. While valuable, this direction primarily maps available tools and functions, offering limited insight into how dyslexic learners themselves describe their needs, frustrations, and lived experiences in naturalistic settings\. The second direction is learner\-centred, examining lived experience and online discourse\. For example, Reddit\-based thematic analysis has shown that online communities can reveal meaningful self\-expressed perspectives from people with dyslexia\(Kok,[2022](https://arxiv.org/html/2606.27619#bib.bib19)\)\. However, this direction focuses more on identity construction and social meaning, rather than AI perceptions and learning\-related use cases\. Thus, prior research explains either the supply side of dyslexia support technologies or selected aspects of dyslexic experience, but provides limited learner\-centred evidence on how AI and related tools are perceived, adopted, and evaluated in practice\. To address this gap, this paper proposes, DysLexLens, an evidence\-traceable framework for analysing Reddit forum data of dyslexic learners and related stakeholders\. DysLexLens helps to investigates how AI tools are discussed in naturalistic online communities, with a focus on learning\-related practices, perceived benefits and limitations, enabling conditions, educational support contexts, and temporal shifts in discourse\. Specifically, in the paper, we will demonstrate how DysLexLens can addresses five research questions:RQ1:What learning\-related use cases for AI tools are described by dyslexic learners in online discussions?RQ2:What benefits and failure modes are reported when AI tools are used for different learning tasks?RQ3:Under what conditions are AI tools perceived as helpful, supportive, or inclusive for dyslexic learners?RQ4:How are AI\-based supports discussed in relation to broader educational challenges, including institutional accommodations and personal coping strategies?RQ5:How have discussions of AI tools among dyslexic learners changed over time in terms of prevalence, tone, and perceived role in learning support? This paper makes the following contributions: - •We propose DysLexLens, designed to be a generalisable, evidence\-traceable framework for analysing low\-resource forum data, where relevant discussions are sparse, noisy, and difficult to identify at scale\. - •We demonstrate the effectiveness of DysLexLens in the domain of dyslexia and AI, using Reddit discussions to examine how dyslexic learners and related stakeholders describe AI\-related learning use cases, benefits, limitations, and support needs\. - •We present an LLM\-based analysis pipeline of DysLexLens that links generated responses to supporting corpus evidence through KG\-based reasoning\. - •We provide a hybrid evaluation approach combining quantitative metrics and human\-grounded validation guidelines to assess response quality, hallucination and evidence alignment\. ## 2\.Related Work This section positions the study at the intersection of dyslexia support, AI\-enabled learning technologies, and LLM\-assisted analysis of online forum data\. We first review research on dyslexic learners and existing support approaches, before discussing the use of LLMs for analysing user\-generated forum data\. Research on Dyslexic Learners:A large body of research has examined dyslexic learners, extending beyond decoding and spelling difficulties\(Snowlinget al\.,[2020](https://arxiv.org/html/2606.27619#bib.bib30)\)to include comprehension\(Georgiouet al\.,[2022](https://arxiv.org/html/2606.27619#bib.bib34)\), written expression\(Grahamet al\.,[2021](https://arxiv.org/html/2606.27619#bib.bib35)\), learner participation\(Nevill and Forsey,[2023](https://arxiv.org/html/2606.27619#bib.bib31)\), learner confidence\(Hamilton Clark,[2024](https://arxiv.org/html/2606.27619#bib.bib33)\), and wider classroom experience\(Ross,[2021](https://arxiv.org/html/2606.27619#bib.bib32)\)\. Existing work on support for dyslexic learners can be grouped into three broad strands: non\-technological pedagogical interventions, non\-AI technological supports, and AI\-specific approaches\. The first strand has examined pedagogical interventions, such as phonics\-based reading instruction and structured spelling programs\. Evidence is not always specific to dyslexia and is sometimes drawn from studies of broader categories, such as reading disabilities\. A meta\-analysis of randomised controlled trials found that phonics was the only reviewed approach that produced statistically significant effects on the reading and spelling ability of children and adolescents with reading disabilities\(Galuschkaet al\.,[2014](https://arxiv.org/html/2606.27619#bib.bib38)\)\. A more recent review also found empirical support for orthographic and morphological instruction\(Galuschkaet al\.,[2020](https://arxiv.org/html/2606.27619#bib.bib36)\)\. However, support for dyslexic learners is not only a matter of literacy outcomes\. Pedagogical choices can also shape whether learners feel understood, whether they feel singled out as different, and the extent to which they have agency over how their dyslexia is supported\(Ross,[2021](https://arxiv.org/html/2606.27619#bib.bib32)\)\. The second strand has focused on non\-AI technological supports, including assistive tools for reading and writing\. Supportive software can improve spelling ability, but it is generally viewed as an aid rather than a substitute for direct instruction\(Galuschkaet al\.,[2020](https://arxiv.org/html/2606.27619#bib.bib36)\)\. A five\-year follow\-up study of nine dyslexic students found that continued use of tools such as audiobooks, text\-to\-speech, and speech\-to\-text depended onschool\-level context, students’ emotional responses in the classroom, and whether students developed meaningful strategies for using the tools\(Almgren Bäcket al\.,[2024](https://arxiv.org/html/2606.27619#bib.bib39)\)\. These findings indicate that technology availability alone is insufficient\. Rather, tools may be difficult to learn, unsuitable for some students’ needs, or hard to integrate into existing learning practices\(Hamilton Clark,[2024](https://arxiv.org/html/2606.27619#bib.bib33)\)\. The third strand has explored AI\-specific interventions\. A review work identified four prominent uses of AI for dyslexic learners: early detection and diagnosis, personalised learning, speech and language processing, and neuroimaging\(Yapet al\.,[2025](https://arxiv.org/html/2606.27619#bib.bib37)\)\. Another review work of AI\-based interventions for learners with learning disabilities, in which dyslexia was the most frequently studied condition, found positive outcomes but also noted risks of bias in the existing evidence base\(Paglialunga and Melogno,[2025](https://arxiv.org/html/2606.27619#bib.bib40)\)\. These strands show that existing research has largely examined either support interventions, available technologies, or the measured effects of AI\-based systems\. Less is known about how dyslexic learners themselves discuss AI tools in naturalistic settings, including what they use them for, where they experience benefits or failures, and how they perceive such tools\. Our study addresses this gap by analysing dyslexia\- and AI\-related Reddit forum data as a low\-resource, user\-generated data setting\. LLM\-Assisted Online Forum Data Analysis:The growing use of LLMs for analysing text data has increased interest in applying them to online forum and social media data\. Existing studies have used LLMs for market sentiment analysis\(Denget al\.,[2023b](https://arxiv.org/html/2606.27619#bib.bib21)\), generating labels for supervised learning models\(Denget al\.,[2023a](https://arxiv.org/html/2606.27619#bib.bib20)\), and analysing public sentiment during social movements\(Sidhartaet al\.,[2025](https://arxiv.org/html/2606.27619#bib.bib22)\)\. A prominent application area is health\-focused online discourse, including Reddit\-based analyis of lupus pain narratives\(Walkeret al\.,[2026](https://arxiv.org/html/2606.27619#bib.bib23)\), dietary and weight\-loss discussions\(Kaloudiset al\.,[2025](https://arxiv.org/html/2606.27619#bib.bib24)\), and eating disorder discourse\(Chopraet al\.,[2024](https://arxiv.org/html/2606.27619#bib.bib29)\)\. Mental health communities have also been a major focus of LLM\-assisted forum analysis\. Existing work has used LLMs to identify latent linguistic dimensions associated with suicidality in Reddit forums and to analyse linguistic characteristics across social networking posts related to mental health\(Baueret al\.,[2024](https://arxiv.org/html/2606.27619#bib.bib25); Kimet al\.,[2023](https://arxiv.org/html/2606.27619#bib.bib26)\)\. Related research also examines online discussions of LLMs themselves, such as ChatGPT, as emerging mental health support tools\(Junget al\.,[2025](https://arxiv.org/html/2606.27619#bib.bib27); Luoet al\.,[2025](https://arxiv.org/html/2606.27619#bib.bib28)\)\. While these studies demonstrate the value of LLMs for analysing online discussions, most focus on broad sentiment, thematic patterns, or mental health signals\. Less attention has been given to low\-resource forum settings where relevant posts are sparse, noisy, and difficult to identify, and where analytical outputs need to be linked back to supporting evidence\. ## 3\.DysLexLens Framework Figure 1\.The overview of DysLexLens frameworkThis section presents DysLexLens, an evidence\-traceable framework for analysing low\-resource forum data\. DysLexLens is designed for settings where relevant discussions are sparse, noisy, and distributed, making it difficult to construct a focused corpus for systematic analysis\. Rather than treating forum data as a directly usable dataset, DysLexLens supports the workflow from targeted data collection to query\-based reasoning and response evaluation\. In this paper, we apply DysLexLens in the domain of dyslexia and AI by analysing Reddit discussions in which dyslexic learners and related stakeholders describe learning experiences, support needs, and perceptions of AI\-supported tools\. As seen in Fig\.[1](https://arxiv.org/html/2606.27619#S3.F1), DysLexLens consists of three layers: data collection, query reasoning, and evaluation\. ### 3\.1\.Data Collection Layer The goal of this layer is to construct a focused corpus from noisy, low\-resource forum data\. In the dyslexia and AI case study, relevant discussions are not concentrated in a single subreddit or labelled dataset; instead, they are scattered across communities related to dyslexia, neurodiversity, education, accessibility, assistive technology, and AI tools\. To address this challenge, DysLexLens adopts a three\-step process: collecting candidate forum data from a broad set of relevant communities, preserving post\-comment discussion structure while removing low\-information records, and applying a concept dictionary\-based filtering method to identify posts most closely aligned with the research objectives\. Collect Initial Subreddit Data:Reddit is selected as the data source because it contains publicly accessible discussions in which individuals with dyslexia, parents, educators, and other stakeholders share experiences, challenges, coping strategies, and views on assistive technologies and AI\-supported learning\. To identify candidate communities, we first identify seed keywords related to dyslexia, neurodiversity, learning support, assistive technology, and AI from literature studies\. Using the Arctic Shift API, posts and discussion threads are collected from 50 screened subreddits222These 50 subreddits are listed in our GitHub repository \(see footnote 1\)\., including communities related to dyslexia and learning differences, broader neurodiversity, education, accessibility, assistive technologies, AI tools, and general user experience discussions\. This broad initial collection is intentional\. In low\-resource forum settings, narrowly querying only a small number of domain\-specific communities risks missing relevant discussions that occur in adjacent communities\. Capture Full Post Structure and Remove Noise Posts:For each subreddit, non\-stickied posts are retrieved to capture ordinary user\-generated discussions rather than moderator announcements or pinned reference material\. For each post, recursive comment extraction is then performed to capture the full discussion structure, including top\-level comments and nested replies\. Then, the extracted posts and comments are merged into a single datafarme, with subreddit labels derived from the source filenames\. The resulting initial corpus contains 23,480 posts and comments, comprising 1,663,250 words across 45 subreddit communities\. suitable for analysing dyslexia\-related experiences, support needs, perceptions of AI and assistive technologies\. To reduce low\-information noise, posts with fewer than three words are excluded\. Concept Dictionary\-based Post Filtering:The initial Reddit corpus contains many posts that are only weakly related to the study focus\. This is a common challenge in low\-resource forum data, where relevant discussions are sparse, noisy, and distributed across multiple communities\. To construct a more focused corpus for downstream analysis, DysLexLens applies a concept dictionary\-based filtering method\. The dictionary is developed using a construct\-driven procedure, aligned with established practices in computational text analysis\(Kuckartz,[2019](https://arxiv.org/html/2606.27619#bib.bib44); Bolden and Moscarola,[2000](https://arxiv.org/html/2606.27619#bib.bib42)\)\. We first define five core concepts aligned with the research objectives:dyslexia,AI,technology,learning support, andperception\. For each concept, we identify observable linguistic indicators from prior literature, domain expertise, and common expressions used in online forum discussions\. For example, thelearning supportconcept includes terms related to speech to text, spell checkers, predictive text, reading aids, note\-taking, and similar supports, while the AI concept includes terms related to AI, large language models, machine learning, ChatGPT, Claude, Gemini, and related technologies\. The seed dictionary is then expanded with related terms and close variants\. The filtering process first applies exact keyword matching across the full post text\. It then uses similarity scoring to identify semantically close variants that may not appear in the original dictionary\. For example,dictation,dictating,diction,voice typingandspeech\-to\-textare added to thelearning supportdictionary as terms related tospeech to text\. Candidate variants are exported for manual review, and validated terms are added back into the dictionary for subsequent filtering and analysis\. This process reduces the candidate corpus to a focused dataset of 319 posts from 27 subreddits\. The resulting dataset is not intended to represent all Reddit discussions about dyslexia\. Rather, it provides a targeted, research question\-aligned \(i\.e\.,RQ1\-RQ5\) corpus for analysing how dyslexia and AI are discussed in low\-resource Reddit forum data\. ### 3\.2\.Query Reasoning Layer This layer enables analysis of the filtered forum corpus\. Its goal is to support analysis by retrieving relevant evidence, reasoning over semantic relations, generating responses to user\-given queries \(or questions\), and linking generated claims back to the original source text\. As a prerequisite to query execution, DysLexLens first constructs a knowledge backbone \(i\.e\. KG\) from the final filtered corpus produced in the previous layer\. This KG is built once and used across all user interactions, rather than being dynamically reconstructed for each query\. This layer consists of three components: KG\-based retrieval, evidence\-grounded response generation, and user\-guided follow\-up analysis\. Prerequisite Knowledge Graph Construction:Before query reasoning is performed, DysLexLens constructs a reusable KG from the final filtered corpus\. We use the Property Graph Index in LlamaIndex to build the graph representation\. The corpus is first divided into text chunks, and an LLM extracts semantic triples from each chunk in the form of subject–predicate–object relations\. Subjects and objects are represented as entity nodes, while predicates are encoded as labelled edges\. The original text chunks are retained as source nodes so that extracted relations can be traced back to their supporting textual evidence\. To support query\-time retrieval, vector embeddings are generated for indexed chunks and nodes using ‘text\-embedding\-3\-small’\. Thus, the KG provides the relational structure, while the embeddings support semantic matching between user queries and relevant corpus evidence\. KG\-based Retrieval:Given a user queryqq, DysLexLens’s retrieval process identifies relevant semantic triples, associated source chunks, and supporting sentences from the corpus\. The triples provide relational context by showing how key concepts, tools, needs, and experiences are connected, while the retrieved source chunks provide textual evidence for answer generation\. The retrieved triples and supporting text are then provided to the LLM to generate a text\-based response\. The response includes inline references to indexed source chunks, allowing users to inspect the evidence behind the generated answer\. In this way, this layer combines semantic interpretation from the KG with source\-grounded response generation\. Evidence Tracing Pipeline:DysLexLens applies a three\-stage evidence\-tracing pipeline after response generation\. First, it examines the generated response to identify claims, defined as sentences that include explicit references to relevant source\-chunk identifiers\. Second, it aggregates the chunk identifiers associated with each identified claim, denoted asC\*\. Third, it uses these identifiers to retrieve the corresponding original Reddit records, including posts or comments, from each linked source chunk\. User\-guided Follow\-up Analysis:DysLexLens supports multi\-turn analysis by allowing users to ask follow\-up questions over the same KG context, retrieved source chunks, and relevant subgraphs\. This enables users to refine, extend, or challenge an initial response without restarting the retrieval process\. Instead, the system preserves the analytical context from the previous query and uses it as the basis for deeper exploration\. This human\-in\-the\-loop capability is important for low\-resource forum analysis, where relevant evidence may be sparse, fragmented, or open to multiple interpretations\. It allows users to examine additional nuances, compare alternative explanations, identify evidence gaps, and surface details that may not be captured in the initial response\.  Figure 2\.An example of query processing workflow in DysLexLens ### 3\.3\.Evaluation Layer This assesses the quality and evidence alignment of DysLexLens outputs using both quantitative metrics and human\-grounded assessment \(using structured guidelines\)\. The quantitative metrics consist of a RAGAS evaluation metric set and query robustness\. The aim is to evaluate not only whether the framework generates relevant responses, but also whether those responses are factually consistent, grounded in retrieved evidence, and interpretable through traceable provenance\. RAGAS Evaluation Metrics:Retrieval and generation quality are evaluated using RAGAS evaluation metrics\. We construct a test set of 30 queries333These queries are found at our Github repository\., consisting of the five research questions \(introduced in Section 1 \(RQ1\-5\), each accompanied by five follow\-up questions\. For each query, DysLexLens generates a response and records the retrieved supporting context, source\-chunk index, and source evidence\. Performance is assessed using four metrics:Answer Relevancymeasuring how well the response addresses the query;Faithfulnessevaluating factual consistency with the retrieved context;Context Relevanceassessing the relevance of retrieved text segments; andResponseGroundednessmeasuring the extent to which generated claims are supported by retrieved evidence\. To ensure fair comparison, all final RAGAS results use a fixed retrieval configuration: similarity top\-k=3 and a 512\-token chunk size\. segmentation\. Query Robustness Analysis:Query robustness analysis evaluates whether DysLexLens produces stable responses when the similar questions are given\. This is important because DysLexLens is designed for user\-guided analysis rather than fixed\-question benchmarking\. In low\-resource forum data, relevant evidence is often sparse, noisy, and limited\. Therefore, small changes in query wording may affect which chunks, subgraphs, or evidence traces are retrieved\. For each research question \(RQ1\-5\), we create three semantically aligned query variants: the original wording, a paraphrased version, and a keyword\-based version\. All variants are processed using the same experimental configuration\. Robustness is assessed by comparing variation in the four RAGAS metrics\. Lower variation indicates stronger robustness, suggesting that DysLexLens can retrieve relevant evidence and generate grounded responses despite differences in query phrasing\. Higher variation, particularly inContext RelevanceorResponse Groundedness, indicates potential evidence mismatch or instability in response generation\. Human Assessment: To complement the quantitative metrics, we conduct a human assessment on a purposive sample of generated responses\. The sample is selected to include both strong and weak RAGAS score patterns\. Humans assess response quality using rubric\-based guideline which is developed for this study444The guidelines are found at our Github repository\.\. The rubrics are designed to support consistent assessment of hallucination risk, evidence alignment, and interpretability\. Each sampled response is decomposed into claims, whereC\*exist\. Each claim is then traced to its identified source chunk, and original Reddit posts\. The human assesses whether evidence is present, whether the evidence sufficiently supports the claim, and whether the provenance trail is interpretable\. Support strength is rated on a 3\-scale:strong, partial, and weak\. This audit helps identify cases where generated responses are well grounded, partially supported, or affected by retrieval mismatch, unsupported inference, or potential hallucination\. For instance, a response may correctly identify that an AI tool supported writing, while overextending the same evidence to claim improved confidence\. The rubric makes this distinction visible through separate assessments of evidence presence, support strength, and interpretative utility\. Illustrative Example Fig\.[2](https://arxiv.org/html/2606.27619#S3.F2)demonstrates an example of the query\-processing workflow in DysLexLens, showing how the framework supports user\-driven query refinement, evidence\-grounded response generation, and evidence tracing\. A user may submit any analytical question over the filtered forum corpus, such asRQ2, which asks what benefits and failure modes dyslexic learners report when using AI tools for different learning tasks\. If the retrieved evidence does not sufficiently support all aspects of the query, the query reasoning layer returns an immediate response indicating insufficient contextual evidence and enables the user to refine the question\. The user can then give a more specific follow\-up question, for example asking how learners describe AI as improving clarity, confidence, or speed when completing learning\-related work\. DysLexLens then generates a response with source\-chunk identifiers such asC1andC3, which make the evidence trail explicit\. Through the evidence\-tracing pipeline, these chunck identifiers allow users to identify claims and trace back to the identified source chunks, and the original Reddit records including posts or comments\. DysLexLens’ workflow thus enables users to move from an initial query to more precise, traceable, and interpretable findings while preserving a link between generated responses and the underlying forum evidence\. ## 4\.Evaluation and Results Our evaluation questions include five research questions and 25 follow\-up questions\. The KG constuction and all responses are generated using gpt\-4o\-mini\. ### 4\.1\.Evaluation using RAGAS Metrics Fig\.[3](https://arxiv.org/html/2606.27619#S4.F3)reports the RAGAS results for responses to the fiveRQs, as well as the mean and standard deviation for theRQs, follow\-up questions, and all queries combined\. Across all 30 responses, DysLexLens achieves a meanAnswer Relevancyscore of 0\.75\. In total, 25 of the 30 responses scored at least 0\.65 forAnswer Relevancy\. This indicates that most generated responses address the intended meaning of the input questions, althoughAnswer Relevancyalone does not show evidential support or factual grounding\. Figure 3\.Quantitative evaluation results: Single scores for individualRQs; Mean and standard deviation for query set variants\.The research question responses achieve strongerAnswer Relevancythan follow\-up responses, with mean scores of 0\.87 and 0\.72, respectively\. This may suggest that DysLexLens performs better when questions include keywords closely aligned with the concept dictionary\. However,Faithfulness,Context Relevance, andResponse Groundednessfor research questions are lower thanAnswer Relevancy\. Across all 30 queries’ responses, the meanFaithfulnessscore is 0\.52, meanContext Relevancyis 0\.40, and meanResponse Groundednessis 0\.43\. The research question scores further show thatAnswer Relevancyalone does not guarantee evidential support\.RQ1,RQ2,RQ3, andRQ5achieve highAnswer Relevancy, but their evidential support levels differ\. For example,RQ3achieves strongFaithfulnessbut aResponse Groundednessscore of zero, suggesting that the response is judged to be factually consistent with the retrieved context, but the generated claims are not explicitly supported by identified evidence\.RQ5achieve the highestAnswer Relevancyscore but aContext Relevancyscore of zero, showing that temporal\-change questions are difficult for the retrieval component\. The evidence\-tracing results further provide an integrated interpretation across both research questions and follow\-up queries\. Of the 30 generated responses, 29 includes at least one chunk identifier\. At the sentence level, the export contained 114 unique claims, of which 96 are linked to at least one source\-chunk identifier and 18 have no source\-chunk identifier\. For identified source chunks, the mean retrieval score is 0\.86, while identified claims have a mean retrieval score of 0\.85, at the response level\. However, three identified claims have an average retrieval score of 0\.50 or below, showing that retrieval similarity alone is not sufficient for strong claim\-level evidential support\. ### 4\.2\.Query Robustness Analysis To evaluate query robustness, all 30 evaluation questions are tested using three semantically aligned query formulations: the original question, a paraphrased version, and a keyword\-perturbed version\. The paraphrased and keyword\-perturbed variants are generated using GPT\-5\.5 Pro\. The original queries achieve the highestAnswer Relevancyscore \(0\.75\), followed by the paraphrased queries \(0\.58\)\. This suggests that DysLexLens remains reasonably stable when questions are reworded but their meaning is preserved\. However, the keyword\-perturbed queries show a clear decrease inAnswer Relevancy\(0\.34\), indicating that the framework is more sensitive when domain\-specific terms are changed\. The highestFaithfulnessscore belongs to the keyword\-perturbed query set \(0\.66\), suggesting that these responses are still aligned with the retrieved evidence, despite their lowerAnswer relevancy\. This suggests that the generated responses remain relatively consistent with the retrieved evidence, but the retrieved evidence may not match the intended meaning of the original question\. In other words, keyword perturbation may lead the system to retrieve narrower or different evidence that can still support the generated response, while reducing alignment with the user’s intended query\. Across all variants,Context RelevancyandResponse Groundednessremain lower thanAnswer RelevancyandFaithfulness\. The original queries achieve the strongest evidential support, withContext Relevancyof 0\.40 andResponse Groundednessof 0\.43, while paraphrased and keyword\-perturbed queries show lower grounding scores\. This indicates that retrieval precision and claim\-level evidential support remain the main limitations of the current pipeline\. Hence, DysLexLens is moderately robust to paraphrasing but less stable under keyword perturbation, highlighting the need for improved evidence ranking, and query optimisation guided by core domain\-related keywords\. Table 1\.Query robustness across query variants based on mean RAGAS scores\.Query variantAnswerRelevancyFaithfulnessContextRelevancyResponseGroundednessOriginal0\.750\.520\.400\.43Paraphrased0\.580\.550\.310\.25Keyword\-perturbed0\.340\.660\.330\.28 ### 4\.3\.Human Assessment A human assessment is conducted in three dimensions following our structured qualitative assessment guideline to assess the interpretability and verifiability of DysLexLens responses beyond automated RAGAS scores\. An evaluation file is automatically generated for this assessment, where multiple rows can appear for a single query response, depending on the number of identified claims, in the response\. Each claim in a query response is supported by aC\*identifier, which links to the corresponding source chunk\. Human assessors review 100 claims, including 51 claims with stronger RAGAS metric scores and 49 claims with lower scores, and examine provenance tracing and alignment between generated claims, identified source chunks, and original Reddit records\. This assessment is conducted through three audit dimensions: \(A\) Evidence Verification, \(B\) Support Strength, and \(C\) Interpretative Utility\. Evidence verification indicates whether the identified chunk appears in the original source\. Support strength rates how strongly the identified chunk supports the claim\. Interpretative utility rates how useful the identified chunk is for interpretation\. Each claim is checked against the exportedsource\_chunk, andfull\_postfields\. As shown in Fig\.[4](https://arxiv.org/html/2606.27619#S4.F4), 39 claims are fully verifiable, 55 were partially verifiable, and 6 are not verifiable due to missing citation or source evidence\. Because evaluator judgements vary in some cases, categorical labels, including evidence verification and interpretative utility, are finalised through adjudication, while support strength is summarised using the median score and checked for consistency\. In the support strength assessment, 21 claims receive strong support, while the most claims are at least partially supported and only 18 claims have weak evidential support\. Interpretative utility followed a similar pattern, with 14 high\-utility, 61 medium\-utility, and 25 low\-utility claims\. The audit also showed that main research\-question responses had stronger provenance than follow\-up responses\. Among the audited main\-response rows, 10 of 11 were fully verifiable\. In contrast, 56 of 89 follow\-up rows were only partially verifiable, mainly because several follow\-up rows exported the full retrieved chunk rather than a short exact evidence phrase\. Thus, the evidence trail remained useful for human inspection, but follow\-up responses often required extra manual checking to identify the precise supporting phrase\. Figure 4\.Human assessment of reviewed claims\. \(A\) Yes: clearly present, Partial: partially present, No: absent or contradictory; \(B\) 1: weak, 2: acceptable, 3: strong; and \(C\) High: explicit and clear, Medium: requiring interpretation, Low: vague or missing link\. ### 4\.4\.Discussion The triangulated evaluation shows that DysLexLens can generate reasonably accurate and inspectable responses from dyslexia\-related Reddit data\. RAGAS benchmarking indicates strongAnswer Relevancy, particularly for the research questions, but mediumContext Relevancywhich is justifiable due to the low\-resource data\. Query robustness analysis shows that the framework is reasonably stable under paraphrasing but more sensitive to keyword perturbation, while the human\-grounded audit shows the risk of hallucination which confirms the need for claim\-level expert verification\. Nevertheless, DysLexLens can still be useful as an evidence\-traceable exploratory analysis tool that supports research on dyslexic learners’ lived experiences with AI, assistive technologies, and learning support, which are often difficult to capture through formal studies alone\. The retrieved subgraphs from the KG suggests that dyslexic learners perceive existing AI tools as useful but still limited, indicating scope for further AI development in this area\. A key value of DysLexLens is that it reduces the burden of exploratory thematic analysis in low\-resource research settings while supporting reproducibility through a transparent evidence\-traceable design\. By separating concept\-based filtering, knowledge\-graph construction, retrieval\-augmented generation, evidence tracing, and provenance auditing, DysLexLens provides a modular and reusable pipeline for analysing low\-resource online discourse\. This modularity supports generalisability, as domain experts can replace the dyslexia\-AI concept dictionary with a context\-specific dictionary in the concept\-based filtering module, redefine the questions, and apply the same retrieval, reasoning, evaluation benchmarks, and audit guidelines to assess whether the framework produces grounded and interpretable outputs in a new setting\. Future versions of DysLexLens should improve synonym\-aware retrieval and evidence ranking to ensure interpretations are consistently supported by precise source passages\. Moreover, expanding query\-variant generation through AI\-based query optimisation, guided by the concept dictionary, could be a useful addition to improve response quality\. ## 5\.Conclusion This paper presented DysLexLens, a low\-resource LLM\-assisted framework for analysing dyslexic learners’ experiences with AI in online forum discussions\. The value of DysLexLens lies in its integration of concept\-dictionary filtering with evidence\-traceable question answering, as demonstrated through the Reddit\-based case study, which provides an explainable and reproducible pipeline\. The knowledge graph supported relational interpretation across AI, technology, support, perception, and educational challenges, while the Q/A pipeline generated source\-grounded responses to user\-defined questions\. The evaluation further shows the importance of combining quantitative metrics with qualitative expert review, as automated scores alone may not fully capture weak grounding, retrieval mismatch, or unsupported interpretation\. This study has limitations\. Reddit discussions do not represent all dyslexic learners, and LLM\-based triple extraction and retrieval may introduce noise\. Therefore, the generated graph should be treated as an interpretive aid rather than a complete representation of learner experience\. Future work will extend DysLexLens beyond this dyslexia\-focused Reddit case study\. By indexing new datasets and developing systematic strategies for building domain\-specific concept dictionaries or alternative filtering methods, this modular framework can be adapted to other Reddit communities and online data sources, such as podcasts and social media, to answer user\-defined research questions\. This positions DysLexLens as a general evidence\-traceable analysis tool for low\-resource online discourse, particularly where researchers need to explore noisy textual data, ask structured and follow\-up questions, inspect supporting evidence, and evaluate response quality\. ## GenAI Usage Disclosure In preparing this paper, GPT\-5\.5 is used for identifying and correcting grammatical errors and typos, as well as for assisting in the generation of illustrative figures\. In accordance with academic integrity, all research content, methods, data analysis, and paper drafting are developed, conducted, and validated by the authors\. ## References - G\. Almgren Bäck, E\. Lindeblad, C\. Elmqvist, and I\. Svensson \(2024\)Dyslexic students’ experiences in using assistive technology to support written language skills: a five\-year follow\-up\.Disability and Rehabilitation: Assistive Technology19\(4\),pp\. 1217–1227\.Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p4.1)\. - B\. Bauer, R\. Norel, A\. Leow, Z\. A\. Rached, B\. Wen, and G\. Cecchi \(2024\)Using large language models to understand suicidality in a social media–based taxonomy of mental health disorders: linguistic analysis of reddit posts\.JMIR mental health11,pp\. e57234\.Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - R\. Bolden and J\. Moscarola \(2000\)Bridging the quantitative\-qualitative divide: the lexical approach to textual data analysis\.Social science computer review18\(4\),pp\. 450–460\.Cited by:[§3\.1](https://arxiv.org/html/2606.27619#S3.SS1.p4.1)\. - H\. W\. Catts, N\. P\. Terry, C\. J\. Lonigan, D\. L\. Compton, R\. K\. Wagner, L\. M\. Steacy, K\. Farquharson, and Y\. Petscher \(2024\)Revisiting the definition of dyslexia\.Annals of Dyslexia74\(3\),pp\. 282–302\.Cited by:[§1](https://arxiv.org/html/2606.27619#S1.p1.1)\. - M\. Chopra, A\. Chatterjee, L\. Dey, and P\. P\. Das \(2024\)Deciphering psycho\-social effects of eating disorder: analysis of reddit posts using large language model \(llm\) s and topic modeling\.InProceedings of the 4th International Conference on Natural Language Processing for Digital Humanities,pp\. 156–164\.Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - X\. Deng, V\. Bashlovkina, F\. Han, S\. Baumgartner, and M\. Bendersky \(2023a\)LLMs to the moon? reddit market sentiment analysis with large language models\.InCompanion Proceedings of the ACM Web Conference 2023,WWW ’23 Companion,New York, NY, USA,pp\. 1014–1019\.External Links:ISBN 9781450394192,[Link](https://doi.org/10.1145/3543873.3587605),[Document](https://dx.doi.org/10.1145/3543873.3587605)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - X\. Deng, V\. Bashlovkina, F\. Han, S\. Baumgartner, and M\. Bendersky \(2023b\)What do llms know about financial markets? a case study on reddit market sentiment analysis\.InCompanion Proceedings of the ACM Web Conference 2023,WWW ’23 Companion,New York, NY, USA,pp\. 107–110\.External Links:ISBN 9781450394192,[Link](https://doi.org/10.1145/3543873.3587324),[Document](https://dx.doi.org/10.1145/3543873.3587324)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - K\. Galuschka, R\. Görgen, J\. Kalmar, S\. Haberstroh, X\. Schmalz, and G\. Schulte\-Körne \(2020\)Effectiveness of spelling interventions for learners with dyslexia: A meta\-analysis and systematic review\.Educational Psychologist55\(1\),pp\. 1–20\.External Links:ISSN 0046\-1520, 1532\-6985,[Document](https://dx.doi.org/10.1080/00461520.2019.1659794)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p3.1),[§2](https://arxiv.org/html/2606.27619#S2.p4.1)\. - K\. Galuschka, E\. Ise, K\. Krick, and G\. Schulte\-Körne \(2014\)Effectiveness of Treatment Approaches for Children and Adolescents with Reading Disabilities: A Meta\-Analysis of Randomized Controlled Trials\.PLoS ONE9\(2\),pp\. e89900\.External Links:ISSN 1932\-6203,[Document](https://dx.doi.org/10.1371/journal.pone.0089900)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p3.1)\. - G\. K\. Georgiou, D\. Martinez, A\. P\. A\. Vieira, A\. Antoniuk, S\. Romero, and K\. Guo \(2022\)A meta\-analytic review of comprehension deficits in students with dyslexia\.Annals of Dyslexia72\(2\),pp\. 204–248\.External Links:ISSN 0736\-9387, 1934\-7243,[Document](https://dx.doi.org/10.1007/s11881-021-00244-y)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p2.1)\. - S\. Graham, A\. A\. Aitken, M\. Hebert, A\. Camping, T\. Santangelo, K\. R\. Harris, K\. Eustice, J\. D\. Sweet, and C\. Ng \(2021\)Do children with reading difficulties experience writing difficulties? A meta\-analysis\.\.Journal of Educational Psychology113\(8\),pp\. 1481–1506\.External Links:ISSN 1939\-2176, 0022\-0663,[Document](https://dx.doi.org/10.1037/edu0000643)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p2.1)\. - C\. H\. Hamilton Clark \(2024\)Dyslexia concealment in higher education: Exploring students’ disclosure decisions in the face ofUKuniversities’ approach to dyslexia\.Journal of Research in Special Educational Needs24\(4\),pp\. 922–935\.External Links:ISSN 1471\-3802, 1471\-3802,[Document](https://dx.doi.org/10.1111/1471-3802.12683)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p2.1),[§2](https://arxiv.org/html/2606.27619#S2.p4.1)\. - K\. Jung, G\. Lee, Y\. Huang, and Y\. Chen \(2025\)’I’ve talked to chatgpt about my issues last night\.’: examining mental health conversations with large language models through reddit analysis\.Proc\. ACM Hum\.\-Comput\. Interact\.9\(7\)\.External Links:[Link](https://doi.org/10.1145/3757537),[Document](https://dx.doi.org/10.1145/3757537)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - E\. Kaloudis, V\. Kouti, F\. Triantafillou, P\. Ventouris, R\. Pavlidis, and V\. Bountziouka \(2025\)AI\-powered analysis of weight loss reports from reddit: unlocking social media’s potential in dietary assessment\.Nutrients17\(5\),pp\. 818\.Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - S\. Kim, J\. Cha, D\. Kim, and E\. Park \(2023\)Understanding mental health issues in different subdomains of social networking services: computational analysis of text\-based reddit posts\.Journal of Medical Internet Research25,pp\. e49074\.Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - L\. T\. Kok \(2022\)Developing a dyslectic identity on reddit a thematic analysis of /r/dyslexia\.Master’s Thesis\.External Links:[Link](http://hdl.handle.net/2105/65001)Cited by:[§1](https://arxiv.org/html/2606.27619#S1.p2.1)\. - U\. Kuckartz \(2019\)Qualitative text analysis: a systematic approach\.InCompendium for early career researchers in mathematics education,pp\. 181–197\.Cited by:[§3\.1](https://arxiv.org/html/2606.27619#S3.SS1.p4.1)\. - R\. Lerga, S\. Candrlic, and A\. Jakupovic \(2021\)A review on assistive technologies for students with dyslexia\.\.CSEDU \(2\),pp\. 64–72\.Cited by:[§1](https://arxiv.org/html/2606.27619#S1.p2.1)\. - X\. Luo, S\. Ghosh, J\. L\. Tilley, P\. Besada, J\. Wang, and Y\. Xiang \(2025\)“Shaping chatgpt into my digital therapist”: a thematic analysis of social media discourse on using generative artificial intelligence for mental health\.Digital health11,pp\. 20552076251351088\.Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - G\. R\. Lyon, S\. E\. Shaywitz, and B\. A\. Shaywitz \(2003\)A definition of dyslexia\.Annals of dyslexia53\(1\),pp\. 1–14\.Cited by:[§1](https://arxiv.org/html/2606.27619#S1.p1.1)\. - T\. Nevill and M\. Forsey \(2023\)The social impact of schooling on students with dyslexia: A systematic review of the qualitative research on the primary and secondary education of dyslexic students\.Educational Research Review38,pp\. 100507\.External Links:ISSN 1747938X,[Document](https://dx.doi.org/10.1016/j.edurev.2022.100507)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p2.1)\. - A\. Paglialunga and S\. Melogno \(2025\)The Effectiveness of Artificial Intelligence\-Based Interventions for Students with Learning Disabilities: A Systematic Review\.Brain Sciences15\(8\),pp\. 806\.External Links:ISSN 2076\-3425,[Document](https://dx.doi.org/10.3390/brainsci15080806)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p5.1)\. - H\. Ross \(2021\)‘I’m Dyslexic but What Does That Even Mean?’: Young People’s Experiences of Dyslexia Support Interventions in Mainstream Classrooms\.Scandinavian Journal of Disability Research23\(1\),pp\. 284–294\.External Links:ISSN 1745\-3011,[Document](https://dx.doi.org/10.16993/sjdr.782)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p2.1),[§2](https://arxiv.org/html/2606.27619#S2.p3.1)\. - S\. Sidharta, H\. Pranoto, F\. M\. Gasa, N\. Kholis, and A\. K\. S\. Ong \(2025\)Analysis of public sentiment on the 17\+8 people’s demands issue using indobert and distilbert with llm\-based data annotation\.In2025 International Conference on Informatics, Multimedia, Cyber and Information System \(ICIMCIS\),Vol\.,pp\. 575–580\.External Links:[Document](https://dx.doi.org/10.1109/ICIMCIS68501.2025.11327025)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - C\. Smith and M\. Hattingh \(2020\)Assistive technologies for students with dyslexia: a systematic literature review\.InInternational Conference on Innovative Technologies and Learning,pp\. 504–513\.Cited by:[§1](https://arxiv.org/html/2606.27619#S1.p2.1)\. - M\. J\. Snowling, C\. Hulme, and K\. Nation \(2020\)Defining and understanding dyslexia: past, present and future\.Oxford Review of Education46\(4\),pp\. 501–513\.External Links:ISSN 0305\-4985, 1465\-3915,[Document](https://dx.doi.org/10.1080/03054985.2020.1765756)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p2.1)\. - R\. K\. Wagner, F\. A\. Zirps, A\. A\. Edwards, S\. G\. Wood, R\. E\. Joyner, B\. J\. Becker, G\. Liu, and B\. Beal \(2020\)The prevalence of dyslexia: a new approach to its estimation\.Journal of learning disabilities53\(5\),pp\. 354–365\.Cited by:[§1](https://arxiv.org/html/2606.27619#S1.p1.1)\. - A\. Walker, J\. Leung, A\. Alagappan, S\. Rajwal, S\. Lakamana, T\. Park, N\. Le, A\. Irani, A\. Sarker, T\. Falasinnu, and S\. Bozkurt \(2026\)Centering patient voices in lupus pain: a biopsychosocial analysis of reddit narratives using large language models\.Arthritis Care & Research78\(1\),pp\. 123–133\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1002/acr.25687),[Link](https://acrjournals.onlinelibrary.wiley.com/doi/abs/10.1002/acr.25687),https://acrjournals\.onlinelibrary\.wiley\.com/doi/pdf/10\.1002/acr\.25687Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p7.1)\. - J\. R\. Yap, T\. Aruthanan, and M\. Chin \(2025\)Artificial Intelligence in Dyslexia Research and Education: A Scoping Review\.IEEE Access13,pp\. 7123–7134\.External Links:ISSN 2169\-3536,[Document](https://dx.doi.org/10.1109/ACCESS.2025.3526189)Cited by:[§2](https://arxiv.org/html/2606.27619#S2.p5.1)\.
Similar Articles
HyperLens: Quantifying Cognitive Effort in LLMs with Fine-grained Confidence Trajectory
This paper introduces HyperLens, a high-resolution probe to quantify cognitive effort in LLMs by tracing fine-grained confidence trajectories across layers. It reveals that complex tasks require higher cognitive effort and demonstrates how Supervised Fine-Tuning can reduce this effort, potentially degrading performance.
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
This tutorial paper provides an overview of building multilingual and multimodal LLMs for low-resource languages, covering data creation, model alignment, fine-tuning, and evaluation, with a focus on practical recipes and hands-on resources.
LEXIC: Lightweight Eye-tracking eXtension via Injected Complexity
This paper introduces LEXIC, a lightweight method that injects precomputed word-level difficulty signals (GPT-2 surprisal, word frequency, word length) into gaze-only models for reading comprehension prediction, achieving slight improvements over baseline on the EyeBench benchmark.
From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages
Presents a systematic methodology for converting Hindi WordNet into 1.25 million instruction-response pairs to fine-tune a 12B-parameter language model using LoRA, demonstrating improved pedagogical effectiveness for specialized conversational systems in low-resource languages.
SkillLens: Adaptive Multi-Granularity Skill Reuse for Cost-Efficient LLM Agents
This paper introduces SkillLens, a hierarchical framework for adaptive multi-granularity skill reuse in LLM agents, demonstrating improved accuracy and cost-efficiency on benchmark tasks.