TRACE-BN: Transferring Bangla-English Tutoring Behavior to a Sub-1B Offline Language Model
Summary
TRACE-BN introduces a curriculum-guided dataset for structured Bangla-English tutoring and demonstrates transferring this behavior to a sub-1B language model using LoRA, achieving significant improvements in tutoring quality for resource-constrained offline environments.
View Cached Full Text
Cached at: 08/18/26, 10:03 AM
# TRACE-BN: Transferring Bangla-English Tutoring Behavior to a Sub-1B Offline Language Model
Source: [https://arxiv.org/html/2608.15223](https://arxiv.org/html/2608.15223)
Khan Raiyan Ibne Reza, Sanjana Aktar Maria, Mohammad Tushar Abdullah, Asfee Bhuiyan Leen, and Sumaiya Tabassum NimiAffiliation:Department of Electrical and Computer Engineering, North South University, Dhaka, Bangladesh Email: \{raiyan\.reza, sanjana\.maria, tushar\.abdullah\.232, asfee\.leen\.242, sumaiya\.nimi\}@northsouth\.edu
###### Abstract
Bangla\-English tutoring requires more than producing a correct translation: learners also need explanations of grammar differences, awareness of their likely errors, and targeted practice\. We presentTRACE\-BN, a curriculum\-guided dataset of structured tutoring traces for Bangla\-speaking learners of English at the CEFR A1–A2 level\. Each trace combines word\-level glosses, literal and natural translations, Bangla grammar explanations, a plausible learner error, and a targeted practice question with its answer\. The traces are generated by Gemini 3\.5 Flash Lite as the teacher model from NCTB Classes 9–10 English curriculum units, then filtered for structural validity, script integrity, and semantic duplication\. We transfer the resulting structured tutoring behavior to Qwen3\-0\.6B using LoRA with 4\-bit quantization for resource\-constrained offline deployment\. On held\-out inputs, schema validity increases from 85\.4% to 95\.8%, while, against teacher\-model references, chrF\+\+ improves from 15\.28 to 34\.77 and BLEU from 4\.52 to 21\.03\. Field\-level evaluation by two independent judges shows improvements across translation, grammar explanation, learner\-error diagnosis, and practice alignment, while a human audit supports the quality of the supervision data\. The results show that curriculum\-guided structured supervision can transfer multi\-component tutoring behavior to a sub\-1B model under these resource constraints\. The dataset, model checkpoints, and code are publicly available athttps://huggingface\.co/datasets/RaiyanKhaan/Trace\-BN\.
###### Index Terms:
small language models, LoRA fine\-tuning, low\-resource NLP, Bangla\-English tutoring, computer\-assisted language learning
## IIntroduction
A Bangla\-English tutoring system should do more than produce a target translation: it should explain relevant grammatical differences, identify likely learner errors, and provide practice on the same pattern\. Large language models can support this interaction more fully than conventional translation systems can, but existing Bangla\-focused models are not designed around a structured tutoring interaction\. Such an interaction requires word\-level glosses, literal and natural translations, Bangla grammar explanations, prediction of likely learner errors, and targeted practice within a single response\. Existing Bangla models such as TigerLLM, BanglaLlama, and TituLLM provide general language capabilities, while recent low\-resource tutoring systems include substantially larger models and broader multimodal settings\[[14](https://arxiv.org/html/2608.15223#bib.bib1),[25](https://arxiv.org/html/2608.15223#bib.bib2),[6](https://arxiv.org/html/2608.15223#bib.bib3),[9](https://arxiv.org/html/2608.15223#bib.bib4),[2](https://arxiv.org/html/2608.15223#bib.bib5)\]\.
We address this problem withTRACE\-BN, a curriculum\-guided dataset of 4,099 structured Bangla\-to\-English tutoring traces\. Each trace contains seven fields: word\-level glosses, literal and natural translations, Bangla grammar notes, a common learner mistake, a practice question, and its answer\. Unlike a conventional parallel corpus, each TRACE\-BN instance encodes a complete instructional sequence rather than a single input\-output pair\. Learner\-facing scaffolding is provided in Bangla, while target\-language content remains in English\. The design draws on second\-language acquisition research emphasizing contrastive explanation and structured practice\[[4](https://arxiv.org/html/2608.15223#bib.bib6),[20](https://arxiv.org/html/2608.15223#bib.bib7),[3](https://arxiv.org/html/2608.15223#bib.bib8),[7](https://arxiv.org/html/2608.15223#bib.bib9)\]\.
We then investigate whether these traces can teach the underlying tutoring behavior to a sub\-1B model\. We fine\-tune Qwen3\-0\.6B with LoRA and evaluate it against the untuned base model and three zero\-shot baselines\. On 432 held\-out examples, the tuned model improves schema validity from 85\.4% to 95\.8% and raises chrF\+\+ from 15\.28 to 34\.77 and BLEU from 4\.52 to 21\.03\. A field\-level evaluation with two independent judges further shows consistent improvements across translation, grammar explanation, learner\-error diagnosis, and practice alignment\.
Our contributions are:
- •TRACE\-BN, a curriculum\-guided bilingual tutoring dataset with structured traces across seven fields spanning translation, explanation, learner\-error modeling, and practice\.
- •A structured tutoring task formulationthat combines translation with contrastive grammar explanation, learner\-error prediction, and targeted practice in a single trace, extending beyond the Bangla\-English pair studied here\.
- •An empirical adaptation studyshowing that Qwen3\-0\.6B can learn to generate the TRACE\-BN structure under LoRA fine\-tuning, together with field\-level evaluation and human validation of the supervision data\.
## IIRelated Work
Fig\. 1:TRACE\-BN construction and adaptation pipeline\. Curriculum\-guided Bangla inputs are used to obtain teacher\-generated tutoring traces, which are filtered into TRACE\-BN and used to adapt Qwen3\-0\.6B with LoRA\.### II\-ABangla NLP and Low\-Resource Language Tutoring
Fig\.[1](https://arxiv.org/html/2608.15223#S2.F1)summarizes the TRACE\-BN construction and adaptation pipeline\. Recent Bangla language models, including TigerLLM, BanglaLlama, and TituLLM, target general Bangla language understanding and generation, while KrishokChat adapts a Bangla LLM to a single domain, agricultural advisory\[[14](https://arxiv.org/html/2608.15223#bib.bib1),[25](https://arxiv.org/html/2608.15223#bib.bib2),[6](https://arxiv.org/html/2608.15223#bib.bib3),[15](https://arxiv.org/html/2608.15223#bib.bib10)\]\. None of these systems combine translation with grammar explanation, error prediction, and practice generation in a single tutoring interaction\.
Low\-resource tutoring systems such as AfriLangTutor, LEARN, CaptainA, and LangLearn support AI\-assisted language learning in other languages, but at a different scale and scope: AfriLangTutor fine\-tunes 8B and 12B models on multi\-turn dialogue across ten African languages, LEARN builds oral proficiency through cartoon\-based visual question answering, CaptainA targets pronunciation practice, and LangLearn targets flashcard\-based vocabulary drills\[[2](https://arxiv.org/html/2608.15223#bib.bib5),[19](https://arxiv.org/html/2608.15223#bib.bib11),[11](https://arxiv.org/html/2608.15223#bib.bib12),[26](https://arxiv.org/html/2608.15223#bib.bib13)\]\. TRACE\-BN instead targets a sub\-1B model with one schema spanning translation, grammar explanation, and error\-aware practice for a single language pair\.
A separate feasibility study concludes that Bangla LLMs are needed but that the field still lacks the high\-quality pretraining and instruction\-tuning data required to build them\[[9](https://arxiv.org/html/2608.15223#bib.bib4)\]\. TRACE\-BN targets this gap for one downstream use: every trace is teacher\-generated and then filtered for structural validity, script integrity, and semantic duplication before it is used for tuning \(Section[III](https://arxiv.org/html/2608.15223#S3)\)\.
### II\-BSmall Models and Structured Generation
KD\-LoRA, DistilQwen2\.5, LoRA\-Gen, and PhoneLM each specializes or deploys smaller language models through a different efficiency strategy\[[1](https://arxiv.org/html/2608.15223#bib.bib14),[21](https://arxiv.org/html/2608.15223#bib.bib15),[22](https://arxiv.org/html/2608.15223#bib.bib16),[24](https://arxiv.org/html/2608.15223#bib.bib17)\]\. TRACE\-BN applies this direction to structured tutoring with a 0\.6B target model\.
Structured generation introduces a separate reliability issue: a response can satisfy a required schema without providing correct or useful content\. Prior work has studied schema adherence and the effects of format constraints on language\-model quality\[[12](https://arxiv.org/html/2608.15223#bib.bib18),[8](https://arxiv.org/html/2608.15223#bib.bib19),[17](https://arxiv.org/html/2608.15223#bib.bib20),[16](https://arxiv.org/html/2608.15223#bib.bib21),[23](https://arxiv.org/html/2608.15223#bib.bib22)\]\. Accordingly, our evaluation separates structural validity from translation and pedagogical quality\.
## IIITRACE\-BN Dataset Construction
Fig\. 2:Example TRACE\-BN tutoring trace\. A Bangla learner input is mapped to word\-level glosses, literal and natural English, a Bangla contrastive explanation, a common learner error, and a short practice exercise with its answer\.TRACE\-BN is based on the NCTB \(National Curriculum and Textbook Board\) English curriculum for Classes 9–10\. We manually selected and reviewed the units, topics, and grammar coverage to define the dataset scope\. Fig\.[2](https://arxiv.org/html/2608.15223#S3.F2)shows a complete trace from the resulting dataset\.
For each selected unit, we specified the target learner level, grammar or language pattern, and an everyday context\. The specifications covered common Bangla–English differences relevant to beginners, including question formation, tense and auxiliary use, prepositions, introductoryitandthere, passive constructions, and idiomatic expressions\.
We used these specifications to prompt Gemini 3\.5 Flash Lite as the teacher model\. Given the curriculum unit, target CEFR \(Common European Framework of Reference\) level, and Bangla input sentence, we instructed the model to generate a strictly valid JSON trace containing the seven fields shown in Fig\.[3](https://arxiv.org/html/2608.15223#S3.F3)\. The prompt also required no free\-form text outside the JSON object\.
\[Teacher Generation Prompt Template\] System:You are an expert English tutor for NCTB Class 9\-\-10 exams\. Input:Unit: \{unit\}, CEFR: \{cefr\}, Bangla: "\{sentence\}" Task:Return strictly valid JSON conforming to: \{ "word\_gloss": \[\{"bn": "\.\.", "en": "\.\.", "pos": "\.\."\}\], "literal\_translation": "Word\-for\-word English", "natural\_translation": "Fluent English target", "grammar\_notes": \["Bangla L1 contrastive notes", "\.\."\], "common\_mistake": "Plausible beginner L1 error", "practice\_question": "Targeted test question", "practice\_answer": "Expected correct answer" \} Constraint:Output NO text outside JSON\. No markdown fences\.
Fig\. 3:Teacher prompt template used for Gemini 3\.5 Flash Lite trace generation\.### III\-ACurriculum Coverage and Language Design
TRACE\-BN contains 4,099 examples spanning 15 CEFR A1–A2 topic clusters and 13 NCTB grammar units\. Topics include family, school, food, health, weather, directions, shopping, and daily routines\. Grammar coverage includes passive voice, verbs and tenses, pronouns, prepositions, modals, tag questions, conditionals, sentence transformation, indirect narration, and introductoryitandthere\.
Learner\-facing scaffolding, including grammar notes and practice questions, is provided in Bangla, while target\-language fields remain in English\.
### III\-BTrace Filtering and Final Dataset
We apply three automatic quality filters before training\. First, structural validation requires valid JSON with all seven fields, non\-empty grammar notes, and a word\-gloss length between0\.5×0\.5\\timesand2\.0×2\.0\\timesthe source sentence length\. Second, Unicode script validation checks the Bangla source and target translation fields to prevent cross\-lingual language\-swapping errors\. Third, semantic deduplication uses sentence embeddings fromparaphrase\-multilingual\-MiniLM\-L12\-v2and removes near\-duplicates with cosine similarity\>0\.92\>0\.92within the same topic cluster\.
From 4,450 teacher\-generated candidates, 351 \(7\.9%\) were removed: 180 for structural errors, 45 for script\-integrity violations, and 126 as semantic duplicates\. The final dataset contains 4,099 traces, with mean Bangla source and English target lengths of 5\.5 and 6\.6 words, respectively\.
The dataset is partitioned into 3,667 training examples and a 432\-example evaluation split\. The split is organized by topic clusters, including held\-out domains such as weather conditions and idiomatic expressions, to test both sentence\-level and topic\-level generalization\. No source\-sentence overlap or detected lexical leakage exists between the evaluation and training sets\.
As a human validation of the supervision signal, three bilingual English/Bangla language educators independently evaluated 100 randomly sampled TRACE\-BN traces across the curriculum\. Across five pedagogical dimensions on a 1–5 Likert scale, the supervision traces achieved mean ratings of4\.81±0\.584\.81\\pm 0\.58for translation quality,4\.81±0\.624\.81\\pm 0\.62for grammar explanations,4\.79±0\.684\.79\\pm 0\.68for learner\-mistake plausibility,4\.44±1\.174\.44\\pm 1\.17for practice\-question alignment, and4\.71±0\.594\.71\\pm 0\.59overall \(95\.7%95\.7\\%rated≥4\\geq 4\)\. No factual or grammatical error reached annotator consensus in the audited sample\. This human validation supports the quality of the supervision data, while the larger dual\-judge evaluation assesses the behavior learned by the adapted model\. Table[I](https://arxiv.org/html/2608.15223#S3.T1)records the full dataset composition\.
Fig\. 4:Qualitative comparison of base vs\. TRACE\-BN\-tuned Qwen3\-0\.6B outputs on representative held\-out Bangla inputs\. The base model produces literal word substitutions or degenerative loops, whereas the tuned model outputs fluent target translations alongside contrastive Bangla grammar explanations\.TABLE I:Key specifications and composition of the TRACE\-BN dataset\.Dataset AttributeSpecification / CountCorpus Scale & PartitionsTotal structured tutoring traces4,099Training split \(DtrainD\_\{\\text\{train\}\}\)3,667 \(89\.5%\)Held\-out evaluation split \(DevalD\_\{\\text\{eval\}\}\)432 \(10\.5%\)Curriculum & Pedagogical ScopeTarget learner levelCEFR A1–A2 \(Beginner\)NCTB curriculum grammar units13 unitsThematic topic clusters15 situational domainsTrace Schema & Sequence LengthsStructured fields per trace7 multi\-task fieldsMean Bangla source length5\.5 wordsMean English natural translation length6\.6 words
## IVModel Adaptation and Evaluation
### IV\-ALoRA Fine\-Tuning
We fine\-tune Qwen3\-0\.6B\[[18](https://arxiv.org/html/2608.15223#bib.bib23)\]using LoRA\[[5](https://arxiv.org/html/2608.15223#bib.bib24)\]with 4\-bit quantization through Unsloth on a single 16 GB Google Colab T4 GPU\. The training input is the Bangla sentence together with the target schema specification; the model is trained to generate the complete seven\-field tutoring trace as a single response\. LoRA uses rankr=16r=16, scaling factorα=32\\alpha=32, zero dropout, and targets theq, k, v, o, gate, up, downprojection layers\. Training runs for a fixed schedule of three epochs \(621 optimizer steps\) with a cosine learning\-rate schedule peaking at2×10−42\\times 10^\{\-4\}; loss is logged at 51\-step intervals \(Fig\.[5](https://arxiv.org/html/2608.15223#S4.F5)\)\. The final\-step checkpoint is evaluated on the 432\-example split without hyperparameter tuning or checkpoint cherry\-picking on evaluation data\.
Fig\. 5:Training loss progression of Qwen3\-0\.6B on TRACE\-BN across three epochs \(621 steps\)\. Loss decreases steadily and stabilizes near training completion\.
### IV\-BBaseline Models
We compare the tuned Qwen3\-0\.6B with its untuned base model and three zero\-shot baselines: Gemma\-4 E2B \(2\.3B\) as a larger\-model reference, TigerLLM\-1B as a Bangla\-specialized baseline, and Llama\-3\.2\-1B\-Instruct as a general multilingual baseline\. All models receive the same task instruction, formatted through each model family’s native chat template, and are evaluated on the same 432 held\-out Bangla inputs; the comparison models are not fine\-tuned\. We also piloted BanglaLlama\-3\.2\-3B on a stratified sample of 18 inputs but excluded it from the full comparison after it produced 0% schema\-valid outputs on that sample; the model could not complete the task in a form our metrics could score\.
Fig\. 6:Empirical evaluation of TRACE\-BN adaptation: \(a\) Benchmark model comparison against zero\-shot baselines on schema validity \(%\), chrF\+\+, and BLEU; \(b\) Field\-level pedagogical ratings on a 0–4 anchored scale evaluated independently by dual automated judges \(GPT\-4o\-mini and DeepSeek\-V4\-Flash\-0731\); \(c\) Trace component ablation demonstrating the impact of multi\-task supervision on translation and schema adherence\.
### IV\-CEvaluation
Schema validity requires a response to be valid JSON with all seven fields present, including a parseable Englishnatural\_translationfield\. For BLEU\[[10](https://arxiv.org/html/2608.15223#bib.bib26)\]and chrF\+\+\[[13](https://arxiv.org/html/2608.15223#bib.bib27)\], scores are computed on extractednatural\_translationstrings when parseable; outputs lacking a valid translation field are scored as empty strings against the reference\. Because the teacher model generates both training traces and silver references, these metrics measure agreement with the teacher reference rather than independently adjudicated translation quality\.
To obtain a more granular assessment of tutoring behavior, we conduct an additional field\-level evaluation on the held\-out split using two independent automated judges, GPT\-4o\-mini and DeepSeek\-V4\-Flash\-0731\. Each judge receives the Bangla source sentence, the corresponding TRACE\-BN reference trace, and the generated response\. It then rates seven dimensions \(word\-level gloss accuracy, literal translation accuracy, natural translation accuracy, grammar explanation accuracy, learner\-error diagnosis, practice alignment, and overall tutoring quality\) on a 0–4 anchored ordinal scale, where 4 denotes fully correct and 0 denotes unusable output\. The reference trace provides the intended instructional target against which the generated response is scored\.
The two judges score the same 432 held\-out outputs\. We report mean field\-level scores, the proportion of outputs receiving a score of at least 3, and inter\-judge agreement via quadratic weighted Cohen’sκ\\kappa\[[27](https://arxiv.org/html/2608.15223#bib.bib25)\]\. This field\-level analysis distinguishes local errors from failures affecting the tutoring response more broadly\.
For metric comparisons, we use paired bootstrap resampling with 1,000 iterations and reportpp\-values for the main comparisons\.
## VResults
We first compare the tuned model with the untuned and zero\-shot baselines, then examine pedagogical quality and failure modes, the contribution of individual trace components, qualitative behavior, and structural reliability\.
### V\-AOverall Performance
TABLE II:Performance of zero\-shot baselines and the TRACE\-BN\-tuned Qwen3\-0\.6B on the shared held\-out set \(n=432n=432\)\. Schema validity requires valid JSON with all seven fields present, including a parseable Englishnatural\_translation; translation metrics use the teacher\-model output as a silver reference; Tok/s is inference throughput on the evaluation set\.ModelParamsSchemachrF\+\+↑\\uparrowBLEU↑\\uparrowTok/sGemma\-4 E2B2\.3B90\.5%50\.9133\.5445\.97TigerLLM\-1B1\.0B67\.1%22\.238\.2148\.79Llama\-3\.2\-1B\-Instruct1\.0B12\.7%13\.211\.9573\.11Qwen3\-0\.6B \(base\)0\.6B85\.4%15\.284\.5290\.89Qwen3\-0\.6B \(ours\)0\.6B95\.8%34\.7721\.0375\.47
Ours improves over the untuned Qwen3\-0\.6B by\+19\.49\+19\.49chrF\+\+ \(paired bootstrap, 1,000 resamples,p<0\.05p<0\.05\)\.
The tuned Qwen3\-0\.6B improves both structural validity and reference\-based translation scores over the untuned model \(Fig\.[6](https://arxiv.org/html/2608.15223#S4.F6)\(a\), Table[II](https://arxiv.org/html/2608.15223#S5.T2)\)\. Schema validity increases from 85\.4% to 95\.8%, while chrF\+\+ and BLEU rise from 15\.28 to 34\.77 and from 4\.52 to 21\.03, respectively\. The chrF\+\+ gain over the base model is\+19\.49\+19\.49\(p<0\.05p<0\.05, paired bootstrap\), while the gain over TigerLLM\-1B is\+12\.54\+12\.54\(p<0\.05p<0\.05\)\. Gemma\-4 E2B remains the strongest translation baseline, with 50\.91 chrF\+\+ and 33\.54 BLEU\. On inference throughput, the tuned model reaches 75\.47 tokens/s, comparable to the untuned base model \(90\.89 tokens/s\) and above Gemma\-4 E2B \(45\.97 tokens/s\) despite Gemma\-4 E2B’s 2\.3B parameters, consistent with the offline, on\-device deployment the model targets\.
### V\-BPedagogical Quality and Failure Analysis
TABLE III:Field\-level pedagogical evaluation of the base and TRACE\-BN\-tuned Qwen3\-0\.6B on the shared held\-out set \(n=432n=432\)\. Scores are means on a 0–4 scale; the percentage in parentheses denotes outputs receiving a score of at least 3\. Inter\-judge agreement is reported using quadratic weighted Cohen’sκ\\kappa\.DimensionBaseOursκ\\kappaWord gloss accuracy0\.10 \(0\.0%\)1\.48\(16\.4%\)0\.788Literal translation0\.21 \(0\.0%\)1\.33\(14\.8%\)0\.754Natural translation0\.26 \(1\.9%\)1\.31\(16\.2%\)0\.803Grammar explanation0\.27 \(0\.0%\)1\.48\(9\.7%\)0\.489Learner\-error diagnosis0\.23 \(0\.0%\)1\.47\(11\.3%\)0\.585Practice alignment0\.16 \(0\.2%\)1\.17\(9\.7%\)0\.529Overall tutoring quality0\.13 \(0\.0%\)1\.14\(7\.5%\)0\.700
Schema validity alone does not establish tutoring quality\. The field\-level evaluation shows that TRACE\-BN fine\-tuning improves every evaluated tutoring dimension relative to the untuned Qwen3\-0\.6B \(Fig\.[6](https://arxiv.org/html/2608.15223#S4.F6)\(b\), Table[III](https://arxiv.org/html/2608.15223#S5.T3)\)\. The largest gains are observed in word\-level gloss accuracy \(\+1\.38\+1\.38\), learner\-error diagnosis \(\+1\.24\+1\.24\), grammar explanation accuracy \(\+1\.21\+1\.21\), and literal translation accuracy \(\+1\.12\+1\.12\)\. The tuned model also improves natural translation quality from 0\.26 to 1\.31 and overall tutoring quality from 0\.13 to 1\.14 on the 0–4 scale\. Although the absolute scores remain below the upper end of the rubric, the consistent improvement across all fields suggests that the effect of fine\-tuning extends beyond output formatting\.
Inter\-judge agreement is substantial for word\-level glosses \(κ=0\.788\\kappa=0\.788\), literal translation \(κ=0\.754\\kappa=0\.754\), natural translation \(κ=0\.803\\kappa=0\.803\), and overall tutoring quality \(κ=0\.700\\kappa=0\.700\), with moderate agreement for grammar explanation \(κ=0\.489\\kappa=0\.489\), learner\-error diagnosis \(κ=0\.585\\kappa=0\.585\), and practice alignment \(κ=0\.529\\kappa=0\.529\)\. These differences are expected because fine\-grained pedagogical judgments involve greater interpretive variation than translation correctness\. Overall, the agreement results support the use of field\-level automated evaluation as a scalable complement to human validation rather than as a substitute for learner studies\.
Qualitative inspection indicates that remaining errors are concentrated mainly in held\-out idiomatic expressions and in fine\-grained word\-level gloss decisions\. These localized failures are distinct from the structural failures captured by schema validity\.
### V\-CTrace Component Ablation
TABLE IV:Ablation of tutoring trace components on Qwen3\-0\.6B \(n=432n=432\)\.Training ConfigurationTarget FieldsSchemachrF\+\+↑\\uparrowTranslation Only1 \(String\)N/A31\.84Translation \+ Explanation2 \(JSON\)91\.2%33\.12Full TRACE\-BN \(Ours\)7 \(JSON\)95\.8%34\.77
We evaluate three supervision configurations \(Fig\.[6](https://arxiv.org/html/2608.15223#S4.F6)\(c\), Table[IV](https://arxiv.org/html/2608.15223#S5.T4)\)\. Translation\-only training reaches 31\.84 chrF\+\+, while adding Bangla explanations increases chrF\+\+ to 33\.12\. The full seven\-field TRACE\-BN configuration reaches 34\.77 chrF\+\+, suggesting that richer structured supervision benefits even the translation component alone, rather than only adding auxiliary output fields\.
### V\-DQualitative Behavior
Beyond schema compliance, the tuned model generates fluent translations with targeted grammar notes \(Fig\.[4](https://arxiv.org/html/2608.15223#S3.F4)\)\. For question formation, it correctly identifies auxiliarydoinsertion rather than translating the Bangla question marker literally\.
### V\-EStructured Reliability
The 95\.8% schema validity rate corresponds to 18 invalid outputs \(n=432n=432\), versus 63 for the base model\. To confirm these results transfer to the target deployment setting, we additionally exported the merged adapter to GGUF format and confirmed matching outputs underllama\.cpplocal inference\.
## VIConclusion
We presented TRACE\-BN, a curriculum\-guided dataset for structured Bangla\-to\-English tutoring, and showed that Qwen3\-0\.6B can learn to generate seven\-field tutoring traces with 95\.8% schema validity and a\+19\.49\+19\.49chrF\+\+ improvement over the base model\. Field\-level evaluation with two independent judges shows consistent improvement across all evaluated tutoring dimensions, with moderate\-to\-substantial inter\-judge agreement, while a 100\-trace human audit supports the quality of the supervision data\. These results show that curriculum\-guided structured supervision can transfer tutoring structure and content to a sub\-1B model, achieving higher schema validity and throughput than the larger Gemma\-4 E2B zero\-shot baseline, while remaining behind it on translation quality\. The same supervision strategy provides a basis for future extensions to other language pairs and curricula; establishing whether such tutors improve learners’ English proficiency will require controlled learner studies\.
## AI Disclosure
LLMs were used for teacher\-trace synthesis and editorial assistance; human authors designed all methodology and verified all results\.
## References
- \[1\]\(2024\)KD\-LoRA: a hybrid approach to efficient fine\-tuning with LoRA and knowledge distillation\.arXiv preprint arXiv:2410\.20777\.Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p1.1)\.
- \[2\]T\. D\. Belay, S\. K\. Nahin, I\. A\. Azime, O\. Monjur, M\. Rei, C\. Biemann, S\. H\. Muhammad, S\. M\. Yimam, and A\. Chhabra\(2026\)AFRILANGTUTOR: advancing language tutoring and culture education in low\-resource languages with large language models\.arXiv preprint arXiv:2604\.20996\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p2.1)\.
- \[3\]G\. Cook\(2010\)Translation in language teaching: an argument for reassessment\.Oxford University Press\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p2.1)\.
- \[4\]R\. Ellis, S\. Loewen, and R\. Erlam\(2006\)Implicit and explicit corrective feedback and the acquisition of L2 grammar\.Studies in Second Language Acquisition28\(2\),pp\. 339–368\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p2.1)\.
- \[5\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2021\)LoRA: low\-rank adaptation of large language models\.arXiv preprint arXiv:2106\.09685\.Cited by:[§IV\-A](https://arxiv.org/html/2608.15223#S4.SS1.p1.1)\.
- \[6\]S\. Kabir Nahin, R\. Nath Nandi, S\. Sarker, Q\. Sarwar Muhtaseem, M\. Kowsher, A\. Chandraw Shill, M\. Ibrahim, M\. H\. Menon, T\. Al Muntasir, and F\. Alam\(2025\)TituLLMs: a family of Bangla LLMs with comprehensive benchmarking\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 24922–24940\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p1.1)\.
- \[7\]V\. Leonardi\(2010\)The role of pedagogical translation in second language acquisition: from theory to practice\.Peter Lang\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p2.1)\.
- \[8\]Y\. Lu, H\. Li, X\. Cong, Z\. Zhang, Y\. Wu, Y\. Lin, Z\. Liu, F\. Liu, and M\. Sun\(2025\)Learning to generate structured output with schema reinforcement learning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 4905–4918\.Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p2.1)\.
- \[9\]T\. Mahfuz, S\. K\. Dey, R\. Naswan, H\. Adil, K\. S\. Sayeed, and H\. S\. Shahgir\(2025\)Too late to train, too early to use? a study on necessity and viability of low\-resource Bengali LLMs\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 1183–1200\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p3.1)\.
- \[10\]K\. Papineni, S\. Roukos, T\. Ward, and W\. Zhu\(2002\)BLEU: a method for automatic evaluation of machine translation\.InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics,pp\. 311–318\.Cited by:[§IV\-C](https://arxiv.org/html/2608.15223#S4.SS3.p1.1)\.
- \[11\]N\. Phan, T\. Grósz, and M\. Kurimo\(2023\)CaptainA—a mobile app for practising Finnish pronunciation\.InProceedings of the 24th Nordic Conference on Computational Linguistics \(NoDaLiDa\),pp\. 265–270\.Cited by:[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p2.1)\.
- \[12\]M\. Pokrass, C\. Colby, M\. Guan, T\. Sanders, and B\. Zhang\(2024\)Introducing structured outputs in the API\.Note:OpenAI Blog,https://openai\.com/index/introducing\-structured\-outputs\-in\-the\-api/Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p2.1)\.
- \[13\]M\. Popović\(2017\)chrF\+\+: words helping character n\-grams\.InProceedings of the Second Conference on Machine Translation,pp\. 612–618\.Cited by:[§IV\-C](https://arxiv.org/html/2608.15223#S4.SS3.p1.1)\.
- \[14\]N\. Raihan and M\. Zampieri\(2025\)TigerLLM—a family of Bangla large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 887–896\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p1.1)\.
- \[15\]K\. R\. I\. Reza and O\. I\. Shahid\(2026\)KrishokChat: a citation\-grounded dataset and benchmark for Bengali agricultural advisory\.arXiv preprint arXiv:2606\.29243\.Cited by:[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p1.1)\.
- \[16\]M\. Schall and G\. De Melo\(2025\)The hidden cost of structure: how constrained decoding affects language model performance\.InProceedings of the 15th International Conference on Recent Advances in Natural Language Processing,pp\. 1074–1084\.Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p2.1)\.
- \[17\]Z\. R\. Tam, C\. Wu, Y\. Tsai, C\. Lin, H\. Lee, and Y\. Chen\(2024\)Let me speak freely? a study on the impact of format restrictions on performance of large language models\.arXiv preprint arXiv:2408\.02442\.Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p2.1)\.
- \[18\]Q\. Teamet al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§IV\-A](https://arxiv.org/html/2608.15223#S4.SS1.p1.1)\.
- \[19\]P\. Tushar, B\. Zhang, I\. Atmosukarto, D\. Soh, R\. Tong, and I\. McLoughlin\(2026\)Personalized AI\-directed tutoring for oral proficiency enhancement in language education\.Applied Sciences16\(5\),pp\. 2379\.Cited by:[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p2.1)\.
- \[20\]B\. VanPatten\(2004\)Processing instruction: theory, research, and commentary\.Routledge\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p2.1)\.
- \[21\]C\. Wang, J\. Yan, Y\. Yue, and J\. Huang\(2025\)DistilQwen2\.5: industrial practices of training distilled open lightweight language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 6: Industry Track\),pp\. 32–42\.Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p1.1)\.
- \[22\]Y\. Xiao, L\. Song, R\. Yang, C\. Cheng, Y\. Ge, X\. Li, and Y\. Shan\(2025\)LoRA\-Gen: specializing large language models via online LoRA generation\.arXiv preprint arXiv:2506\.11638\.Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p1.1)\.
- \[23\]J\. Yang, D\. Jiang, L\. He, S\. Siu, Y\. Zhang, D\. Liao, Z\. Li, H\. Zeng, Y\. Jia, H\. Wang,et al\.\(2025\)StructEval: benchmarking LLMs’ capabilities to generate structural outputs\.arXiv preprint arXiv:2505\.20139\.Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p2.1)\.
- \[24\]R\. Yi, X\. Li, W\. Xie, Z\. Lu, C\. Wang, A\. Zhou, S\. Wang, X\. Zhang, and M\. Xu\(2024\)PhoneLM: an efficient and capable small language model family through principled pre\-training\.arXiv preprint arXiv:2411\.05046\.Cited by:[§II\-B](https://arxiv.org/html/2608.15223#S2.SS2.p1.1)\.
- \[25\]A\. K\. Zehady, S\. R\. Dipta, N\. Islam, S\. Al Mamun, and S\. Karmaker\(2026\)BanglaLlama: LLaMA for Bangla language\.InProceedings of the Second Workshop on Language Models for Low\-Resource Languages \(LoResLM 2026\),pp\. 73–89\.Cited by:[§I](https://arxiv.org/html/2608.15223#S1.p1.1),[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p1.1)\.
- \[26\]T\. Zhang, L\. Yang, E\. Chen, K\. Riani, J\. Zipf, M\. Shimabukuro, and E\. A\. Lee\(2025\)Learning low\-resource languages through NLP\-driven flashcards: a case study of Hokkien in language learning applications\.InProceedings of NAACL 2025 \(System Demonstrations\),pp\. 303–312\.Cited by:[§II\-A](https://arxiv.org/html/2608.15223#S2.SS1.p2.1)\.
- \[27\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.\(2023\)Judging LLM\-as\-a\-judge with MT\-Bench and chatbot arena\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 46595–46623\.Cited by:[§IV\-C](https://arxiv.org/html/2608.15223#S4.SS3.p3.1)\.Similar Articles
Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation
This paper quantifies the magnitude of subliminal behavioral transfer in language model distillation, showing that undesirable traits can transfer robustly from teacher to student models even with benign training data, and that transfer scales differently across model families.
CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
CSTutorBench is a benchmark for evaluating small language models as tutors for block-based programming, focusing on pedagogical behaviors. Preliminary results show models struggle with deeper tutoring skills like avoiding answer leakage, and prompt revision improves scores.
Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR
This paper proposes a tokenizer transplantation pipeline for lightweight ASR models like Moonshine to address autoregressive collapse in Bengali. By replacing the English-centric tokenizer with a BanglaBERT WordPiece vocabulary, token fertility drops from 9.16 to 1.30 and sequence length by 85.8%, achieving 21.54% WER on the Lipi-Ghor dataset.
Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generation
This paper introduces BLADE, a culturally aligned instruction-tuning dataset of 4,196 interaction pairs for fixing honorific failures and pragmatic gaps in multilingual Bangla generation. Fine-tuning models like DeepSeek-8B and LLaMA-3.2-3B on this dataset yields substantial improvements in structural fidelity and honorific alignment.
Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection
This paper investigates why LLMs underperform in Arabic medical tasks, showing via mechanistic analysis that knowledge exists internally but fails to surface, then proposes TLoRA, a targeted low-rank adaptation method that outperforms full-network LoRA on medical QA and introduces a new Arabic clinical dialogue benchmark.