Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA

arXiv cs.LG Papers

Summary

This paper proposes a Damage Cause Encoder for bridge diagnostic agents that uses knowledge triple extraction from manuals and retrieval-augmented fine-tuning with QLoRA, achieving high accuracy with lower memory usage.

arXiv:2607.21680v1 Announce Type: new Abstract: Bridge infrastructure deteriorates gradually, yet its root causes---salt intrusion, freezing, fatigue cracking, and others---remain invisible to the naked eye. Expert diagnosis relies on tacit knowledge built over years of practice. We address the challenge of automating this latent causal reasoning by proposing a Damage Cause Encoder that classifies 10-class damage causes from visible damage descriptions $S_i$ for use in autonomous bridge diagnostic agents. Our approach chains three components: (i)Knowledge Triple Extraction---a large language model extracts causal triples of the form (damage $\xrightarrow{\mathtt{caused\_by}}$ cause) from 15--35 diagnostic PDF manuals and indexes them in a FAISS vector store; (ii)Retrieval-Augmented Context---at training and inference time, relevant causal triples $\mathcal{C}_i$ are retrieved and concatenated with $S_i$, converting implicit domain knowledge into explicit Encoder context; (iii)Systematic Fine-tuning Comparison---we conduct a rigorous comparison of LoRA, QLoRA, and QA-LoRA on a fixed Golden Testset (116 stratified samples), demonstrating that QLoRA achieves the optimal trade-off: identical test accuracy (87.07%) to full-precision LoRA, 11% faster inference, 72% lower GPU memory, and superior generalization across diverse unseen inputs. A controlled Golden Testset---stratified, deduplicated, and difficulty-tagged---is introduced as a reusable benchmark contribution. QLoRA further outperforms LoRA by 13 percentage points on a 100-sample diverse evaluation spanning all 10 damage cause classes.These findings enable memory-efficient, high-accuracy diagnostic agents on consumer-grade hardware for edge deployment.
Original Article
View Cached Full Text

Cached at: 07/27/26, 07:41 AM

# Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA
Source: [https://arxiv.org/html/2607.21680](https://arxiv.org/html/2607.21680)
###### Abstract

Bridge infrastructure deteriorates gradually, yet its root causes—salt intrusion, freezing, fatigue cracking, and others—remain*invisible*to the naked eye\. Expert diagnosis relies on tacit knowledge built over years of practice\. We address the challenge of automating this latent causal reasoning by proposing aDamage Cause Encoderthat classifies 10\-class damage causes from visible damage descriptionsSiS\_\{i\}for use in autonomous bridge diagnostic agents\.

Our approach chains three components: \(i\)*Knowledge Triple Extraction*—a large language model extracts causal triples of the form \(damage→𝚌𝚊𝚞𝚜𝚎𝚍​\_​𝚋𝚢\\xrightarrow\{\\mathtt\{caused\\\_by\}\}cause\) from 15–35 diagnostic PDF manuals and indexes them in a FAISS vector store; \(ii\)*Retrieval\-Augmented Context*—at training and inference time, relevant causal triples𝒞i\\mathcal\{C\}\_\{i\}are retrieved and concatenated withSiS\_\{i\}, converting implicit domain knowledge into explicit Encoder context; \(iii\)*Systematic Fine\-tuning Comparison*—we conduct a rigorous comparison of LoRA, QLoRA, and QA\-LoRA on a fixed Golden Testset \(116 stratified samples\), demonstrating that QLoRA achieves the optimal trade\-off: identical test accuracy \(87\.07%\) to full\-precision LoRA, 11% faster inference, 72% lower GPU memory, and superior generalization across diverse unseen inputs\.

A controlled*Golden Testset*—stratified, deduplicated, and difficulty\-tagged—is introduced as a reusable benchmark contribution\. QLoRA further outperforms LoRA by 13 percentage points on a 100\-sample diverse evaluation spanning all 10 damage cause classes\. We identify persistent failure modes for Water Accumulation, Soil Liquefaction, and ASR classes, providing actionable guidance for targeted data augmentation\. These findings enable memory\-efficient, high\-accuracy diagnostic agents on consumer\-grade hardware for edge deployment\.

Keywords:Bridge Inspection, Damage Cause Encoding, Knowledge Triples, QLoRA, LoRA Comparison, Golden Testset, Retrieval\-Augmented Generation, Diagnostic Agent, Tacit Knowledge\.

## 1Introduction

### 1\.1Background and Motivation

Japan operates more than 730,000 road bridges, the majority of which were built during the high\-growth era of the 1960–70s and are now exceeding their 50\-year design life\[[17](https://arxiv.org/html/2607.21680#bib.bib10)\]\. Periodic visual inspection is legally mandated under the Road Act, yet the inspection backlog is widening as the number of structures outpaces the number of qualified engineers\.

A fundamental asymmetry underlies the inspection process:damage is visible but its cause is not\. An inspector can observe surface cracking, spalling, or efflorescence, yet correctly diagnosing the root cause—whether salt\-induced rebar corrosion, freeze–thaw cycling, alkali–silica reaction \(ASR\), or fatigue from traffic loading—demands years of domain expertise encoded as tacit knowledge\. This invisible causal layer is precisely what impedes automation: off\-the\-shelf image classifiers label damage types but cannot explain*why*the damage occurred\.

We argue that this tacit causal knowledge can beencodedfrom existing diagnostic manuals in the form of structured causal triples and subsequently injected into a pre\-trained language model via retrieval\-augmented fine\-tuning\.

### 1\.2Problem Statement

Given a damage descriptionSiS\_\{i\}\(a natural\-language text observing structural anomalies\), predict the damage cause categoryc^i∈\{0,…,9\}\\hat\{c\}\_\{i\}\\in\\\{0,\\ldots,9\\\}from a closed set of 10 domain\-specific labels \(Table[1](https://arxiv.org/html/2607.21680#S3.T1)\)\. The key challenge is thatSiS\_\{i\}alone is often insufficient; cause inference requires knowledge of deterioration mechanisms, material properties, and environmental exposure—all encapsulated in diagnostic manuals as implicit expert knowledge\.

### 1\.3Contributions

This paper makes four contributions:

1. 1\.Triple\-Guided Retrieval\-Augmented Encoding: We convert tacit causal knowledge from 15–35 PDF diagnostic manuals into 6,745 structured triples indexed in FAISS, enabling context\-aware damage cause classification at training and inference time\.
2. 2\.Golden Testset Construction: We introduce a stratified, deduplicated, and difficulty\-tagged benchmark of 116 samples spanning all 10 damage cause classes \(Section[5\.6](https://arxiv.org/html/2607.21680#S5.SS6)\), enabling reproducible comparison across fine\-tuning methods\.
3. 3\.QLoRA as Optimal Fine\-tuning Strategy: A controlled comparison of LoRA, QLoRA, and QA\-LoRA on the Golden Testset demonstrates that QLoRA achieves identical test accuracy \(87\.07%\) to full\-precision LoRA while delivering 11% faster inference, 72% lower GPU memory \(1\.45 GB→\\to0\.40 GB\), and superior generalization on 100 diverse unseen samples \(47\.0% vs\. 34\.0%\)\.
4. 4\.Class\-wise Failure Mode Analysis: A 100\-sample diverse evaluation across all 10 classes reveals systematic failure classes \(Water Accumulation: 0%, Soil Liquefaction:≤\\leq10%, ASR:≤\\leq20%\), identifying where targeted data augmentation is most needed\.

## 2Related Work

### 2\.1BERT Fine\-tuning for Sequence Classification

BERT\[[6](https://arxiv.org/html/2607.21680#bib.bib1)\]established the paradigm of pre\-training large Transformer models and fine\-tuning them on downstream tasks\. For classification, a linear head is appended to the\[CLS\]token representation\. Japanese\-specific models such ascl\-tohoku/bert\-large\-japanese\-v2\[[21](https://arxiv.org/html/2607.21680#bib.bib9)\]extend this paradigm to morphologically rich Japanese text using character\-level tokenization, making them well\-suited for domain\-specific inspection texts that frequently include technical compound nouns\.

### 2\.2Parameter\-Efficient Fine\-Tuning and Quantization

LoRA\[[10](https://arxiv.org/html/2607.21680#bib.bib2)\]inserts low\-rank adapters into pre\-trained weight matrices, enabling efficient fine\-tuning with only 0\.1–2% of total parameters active\. QLoRA\[[5](https://arxiv.org/html/2607.21680#bib.bib3)\]extends LoRA by 4\-bit quantizing the frozen backbone \(NF4 format\) while keeping adapters in BFloat16, reducing memory from∼\\sim28 GB to∼\\sim6 GB for 7B\-parameter models\.

QA\-LoRA\[[24](https://arxiv.org/html/2607.21680#bib.bib4)\]further introduces quantization\-aware initialization of LoRA adapters, allowing them to compensate for quantization error during training rather than only during inference\. Our work applies these principles to encoder\-only BERT models for sequence classification—a setting not addressed in prior work—and uncovers a compatibility issue with task\-specific classifier layers that requires a dedicated solution \(Section[4\.3\.2](https://arxiv.org/html/2607.21680#S4.SS3.SSS2)\)\.

### 2\.3Retrieval\-Augmented Generation

Lewis et al\.\[[15](https://arxiv.org/html/2607.21680#bib.bib5)\]demonstrated that augmenting language model inputs with retrieved passages from a non\-parametric memory significantly improves performance on knowledge\-intensive tasks\. Dense Passage Retrieval\[[13](https://arxiv.org/html/2607.21680#bib.bib11)\]showed that FAISS\-indexed dense embeddings outperform BM25 for open\-domain question answering\. Graph\-based RAG\[[7](https://arxiv.org/html/2607.21680#bib.bib14)\]further structures the retrieved knowledge as entity\-relation graphs, analogous to our triple\-based approach\.

Our work differs in that retrieval serves a*discriminative classification*task rather than generative text production, and the retrieved knowledge is*causal*—linking damage observations to deterioration mechanisms rather than factual passages\.

### 2\.4Bridge Inspection and Damage Assessment

Computer vision approaches to bridge damage detection have advanced significantly, with CNN\-based crack detectors\[[14](https://arxiv.org/html/2607.21680#bib.bib13)\]and point\-cloud segmentation methods achieving practical deployment\. Chen et al\.\[[3](https://arxiv.org/html/2607.21680#bib.bib12)\]applied deep learning to automated surface crack inspection\. However, existing systems focus predominantly on damage*detection*and*localization*rather than*cause inference*—the gap our work addresses\. Yasuno\[[25](https://arxiv.org/html/2607.21680#bib.bib17)\]explored quantized VLMs for damage description quality, providing complementary evidence on quantization trade\-offs in infrastructure inspection for diagnostic purposes\.

### 2\.5Agentic AI for Infrastructure and Industrial Inspection

Recent advances in LLM\-based autonomous agents\[[22](https://arxiv.org/html/2607.21680#bib.bib18),[23](https://arxiv.org/html/2607.21680#bib.bib19)\]have opened new possibilities for intelligent inspection systems\. Bommasani et al\.\[[1](https://arxiv.org/html/2607.21680#bib.bib25)\]argued that foundation models can be adapted to specialized downstream tasks via fine\-tuning, motivating the use of domain\-specific encoder modules within broader agentic pipelines\. Shen et al\.\[[20](https://arxiv.org/html/2607.21680#bib.bib26)\]demonstrated that a planning LLM can orchestrate specialized task\-specific models as tools—precisely the role our Damage Cause Encoder is designed to fill in a diagnostic agent\. Park et al\.\[[18](https://arxiv.org/html/2607.21680#bib.bib27)\]further showed that agents equipped with specialized memory and reasoning tools exhibit emergent planning behaviors\.

Condition monitoring and fault diagnosis represent natural application domains for such architectures\. Farrar and Worden\[[8](https://arxiv.org/html/2607.21680#bib.bib20)\]established structural health monitoring \(SHM\) as a systematic framework for damage detection, localization, and prognosis across civil infrastructure\. Zhao et al\.\[[26](https://arxiv.org/html/2607.21680#bib.bib21)\]demonstrated deep learning’s applicability to machine health monitoring across rotating machinery, bearings, and gearboxes\. Jardine et al\.\[[12](https://arxiv.org/html/2607.21680#bib.bib24)\]surveyed condition\-based maintenance with diagnostic and prognostic algorithms covering multiple industrial domains\. Cha et al\.\[[2](https://arxiv.org/html/2607.21680#bib.bib22)\]and Maeda et al\.\[[16](https://arxiv.org/html/2607.21680#bib.bib23)\]applied deep neural networks to crack detection and road damage classification, respectively—both demonstrating that modest\-scale domain\-specific training data enables practical inspection automation\.

Our work complements these efforts by targeting the*causal reasoning*gap: existing inspection AI detects*what*damage is present, while the Damage Cause Encoder infers*why*, providing the root\-cause reasoning module that agentic diagnostic systems require to identify root causes\.

## 3Problem Formulation

### 3\.1Definitions

Let𝒟=\{\(Si,ci\)\}i=1N\\mathcal\{D\}=\\\{\(S\_\{i\},c\_\{i\}\)\\\}\_\{i=1\}^\{N\}be a labeled dataset whereSiS\_\{i\}is a natural\-language damage description \(Japanese text\) andci∈\{0,1,…,9\}c\_\{i\}\\in\\\{0,1,\\ldots,9\\\}is the ground\-truth damage cause label\. A*causal triple*is a structured tuple:

τ=\(subject,𝚛𝚎𝚕𝚊𝚝𝚒𝚘𝚗,object\)\\tau=\(\\text\{subject\},~\\mathtt\{relation\},~\\text\{object\}\)\(1\)where𝚛𝚎𝚕𝚊𝚝𝚒𝚘𝚗∈\{𝚌𝚊𝚞𝚜𝚎𝚍\_𝚋𝚢,𝚊𝚌𝚌𝚎𝚕𝚎𝚛𝚊𝚝𝚎𝚍\_𝚋𝚢,\\mathtt\{relation\}\\in\\\{\\mathtt\{caused\\\_by\},\\,\\mathtt\{accelerated\\\_by\},\\,𝚛𝚎𝚕𝚊𝚝𝚎𝚍\_𝚝𝚘\}\\mathtt\{related\\\_to\}\\\}\. For example: \(rebar corrosion,caused\_by,chloride ion concentration exceeding 1\.2 kg/m3\)\.

Let𝒯=\{τ1,…,τM\}\\mathcal\{T\}=\\\{\\tau\_\{1\},\\ldots,\\tau\_\{M\}\\\}be the set of all extracted triples indexed in a FAISS vector store\. Given a querySiS\_\{i\}, a retrieval functionRet​\(Si,𝒯,k\)\\text\{Ret\}\(S\_\{i\},\\mathcal\{T\},k\)returns the top\-kkmost similar triples by cosine similarity in the embedding space\. After filtering by an LLM relevance judge, the surviving context is:

𝒞i=\{τ∈Ret​\(Si,𝒯,k\)∣LLM\-relevant​\(Si,τ\)=YES\}\\mathcal\{C\}\_\{i\}=\\\{\\tau\\in\\text\{Ret\}\(S\_\{i\},\\mathcal\{T\},k\)\\mid\\text\{LLM\-relevant\}\(S\_\{i\},\\tau\)=\\text\{YES\}\\\}\(2\)

### 3\.2Classification Task

The augmented input to the encoder is:

xi=\[CLS\]​Si​\[SEP\]​o1⋅o2⋅⋯⋅o\|𝒞i\|​\[SEP\]x\_\{i\}=\\texttt\{\[CLS\]\}\\;S\_\{i\}\\;\\texttt\{\[SEP\]\}\\;o\_\{1\}\\,\\cdot\\,o\_\{2\}\\,\\cdot\\,\\cdots\\,\\cdot\\,o\_\{\|\\mathcal\{C\}\_\{i\}\|\}\\;\\texttt\{\[SEP\]\}\(3\)whereojo\_\{j\}is the*object*field of tripleτj∈𝒞i\\tau\_\{j\}\\in\\mathcal\{C\}\_\{i\}\(the causal explanation text\), and “·” denotes the Japanese sentence period\. The fine\-tuned encoderfθf\_\{\\theta\}mapsxix\_\{i\}to a probability distribution over 10 classes:

p^i=softmax​\(Wc⋅fθ​\(xi\)​\[CLS\]\)∈ℝ10\\hat\{p\}\_\{i\}=\\text\{softmax\}\(W\_\{c\}\\cdot f\_\{\\theta\}\(x\_\{i\}\)\[\\texttt\{CLS\}\]\)\\in\\mathbb\{R\}^\{10\}\(4\)and the predicted label isc^i=arg⁡max⁡p^i\\hat\{c\}\_\{i\}=\\arg\\max\\hat\{p\}\_\{i\}\.

### 3\.3Damage Cause Categories

The 10\-class taxonomyC2C\_\{2\}covers the principal deterioration mechanisms identified in Japanese bridge diagnostic manuals \(Table[1](https://arxiv.org/html/2607.21680#S3.T1)\)\. Categories span chemical \(salt, ASR\), physical \(frost, fatigue\), and biological/hydrological \(water retention, sanding\) mechanisms, as well as structural causes \(void tube subsidence, connected girder issues\)\.

Table 1:Damage Cause Categories \(C2C\_\{2\}\)
### 3\.4Domain\-Agnostic Generalization

Although the instantiation above targets bridge inspection, the four\-phase pipeline is parameterized by two domain\-specific components that can be replaced independently:

1. 1\.Document corpus\{Dk\}k=1K\\\{D\_\{k\}\\\}\_\{k=1\}^\{K\}: any collection of expert PDF documents encoding domain causal knowledge \(diagnostic manuals, maintenance guidelines, clinical protocols, etc\.\)\.
2. 2\.Label taxonomyℒ=\{l0,…,lL−1\}\\mathcal\{L\}=\\\{l\_\{0\},\\ldots,l\_\{L\-1\}\\\}: a closed set ofLLcause categories appropriate to the target domain that represent all possible causes\.

Formally, let𝒳\\mathcal\{X\}be the space of observable symptom descriptions and𝒴=\{0,…,L−1\}\\mathcal\{Y\}=\\\{0,\\ldots,L\-1\\\}the cause label space\. The general encoder learns a mapping:

fθ:𝒳×2𝒯→𝒴f\_\{\\theta\}:\\mathcal\{X\}\\times 2^\{\\mathcal\{T\}\}\\rightarrow\\mathcal\{Y\}\(5\)where2𝒯2^\{\\mathcal\{T\}\}denotes the power set of retrieved causal triples\. Sufficient conditions for successful horizontal transfer are:

1. \(i\)Expert diagnostic knowledge exists in text\-parseable document form\.
2. \(ii\)Observable symptoms and latent causes are linguistically separable\.
3. \(iii\)Labeled training data are available at modest scale \(tens to hundreds of samples suffice\[[8](https://arxiv.org/html/2607.21680#bib.bib20),[12](https://arxiv.org/html/2607.21680#bib.bib24)\]\)\.
4. \(iv\)Memory\-efficient deployment is preferred; the proposed QLoRA configuration targets≤\\leq16 GB consumer\-grade GPU, enabling on\-premise deployment without cloud infrastructure\.

Candidate domains include tunnel lining diagnosis, pavement deterioration assessment\[[16](https://arxiv.org/html/2607.21680#bib.bib23)\], rotating machinery fault detection\[[26](https://arxiv.org/html/2607.21680#bib.bib21),[12](https://arxiv.org/html/2607.21680#bib.bib24)\], structural health monitoring\[[8](https://arxiv.org/html/2607.21680#bib.bib20)\], and medical symptom\-to\-diagnosis reasoning\. The agentic AI perspective—equipping autonomous diagnostic agents with specialized encoder tools\[[22](https://arxiv.org/html/2607.21680#bib.bib18),[20](https://arxiv.org/html/2607.21680#bib.bib26),[23](https://arxiv.org/html/2607.21680#bib.bib19)\]—further motivates the design: a lightweight QLoRA\-fine\-tuned encoder can serve as a plug\-in causal reasoning module within LLM\-based inspection agents\[[18](https://arxiv.org/html/2607.21680#bib.bib27),[1](https://arxiv.org/html/2607.21680#bib.bib25)\]\.

## 4Methodology

The proposed pipeline comprises four phases \(Figures[1](https://arxiv.org/html/2607.21680#S4.F1)and[2](https://arxiv.org/html/2607.21680#S4.F2)\)\. Phase A builds a domain knowledge triple index from PDF manuals\. Phase B generates a training dataset by augmenting damage blocks with retrieved causal context\. Phase C fine\-tunes a quantized Encoder model with LoRA adapters\. Phase D executes the same retrieval\-augmented encoding at inference for damage cause reasoning\.

Phase A: Knowledge Triple PipelinePhase B: Training Dataset GenerationPDF Corpus\(15/35 PDFs\)pypdfium2text OCRPaddleOCRfig/tab OCRText\+Fig Blocks\(872 / 1721\)Triple ExtractionQwen2\.5\-7BCausal Triples\(4,186 / 6,745\)FAISS Index1024\-dim cosineLabeled BlocksSiS\_\{i\}\+ labelccDense Retrievaltop\-kktriplesLLM YES/NOFilterContext𝒞i\\mathcal\{C\}\_\{i\}\(relevant triples\)Balanced Samplingmin=14, max=100Training Dataset\(428 / 388 / 642\)⟶\\longrightarrowPhase C\(Fine\-tuning\)FAISS index

Figure 1:Pipeline Phases A–B\.Phase A: PDF manuals are processed by two OCR engines \(pypdfium2 for text, PaddleOCR for figures/tables\) to extract blocks; Qwen2\.5\-7B then extracts causal triples\(σj,rj,oj\)\(\\sigma\_\{j\},r\_\{j\},o\_\{j\}\)indexed in a 1024\-dimensional FAISS store\.Phase B: For each labeled block\(Si,ci\)\(S\_\{i\},c\_\{i\}\), the FAISS index is queried to retrieve top\-kktriples, filtered by an LLM relevance judge, forming context𝒞i\\mathcal\{C\}\_\{i\}\. Balanced sampling \(min=14, max=100 per class\) yields the training dataset\.Phase C: QLoRA Fine\-tuningPhase D: Inference←\\leftarrowDataset \(Phase B\)BERT\-Large\-Ja344M params4\-bit NF4 Quantdouble quantFP16 Classifiermanual replaceprepare\_kbittrainingLoRA Adaptersr=16r\{=\}16,α=32\\alpha\{=\}32Weighted Training20 epochs, lr=2e\-4Fine\-tuned Model7\.1M params \(2\.07%\)NewSiS\_\{i\}\(unseen damage\)Dense Retrievaltop\-kktriplesLLM YES/NOFilterContext𝒞^i\\hat\{\\mathcal\{C\}\}\_\{i\}Encode:\[Si;𝒞i\]\[S\_\{i\};\\,\\mathcal\{C\}\_\{i\}\]C^2\\hat\{C\}\_\{2\}Prediction10\-class softmaxTop\-3 causes\+ confidenceFAISS Index\(Phase A\)weights

Figure 2:Pipeline Phases C–D\.Phase C: BERT\-Large\-Ja \(344M params\) is fine\-tuned with QLoRA \(recommended\); the highlightedFP16 Classifier Replacementresolves the BitsAndBytes incompatibility \(Section[4\.3\.2](https://arxiv.org/html/2607.21680#S4.SS3.SSS2)\), followed by 4\-bit NF4 quantization, LoRA adapter injection \(r=16r\{=\}16,α=32\\alpha\{=\}32\), and weighted cross\-entropy training\.Phase D: At inference, the same retrieval pipeline augments a new damage descriptionSiS\_\{i\}with context𝒞^i\\hat\{\\mathcal\{C\}\}\_\{i\}from the Phase A FAISS index; the fine\-tuned encoder predicts the top\-3 damage causes with confidence scores\.### 4\.1Phase A: Knowledge Triple Extraction

PDF OCR\.Two complementary engines extract content from diagnostic manuals:pypdfium2for text\-layer extraction \(primary, fast\) and PaddleOCR 2\.7 in full\-scan mode for figure and table regions \(fallback, CPU mode, DPI = 200\)\. For 35 PDFs, this yields 870 text blocks and 851 figure/table blocks\.

Triple Extraction\.Each block is passed to Qwen2\.5\-7B\-Instruct\[[19](https://arxiv.org/html/2607.21680#bib.bib8)\]running via Ollama with a structured prompt:

> “Extract causal triples from the following text\. Each triple must have fields: subject, relation \(one of: caused\_by / accelerated\_by / related\_to\), object, stage, damage\_type, bridge\_type, cause\_label\. Return JSON\.”

From 35 PDFs, 6,745 triples are extracted \(relation distribution:caused\_by61\.9%,accelerated\_by18\.0%,related\_to20\.1%\)\.

FAISS Indexing\.Triple texts are embedded withhotchpotch/static\-embedding\-japanese\[[9](https://arxiv.org/html/2607.21680#bib.bib16)\]\(1024\-dimensional static embeddings, 109 it/s on GPU\) and stored in a FAISSIndexFlatIPindex for exact cosine similarity search\.

Formal Triple Extraction\.The extraction step is formalized asℰ:ℬ→2𝒯\\mathcal\{E\}:\\mathcal\{B\}\\rightarrow 2^\{\\mathcal\{T\}\}, whereℬ\\mathcal\{B\}is the set of document blocks\. For each blockb∈ℬb\\in\\mathcal\{B\}, the LLMfΘf\_\{\\Theta\}generates a set of typed causal triples:

ℰ\(b\)=\{\(σj,rj,oj\)∣rj∈ℛ,\(σj,rj,oj\)∈fΘ\(prompt\(b\)\)\}\\mathcal\{E\}\(b\)=\\bigl\\\{\(\\sigma\_\{j\},\\,r\_\{j\},\\,o\_\{j\}\)\\mid r\_\{j\}\\in\\mathcal\{R\},\\\\ \\quad\(\\sigma\_\{j\},r\_\{j\},o\_\{j\}\)\\in f\_\{\\Theta\}\(\\mathrm\{prompt\}\(b\)\)\\bigr\\\}\(6\)whereℛ=\{𝚌𝚊𝚞𝚜𝚎𝚍​\_​𝚋𝚢,𝚊𝚌𝚌𝚎𝚕𝚎𝚛𝚊𝚝𝚎𝚍​\_​𝚋𝚢,𝚛𝚎𝚕𝚊𝚝𝚎𝚍​\_​𝚝𝚘\}\\mathcal\{R\}=\\\{\\mathtt\{caused\\\_by\},\\mathtt\{accelerated\\\_by\},\\mathtt\{related\\\_to\}\\\}andσj\\sigma\_\{j\},ojo\_\{j\}are subject and object texts\. The full knowledge index is𝒯=⋃b∈ℬℰ​\(b\)\\mathcal\{T\}=\\bigcup\_\{b\\in\\mathcal\{B\}\}\\mathcal\{E\}\(b\)\.

### 4\.2Phase B: Training Dataset Generation

Retrieval and Filtering\.For each labeled block\(Si,ci\)\(S\_\{i\},c\_\{i\}\), the top\-k=10k=10triples are retrieved from FAISS with minimum cosine score≥0\.5\\geq 0\.5\. A Qwen2\.5\-7B YES/NO judge then filters to retain only triples causally relevant toSiS\_\{i\}, forming context𝒞i\\mathcal\{C\}\_\{i\}\.

Formal Retrieval Scoring\.Embeddings are computed by encoderϕ:𝒳∪𝒯→ℝd\\phi:\\mathcal\{X\}\\cup\\mathcal\{T\}\\rightarrow\\mathbb\{R\}^\{d\}\(1024\-dimensional static embeddings\[[9](https://arxiv.org/html/2607.21680#bib.bib16)\]\)\. The cosine similarity between a query and a triple is:

sim​\(Si,τj\)=ϕ​\(Si\)⊤​ϕ​\(text​\(τj\)\)‖ϕ​\(Si\)‖​‖ϕ​\(text​\(τj\)\)‖\\mathrm\{sim\}\(S\_\{i\},\\tau\_\{j\}\)=\\frac\{\\phi\(S\_\{i\}\)^\{\\top\}\\,\\phi\(\\mathrm\{text\}\(\\tau\_\{j\}\)\)\}\{\\\|\\phi\(S\_\{i\}\)\\\|\\;\\\|\\phi\(\\mathrm\{text\}\(\\tau\_\{j\}\)\)\\\|\}\(7\)The filtered candidate set is𝒞^i=arg⁡topk​\{τ∈𝒯:sim​\(Si,τ\)≥δ\}\\hat\{\\mathcal\{C\}\}\_\{i\}=\\arg\\mathrm\{top\}\_\{k\}\\\{\\tau\\in\\mathcal\{T\}:\\mathrm\{sim\}\(S\_\{i\},\\tau\)\\geq\\delta\\\}with thresholdδ=0\.5\\delta=0\.5, from which LLM relevance filtering yields the final context𝒞i\\mathcal\{C\}\_\{i\}\.

Balanced Sampling Strategy\.To control class imbalance, we apply minority\-preserving sampling:

ntarget​\(c\)=\{ncif​nc≤nmaxnmaxif​nc\>nmaxn\_\{\\text\{target\}\}\(c\)=\\begin\{cases\}n\_\{c\}&\\text\{if \}n\_\{c\}\\leq n\_\{\\max\}\\\\ n\_\{\\max\}&\\text\{if \}n\_\{c\}\>n\_\{\\max\}\\end\{cases\}\(8\)wherencn\_\{c\}is the available block count for classccandnmaxn\_\{\\max\}is a per\-experiment cap\. For v0\.3 we usenmax=100n\_\{\\max\}=100, reducing the Imbalance Ratio \(abbreviated IR\) from 35\.1 to 7\.7\.

### 4\.3Phase C: Fine\-tuning Strategy Comparison

#### 4\.3\.1Base Model Configuration

We usecl\-tohoku/bert\-large\-japanese\-v2\[[21](https://arxiv.org/html/2607.21680#bib.bib9)\]\(BERT\-Large, 344M parameters, character\-level tokenization\) extended with a 10\-class linear classifier head\. LoRA adapters are applied to \{query, key, value, dense\} attention layers \(r=16r=16,α=32\\alpha=32, dropout = 0\.1\), yielding 7\.1M trainable parameters \(2\.07% of 344M total\)\. Training uses aWeightedLossTrainerwith inverse\-frequency class weights as follows:

wc=N\|ℒ\|⋅nc,IR=maxc⁡ncminc⁡ncw\_\{c\}=\\frac\{N\}\{\|\\mathcal\{L\}\|\\cdot n\_\{c\}\},\\quad\\mathrm\{IR\}=\\frac\{\\max\_\{c\}n\_\{c\}\}\{\\min\_\{c\}n\_\{c\}\}\(9\)whereN=∑cncN=\\sum\_\{c\}n\_\{c\}is total training samples and\|ℒ\|=10\|\\mathcal\{L\}\|=10\. The weighted cross\-entropy objective is:

ℒCE=−1N​∑i=1Nwci​log⁡p^i​\[ci\]\\mathcal\{L\}\_\{\\mathrm\{CE\}\}=\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}w\_\{c\_\{i\}\}\\,\\log\\,\\hat\{p\}\_\{i\}\[c\_\{i\}\]\(10\)wherep^i​\[ci\]\\hat\{p\}\_\{i\}\[c\_\{i\}\]is the predicted probability for the true labelcic\_\{i\}\. Hyperparameters: batch = 8, gradient accumulation = 4 \(effective batch = 32\), lr = 2e\-4, 20 epochs, AdamW optimizer, warmup ratio = 0\.1\.

For QLoRA and QA\-LoRA, the backbone is 4\-bit quantized with BitsAndBytes\[[4](https://arxiv.org/html/2607.21680#bib.bib6)\]using NF4 format and double quantization \(s1,s2s\_\{1\},s\_\{2\}scale factors\), with bfloat16 compute dtype\.

#### 4\.3\.2BitsAndBytes Incompatibility Fix

When applying 4\-bit quantization toBertForSequenceClassification, anAssertionErrorarises atbitsandbytes/nn/modules\.py:415because the classifier head\(dhidden,nclass\)=\(1024,10\)\(d\_\{\\text\{hidden\}\},\\,n\_\{\\text\{class\}\}\)=\(1024,10\)violates the weight\-shape assumption for 4\-bit kernels\. This issue is not specific to bridge inspection—it surfaces whenever a small\-output linear head is attached to a large quantized backbone, and the same fix applies in any domain\.

The resolution follows a four\-step procedure, using BitsAndBytes\[[4](https://arxiv.org/html/2607.21680#bib.bib6)\]as the quantization backend:

1. 1\.Identify the incompatible layer\.Inspect the model graph to locate anynn\.Linearlayer whose weight shape is incompatible with 4\-bit kernels \(typically a task\-specific classification head with small output dimension, here\(1024,10\)\(1024,10\)\)\.
2. 2\.Replace with a full\-precision head\.Before quantization setup, substitute the incompatible layer with a standardnn\.Linearcast tofloat16\. This preserves classifier capacity while keeping it outside the 4\-bit quantization scope\.
3. 3\.Prepare the backbone for quantized training\.Callprepare\_model\_for\_kbit\_training\(\)from the BitsAndBytes / PEFT library\[[11](https://arxiv.org/html/2607.21680#bib.bib15)\]on the model after the head replacement\. This step freezes the quantized backbone weights and enables gradient checkpointing for the LoRA adapters only\.
4. 4\.Attach LoRA adapters\.Applyget\_peft\_model\(\)withLoraConfig\(r=16r\{=\}16,α=32\\alpha\{=\}32, dropout=0\.1\\,\{=\}\\,0\.1, target modules:query,key,value,dense\) to inject low\-rank trainable adapters into the frozen quantized backbone\.

The resulting model has a fully quantized \(4\-bit NF4\) backbone, FP16 classifier head, and BFloat16 LoRA adapters—satisfying BitsAndBytes constraints while maintaining task\-specific classification\. This procedure generalizes to any encoder\-only or decoder\-only model where a small\-dimension head would trigger the weight\-shape assertion\.

#### 4\.3\.3Three Fine\-tuning Variants

We compare three variants under identical LoRA configuration and training hyperparameters:

\(1\) LoRA\(full\-precision baseline\): Standard LoRA\[[10](https://arxiv.org/html/2607.21680#bib.bib2)\]with FP32 backbone weights\. No quantization is applied\.

\(2\) QLoRA\(recommended\): QLoRA\[[5](https://arxiv.org/html/2607.21680#bib.bib3)\]applies 4\-bit NF4 quantization to the frozen backbone, with LoRA adapters in bfloat16\. The FP16 classifier replacement \(above\) is applied\.

\(3\) QA\-LoRA\(error correction variant\): We implement a practical approximation of QA\-LoRA\[[24](https://arxiv.org/html/2607.21680#bib.bib4)\]by adding an L2 regularization loss on LoRA adapter weights to compensate for quantization error, without LoftQ initialization \(which proved incompatible with our manual classifier replacement\):

ℒtotal=ℒCE\+λ⋅1\|ΘLoRA\|​∑𝐀,𝐁∈ΘLoRA\(‖𝐀‖F2\+‖𝐁‖F2\)\\mathcal\{L\}\_\{\\text\{total\}\}=\\mathcal\{L\}\_\{\\text\{CE\}\}\+\\lambda\\cdot\\frac\{1\}\{\|\\Theta\_\{\\text\{LoRA\}\}\|\}\\sum\_\{\\mathbf\{A\},\\mathbf\{B\}\\in\\Theta\_\{\\text\{LoRA\}\}\}\\bigl\(\\\|\\mathbf\{A\}\\\|\_\{F\}^\{2\}\+\\\|\\mathbf\{B\}\\\|\_\{F\}^\{2\}\\bigr\)\(11\)whereλ=0\.01\\lambda=0\.01andΘLoRA\\Theta\_\{\\text\{LoRA\}\}denotes all LoRA adapter matrices \(lora\_A,lora\_B\)\.

#### 4\.3\.4Comparison Results on Golden Testset

Table[2](https://arxiv.org/html/2607.21680#S4.T2)reports evaluation on the Golden Testset \(116 samples; Section[5\.6](https://arxiv.org/html/2607.21680#S5.SS6)\)\.

Table 2:Fine\-tuning method comparison on Golden Testset \(116 samples\)\. Best values inbold\.Key findings:

- •QLoRA matches LoRA in accuracy\(87\.07% vs\. 87\.07%\) while reducing GPU memory by 72% and inference latency by 11%\.
- •QA\-LoRA underperformsby 1\.73% relative to LoRA\. The L2 regularization atλ=0\.01\\lambda=0\.01over\-constrains the adapters, suppressing the model’s ability to compensate for quantization error\. Smallerλ\\lambda\(0\.001–0\.005\) may improve results; LoftQ initialization was not applied due to incompatibility with the manual FP16 replacement\.
- •QLoRA generalizes better: on 100 diverse unseen inputs, QLoRA achieves 47\.0% vs\. 34\.0% for LoRA—a 13\-point gap suggesting that 4\-bit quantization noise acts as implicit regularization on the training distribution \(Section[5\.8](https://arxiv.org/html/2607.21680#S5.SS8)\)\.

Recommendation:QLoRA is the preferred fine\-tuning strategy for production deployment, offering the best balance of accuracy, speed, memory, and generalization for diagnostic purposes\.

### 4\.4Phase D: Inference

At inference, new damage textSitestS\_\{i\}^\{\\text\{test\}\}undergoes the same retrieval pipeline \(FAISS top\-kk\+ LLM YES/NO filter\) to construct context𝒞^i\\hat\{\\mathcal\{C\}\}\_\{i\}\. The augmented inputxix\_\{i\}is passed through the fine\-tuned BERT\+LoRA model to predictc^i\\hat\{c\}\_\{i\}with confidence scores, returning the top\-3 most probable causes\.

## 5Experiments

### 5\.1Experimental Setup

Hardware\.Ollama with Qwen2\.5\-7B occupies 4\.7 GB VRAM during triple extraction and LLM filtering in our pipeline\.

Software\.Python 3\.12\.10 for training \(PyTorch 2\.6\.0\+cu124, transformers, peft, bitsandbytes\); Python 3\.10\.11 for OCR \(PaddleOCR 2\.7\.0, pypdfium2 5\.12\.1, NumPy 1\.25\.2\)\.

Evaluation Metrics\.Validation accuracy and weighted F1 score, evaluated on a 20% stratified held\-out split\.

### 5\.2Datasets

Three progressive dataset versions are used \(Table[3](https://arxiv.org/html/2607.21680#S5.T3)\):

Table 3:Dataset VersionsOriginal, Imbalanced\(v0\.2\):15 PDFs, text blocks only, no sampling constraint\. The Salt class dominates at 65\.7% \(281/428 samples\); Void tube and Fatigue classes have fewer than 10 samples each \(IR = 35\.1\)\. Figure[3](https://arxiv.org/html/2607.21680#S5.F3)illustrates the dramatic imbalance\.

Balanced\(v0\.3\):Same 15 PDFs but text \+ figure/table blocks combined \(872 blocks total, 4,186 triples\), with balanced sampling \(nmax=100n\_\{\\max\}=100\)\. Salt reduced to 25\.8%; all minority classes reach≥\\geq13 samples \(IR = 7\.7\)\.

Scale\-up\(v0\.4\):35 PDFs spanning additional bridge types—culvert bridges \(dobashi\), arches \(arch\), foundations \(kiso\), bearings \(shisho\), abutments and piers \(kyodai/kyakyu\)—yielding 6,745 triples and 642 balanced training samples \(IR = 7\.69\)\. Class distribution: Salt 100, Frost 92, Sanding 98, Fatigue 99, Rebar 27, ASR 82, Void 13, Water 24, Girder 16, Other 91\. Class weights span 0\.642–4\.938 \(7\.7×\\times\), consistent with v0\.3 balance\.

SaltFrostSandFatigueRebarASRVoidWaterGirderOther05050100100150150200200250250300300Number of samplesv0\.2 Original \(IR=35\.1=35\.1\)v0\.3 Balanced \(IR=7\.7=7\.7\)v0\.4 Scale\-up \(IR=7\.69=7\.69\)Figure 3:Class distribution comparison among v0\.2 \(original, imbalanced\), v0\.3 \(balanced\), and v0\.4 \(scale\-up, IR = 7\.69\)\. Salt class is reduced from 281 to 100 samples \(65\.7%→\\to25\.8%\); v0\.4 further expands minority classes \(Frost 92, Fatigue 99, ASR 82, Other 91\) while maintaining the same balanced cap for Salt, Sand, Rebar, Void, Water, and Girder\.
### 5\.3Main Results

Table[4](https://arxiv.org/html/2607.21680#S5.T4)reports progressive dataset results using QLoRA \(our recommended configuration\)\. The Golden Testset comparison \(LoRA vs\. QLoRA vs\. QA\-LoRA\) is reported in Section[4\.3](https://arxiv.org/html/2607.21680#S4.SS3), Table[2](https://arxiv.org/html/2607.21680#S4.T2)\.

Table 4:Progressive dataset results with QLoRA\. Best accuracy per dataset inbold\. All use 4\-bit NF4 quantization with FP16 classifier replacement \(LoRA rows use full\-precision weights\)\.Key observations:

- •QLoRA matches LoRA on Golden Testset\(87\.07% each\) while reducing GPU memory by 72% and training time by 29%\.
- •Balanced data \(v0\.3\):QLoRA*outperforms*LoRA \(85\.90% vs\. 83\.33%,Δ=\+2\.57%\\Delta=\+2\.57\\%\), confirming that 4\-bit quantization noise acts as implicit regularization when class imbalance is low \(IR<<10\)\.
- •Scale\-up \(v0\.4\):QLoRA on 35 PDFs achieves 91\.47%—recovering v0\.2 accuracy while maintaining full memory efficiency \(0\.40 GB\)\.
- •QA\-LoRA underperformson the Golden Testset \(85\.34%\), suggesting thatλ=0\.01\\lambda=0\.01L2 regularization over\-constrains the LoRA adapters\.

### 5\.4Learning Curves

Figure[4](https://arxiv.org/html/2607.21680#S5.F4)shows validation accuracy across training epochs for all four configurations\. On imbalanced data \(v0\.2\), both models converge by epoch 14–16, with LoRA maintaining a consistent∼\\sim1% lead throughout\. On balanced data \(v0\.3\), QLoRA initially lags \(14\.1% at epoch 1 vs\. 18% for LoRA\) but surpasses LoRA around epoch 10 and peaks at epoch 13 \(85\.90%\), demonstrating that quantization regularization requires slightly longer training to manifest\.

11551010151520202020404060608080909010010091\.8690\.7085\.9083\.3391\.47EpochValidation Accuracy \(%\)v0\.2 No Quant \(91\.86%\)v0\.2 QLoRA 4\-bit \(90\.70%\)v0\.3 No Quant \(83\.33%\)v0\.3 QLoRA 4\-bit \(85\.90%\)v0\.4 QLoRA 4\-bit \(91\.47%\)Figure 4:Validation accuracy vs\. epoch for four experimental configurations\. Filled marks indicate the best epoch used for evaluation\. On balanced data \(v0\.3\), QLoRA overtakes LoRA after epoch 10\.
### 5\.5Scale\-up Results

Expanding the corpus to 35 PDFs yields 642 balanced training samples \(IR = 7\.69\), a 2\.3×\\timesscale\-up in data volume\. With QLoRA \(per the guideline in Section[6\.2](https://arxiv.org/html/2607.21680#S6.SS2)\), the model achieves91\.47% accuracy and 91\.39% weighted F1on the 129\-sample validation set—nearly recovering v0\.2’s imbalanced\-data peak \(91\.86%\) but with dramatically improved minority\-class coverage and balanced training\.

Key learning curve characteristics: the model reaches its peak at epoch 6 \(91\.47%\), then plateaus at 90\.70% for epochs 7–20, indicating early convergence driven by the quantization regularization effect on balanced data \(IR = 7\.69<<10\)\. Training completed in 16m 16s \(976\.4 s\) with GPU memory of 0\.40 GB—consistent with all QLoRA configurations regardless of dataset size\.

Summary:corpus scaling \(15→\\to35 PDFs, 388→\\to642 samples\) is the primary lever for recovering accuracy, while QLoRA maintains memory efficiency\.

### 5\.6Golden Testset Construction

To enable reproducible and fair comparison across fine\-tuning methods, we construct a*Golden Testset*using a stratified, deduplicated split of the merged v0\.4 corpus \(767 samples total\)\.

Construction procedure:

- •Deduplication: fingerprint viasource\_block\_id\(if available\) or SHA\-1 of the first 300 characters; v0\.4 samples overwrite v0\.3 duplicates to prioritize richer triple context\.
- •Stratified split: 70/15/15 \(train/val/test\) withseed=42, yielding 536 / 115 / 116 samples\.
- •Difficulty tagging: three levels including Easy/Medium/Hard based on text length, number of retrieved triplesnc​in\_\{ci\}, training stage, and membership in rare classes \(Void Floating, Water Accumulation, Continuous Girder\)\.
- •Canonical labels: label IDs mapped to a fixed class vocabulary to absorb variation across prepared dataset versions\.

Table 5:Golden Testset statistics by split and difficulty\.The fixed seed \(42\) guarantees that the 116\-sample test partition is invariant across all experimental iterations, enabling fair comparison of LoRA, QLoRA, and QA\-LoRA reported in Table[2](https://arxiv.org/html/2607.21680#S4.T2)\.

### 5\.7Confusion Matrix Analysis

Figure[5](https://arxiv.org/html/2607.21680#S5.F5)shows normalized confusion matrices for all three fine\-tuning methods evaluated on the Golden Testset\.

![Refer to caption](https://arxiv.org/html/2607.21680v1/figures/confusion_matrices.png)Figure 5:Confusion matrices for LoRA, QLoRA, and QA\-LoRA on the Golden Testset \(116 samples\)\. Predicted labels \(columns\) vs\. true labels \(rows\); darker shading indicates higher count\. The diagonal represents correct predictions\.Key patterns across all three methods:

- •Common failure pairs:Freeze\-Thaw↔\\leftrightarrowRebar Corrosion \(4 cases each\), Salt Damage↔\\leftrightarrowOthers, and Freeze\-Thaw↔\\leftrightarrowASR exhibit systematic misclassification, likely due to overlapping surface crack patterns in training descriptions\.
- •QLoRA vs\. LoRA:Both achieve 87\.07% accuracy with nearly identical confusion patterns, confirming that 4\-bit quantization preserves decision boundaries\.
- •QA\-LoRA:Shows increased confusion toward Continuous Girder, suggesting that the L2 regularization biases predictions toward structurally\-specific language patterns\.

### 5\.8Qualitative Evaluation: Unseen Examples

To evaluate generalization beyond the Golden Testset distribution, we construct 100 diverse inputs \(10 per class\) that systematically cover different damage locations, mechanisms, and structural elements not exhaustively present in the training data\.

Figure[6](https://arxiv.org/html/2607.21680#S5.F6)reports class\-wise accuracy for all three methods across these 100 samples\.

![Refer to caption](https://arxiv.org/html/2607.21680v1/figures/classwise_accuracy_comparison.png)Figure 6:Class\-wise accuracy on 100 diverse samples \(10 per class\)\. Dashed lines indicate overall accuracy: QLoRA 47\.0%, LoRA 34\.0%, QA\-LoRA 30\.0%\.Overall results:QLoRA achieves 47\.0% \(47/100\), LoRA 34\.0% \(34/100\), and QA\-LoRA 30\.0% \(30/100\)\. The 13\-point gap between QLoRA and LoRA despite identical Golden Testset accuracy strongly suggests that 4\-bit quantization provides implicit regularization against distribution shift\.

Class\-wise findings:

- •High accuracy \(all models\):Freeze\-Thaw \(90% each\), Fatigue \(QLoRA: 100%, LoRA: 80%, QA\-LoRA: 60%\)\.
- •QLoRA\-exclusive advantage:Salt Damage \(80% vs\. 30%/20%\), Void Floating \(70% vs\. 20%/0%\), Others \(50% vs\. 20%/0%\)\.
- •QA\-LoRA anomaly:Continuous Girder achieves 100% with QA\-LoRA but 0% with QLoRA, suggesting overfitting to specific structural terminology under L2 regularization\.
- •Persistent failures \(all models\):Water Accumulation \(0%/0%/0%\), Soil Liquefaction \(0%/10%/0%\), and ASR \(10%/20%/10%\) remain near\-zero, pointing to insufficient discriminative training examples for these classes\.

These failure patterns provide actionable targets for future data augmentation: Water Accumulation, Soil Liquefaction, and ASR collectively account for the largest performance gap between the Golden Testset \(balanced\) and diverse real\-world inputs\.

## 6Discussion

### 6\.1Quantization as a Regularizer

The\+\+2\.57% accuracy improvement of QLoRA over LoRA on balanced data reveals a previously underappreciated role of quantization beyond memory compression\. We hypothesize the following mechanism:

Reduced weight expressivity as implicit regularization\.4\-bit NF4 quantization constrains each weight to one of 16 discrete values, compared to 65,536 values in FP16\. This 4,096×\\timesreduction in representational capacity forces the model to learn more coarse\-grained, generalizable weight patterns rather than memorizing training\-set idiosyncrasies—an effect analogous to dropout regularization, where reduced model capacity encourages generalization over memorization\.

Interaction with class weights\.On the severely imbalanced v0\.2 dataset, inverse\-frequency class weights range from 0\.152 \(Salt, 65\.7%\) to 6\.114 \(Void, 1\.6%\), a 40\.2×\\timesspan\. These extreme weights demand high representational capacity to simultaneously learn the dominant class pattern and the subtle minority class features; quantization’s capacity constraint partially frustrates this, yielding the observed 1\.16% accuracy drop\. On the balanced v0\.3 dataset, class weights span only 0\.388–2\.985 \(7\.7×\\times\), making the capacity constraint harmless—and its regularization effect beneficial\. This is further corroborated by the 100\-sample diverse test \(Section[5\.8](https://arxiv.org/html/2607.21680#S5.SS8)\), where QLoRA outperforms LoRA by 13 points despite identical Golden Testset accuracy\.

### 6\.2Imbalance Ratio Guideline for Quantization Selection

Figure[7](https://arxiv.org/html/2607.21680#S6.F7)visualizes the accuracy differenceΔ​Acc=AccQLoRA−AccLoRA\\Delta\\text\{Acc\}=\\text\{Acc\}\_\{\\text\{QLoRA\}\}\-\\text\{Acc\}\_\{\\text\{LoRA\}\}as a function of Imbalance Ratio\. Based on experimental data points and the theoretical argument above, we propose the following practical guideline:

Table 6:Quantization Strategy vs\. Imbalance Ratio0551010151520202525303035354040−2\-2−1\-10112233Quant helpfulQuant harmfulv0\.3 \(IR=7\.7\)v0\.4 \(IR=7\.69\)v0\.2 \(IR=35\.1\)Imbalance Ratio \(IR = max / min class count\)Δ\\DeltaAccuracy \(%\) = QLoRA−\-No QuantMeasuredΔ\\DeltaAccv0\.4 Scale\-upTrendFigure 7:Accuracy gain of QLoRA over LoRA as a function of Imbalance Ratio\. The green region \(IR<<12\) indicates where 4\-bit quantization acts as a regularizer and improves accuracy; the red region \(IR\>\>12\) indicates degradation\.These thresholds are heuristic estimates and should be validated on additional datasets\. The v0\.4 result \(IR = 7\.69, QLoRA91\.47%91\.47\\%vs\. v0\.3 QLoRA85\.90%85\.90\\%,Δ=\+5\.57%\\Delta=\+5\.57\\%\) confirms the guideline: corpus scaling from 388 to 642 samples recovers accuracy from 85\.90% to 91\.47%, validating the generalizability of the IR<<10 regime\.

### 6\.3Generalization and Horizontal Transferability

The proposed pipeline is not inherently specific to bridge inspection\. The four\-phase architecture—\(A\) extract causal triples from domain PDF documents, \(B\) build a FAISS retrieval store, \(C\) fine\-tune with QLoRA using retrieved causal context, \(D\) infer with retrieval augmentation—is applicable to another domain that meets the following criteria:

1. 1\.Expert diagnostic knowledge exists in PDF/document form,
2. 2\.Symptoms are observable but causes require inference,
3. 3\.Labeled training data are limited \(tens to hundreds of samples\),
4. 4\.Deployment hardware is resource\-constrained\.

Candidate domains include tunnel lining defect diagnosis, pavement deterioration classification, mechanical component fault analysis, and medical symptom\-to\-diagnosis reasoning\.

### 6\.4Limitations

Triple extraction quality\.Qwen2\.5\-7B occasionally produces malformed JSON or semantically imprecise triples, requiring post\-hoc filtering\. The 10\.6% of blocks producing no triples \(in the 35\-PDF corpus\) represents information loss\.

Persistent failure classes\.Water Accumulation, Soil Liquefaction, and ASR score near\-zero on diverse inputs across all fine\-tuning methods, indicating that the current training data—even with retrieval augmentation—lacks sufficient discriminative examples for these classes\. Targeted data collection and augmentation are the highest\-priority next steps\.

Language and domain coverage\.The current models are Japanese\-only\. Extending to multilingual or multimodal \(image \+ text\) inputs is a straightforward extension\.

## 7Conclusion

We have presented theDamage Cause Encoder, a retrieval\-augmented fine\-tuning framework that encodes invisible causal knowledge from bridge diagnostic manuals into a BERT\-based 10\-class classifier\.

The four principal findings are:

1. 1\.Triple\-Guided RAGenables a BERT encoder to leverage domain causal knowledge that would otherwise require years of expert tacit knowledge accumulation\. Scaling from 15 to 35 PDFs \(4,186 to 6,745 triples\) expands vocabulary coverage across additional bridge structure types, recovering accuracy from 85\.90% \(v0\.3\) to91\.47%\(v0\.4\) while preserving balanced class distribution\.
2. 2\.Golden Testset as Benchmark Contribution: The stratified, deduplicated, difficulty\-tagged 116\-sample testset enables reproducible comparison of fine\-tuning methods and provides a reusable evaluation fixture for future work\.
3. 3\.QLoRA is the Optimal Fine\-tuning Strategy: On the Golden Testset, QLoRA matches LoRA accuracy \(87\.07%\) while being 11% faster and requiring 72% less GPU memory\. Critically, QLoRA outperforms LoRA by 13 percentage points on 100 diverse unseen inputs \(47\.0% vs\. 34\.0%\), confirming that 4\-bit quantization noise provides beneficial implicit regularization\.
4. 4\.Persistent Failure Classes: Water Accumulation, Soil Liquefaction, and ASR consistently score near\-zero across all methods on diverse inputs, identifying the highest\-priority targets for data augmentation in future iterations\.

These results collectively demonstrate that memory\-efficient, high\-accuracy damage cause diagnosis is achievable on consumer\-grade hardware \(RTX 4060 Ti, 16 GB\), lowering the barrier for real\-world deployment of diagnostic agents in infrastructure maintenance organizations that cannot afford cloud GPU infrastructure\.

Future work includes: replacing the LLM YES/NO filter with a learned relevance classifier for faster inference; exploring multimodal fusion \(vision \+ text\) using figure and table blocks directly; augmenting training data for Water Accumulation, Soil Liquefaction, and ASR to close the identified failure gaps; extending the framework to additional structure types \(tunnels, retaining walls\) and languages; and validating the IR\-based quantization guideline on a wider range of datasets\.

## Acknowledgments

The author developed this work as part of applied research in infrastructure maintenance\. Ollama, Hugging Face, and their open\-source communities are gratefully acknowledged\.

## References

- \[1\]R\. Bommasani, D\. A\. Hudson, E\. Aditi,et al\.\(2021\)On the opportunities and risks of foundation models\.arXiv preprint arXiv:2108\.07258\.Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p1.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.
- \[2\]Y\. Cha, W\. Choi, and O\. Büyüköztürk\(2017\)Deep learning\-based crack damage detection using convolutional neural networks\.Computer\-Aided Civil and Infrastructure Engineering32\(5\),pp\. 361–378\.Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p2.1)\.
- \[3\]K\. Chen, G\. Reichard, A\. Akanmu, and X\. Xu\(2022\)Automated visual inspection of bridge surface cracks using deep learning\.Computer\-Aided Civil and Infrastructure Engineering37\(2\),pp\. 193–212\.Cited by:[§2\.4](https://arxiv.org/html/2607.21680#S2.SS4.p1.1)\.
- \[4\]T\. Dettmers, M\. Lewis, Y\. Belkada, and L\. Zettlemoyer\(2022\)LLM\.int8\(\): 8\-bit matrix multiplication for transformers at scale\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§4\.3\.1](https://arxiv.org/html/2607.21680#S4.SS3.SSS1.p2.1),[§4\.3\.2](https://arxiv.org/html/2607.21680#S4.SS3.SSS2.p2.1)\.
- \[5\]T\. Dettmers, A\. Pagnoni, A\. Holtzman, and L\. Zettlemoyer\(2023\)QLoRA: efficient finetuning of quantized LLMs\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2\.2](https://arxiv.org/html/2607.21680#S2.SS2.p1.2),[§4\.3\.3](https://arxiv.org/html/2607.21680#S4.SS3.SSS3.p3.1)\.
- \[6\]J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova\(2019\)BERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of NAACL\-HLT,pp\. 4171–4186\.Cited by:[§2\.1](https://arxiv.org/html/2607.21680#S2.SS1.p1.1)\.
- \[7\]D\. Edge, H\. Trinh, N\. Cheng, J\. Bradley, A\. Chao, A\. Mody, S\. Truitt, and J\. Larson\(2024\)From local to global: a graph RAG approach to query\-focused summarization\.arXiv preprint arXiv:2404\.16130\.Cited by:[§2\.3](https://arxiv.org/html/2607.21680#S2.SS3.p1.1)\.
- \[8\]C\. R\. Farrar and K\. Worden\(2007\)An introduction to structural health monitoring\.Philosophical Transactions of the Royal Society A365\(1851\),pp\. 303–315\.Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p2.1),[item \(iii\)](https://arxiv.org/html/2607.21680#S3.I2.i3.p1.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.
- \[9\]Hotchpotch\(2024\)Static embedding japanese\.Note:[https://huggingface\.co/hotchpotch/static\-embedding\-japanese](https://huggingface.co/hotchpotch/static-embedding-japanese)Accessed: 2026Cited by:[§4\.1](https://arxiv.org/html/2607.21680#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2607.21680#S4.SS2.p2.1)\.
- \[10\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.2](https://arxiv.org/html/2607.21680#S2.SS2.p1.2),[§4\.3\.3](https://arxiv.org/html/2607.21680#S4.SS3.SSS3.p2.1)\.
- \[11\]Hugging Face\(2023\)PEFT: state\-of\-the\-art parameter\-efficient fine\-tuning\.Note:[https://github\.com/huggingface/peft](https://github.com/huggingface/peft)Accessed: 2026Cited by:[item 3](https://arxiv.org/html/2607.21680#S4.I1.i3.p1.1)\.
- \[12\]A\. K\. S\. Jardine, D\. Lin, and D\. Banjevic\(2006\)A review on machinery diagnostics and prognostics implementing condition\-based maintenance\.Mechanical Systems and Signal Processing20\(7\),pp\. 1483–1510\.Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p2.1),[item \(iii\)](https://arxiv.org/html/2607.21680#S3.I2.i3.p1.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.
- \[13\]V\. Karpukhin, B\. Oğuz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. Yih\(2020\)Dense passage retrieval for open\-domain question answering\.InProceedings of EMNLP,Cited by:[§2\.3](https://arxiv.org/html/2607.21680#S2.SS3.p1.1)\.
- \[14\]B\. Kim and S\. Cho\(2023\)Bridge crack detection and classification using a two\-stage deep convolutional neural network\.Automation in Construction150,pp\. 104821\.Cited by:[§2\.4](https://arxiv.org/html/2607.21680#S2.SS4.p1.1)\.
- \[15\]P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni,et al\.\(2020\)Retrieval\-augmented generation for knowledge\-intensive NLP tasks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2\.3](https://arxiv.org/html/2607.21680#S2.SS3.p1.1)\.
- \[16\]H\. Maeda, Y\. Sekimoto, T\. Seto, T\. Kashiyama, and H\. Omata\(2018\)Road damage detection and classification using deep neural networks with smartphone images\.Computer\-Aided Civil and Infrastructure Engineering33\(12\),pp\. 1127–1141\.Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p2.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.
- \[17\]Ministry of Land, Infrastructure, Transport and Tourism, Japan\(2023\)Bridge and tunnel inspection status \(in japanese\)\.Note:[https://www\.mlit\.go\.jp/road/road/traffic/hyoushiki/index\.html](https://www.mlit.go.jp/road/road/traffic/hyoushiki/index.html)Accessed: 2026Cited by:[§1\.1](https://arxiv.org/html/2607.21680#S1.SS1.p1.1)\.
- \[18\]J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. Bernstein\(2023\)Generative agents: interactive simulacra of human behavior\.InProceedings of the ACM Symposium on User Interface Software and Technology \(UIST\),Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p1.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.
- \[19\]Qwen Team\(2025\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§4\.1](https://arxiv.org/html/2607.21680#S4.SS1.p2.1)\.
- \[20\]Y\. Shen, K\. Song, X\. Tan, D\. Li, W\. Lu, and Y\. Zhuang\(2023\)HuggingGPT: solving AI tasks with ChatGPT and its friends in HuggingFace\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p1.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.
- \[21\]Tohoku NLP Group\(2023\)Japanese BERT\-large \(cl\-tohoku/bert\-large\-japanese\-v2\)\.Note:[https://huggingface\.co/cl\-tohoku/bert\-large\-japanese\-v2](https://huggingface.co/cl-tohoku/bert-large-japanese-v2)Accessed: 2026Cited by:[§2\.1](https://arxiv.org/html/2607.21680#S2.SS1.p1.1),[§4\.3\.1](https://arxiv.org/html/2607.21680#S4.SS3.SSS1.p1.2)\.
- \[22\]L\. Wang, C\. Ma, X\. Feng, Z\. Zhang,et al\.\(2024\)A survey on large language model based autonomous agents\.Frontiers of Computer Science18\(6\),pp\. 186345\.Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p1.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.
- \[23\]Z\. Xi, W\. Chen, X\. Guo, W\. He,et al\.\(2023\)The rise and potential of large language model based agents: a survey\.arXiv preprint arXiv:2309\.07864\.Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p1.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.
- \[24\]Y\. Xu, L\. Xie, X\. Gu, X\. Chen, H\. Chang, H\. Zhang, Z\. Chen, X\. Zhang, and Q\. Tian\(2023\)QA\-LoRA: quantization\-aware low\-rank adaptation of large language models\.arXiv preprint arXiv:2309\.14717\.Cited by:[§2\.2](https://arxiv.org/html/2607.21680#S2.SS2.p2.1),[§4\.3\.3](https://arxiv.org/html/2607.21680#S4.SS3.SSS3.p4.3)\.
- \[25\]T\. Yasuno\(2026\)Quantized vision\-language models for damage assessment: a comparative study of LLaVA\-1\.5\-7B quantization levels\.Note:arXiv preprintCited by:[§2\.4](https://arxiv.org/html/2607.21680#S2.SS4.p1.1)\.
- \[26\]R\. Zhao, R\. Yan, Z\. Chen, K\. Mao, P\. Wang, and R\. X\. Gao\(2019\)Deep learning and its applications to machine health monitoring\.Mechanical Systems and Signal Processing115,pp\. 213–237\.Cited by:[§2\.5](https://arxiv.org/html/2607.21680#S2.SS5.p2.1),[§3\.4](https://arxiv.org/html/2607.21680#S3.SS4.p5.1)\.

Similar Articles

Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis

arXiv cs.LG

This paper introduces the Diagnostic Evidence Network (DENet), a multi-task framework that extends AI-based bearing fault diagnosis to produce physically verifiable evidence, such as predicted characteristic frequencies and temporal localization of impulses, while using a QLoRA-adapted language model to generate constrained diagnostic reports that reduce hallucinated content.

Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning

arXiv cs.CL

This paper investigates why instruction-tuned language models give different answers to causal reasoning questions when variable names are replaced with placeholders, finding that the issue stems from representational misalignment rather than information loss. The authors introduce Vernier, a method using paired-view weight updates and mechanism inspection to reveal that answer-relevant content is still present in the placeholder view but misaligned.