MiLSD: A Micro Line-Segment Detector for Resource-Constrained Devices
Summary
This paper presents MiLSD, a micro line-segment detector designed for resource-constrained devices like microcontrollers. It systematically compares output representations within a compact fully-convolutional backbone and shows that 8-bit quantization preserves full-precision performance, while 4-bit quantization causes degradation, achieving improved accuracy within a 1 MB activation budget.
View Cached Full Text
Cached at: 07/09/26, 07:57 AM
# MiLSD: A Micro Line-Segment Detector for Resource-Constrained Devices
Source: [https://arxiv.org/html/2607.06600](https://arxiv.org/html/2607.06600)
Parsa Hassani Shariat Panahiπ[https://orcid.org/0009-0005-2912-3754](https://orcid.org/0009-0005-2912-3754), Amir Hossein Jalilvandπ[https://orcid.org/0000-0002-7641-6606](https://orcid.org/0000-0002-7641-6606), and M\. Hassan Najafi\+[https://orcid.org/0000-0002-4655-6229](https://orcid.org/0000-0002-4655-6229) πSchool of Computer Engineering, Iran University of Science and Technology, Tehran, Iran \+Electrical, Computer, and Systems Engineering Department, Case Western Reserve University, OH, USA
###### Abstract
Line segment detection is a key building block in visual SLAM, 3D reconstruction, and industrial inspection\. Recent deep learning methods have greatly improved accuracy, yet even the smallest models require several megabytes of memory, exceeding low\-cost MCU capacity\. This work investigates the maximum achievable accuracy under a sub\-megabyte budget\. We propose MiLSD, a detector tailored for MCU\-level constraints, and systematically compare three output representations within a compact fully\-convolutional backbone\.
Our study shows that the proposed F\-Clip center\-with\-length\-and\-angle formulation learns most effectively at small model sizes\. We find that 8\-bit quantization preserves full\-precision performance, while 4\-bit quantization causes significant degradation, particularly in angle regression, with quantization\-aware training recovering only part of the loss\. With a one\-megabyte activation budget and inference enhancements including sub\-pixel decoding, test\-time augmentation, and a lightweight verifier, MiLSD improvessAP10\\mathrm\{sAP\}^\{10\}on ShanghaiTech Wireframe from10\.610\.6\(25k parameters,0\.250\.25MB\) to24\.124\.1within11MB\. Rather than competing with GPU\-scale parsers, we map the accuracy–memory trade\-off across representations, bit\-widths, capacities, and post\-processing strategies for embedded vision systems\.
## IIntroduction
Line segments are a primitive structural feature in computer vision: the straight edges of walls, doors, buildings, and machined parts\[[27](https://arxiv.org/html/2607.06600#bib.bib3)\]\. They support SLAM, structure\-from\-motion, vanishing\-point estimation, lane and power\-line detection, and industrial inspection\. While much recent progress has targeted GPU\- or cloud\-based platforms, this work focuses on detection under the tight memory and compute constraints of real\-time embedded hardware\.
Classical detectors such as LSD\[[27](https://arxiv.org/html/2607.06600#bib.bib3)\]\(Line Segment Detector\) and EDLines\[[1](https://arxiv.org/html/2607.06600#bib.bib4)\]run on a CPU but find*all*edges, while modern learned wireframe parsers\[[35](https://arxiv.org/html/2607.06600#bib.bib8),[31](https://arxiv.org/html/2607.06600#bib.bib10),[4](https://arxiv.org/html/2607.06600#bib.bib15),[5](https://arxiv.org/html/2607.06600#bib.bib17)\]recover only salient segments but require GPU\- or phone\-class compute\. Table[I](https://arxiv.org/html/2607.06600#S1.T1)situates representative methods across this spectrum\. Classical detectors grow line\-support regions from local gradients and validate them statistically\. LSD groups pixels with consistent gradient orientation and accepts a segment if its number of false alarms is below one\. EDLines reaches comparable quality faster by chaining edge pixels into clean chains\. ELSED\[[25](https://arxiv.org/html/2607.06600#bib.bib5)\]targets embedded CPUs for high frame rates\. Their shared weakness: accuracy degrades under blur, low contrast, and clutter, and runtime is content\-dependent\. At its evaluated640×480640\\times 480resolution, ELSED also requires several full\-frame gradient and edge buffers \(∼\\sim1\.5–2 MB\) that exceed even the 1 MB SRAM of an STM32H7, and since its edge walk is global and data\-dependent, it cannot be tiled and admits no static worst\-case memory bound, unlike a fixed\-cost CNN\. Classical detectors have also been mapped to FPGAs and ASICs for deterministic latency, but these implementations accelerate hand\-designed gradient logic, not neural networks on general\-purpose microcontrollers\.
ShanghaiTech Wireframe\[[9](https://arxiv.org/html/2607.06600#bib.bib7)\]reframed line detection as a learning problem\. L\-CNN\[[35](https://arxiv.org/html/2607.06600#bib.bib8)\]proposed junctions and verified candidate lines\. AFM\[[30](https://arxiv.org/html/2607.06600#bib.bib9)\]used attraction fields, while HAWP\[[31](https://arxiv.org/html/2607.06600#bib.bib10)\]combined holistic fields with endpoint verification\. ULSD\[[15](https://arxiv.org/html/2607.06600#bib.bib13)\]generalized across pinhole, fisheye, and spherical cameras, and LETR\[[29](https://arxiv.org/html/2607.06600#bib.bib12)\]uses transformers for direct line detection\. L\-CNN, HAWP, ULSD, and LETR achievesAP10≈63\\mathrm\{sAP\}^\{10\}\\approx 63–7070\(AFM, an earlier field\-based method, scores≈24\\approx 24\), but assume workstation\-class memory\. The unifying observation is that classical detectors are light but find every edge, while learned parsers are accurate but GPU\-bound; no learned method yet occupies the MCU column of Table[I](https://arxiv.org/html/2607.06600#S1.T1)\.
Even the lightest learned detector, M\-LSD\-tiny\[[5](https://arxiv.org/html/2607.06600#bib.bib17)\], requires at least 78 MB of runtime memory, orders of magnitude beyond the SRAM budget of typical microcontrollers\. For context, the STM32F746 provides only 320 KB of SRAM and 1 MB of flash\[[23](https://arxiv.org/html/2607.06600#bib.bib35)\]\. On such devices, the bottleneck is peak activation memory, not parameter storage\. Prior work differs in how segments are encoded at the network output\. Heatmap methods require a separate linker\. Tri\-point and center\-with\-displacement designs \(TP\-LSD\[[10](https://arxiv.org/html/2607.06600#bib.bib14)\], M\-LSD\[[5](https://arxiv.org/html/2607.06600#bib.bib17)\]\) target mobile inference\. F\-Clip\[[4](https://arxiv.org/html/2607.06600#bib.bib15)\]compresses each segment to center, length, and angle; we encode that angle as a double\-angle\(cos2θ,sin2θ\)\(\\cos 2\\theta,\\sin 2\\theta\)\. On microcontrollers, activations dominate SRAM usage\[[16](https://arxiv.org/html/2607.06600#bib.bib21),[2](https://arxiv.org/html/2607.06600#bib.bib22)\]\. MCUNet\[[17](https://arxiv.org/html/2607.06600#bib.bib20)\]uses hardware\-aware search, and MCUNetV2\[[16](https://arxiv.org/html/2607.06600#bib.bib21)\]adds patch\-based inference\. Integer quantization with straight\-through gradients\[[11](https://arxiv.org/html/2607.06600#bib.bib25),[3](https://arxiv.org/html/2607.06600#bib.bib26)\]underpins PTQ and QAT\. Int8 is generally safe; sub\-8\-bit demands care\[[22](https://arxiv.org/html/2607.06600#bib.bib29),[19](https://arxiv.org/html/2607.06600#bib.bib30)\]\. To our knowledge, no prior work combines these threads for MCU\-scale line segment detection\.
This work investigates the maximum achievable accuracy under a sub\-megabyte memory budget\. We study three axes: \(i\) output representation: heatmap, center\-with\-displacement, and F\-Clip\-style center\-with\-length\-and\-angle; \(ii\) quantization: full\-precision, 8\-bit, and 4\-bit; and \(iii\) inference enhancements: sub\-pixel decoding, test\-time augmentation, and a lightweight verifier\.
We propose MiLSD, a detector designed for MCU\-scale memory\. With a 1 MB activation budget, MiLSD achievessAP10=24\.1\\mathrm\{sAP\}^\{10\}=24\.1on ShanghaiTech Wireframe, improving over a0\.250\.25MB baseline at10\.610\.6\. Our quantization study reveals that 8\-bit inference preserves full\-precision performance, while 4\-bit quantization causes significant degradation, particularly in angle regression, where quantization\-aware training recovers only part of the loss\. This sensitivity has not been reported in prior work\.
The main contributions are:
- •A comparison of three output representations under extreme memory constraints, identifying F\-Clip\-style as the most effective at small model sizes\.
- •A quantization study revealing angle regression in the\(cos2θ,sin2θ\)\(\\cos 2\\theta,\\sin 2\\theta\)space as the most sensitive component to bit\-width reduction\.
- •MiLSD, operating within 1 MB memory while achievingsAP10=24\.1\\mathrm\{sAP\}^\{10\}=24\.1on ShanghaiTech Wireframe\.
- •An accuracy–resource frontier characterizing trade\-offs among capacity, quantization, and post\-processing\.
Targeting MCU\-scale detection, where memory is the binding constraint, we show that meaningful wireframe quality is achievable within 1 MB despite GPU\-class parsers remaining out of reach\. Our study maps the accuracy–resource trade\-offs across representations, quantization, and post\-processing\.
The rest of the paper is organized as follows\. Section[II](https://arxiv.org/html/2607.06600#S2)details the proposed architecture, including the three output representations, the compact backbone, and the quantization\-aware training pipeline\. Section[III](https://arxiv.org/html/2607.06600#S3)describes the experimental protocol, dataset, and training hyperparameters\. Section[IV](https://arxiv.org/html/2607.06600#S4)presents our empirical findings on representation selection, quantization sensitivity, capacity scaling, and comparisons to prior art\. Section[V](https://arxiv.org/html/2607.06600#S5)introduces the full MiLSD system on the STM32H7, incorporating capacity scaling, sub\-pixel decoding, test\-time augmentation, and a learned verification head\. Section[VI](https://arxiv.org/html/2607.06600#S6)summarizes our contributions and discusses directions for future work\.
TABLE I:Representative line segment detectors across classical, learned, and efficient regimes \(core ideas and platforms per the respective papers\)\. The*On MCU?*column highlights that prior learned methods target GPU, phone, or FPGA platforms\.
## IIProposed MiLSD design
This section describes the design of MiLSD, a line\-segment detector optimized for microcontroller\-scale memory\. We first present three output representations for encoding segments on a fixed grid, then describe the compact backbone shared across all variants, followed by the quantization strategy that enables int8 deployment, and finally the training pipeline and deployment flow\.
### II\-AOutput Representations
A key design decision for any learned line\-segment detector is how to encode the continuous geometry of a segment into discrete network outputs\. We study three representations on a128×128128\\times 128output grid \(Fig\.[1](https://arxiv.org/html/2607.06600#S2.F1)\)\.
*\(i\) Heatmap:*A per\-pixel binary classification map indicating whether a pixel lies on a line\. This is the most direct representation but requires an external post\-processing linker to assemble pixels into continuous segments\. While conceptually simple, the linker introduces additional computational overhead and hyperparameters, and the network itself does not produce geometric primitives\.
*\(ii\) Center\-with\-Displacement:*Inspired by TP\-LSD and M\-LSD\[[10](https://arxiv.org/html/2607.06600#bib.bib14),[5](https://arxiv.org/html/2607.06600#bib.bib17)\], this representation predicts a center confidence map alongside displacement vectors from each center pixel to the two endpoints\. Each segment is thus encoded as a center point plus two offset vectors\. This formulation is single\-stage and does not require an external linker, but the network must learn to regress four continuous values \(two displacements\) per detected segment\.
*\(iii\) F\-Clip:*Following Dai et al\.\[[4](https://arxiv.org/html/2607.06600#bib.bib15)\], this representation encodes each segment as a center confidence map, a lengthℓ\\ell, and an orientation\. Whereas F\-Clip regresses the angle directly as a scalar, we encode it as\(cos2θ,sin2θ\)\(\\cos 2\\theta,\\sin 2\\theta\); this double\-angle encoding resolves the180∘180^\{\\circ\}ambiguity inherent to undirected line segments, making it uniquely defined for any line orientation\. For a ground\-truth segment with endpoints𝐩1,𝐩2\\mathbf\{p\}\_\{1\},\\mathbf\{p\}\_\{2\}, the targets at the center cell are:
ℓ=∥𝐩2−𝐩1∥,θ=atan2\(y2−y1,x2−x1\)\\ell=\\lVert\\mathbf\{p\}\_\{2\}\-\\mathbf\{p\}\_\{1\}\\rVert,\\quad\\theta=\\operatorname\{atan2\}\(y\_\{2\}\-y\_\{1\},x\_\{2\}\-x\_\{1\}\)During inference, decoding inverts this representation: for each detected center peak above a confidence threshold, the segment is reconstructed as a line of lengthℓ\\elloriented atθ\\theta, centered at the detected location\. This compact encoding requires only four output channels \(center, length,cos2θ\\cos 2\\theta,sin2θ\\sin 2\\theta\), making it particularly attractive for memory\-constrained deployment\. Section[IV\-A](https://arxiv.org/html/2607.06600#S4.SS1)compares these three representations under identical memory budgets\.
Figure 1:Three output encodings as dense per\-pixel maps \(not explicit segments\)\.v1predicts per\-pixel line probability and requires an external linker\.v2Center \+ displacement\[[10](https://arxiv.org/html/2607.06600#bib.bib14),[5](https://arxiv.org/html/2607.06600#bib.bib17)\]: center confidence with two endpoint offset vectors \(four channels; int4\-fragile\)\.v6F\-Clip\[[4](https://arxiv.org/html/2607.06600#bib.bib15)\]: center plus length and angle\(cos2θ,sin2θ\)\(\\cos 2\\theta,\\sin 2\\theta\)\.
### II\-BBackbone Architecture
Figure 2:Backbone and output head architecture\. A256×256256\\times 256grayscale input is encoded through five strided convolutions, reducing spatial resolution to64×6464\\times 64and expanding channels to3232\. A1×11\\times 1projection and nearest\-neighbor upsampling restore the resolution to128×128128\\times 128, followed by a3×33\\times 3output head producing the four prediction maps\.All three output heads share a common compact backbone designed to minimize activation memory while preserving sufficient spatial resolution for accurate line localization\. The architecture \(Fig\.[2](https://arxiv.org/html/2607.06600#S2.F2)\) consists of:
1. 1\.A strided fully\-convolutional encoder with five convolutional layers, channel widths8→16→32→32→328\\to 16\\to 32\\to 32\\to 32, and stride\-2 downsampling that reduces the input resolution from256×256256\\times 256to64×6464\\times 64\.
2. 2\.A1×11\\times 1convolutional reduction layer that projects features to a compact representation\.
3. 3\.A nearest\-neighbor upsampling layer that restores the spatial resolution to128×128128\\times 128\.
4. 4\.A3×33\\times 3output head that produces the final prediction maps \(center confidence, length, and orientation for F\-Clip; center and displacements for endpoint representation; or a single heatmap\)\.
The total parameter count is approximately2525k at the smallest width\. Section[IV\-D](https://arxiv.org/html/2607.06600#S4.SS4)demonstrates that parameter capacity is not the primary bottleneck in this regime; rather, input resolution and activation memory constrain performance\.
### II\-CQuantization for MCU Deployment
Figure 3:Quantization scheme\. Continuous weights are snapped to discrete levels: int8 provides 256 levels \(fine grid\), while int4 provides only 16 levels \(coarse grid\)\. Quantization is performed per\-tensor and symmetric\.To fit the model within microcontroller SRAM and use optimized integer inference kernels such as CMSIS\-NN\[[14](https://arxiv.org/html/2607.06600#bib.bib24)\], we quantize weights and activations to integer precision\. We adopt per\-tensor symmetric quantization, implemented with fake\-quantization in the forward pass and the straight\-through estimator \(STE\) for gradient propagation during backpropagation\[[11](https://arxiv.org/html/2607.06600#bib.bib25),[3](https://arxiv.org/html/2607.06600#bib.bib26)\]\.
For a weight tensorwwand target bit\-widthbb, the quantization scale is:
s=max\|w\|2b−1−1s=\\frac\{\\max\|w\|\}\{2^\{b\-1\}\-1\}The quantized weight is computed as:
w^=s⋅round\(ws\)\\hat\{w\}=s\\cdot\\operatorname\{round\}\\left\(\\frac\{w\}\{s\}\\right\)This operation snaps continuous values onto a discrete grid of2b2^\{b\}levels\. As illustrated in Fig\.[3](https://arxiv.org/html/2607.06600#S2.F3), 8\-bit quantization provides 256 levels, offering fine granularity, while 4\-bit quantization reduces this to only 16 levels, introducing substantial rounding error\.
We evaluate two quantization strategies:
- •*Post\-training quantization \(PTQ\):*The model is first trained in full precision, then weights and activations are quantized once using a small calibration set\. This is computationally efficient but can suffer from accuracy degradation, particularly at low bit\-widths\.
- •*Quantization\-aware training \(QAT\):*The quantization operation is simulated during training, allowing the model to learn weights that are robust to quantization error\. While more expensive, QAT often recovers some of the accuracy lost in PTQ\.
Our deployed model uses int8 quantization with QAT, achieving performance comparable to full precision as shown in Section[IV\-B](https://arxiv.org/html/2607.06600#S4.SS2)\. We also investigate 4\-bit quantization to understand the limits of aggressive compression, revealing that angle regression is particularly sensitive to bit\-width reduction\.
### II\-DTraining Pipeline and Deployment Flow
Training is performed entirely off\-device on a GPU workstation; the microcontroller executes only int8 inference\. This train\-off / infer\-on split is the defining premise of our deployment strategy and is common practice in TinyML systems\.
The training pipeline proceeds as follows:
1. 1\.A256×256256\\times 256grayscale image is fed into the backbone\.
2. 2\.The network produces prediction maps \(center, length,cos2θ\\cos 2\\theta,sin2θ\\sin 2\\thetafor F\-Clip\) on a128×128128\\times 128grid\.
3. 3\.Loss is computed against ground\-truth segments encoded in the same representation, using a combination of binary cross\-entropy for center confidence and smooth L1 loss for geometric attributes\.
4. 4\.For QAT, quantization simulation is enabled during training with STE gradient propagation\.
For deployment, the trained model is exported through X\-CUBE\-AI, STMicroelectronics’ neural network inference library for STM32 microcontrollers\. The export process generates optimized C code that runs on the Arm Cortex\-M7 core, using CMSIS\-NN for efficient integer arithmetic\. The inference pipeline on the MCU \(Fig\.[4](https://arxiv.org/html/2607.06600#S2.F4)\) consists of:
1. 1\.Input image capture \(grayscale,256×256256\\times 256\)\.
2. 2\.int8 inference through the quantized network\.
3. 3\.Decoding of output maps into line segments \(center detection, length and angle extraction, endpoint computation\)\.
4. 4\.Optional post\-processing: Line\-of\-Interest verification and non\-maximum suppression\.
The entire inference pipeline is designed to operate within the 320 KB SRAM budget of the STM32F746, with peak activation memory as the primary constraint rather than parameter storage\.
Figure 4:Off\-device training and on\-MCU inference pipeline\. The model is trained on GPU with quantization simulation, then exported through X\-CUBE\-AI for deployment on the STM32F746\. The MCU executes int8 inference only\.
## IIIExperimental Setup
#### Dataset and metric
We train and evaluate on the ShanghaiTech Wireframe benchmark\[[9](https://arxiv.org/html/2607.06600#bib.bib7)\]\(5,000 training and 462 held\-out evaluation images,∼\\sim74 segments per image\)\. Accuracy is structural average precisionsAPt\\mathrm\{sAP\}^\{t\}at squared\-endpoint\-distance thresholdst∈\{5,10,15\}t\\in\\\{5,10,15\\\}in the1282128^\{2\}output space; we headlinesAP10\\mathrm\{sAP\}^\{10\}\. For comparability with the FPGA predecessor\[[20](https://arxiv.org/html/2607.06600#bib.bib2)\]we also reference its Q1 \(coverage\) and Q2 \(noise\-suppression\) measures\.
#### Implementation
Table[II](https://arxiv.org/html/2607.06600#S3.T2)lists the training configuration\. Training is in PyTorch on a GPU; versions v1–v6 of the design search\[[7](https://arxiv.org/html/2607.06600#bib.bib38)\]share the backbone of Section[II](https://arxiv.org/html/2607.06600#S2)\. All models are trained for 300 epochs with a batch size of 32, using the Adam optimizer and a cosine annealing learning rate schedule starting from10−310^\{\-3\}\. Data augmentation is limited to random horizontal and vertical flips\. The loss function combines binary cross\-entropy for center classification with smooth L1 loss for length and angle regression, weighted by a factor of 2\.0 for the geometric terms and masked to ground\-truth segment locations\. We evaluate both PTQ and QAT at 8 and 4 bits\.
TABLE II:Training hyperparameters \(deployed F\-Clip model\)\.
## IVResults
### IV\-ARepresentation Comparison: The Climb
Figure 5:From baseline to MiLSD:sAP10\\mathrm\{sAP\}^\{10\}on Wireframe at each step\. The output representation drives the first gains \(heatmap v1→\\toendpoint v2→\\toF\-Clip\), reaching10\.610\.6for the 25k\-parameter F\-Clip model; scaling capacity to the width\-4 MiLSD backbone \(0\.390\.39M parameters\) lifts this to17\.817\.8, and the inference\-time stages \(test\-time augmentation then the Line\-of\-Interest verification head\) carry it tosAP10=24\.1\\mathrm\{sAP\}^\{10\}=24\.1\(purple\)\.Fig\.[5](https://arxiv.org/html/2607.06600#S4.F5)traces the evolution ofsAP10\\mathrm\{sAP\}^\{10\}across the successive versions v1–v6 of our design search, providing a step\-by\-step account of how accuracy accumulates as each design choice is introduced\. At the lowest rung of this progression, the heatmap baseline proves fundamentally inadequate for producing clean, discrete segments, registering onlysAP10=0\.3\\mathrm\{sAP\}^\{10\}=0\.3\. Endpoint regression constitutes the first formulation capable of yielding a meaningful structural score, reachingsAP10=3\.6\\mathrm\{sAP\}^\{10\}=3\.6, yet it remains limited in its ability to recover coherent geometry\. The decisive inflection occurs with the adoption of the F\-Clip representation, which at the identical 25k\-parameter budget more than doubles performance tosAP10=7\.2\\mathrm\{sAP\}^\{10\}=7\.2, establishing that the output encoding itself, rather than model size, is the dominant factor at this scale\. Subsequent refinements complete the climb in a more incremental fashion: aligning the output grid with the label resolution at256256px input lifts accuracy to8\.68\.6; incorporating the full training set of 5,000 images together with flip augmentation raises it further to9\.39\.3; and extending the training schedule to 300 epochs yields the final deployed score ofsAP10=10\.6\\mathrm\{sAP\}^\{10\}=10\.6\. Taken together, these results demonstrate unambiguously that the largest gains originate from the choice of*representation*, not from additional capacity\. This finding carries particular significance for microcontroller\-scale design: a geometric encoding that explicitly parameterizes each segment by its center, length, and orientation equips even a 25k\-parameter network with sufficient inductive structure to learn meaningful segment hypotheses, whereas the heatmap and endpoint alternatives remain unable to assemble coherent geometric predictions under the same severe parameter constraint\.
### IV\-BQuantization
Figure 6:sAP10\\mathrm\{sAP\}^\{10\}versus bit\-width for the deployed F\-Clip model on Wireframe\. The int8 point overlaps fp32; PTQ at 4 bits fails while QAT partially recovers\.Fig\.[6](https://arxiv.org/html/2607.06600#S4.F6)and Table[III](https://arxiv.org/html/2607.06600#S4.T3)present the results of our systematic quantization study, evaluated across the three principal output representations corresponding to versions v1, v2, and v6 of the design search\. For the deployed F\-Clip model, the transition from full\-precision fp32 inference to 8\-bit integer quantization incurs a degradation of only0\.60\.6sAP10\\mathrm\{sAP\}^\{10\}points \(10\.6→10\.010\.6\\rightarrow 10\.0\), indicating that int8 arithmetic is sufficient for on\-device deployment with negligible impact on structural accuracy\. By contrast, post\-training quantization at 4 bits proves catastrophic, collapsing performance tosAP10=0\.7\\mathrm\{sAP\}^\{10\}=0\.7; quantization\-aware training partially mitigates this failure, recovering the score to6\.96\.9, which corresponds to approximately60%60\\%of the int4\-induced gap relative to the fp32 baseline\. Inspection of the per\-head errors reveals that the degradation is concentrated almost entirely in the\(cos2θ,sin2θ\)\(\\cos 2\\theta,\\sin 2\\theta\)angle regression branch, whose inherently narrow dynamic range is poorly served by the coarse 4\-bit quantization grid\. On this basis, the deployed model adopts int8 throughout\. More broadly, these results suggest that pushing quantization beyond 8 bits is unlikely to remain viable for geometric regression heads of this kind without substantial architectural or training modifications\.
TABLE III:Quantization results,sAP10\\mathrm\{sAP\}^\{10\}on Wireframe\.∗Heatmap is not built for sAP; on its own terms Q2=0\.86=0\.86, recall=0\.44=0\.44\.
### IV\-CResolution–Memory Trade\-off
Figure 7:Input resolution sets a genuine accuracy–memory trade\-off\. Accuracy \(blue, left axis\) climbs as resolution rises, but only until the128128output grid reaches the128128\-px label resolution at256256px input; past that the output is finer than the labels and accuracy saturates, while peak SRAM \(coral, right axis\) keeps growing and crosses the320320KB budget\.256256px is therefore the operating point\.Fig\.[7](https://arxiv.org/html/2607.06600#S4.F7)plots bothsAP10\\mathrm\{sAP\}^\{10\}and peak SRAM consumption as functions of input resolution, enabling a joint assessment of the accuracy–memory trade\-off that governs operating\-point selection on resource\-constrained hardware\. The chosen configuration of256256px input is selected at the point where the128×128128\\times 128output grid aligns with the native label resolution, and where peak SRAM remains within the 320 KB ceiling imposed by the STM32F746\. This analysis reveals a pronounced and interpretable trade\-off: as input resolution increases, structural accuracy improves steadily until the output grid reaches parity with the label resolution, at which point further resolution gains yield diminishing or negligible returns\. Beyond this saturation point, peak SRAM continues to grow without a commensurate accuracy benefit\. The256256px operating point therefore represents the optimal balance between detection quality and memory footprint for the F746 deployment target\.
### IV\-DCapacity and Overfitting
Figure 8:Training the F\-Clip model over 300 epochs\. Train and held\-out loss track together throughout, with a final gap of≈0\.01\\approx 0\.01: the 25k\-parameter model shows no overfitting despite only5,0005\{,\}000training images\.To test whether the 25k\-parameter backbone is capacity\-limited, we swept backbone width through three operating points \(∼\\sim25k,∼\\sim98k, and∼\\sim209k parameters\) while holding all other training settings fixed, and recorded the held\-out loss floor at convergence\. The result is nearly flat: the smallest model settles at1\.741\.74, while the intermediate and largest variants both reach1\.731\.73; an eightfold increase in parameters yields only a0\.010\.01reduction in held\-out loss\. This pattern is consistent with the capacity\-gap effect documented in the knowledge\-distillation literature\[[18](https://arxiv.org/html/2607.06600#bib.bib31)\], and implies limited headroom for naive distillation\-based improvement at this scale\. The flatness further suggests that the model is already operating near the information\-theoretic limit imposed by the dataset and the chosen input resolution, such that additional parameters are unlikely to translate into measurable gains in structural accuracy\. Complementing this capacity analysis, Fig\.[8](https://arxiv.org/html/2607.06600#S4.F8)plots the training and held\-out loss trajectories over the full 300\-epoch schedule\. The two curves remain closely aligned throughout training, converging to a final gap of approximately0\.010\.01, indicating the absence of overfitting despite the severely constrained 25k\-parameter architecture and the modest size of the training set\. Together, these observations support the conclusion that the fixed small\-width design is both memory\-efficient and well\-regularized by its architectural constraints\.
### IV\-EComparison with Prior Work
Figure 9:Accuracy vs\. parameter budget on the Wireframe benchmark \(log\-scalexx\)\. Our two operating points sit at the extreme low\-resource end: the 25k\-parameter F\-Clip model on the STM32F746 \(sAP10=10\.6\\mathrm\{sAP\}^\{10\}=10\.6\) and MiLSD on the STM32H7 \(0\.390\.39M parameters,sAP10=24\.1\\mathrm\{sAP\}^\{10\}=24\.1, purple\)\. Related learned detectors use24×24\\timesto1,600×1\{,\}600\\timesmore parameters and assume mobile or GPU compute; their figures are from the respective papers \(Table[IV](https://arxiv.org/html/2607.06600#S4.T4)\)\[[32](https://arxiv.org/html/2607.06600#bib.bib11),[4](https://arxiv.org/html/2607.06600#bib.bib15),[31](https://arxiv.org/html/2607.06600#bib.bib10),[35](https://arxiv.org/html/2607.06600#bib.bib8),[5](https://arxiv.org/html/2607.06600#bib.bib17),[30](https://arxiv.org/html/2607.06600#bib.bib9),[9](https://arxiv.org/html/2607.06600#bib.bib7)\]\.Table[IV](https://arxiv.org/html/2607.06600#S4.T4)and Fig\.[9](https://arxiv.org/html/2607.06600#S4.F9)jointly situate our two operating points within the broader accuracy–resource frontier of the learned line\-segment detection literature\. At the low\-resource end of this range, the 25k\-parameter F746 model achievessAP10=10\.6\\mathrm\{sAP\}^\{10\}=10\.6within a 0\.25 MB activation footprint, while MiLSD \(Section[V](https://arxiv.org/html/2607.06600#S5)\) extends this capability tosAP10=24\.1\\mathrm\{sAP\}^\{10\}=24\.1under the expanded 1 MB SRAM budget of the STM32H7\. In absolute accuracy terms, both models remain substantially below the performance of contemporary transformer\-era parsers, including DT\-LSD at71\.771\.7\[[12](https://arxiv.org/html/2607.06600#bib.bib42)\]and LINEA at65\.065\.0–67\.967\.9\[[13](https://arxiv.org/html/2607.06600#bib.bib40)\], as well as compact GPU\-oriented designs such as EM\-LSD, which attains63\.263\.2with1\.11\.1M parameters\[[8](https://arxiv.org/html/2607.06600#bib.bib41)\]\. This disparity is an expected consequence of operating in a memory regime where peak SRAM is measured in kilobytes rather than megabytes\. On the resource axis, however, our position is distinctive: Table[IV](https://arxiv.org/html/2607.06600#S4.T4)lists no other learned detector with an affirmative*On MCU?*entry or an accompanying int8/int4 quantization study\. To our knowledge, this is the first learned line\-segment detector designed and evaluated under sub\-megabyte MCU SRAM budgets, occupying a previously empty region between high\-accuracy GPU\-based parsers and classical lightweight detectors\.
TABLE IV:Accuracy vs\. resources on Wireframe\. PriorsAP10\\mathrm\{sAP\}^\{10\}figures are from the respective papers or, for methods predating the sAP metric, from later re\-evaluations\[[12](https://arxiv.org/html/2607.06600#bib.bib42),[32](https://arxiv.org/html/2607.06600#bib.bib11),[13](https://arxiv.org/html/2607.06600#bib.bib40),[4](https://arxiv.org/html/2607.06600#bib.bib15),[31](https://arxiv.org/html/2607.06600#bib.bib10),[8](https://arxiv.org/html/2607.06600#bib.bib41),[35](https://arxiv.org/html/2607.06600#bib.bib8),[5](https://arxiv.org/html/2607.06600#bib.bib17),[30](https://arxiv.org/html/2607.06600#bib.bib9),[9](https://arxiv.org/html/2607.06600#bib.bib7),[27](https://arxiv.org/html/2607.06600#bib.bib3)\]; parameter counts are largely as tabulated by LINEA\[[13](https://arxiv.org/html/2607.06600#bib.bib40)\]\. LINEA reports the 462\-image Wireframe validation split\.MethodsAP10\\mathrm\{sAP\}^\{10\}ParamsOn MCU?DT\-LSD\[[12](https://arxiv.org/html/2607.06600#bib.bib42)\]71\.7217 MnoHAWPv2\[[32](https://arxiv.org/html/2607.06600#bib.bib11)\]69\.7∼\\sim11 MnoLINEA\-L\[[13](https://arxiv.org/html/2607.06600#bib.bib40)\]67\.925 MnoF\-Clip\[[4](https://arxiv.org/html/2607.06600#bib.bib15)\]67\.4∼\\sim28 MnoHAWP\[[31](https://arxiv.org/html/2607.06600#bib.bib10)\]66\.5∼\\sim10 MnoLINEA\-N\[[13](https://arxiv.org/html/2607.06600#bib.bib40)\]65\.03\.9 MnoEM\-LSD\[[8](https://arxiv.org/html/2607.06600#bib.bib41)\]63\.21\.1 MnoL\-CNN\[[35](https://arxiv.org/html/2607.06600#bib.bib8)\]62\.9∼\\sim9\.7 MnoM\-LSD\[[5](https://arxiv.org/html/2607.06600#bib.bib17)\]62\.11\.5 MnoM\-LSD\-tiny\[[5](https://arxiv.org/html/2607.06600#bib.bib17)\]58\.00\.6 M \(≥\\geq78 MB\)noAFM\[[30](https://arxiv.org/html/2607.06600#bib.bib9)\]24\.4∼\\sim43 MnoWireframe\-CNN\[[9](https://arxiv.org/html/2607.06600#bib.bib7)\]5\.1∼\\sim30 MnoClassical LSD\[[27](https://arxiv.org/html/2607.06600#bib.bib3)\]≈\\approx0n/aCPUOurs: F\-Clip int8 \(F746\)10\.60\.025 M \(0\.25 MB\)yesOurs: MiLSD \(H7\)24\.10\.39 M \(∼\\sim1 MB\)yes
### IV\-FQualitative Results
Figure 10:Detections on two Wireframe\-val images for v1 \(heatmap\), v2 \(endpoint\), and v3–v6 \(F\-Clip progression\); layout described in Section[IV\-F](https://arxiv.org/html/2607.06600#S4.SS6)\. The dominant visual step is v2→\\rightarrowv3; later versions refine segment placement incrementally\.Fig\.[10](https://arxiv.org/html/2607.06600#S4.F10)provides a qualitative side\-by\-side comparison of versions v1 through v6 on two representative Wireframe\-val images\[[7](https://arxiv.org/html/2607.06600#bib.bib38)\], offering visual corroboration of the quantitative trends reported above\. In the heatmap formulation \(v1\), predictions remain densely distributed across the image without resolving into discrete, well\-formed segments\. Endpoint regression \(v2\) produces scattered short segments that fail to reconstruct the underlying room structure\. Beginning with F\-Clip at v3, the detections progressively recover coherent architectural geometry, a visual pattern that directly mirrors the quantitativesAP10\\mathrm\{sAP\}^\{10\}jump at v3 \(3\.6→7\.23\.6\\rightarrow 7\.2\) and the more incremental refinement observed across v3–v6 \(7\.2→10\.67\.2\\rightarrow 10\.6\)\. The qualitative contrast is visually striking: F\-Clip yields coherent, well\-localized segments that align closely with salient architectural edges, whereas the heatmap and endpoint alternatives continue to produce noisy, fragmented, or incomplete detections that lack geometric consistency\.
## VMiLSD: Verification\-Augmented Detection on the STM32H7
Figure 11:MiLSD inference pipeline on the STM32H7\. An int8 fully\-convolutional backbone predicts F\-Clip center/length/angle maps; candidate segments are decoded by3×33\{\\times\}3peak non\-maximum suppression with sub\-pixel refinement; optional test\-time flip augmentation averages the predicted maps; and a small Line\-of\-Interest \(LoI\) head pools features along each candidate and re\-scores it\. All stages reuse a single≈1\{\\approx\}1MB activation arena\.The 320 KB SRAM budget of the STM32F746 imposes a hard ceiling on the detector developed in Section[II](https://arxiv.org/html/2607.06600#S2), restricting it to approximately 25k parameters and a correspondingly minimal activation footprint\. A larger yet still microcontroller\-class platform, the STM32H7, which provides 1 MB of SRAM and is executed via CMSIS\-NN\[[14](https://arxiv.org/html/2607.06600#bib.bib24)\]on a Cortex\-M7 core, affords sufficient headroom to accommodate both a more capable backbone and a substantially richer inference pipeline\. We designate the resulting system*MiLSD*\(Micro Line\-Segment Detector\)\. MiLSD preserves the F\-Clip center–length–angle output representation\[[4](https://arxiv.org/html/2607.06600#bib.bib15)\]\(Section[II\-A](https://arxiv.org/html/2607.06600#S2.SS1)\) introduced earlier and extends it along two complementary axes: first, a capacity\-scaled int8 backbone whose activation arena is deliberately sized to approach, but not exceed, the 1 MB SRAM budget; and second, a sequence of inference\-time refinement stages, comprising sub\-pixel decoding, test\-time augmentation, and a learned verification head, each of which contributes additional accuracy without increasing the size of the trained network\. The complete pipeline is illustrated in Fig\.[11](https://arxiv.org/html/2607.06600#S5.F11)\. In contrast to the F746 model, whose performance is bounded by its severely constrained 25k\-parameter budget, MiLSD exploits the expanded memory envelope to scale model capacity while preserving a lightweight, inference\-efficient pipeline appropriate for real\-time embedded deployment\.
### V\-ACapacity\-Scaled Backbone
With eight times the SRAM of the F746, the backbone width is increased until the int8 activation arena approaches, but does not exceed, 1 MB, yielding a approximately 0\.39 M\-parameter model \(peak arena approximately 1 MB at 256 px input; a 0\.7 MB fallback configuration is also retained for devices with tighter memory\)\. Trained as in Section[II](https://arxiv.org/html/2607.06600#S2)for 300 epochs, this model attainssAP10=17\.8\\mathrm\{sAP\}^\{10\}=17\.8on Wireframe\-val, versus10\.610\.6for the 25k F746 model, confirming that, beyond the extreme F746 regime, capacity is a genuine lever\. This 68% relative improvement demonstrates that the additional parameters are effectively utilized when the memory budget permits, validating our decision to scale the backbone for the H7 platform\.
### V\-BSub\-Pixel Decoding and Test\-Time Augmentation
The center peaks are refined to sub\-pixel accuracy by a one\-dimensional parabolic fit over each peak’s neighbourhood, directly targeting the endpoint precision that structural AP rewards \(sAP10:17\.8→18\.1\\mathrm\{sAP\}^\{10\}:17\.8\\rightarrow 18\.1\)\. This refinement is particularly beneficial for structural AP, which penalizes endpoint localization errors quadratically\. Averaging the predicted maps over the image and its horizontal, vertical, and diagonal flips \(test\-time augmentation, TTA\) further improves robustness \(sAP10:18\.1→21\.0\\mathrm\{sAP\}^\{10\}:18\.1\\rightarrow 21\.0\), a gain of nearly 3 points from the four\-view ensemble\. TTA multiplies inference latency by the number of views but reuses the same activation arena, so it does not raise peak memory, making it a memory\-free accuracy boost at the cost of increased latency\.
### V\-CLine\-of\-Interest Verification Head
The largest gain comes from a trained verifier in the spirit of L\-CNN\[[35](https://arxiv.org/html/2607.06600#bib.bib8)\]and HAWP\[[31](https://arxiv.org/html/2607.06600#bib.bib10)\], scaled to the microcontroller\. With the backbone*frozen*, a small multilayer perceptron takes, for each candidate segment, features bilinearly pooled at 32 points along the line from the backbone feature map and the output maps, together with a short geometric descriptor, and predicts a verification score \(real vs\. spurious\), trained with one\-to\-one matched labels\. Re\-ranking by verification×\\timescenter score reaches the MiLSD row of Table[V](https://arxiv.org/html/2607.06600#S5.T5)\. The decode and post\-filtering design were obtained by analyzing the inference stages of the principal wireframe parsers\[[35](https://arxiv.org/html/2607.06600#bib.bib8),[31](https://arxiv.org/html/2607.06600#bib.bib10),[4](https://arxiv.org/html/2607.06600#bib.bib15),[33](https://arxiv.org/html/2607.06600#bib.bib36),[34](https://arxiv.org/html/2607.06600#bib.bib37)\]; across that study, weight\-free post\-processing saturates nearsAP10=21\\mathrm\{sAP\}^\{10\}=21, and only the trained verifier advances beyond it, contributing an additional 3\.1 points to reach 24\.1\.
TABLE V:MiLSD on Wireframe\-val \(STM32H7 model\)\. Each stage is inference\-time only; the trained network is unchanged\. The oracle ranks the same candidates by their true labels and is the recall ceiling of the candidate set\.The oracle row \(sAP10=37\.4\\mathrm\{sAP\}^\{10\}=37\.4\) bounds what the candidate set can deliver; the LoI head recovers about 55% of the gap from TTA\. The remainder reflects unannotated image edges and recall limits; junction\-based candidate generation\[[35](https://arxiv.org/html/2607.06600#bib.bib8)\]could help but exceeds the 1 MB budget\.
### V\-DQualitative Results
Fig\.[12](https://arxiv.org/html/2607.06600#S5.F12)shows MiLSD on held\-out scenes\. The detector is complete on the salient structure \(cabinetry, counters, window mullions, and architectural edges are recovered at correct orientation and length\)\. The dense scene of example 2 illustrates the verifier suppressing the redundant ridge detections that a raw center\-heatmap emits, demonstrating the effectiveness of the learned verification head\. Example 3 shows that the remaining stray predictions are real but unannotated edges \(reflections, texture seams\), consistent with the oracle analysis above, which identified that many false positives are actually unlabeled ground\-truth edges\. These qualitative results validate that MiLSD achieves its accuracy gains through meaningful structural understanding rather than overfitting to the training set\.
Figure 12:Four Wireframe\-val scenes: ground truth \(left, green\) and MiLSD output \(right, yellow\) after LoI verification; line opacity encodes verifier confidence\.
### V\-EDeployment Budget
The SRAM ceiling is set by the backbone’s int8 activation arena \(approximately 1 MB\); the LoI head’s pooled feature \(<<0\.5 MB\) fits within that peak and TTA reuses it, so neither raises it\. Weights total approximately 0\.54 MB \(backbone plus head\), well within flash\. Latency is dominated by the int8 convolutions and scales with the number of TTA views; the LoI head adds a small per\-line cost\. On\-device arena and latency must be confirmed with ST Edge AI\[[24](https://arxiv.org/html/2607.06600#bib.bib34)\]; the figures here are software\-measured\. The memory\-efficient design ensures that all inference stages operate within the 1 MB SRAM budget, making MiLSD deployable on the STM32H7 without requiring external memory\. Source code is released for reproducibility\[[6](https://arxiv.org/html/2607.06600#bib.bib39)\]\.
## VIConclusion
This paper presented MiLSD, a line\-segment detector explicitly designed for microcontroller\-scale memory, alongside a systematic study of the accuracy–resource trade\-off under extreme memory constraints\. We compared three output representations and found that the F\-Clip center\-with\-length\-and\-angle encoding learns most effectively at small model sizes, achievingsAP10=10\.6\\mathrm\{sAP\}^\{10\}=10\.6with only 25k parameters\. Our quantization study revealed that 8\-bit weights preserve full\-precision accuracy, while 4\-bit quantization collapses, particularly in the\(cos2θ,sin2θ\)\(\\cos 2\\theta,\\sin 2\\theta\)angle regression, with quantization\-aware training recovering only part of the loss\. By scaling the backbone to a 1 MB activation budget and adding inference\-time enhancements including sub\-pixel decoding, test\-time augmentation, and a Line\-of\-Interest verification head, MiLSD improvessAP10\\mathrm\{sAP\}^\{10\}from10\.610\.6at 0\.25 MB to24\.124\.1within 1 MB on the ShanghaiTech Wireframe benchmark\.
To our knowledge, no prior work characterizes the joint trade\-off among representation choice, quantization bit\-width, and on\-device post\-processing at this memory scale\. F\-Clip wins at int8 but its angle head is int4\-fragile, whereas heatmaps tolerate lower precision yet require an external linker, a finding that informs future TinyML geometric\-vision designs\. As expected, Wireframe accuracy remains far below GPU\-class parsers given the SRAM envelope we target; that gap is inherent to the memory regime rather than a limitation we aim to overcome\.
## References
- \[1\]\(2011\)EDLines: a real\-time line segment detector with a false detection control\.Pattern Recognition Letters32\(13\),pp\. 1633–1642\.Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.8.2.1.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p2.2)\.
- \[2\]C\. Banbury, C\. Zhou, I\. Fedorov, R\. Matas Navarro, U\. Thakker, D\. Gope, V\. Janapa Reddi, M\. Mattina, and P\. N\. Whatmough\(2021\)MicroNets: neural network architectures for deploying tinyml applications on commodity microcontrollers\.Proc\. of Machine Learning and Systems \(MLSys\)\.Note:arXiv:2010\.11267Cited by:[§I](https://arxiv.org/html/2607.06600#S1.p4.1)\.
- \[3\]Y\. Bengio, N\. Léonard, and A\. Courville\(2013\)Estimating or propagating gradients through stochastic neurons for conditional computation\.arXiv preprint arXiv:1308\.3432\.Note:Straight\-through estimatorCited by:[§I](https://arxiv.org/html/2607.06600#S1.p4.1),[§II\-C](https://arxiv.org/html/2607.06600#S2.SS3.p1.1)\.
- \[4\]X\. Dai, H\. Gong, S\. Wu, X\. Yuan, and Y\. Ma\(2022\)Fully convolutional line parsing\.Neurocomputing506,pp\. 1–11\.Note:F\-Clip\. arXiv:2104\.11207External Links:[Document](https://dx.doi.org/10.1016/j.neucom.2022.07.026)Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.4.4.3.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p2.2),[§I](https://arxiv.org/html/2607.06600#S1.p4.1),[Figure 1](https://arxiv.org/html/2607.06600#S2.F1),[§II\-A](https://arxiv.org/html/2607.06600#S2.SS1.p4.4),[Figure 9](https://arxiv.org/html/2607.06600#S4.F9),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.5.3.2.1.1),[§V\-C](https://arxiv.org/html/2607.06600#S5.SS3.p1.2),[§V](https://arxiv.org/html/2607.06600#S5.p1.1)\.
- \[5\]G\. Gu, B\. Ko, S\. Go, S\. Lee, J\. Lee, and M\. Shin\(2022\)Towards light\-weight and real\-time line segment detection\.InAAAI Conf\. on Artificial Intelligence,Note:M\-LSD / M\-LSD\-tiny; MobileNetV2, center\+displacement\. arXiv:2106\.00186Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.5.2.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p2.2),[§I](https://arxiv.org/html/2607.06600#S1.p4.1),[Figure 1](https://arxiv.org/html/2607.06600#S2.F1),[§II\-A](https://arxiv.org/html/2607.06600#S2.SS1.p3.1),[Figure 9](https://arxiv.org/html/2607.06600#S4.F9),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.12.15.5.1.1.1),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.8.6.2.1.1)\.
- \[6\]P\. Hassani Shariat Panahi, A\. H\. Jalilvand, and M\. H\. Najafi\(2026\)MiLSD: micro line\-segment detector\.Note:[https://github\.com/F4RAN/MiLSD](https://github.com/F4RAN/MiLSD)Source code, trained models, and evaluation scripts for the MiLSD implementation\.Cited by:[§V\-E](https://arxiv.org/html/2607.06600#S5.SS5.p1.1)\.
- \[7\]P\. Hassani Shariat Panahi\(2026\)LSD\-TML\-All: training and evaluation notebooks for versions v1–v6\.Note:[https://www\.kaggle\.com/code/parsahshariatpanahi/lsd\-tml\-all](https://www.kaggle.com/code/parsahshariatpanahi/lsd-tml-all)Kaggle notebooks for the representation and quantization study \(v1–v6\)\.Cited by:[§III](https://arxiv.org/html/2607.06600#S3.SS0.SSS0.Px2.p1.1),[§IV\-F](https://arxiv.org/html/2607.06600#S4.SS6.p1.3)\.
- \[8\]S\. Hu, L\. Zhao, and Q\. Wang\(2026\)EM\-lsd: a lightweight and efficient model for multi\-scale line segment detection\.Robotics and Autonomous Systems195,pp\. 105192\.External Links:ISSN 0921\-8890,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.robot.2025.105192),[Link](https://www.sciencedirect.com/science/article/pii/S0921889025002891)Cited by:[§IV\-E](https://arxiv.org/html/2607.06600#S4.SS5.p1.7),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.12.14.4.1.1.1)\.
- \[9\]K\. Huang, Y\. Wang, Z\. Zhou, T\. Ding, S\. Gao, and Y\. Ma\(2018\)Learning to parse wireframes in images of man\-made environments\.InIEEE/CVF Conf\. on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.1.1.2.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p3.3),[§III](https://arxiv.org/html/2607.06600#S3.SS0.SSS0.Px1.p1.5),[Figure 9](https://arxiv.org/html/2607.06600#S4.F9),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.10.8.2.1.1)\.
- \[10\]S\. Huang, F\. Qin, P\. Xiong, N\. Ding, Y\. He, and X\. Liu\(2020\)TP\-LSD: tri\-points based line segment detector\.InEuropean Conf\. on Computer Vision \(ECCV\),Note:arXiv:2009\.05505Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.2.2.2.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p4.1),[Figure 1](https://arxiv.org/html/2607.06600#S2.F1),[§II\-A](https://arxiv.org/html/2607.06600#S2.SS1.p3.1)\.
- \[11\]B\. Jacob, S\. Kligys, B\. Chen, M\. Zhu, M\. Tang, A\. Howard, H\. Adam, and D\. Kalenichenko\(2018\)Quantization and training of neural networks for efficient integer\-arithmetic\-only inference\.InIEEE/CVF Conf\. on Computer Vision and Pattern Recognition \(CVPR\),Note:arXiv:1712\.05877Cited by:[§I](https://arxiv.org/html/2607.06600#S1.p4.1),[§II\-C](https://arxiv.org/html/2607.06600#S2.SS3.p1.1)\.
- \[12\]S\. Janampa and M\. Pattichis\(2025\-02\)DT\-lsd: deformable transformer\-based line segment detection\.InProceedings of the Winter Conference on Applications of Computer Vision \(WACV\),pp\. 3477–3486\.Cited by:[§IV\-E](https://arxiv.org/html/2607.06600#S4.SS5.p1.7),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.12.11.1.1.1.1)\.
- \[13\]S\. Janampa and M\. Pattichis\(2025\)LINEA: fast and accurate line detection using scalable transformers\.External Links:2505\.16264,[Link](https://arxiv.org/abs/2505.16264)Cited by:[§IV\-E](https://arxiv.org/html/2607.06600#S4.SS5.p1.7),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.12.12.2.1.1.1),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.12.13.3.1.1.1)\.
- \[14\]L\. Lai, N\. Suda, and V\. Chandra\(2018\)CMSIS\-NN: efficient neural network kernels for arm cortex\-m CPUs\.arXiv preprint arXiv:1801\.06601\.Cited by:[§II\-C](https://arxiv.org/html/2607.06600#S2.SS3.p1.1),[§V](https://arxiv.org/html/2607.06600#S5.p1.1)\.
- \[15\]H\. Li, H\. Yu, J\. Wang, W\. Yang, L\. Yu, and S\. Scherer\(2021\)ULSD: unified line segment detection across pinhole, fisheye, and spherical cameras\.ISPRS J\. of Photogrammetry and Remote Sensing\.Note:arXiv:2011\.03174Cited by:[§I](https://arxiv.org/html/2607.06600#S1.p3.3)\.
- \[16\]J\. Lin, W\. Chen, H\. Cai, C\. Gan, and S\. Han\(2021\)MCUNetV2: memory\-efficient patch\-based inference for tiny deep learning\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2110\.15352Cited by:[§I](https://arxiv.org/html/2607.06600#S1.p4.1)\.
- \[17\]J\. Lin, W\. Chen, Y\. Lin, J\. Cohn, C\. Gan, and S\. Han\(2020\)MCUNet: tiny deep learning on IoT devices\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2007\.10319Cited by:[§I](https://arxiv.org/html/2607.06600#S1.p4.1)\.
- \[18\]S\. I\. Mirzadeh, M\. Farajtabar, A\. Li, N\. Levine, A\. Matsukawa, and H\. Ghasemzadeh\(2020\)Improved knowledge distillation via teacher assistant\.InAAAI Conf\. on Artificial Intelligence,Note:capacity\-gap effectCited by:[§IV\-D](https://arxiv.org/html/2607.06600#S4.SS4.p1.7)\.
- \[19\]P\. Novac, G\. Boukli Hacene, A\. Pegatoquet, B\. Miramond, and V\. Gripon\(2021\)Quantization and deployment of deep neural networks on microcontrollers\.Sensors21\(9\),pp\. 2984\.Cited by:[§I](https://arxiv.org/html/2607.06600#S1.p4.1)\.
- \[20\]C\. Ossimitz and N\. Taherinejad\(2021\-05\)A fast line segment detector using approximate computing\.pp\. 1–5\.External Links:[Document](https://dx.doi.org/10.1109/ISCAS51556.2021.9401660)Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.17.11.1.1.1),[§III](https://arxiv.org/html/2607.06600#S3.SS0.SSS0.Px1.p1.5)\.
- \[21\]R\. Pautrat, D\. Barath, V\. Larsson, M\. R\. Oswald, and M\. Pollefeys\(2023\)DeepLSD: line segment detection and refinement with deep image gradients\.InIEEE/CVF Conf\. on Computer Vision and Pattern Recognition \(CVPR\),Note:Line attraction field\. arXiv:2212\.07766Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.13.7.1.1.1)\.
- \[22\]M\. Rusci, A\. Capotondi, and L\. Benini\(2019\)Memory\-driven mixed low precision quantization for enabling deep network inference on microcontrollers\.arXiv preprint arXiv:1905\.13082\.Cited by:[§I](https://arxiv.org/html/2607.06600#S1.p4.1)\.
- \[23\]STMicroelectronicsSTM32F746xx arm cortex\-m7 microcontroller datasheet\.Note:[https://www\.st\.com/en/microcontrollers\-microprocessors/stm32f746ng\.html](https://www.st.com/en/microcontrollers-microprocessors/stm32f746ng.html)320 KB SRAM, 1 MB flash, 216 MHz Cortex\-M7\.Cited by:[§I](https://arxiv.org/html/2607.06600#S1.p4.1)\.
- \[24\]STMicroelectronics\(2025\)ST Edge AI Suite / X\-CUBE\-AI / ST Edge AI Developer Cloud\.Note:[https://www\.st\.com/en/embedded\-software/x\-cube\-ai\.html](https://www.st.com/en/embedded-software/x-cube-ai.html)Model analysis, int8 code generation, and on\-target benchmarking for STM32\.Cited by:[§V\-E](https://arxiv.org/html/2607.06600#S5.SS5.p1.1)\.
- \[25\]I\. Suárez, J\. M\. Buenaposada, and L\. Baumela\(2022\)ELSED: enhanced line segment drawing\.Pattern Recognition127,pp\. 108619\.Note:arXiv:2108\.03144External Links:[Document](https://dx.doi.org/10.1016/j.patcog.2022.108619)Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.9.3.1.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p2.2)\.
- \[26\]L\. Teplyakov, L\. Erlygin, and E\. Shvets\(2022\)LSDNet: trainable modification of LSD algorithm for real\-time line segment detection\.IEEE Access10,pp\. 45256–45265\.Note:arXiv:2209\.04642External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2022.3169177)Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.15.9.1.1.1)\.
- \[27\]R\. G\. von Gioi, J\. Jakubowicz, J\. Morel, and G\. Randall\(2010\)LSD: a fast line segment detector with a false detection control\.IEEE Trans\. on Pattern Analysis and Machine Intelligence \(TPAMI\)32\(4\),pp\. 722–732\.Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.7.1.1.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p1.1),[§I](https://arxiv.org/html/2607.06600#S1.p2.2),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.11.9.2.1.1)\.
- \[28\]Wanget al\.\(2024\)EvLSD\-IED: event\-based line segment detection with image\-to\-event distillation\.IEEE Trans\. on Instrumentation and Measurement \(TIM\)\.Note:Knowledge\-distillation precedent for LSD\.Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.16.10.1.1.1)\.
- \[29\]Y\. Xu, W\. Xu, D\. Cheung, and Z\. Tu\(2021\)Line segment detection using transformers without edges\.InIEEE/CVF Conf\. on Computer Vision and Pattern Recognition \(CVPR\),Note:LETR\. arXiv:2101\.01909Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.14.8.1.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p3.3)\.
- \[30\]N\. Xue, S\. Bai, F\. Wang, G\. Xia, T\. Wu, and L\. Zhang\(2019\)Learning attraction field representation for robust line segment detection\.InIEEE/CVF Conf\. on Computer Vision and Pattern Recognition \(CVPR\),Note:AFM\. arXiv:1812\.02122Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.11.5.1.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p3.3),[Figure 9](https://arxiv.org/html/2607.06600#S4.F9),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.9.7.2.1.1)\.
- \[31\]N\. Xue, T\. Wu, S\. Bai, F\. Wang, G\. Xia, L\. Zhang, and P\. H\.S\. Torr\(2020\)Holistically\-attracted wireframe parsing\.InIEEE/CVF Conf\. on Computer Vision and Pattern Recognition \(CVPR\),Note:HAWP\. arXiv:2003\.01663Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.12.6.1.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p2.2),[§I](https://arxiv.org/html/2607.06600#S1.p3.3),[Figure 9](https://arxiv.org/html/2607.06600#S4.F9),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.6.4.2.1.1),[§V\-C](https://arxiv.org/html/2607.06600#S5.SS3.p1.2)\.
- \[32\]N\. Xue, T\. Wu, S\. Bai, F\. Wang, G\. Xia, L\. Zhang, and P\. H\.S\. Torr\(2023\)Holistically\-attracted wireframe parsing: from supervised to self\-supervised learning\.IEEE Trans\. on Pattern Analysis and Machine Intelligence \(TPAMI\)\.Note:HAWPv2/v3\. arXiv:2210\.12971Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.12.6.1.1.1),[Figure 9](https://arxiv.org/html/2607.06600#S4.F9),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.4.2.2.1.1)\.
- \[33\]Z\. Zhang, Z\. Li, N\. Bi, J\. Zheng, J\. Wang, K\. Huang, W\. Luo, Y\. Xu, and S\. Gao\(2019\)PPGNet: learning point\-pair graph for line segment detection\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 7105–7114\.Cited by:[§V\-C](https://arxiv.org/html/2607.06600#S5.SS3.p1.2)\.
- \[34\]K\. Zhao, Q\. Han, C\. Zhang, J\. Xu, and M\. Cheng\(2020\)Deep Hough transform for semantic line detection\.InProceedings of the European Conference on Computer Vision \(ECCV\),Cited by:[§V\-C](https://arxiv.org/html/2607.06600#S5.SS3.p1.2)\.
- \[35\]Y\. Zhou, H\. Qi, and Y\. Ma\(2019\)End\-to\-end wireframe parsing\.InIEEE/CVF Int\. Conf\. on Computer Vision \(ICCV\),Note:L\-CNN\. arXiv:1905\.03246Cited by:[TABLE I](https://arxiv.org/html/2607.06600#S1.T1.5.10.4.1.1.1),[§I](https://arxiv.org/html/2607.06600#S1.p2.2),[§I](https://arxiv.org/html/2607.06600#S1.p3.3),[Figure 9](https://arxiv.org/html/2607.06600#S4.F9),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4),[TABLE IV](https://arxiv.org/html/2607.06600#S4.T4.7.5.2.1.1),[§V\-C](https://arxiv.org/html/2607.06600#S5.SS3.p1.2),[§V\-C](https://arxiv.org/html/2607.06600#S5.SS3.p2.1)\.
![[Uncaptioned image]](https://arxiv.org/html/2607.06600v1/author1.png)Parsa Hassani Shariat Panahireceived his B\.Sc\. degree in Computer Engineering from Azad University South Tehran Branch, Tehran, Iran, and the M\.Sc\. degree in Computer Engineering \- Computer Networks from Iran University of Science and Technology, Tehran, Iran\. His research interests include cellular networks, QoE assessment, telecommunication networks, wireless communication, and machine learning\. He can be reached at parsa\_hassani@comp\.iust\.ac\.ir\.![[Uncaptioned image]](https://arxiv.org/html/2607.06600v1/author2.png)Amir Hossein Jalilvandreceived his B\.Sc\. degree in Computer Engineering from Bu\-Ali Sina University, Hamadan, Iran, and the M\.Sc\. degree in Computer Engineering \- Computer Architecture from Iran University of Science and Technology, Tehran, Iran\. He is currently pursuing his Ph\.D\. in Computer Engineering\. His research interests include cellular networks, stochastic and unary computing, computer architecture, fuzzy logic, and machine learning\. Mr\. Jalilvand has authored several publications in these fields\. He can be reached at jalilvand\_a@comp\.iust\.ac\.ir\.![[Uncaptioned image]](https://arxiv.org/html/2607.06600v1/author3.png)M\. Hassan Najafireceived his Ph\.D\. in electrical and electronics engineering from the University of Minnesota\-Twin Cities, Minneapolis, MN, USA, in 2018\. He is currently an Associate Professor at the Electrical, Computer, and Systems Engineering Department at Case Western Reserve University\. His research interests include stochastic and approximate computing, unary processing, in\-memory computing, and hyperdimensional computing\. He has authored/coauthored more than 120 peer\-reviewed papers and has been granted 12 U\.S\. patents with more pending\. Dr\. Najafi received the NSF CAREER Award in 2024, the Best Paper Award at GLSVLSI’23 and ICCD’17, and the 2018 EDAA Outstanding Dissertation Award\. Dr\. Najafi is a senior member of IEEE and a senior member of the U\.S\. National Academy of Inventors \(NAI\)\. He can be reached at najafi@case\.edu\.Similar Articles
LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
LightMIS presents a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation, achieving high accuracy with reduced parameters and GFLOPs, suitable for on-device execution.
Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment
This paper systematically evaluates component-wise quantization of small vision-language models on Jetson edge devices, finding that model architecture (MoE vs dense) significantly affects quantization sensitivity and that quantization errors are largely additive except along modality-alignment paths.
Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs
This paper investigates whether open-source quantized LLMs encode a linearly separable truthfulness signal in their hidden states. Across three 7B-8B instruction-tuned models, a linear probe on a single mid-network layer achieves 0.904-1.000 AUROC on hallucination detection benchmarks, outperforming sampling-based methods.
Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
This paper presents a framework for quantizing vision-language models to 2.7 bits per parameter, enabling efficient mobile deployment by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB while preserving performance on visual QA tasks.
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
This paper proposes a multi-pass prompt verification method to improve the performance of quantized LLMs (LLaMA-3.1 8B) in qualitative analysis, reducing hallucinations and increasing stability across different quantization levels (8-bit, 4-bit, 3-bit, 2-bit).