Low-Overhead Error-Corrected QCNNs Using Bivariate Bicycle Codes

arXiv cs.LG Papers

Summary

Proposes a low-overhead error correction technique for quantum convolutional neural networks using bivariate bicycle codes, demonstrating improved learning under realistic noise compared to unprotected QCNNs.

arXiv:2607.05724v1 Announce Type: new Abstract: Quantum convolutional neural networks (QCNNs) combine the power of quantum computing and classical CNN for computational speedup in classification tasks. However, noise levels on state-of-the-art quantum devices remain too high for practical QCNN execution. In addition, despite the reliable surface code providing a method for error rates below a threshold value, they have a prohibitively large qubit cost. Recently introduced bivariate bicycle (BB) codes are of particular interest for their high error threshold, constant encoding rate, and linear code distance. Through simulation with realistic hardware noise sources, we demonstrate that a 4-qubit unprotected QCNN fails to converge and exhibits a worse learning rate compared to numerical simulations. Addressing both limitations, we propose a distance-4 BB quantum error-correction (QEC) technique for QCNNs. In doing so, we validate that our low-overhead QEC technique for QCNNS represents a step toward practical QCNNs.
Original Article
View Cached Full Text

Cached at: 07/08/26, 04:44 AM

# Low-Overhead Error-Corrected QCNNs Using Bivariate Bicycle Codes
Source: [https://arxiv.org/html/2607.05724](https://arxiv.org/html/2607.05724)
###### Abstract

Quantum convolutional neural networks \(QCNNs\) combine the power of quantum computing and classical CNN for computational speedup in classification tasks\. However, noise levels on state\-of\-the\-art quantum devices remain too high for practical QCNN execution\. In addition, despite the reliable surface code providing a method for error rates below a threshold value, they have a prohibitively large qubit cost\. Recently introduced bivariate bicycle \(BB\) codes are of particular interest for their high error threshold, constant encoding rate, and linear code distance\. Through simulation with realistic hardware noise sources, we demonstrate that a 4\-qubit unprotected QCNN fails to converge and exhibits a worse learning rate compared to numerical simulations\. Addressing both limitations, we propose a distance\-4 BB quantum error\-correction \(QEC\) technique for QCNNs\. In doing so, we validate that our low\-overhead QEC technique for QCNNS represents a step toward practical QCNNs\.

## IIntroduction

Advancements in machine learning \(ML\) have made it a crucial technology across various domains, from signal processing to healthcare, with applications such as speech recognition, computer vision, and drug discovery\. However, ML procedures suffer from gradient vanishing in high\-dimensional parameter spaces\[[16](https://arxiv.org/html/2607.05724#bib.bib1),[27](https://arxiv.org/html/2607.05724#bib.bib2)\]and quadratic sample sizes and training time for certain tasks\[[17](https://arxiv.org/html/2607.05724#bib.bib40)\]\. On the other hand, quantum computing \(QC\) shows exponential computational speedups\[[20](https://arxiv.org/html/2607.05724#bib.bib6),[26](https://arxiv.org/html/2607.05724#bib.bib10),[15](https://arxiv.org/html/2607.05724#bib.bib12)\]\. The combination of these two technologies, termed quantum machine learning \(QML\), promises processing of data in exponentially large feature spaces\[[4](https://arxiv.org/html/2607.05724#bib.bib3),[13](https://arxiv.org/html/2607.05724#bib.bib4)\]by leveraging quantum effects like superposition and entanglement, and more efficient navigation of high\-dimensional optimization landscapes via quantum\-enhanced optimization\[[24](https://arxiv.org/html/2607.05724#bib.bib5)\]\. Although QML holds significant promise, the limited qubit counts of current noisy intermediate\-scale quantum \(NISQ\) devices pose substantial challenges\. Furthermore, these devices suffer from high physical error rates introduced by noise sources including stray electromagnetic fields, cosmic rays, and thermal and temporal decoherence of quantum states\[[9](https://arxiv.org/html/2607.05724#bib.bib16)\]\.

Quantum error correction \(QEC\) is a method for protecting fragile quantum information while controlling qubits to induce a desired computation\. For over two decades, the topological toric code has served as the canonical model for topological quantum error correction\.\[[19](https://arxiv.org/html/2607.05724#bib.bib22)\]\. The code encodes two logical qubits into ad×dd\\times dlattice of physical qubits, with the total qubit count scaling asn=2​d2n=2d^\{2\}, whereddis the code distance\. With minimum distancedd, the code can correct up to⌊d−12⌋\\left\\lfloor\\frac\{d\-1\}\{2\}\\right\\rfloorarbitrary single\-qubit errors\. The toric code thus provides reliable protection of quantum information, however it suffers from inefficient encoding rates that makes scaling to hundreds of logical qubits prohibitively resource\-intensive\. Recent advances in quantum Low\-Density Parity\-Check \(qLDPC\), most notably the bivariate bicycle \(BB\) code\[[5](https://arxiv.org/html/2607.05724#bib.bib18)\], have emerged as a promising alternative\. The constant encoding rate of BB codes enables fault\-tolerant quantum memory with constant space overhead\[[5](https://arxiv.org/html/2607.05724#bib.bib18),[12](https://arxiv.org/html/2607.05724#bib.bib30)\], making them an attractive candidate for scalable architectures\. However, a key open problem remains for the BB code is that it has yet to support fault\-tolerant computation, particularly for deep\-circuit algorithms and poses an obstacle to real\-world QML use cases\.

To this end, we introduce a constant\-overhead QEC protocol that integrates with QML, particularly, quantum convolutional neural networks \(QCNN\)\. QCNNs are the quantum counterpart of classical convolutional neural networks \(CNNs\)\. These classical supervised learning architecture trains on labeled data to learn a mapping between inputs and their respective labels\. QCNNs extend this framework by leveraging quantum properties such as quantum parallelism, to accelerate computation\. However, practical QCNN execution on NISQ hardware is severely hindered by quantum decoherence and gate noise, which degrades the loss landscape and therefore compromise training and inference\. The proposed constant\-overhead QEC protocol directly mitigates these impairments, enabling more reliable and practical QCNN deployment on near\-term quantum devices\.

### I\-ARelated Work

In\[[7](https://arxiv.org/html/2607.05724#bib.bib7)\], the authors introduced a QCNN architecture that incorporated the multi\-scale entanglement renormalization ansatz \(MERA\) with QEC, demonstrating QCNNs for quantum phase recognition \(QPR\) of11D symmetry\-protected topological \(SPT\) phases and designing a QEC scheme\. Their QCNN uses𝒪​\(log⁡N\)\\mathcal\{O\}\(\\log N\)variational parameters for input sizes ofNNqubits\. However, there is still a gap for QEC codes for error correction on deep circuits, like the QCNN, while being practical\.

In\[[5](https://arxiv.org/html/2607.05724#bib.bib18)\], the authors proposed the BB code that requiresnnancillary qubits, a depth\-77syndrome measurement circuit, and a degree\-66qubit connectivity graph, enabling efficient error correction with minimal overhead\. Their code demonstrated reduced hardware requirements compared to the surface code and making fault\-tolerant quantum memory feasible for near\-term quantum processors\. The application of the BB code for quantum computation is a high barrier, given the BB code works for Clifford gates only, unlike the Non\-Clifford gates in the QCNN\. In\[[32](https://arxiv.org/html/2607.05724#bib.bib14)\], the authors demonstrated low\-overhead qLDPC codes, a distance\-33qLDPC code and a distance\-44BB code utilizing periodic execution of the syndrome extraction circuit\.

The work of\[[3](https://arxiv.org/html/2607.05724#bib.bib21)\]addresses the classical simulability and computational hardness of the functions of quantum models\. The key distinction is that our research addresses trainability on physical hardware and how hardware noise introduces a stochastic, non\-unitary perturbation at each circuit evaluation, degrading the fidelity of gradient estimates and making gradient estimates unreliable\. This is orthogonal to questions of whether the learned function is classically hard to simulate or evaluate\.

### I\-BContributions

Motivated by the the application of low\-overhead BB code on a real3232long\-range\-coupled NISQ transmon circuit\[[32](https://arxiv.org/html/2607.05724#bib.bib14)\], and the computational speed of machine learning training by the QCNN\[[7](https://arxiv.org/html/2607.05724#bib.bib7)\], we apply the distance44qLDPC codes for the QCNN\. Specifically, we propose a distance\-44BB code\[[5](https://arxiv.org/html/2607.05724#bib.bib18)\]combined with a Feed\-Forward Neural Network \(FFNN\) as an error correction technique for the44\-qubit QCNN\[[7](https://arxiv.org/html/2607.05724#bib.bib7)\]\. We integrate a constant encoding rate, linear code distance BB code that offers an error threshold of 0\.3%\\%for the standard circuit\-based noise model to allow for low\-overhead scaling to large QCNNs\. Next, we compare the QCNN with the new QEC method against an unprotected 4 qubit QCNN in simulation at various NISQ error rates\. We observed that the QCNN with the BB code achieves satisfactory results, which shows the implementation achieves a constant encoding rate and linear code distance offered by the BB code, with low qubit, online overhead for error correction\. We further show that the QEC protocol can be sustained under low\-overhead additional qubit resource requirement while reducing learning loss and improving convergence\.\.

The rest of the paper is organized as follows\. Section II presents the key background concepts described in our model implementation and the experiments conducted\. Section III presents the proposed BB coded QCNN\. Section IV present and discuss the results, comparing them to a plain QCNN architecture, and Section V concludes the paper along with some directions for future research\.

## IIBackground

![Refer to caption](https://arxiv.org/html/2607.05724v1/0p_results_initial_params.png)\(a\)
![Refer to caption](https://arxiv.org/html/2607.05724v1/0p_results_optimal_params.png)\(b\)

Figure 1:\(a\) Points show a64×6464\\times 64test set of ground states overh1h\_\{1\}andh2h\_\{2\}for a Hamiltonian withJ=1J=1\. Phase\-boundary points \(blue and red diamonds\) come from infinite\-size density\-matrix renormalization group \(DMRG\)\. Colors indicate the circuit’s expectation values found via evolutionary search forN=4N=4spins using initial, untrained weights\. \(An artifact can be seen, possibly because of the low qubit spin state\.\) \(b\) Uses the same test set and with trained weights after 100 iterations\.### II\-AQuantum Convolutional Neural Network

Classical CNNs has provided a successful ML framework for image recognition and are widely used to approximate functions due to their ability to capture hierarchical features\[[20](https://arxiv.org/html/2607.05724#bib.bib6),[21](https://arxiv.org/html/2607.05724#bib.bib28)\]\. A QCNN is a quantum counterpart to the classical CNN that exploits quantum properties, such as quantum parallelism, to optimize computation\. The QCNN architecture efficiently implements a hierarchical structure of quantum many\-body states represented as an inverse MERA\[[7](https://arxiv.org/html/2607.05724#bib.bib7)\]\.

The MERA is a tensor network that is designed to efficiently represent quantum many\-body states using a hierarchical and multi\-layer structure\. It can compute low energy properties of the system, described as the ground state\|ΨGS⟩∈ℂ⊗N⊆ℋ\\left\|\\Psi\_\{\\mathrm\{GS\}\}\\right\\rangle\\in\\mathbb\{C\}^\{\\otimes N\}\\subseteq\\mathcal\{H\}, whereℋ\\mathcal\{H\}denotes the Hilbert space of aNN\-qubit system\. Each layerτ\\tauof the MERA consists of disentanglersUτ=⨂i=1NτuiU\_\{\\tau\}=\\bigotimes\_\{i=1\}^\{N\_\{\\tau\}\}u\_\{i\}and isometriesWτ=⨂i=1NτwiW\_\{\\tau\}=\\bigotimes\_\{i=1\}^\{N\_\{\\tau\}\}w\_\{i\}, whereNτN\_\{\\tau\}denotes the number of disentanglers and isometries in that layer,uiu\_\{i\}denotes the parameterized two\-qubit unitary operator applied at theiith convolutional unit of layerτ\\tau, andwiw\_\{i\}denotes the corresponding isometric tensor at the reduced pooling subsystem\.NτN\_\{\\tau\}decreases withτ\\tauas sites are progressively coarse\-grained\. The transformation from a coarse\-grained level to a finer\-grained one is expressed as

Vτ=Uτ∘Wτ:ℋℒτ⟶ℋℒτ−1,V\_\{\\tau\}=U\_\{\\tau\}\\circ W\_\{\\tau\}:\\mathcal\{H\}\_\{\\mathcal\{L\}\_\{\\tau\}\}\\longrightarrow\\mathcal\{H\}\_\{\\mathcal\{L\}\_\{\\tau\-1\}\},whereℋℒτ=⨂s∈ℒτℋs\\mathcal\{H\}\_\{\\mathcal\{L\}\_\{\\tau\}\}=\\bigotimes\_\{s\\in\\mathcal\{L\}\_\{\\tau\}\}\\mathcal\{H\}\_\{s\}is the Hilbert space of the coarse\-grained lattice at layerτ\\tau, with\|ℒτ\|<\|ℒτ−1\|\\lvert\\mathcal\{L\}\_\{\\tau\}\\rvert<\\lvert\\mathcal\{L\}\_\{\\tau\-1\}\\rvertas sites are progressively reduced\. The operator∘\\circdenotes function composition, andVτV\_\{\\tau\}is the output unitary fed into the next layer\[[28](https://arxiv.org/html/2607.05724#bib.bib8)\]\.

A QCNN can be thought of as a circuit\-based ansatz of an inverse MERA with nested QEC\. Learning occurs through gradual tuning of its initial unitary operations\[[7](https://arxiv.org/html/2607.05724#bib.bib7)\]\. The QCNN acts on input states\|ψα⟩∈ℋℒ\|\\psi\_\{\\alpha\}\\rangle\\in\\mathcal\{H\}\_\{\\mathcal\{L\}\}, whereα=1,…,M\\alpha=1,\\ldots,Mindexes the training samples andℋℒ≡ℋℒ0\\mathcal\{H\}\_\{\\mathcal\{L\}\}\\equiv\\mathcal\{H\}\_\{\\mathcal\{L\}\_\{0\}\}denotes the full input Hilbert space before any coarse\-graining,

ℋℒ=⨂ℓ∈ℒℋℓ,\\mathcal\{H\}\_\{\\mathcal\{L\}\}=\\bigotimes\_\{\\ell\\in\\mathcal\{L\}\}\\mathcal\{H\}\_\{\\ell\},\(1\)andℋℓ≅ℂ2\\mathcal\{H\}\_\{\\ell\}\\cong\\mathbb\{C\}^\{2\}is the local qubit Hilbert space at siteℓ\\ell\.

The state of an individual siteℓ\\ellin any given step of the process, can be represented by its reduced density matrixρ\[ℓ\]=trℓ¯⁡\(\|Ψ⟩​⟨Ψ\|\)\\rho^\{\[\\ell\]\}=\\operatorname\{tr\}\_\{\\bar\{\\ell\}\}\(\|\\Psi\\rangle\\langle\\Psi\|\)and obtained by tracing out all other sitesℓ¯\\bar\{\\ell\}exceptℓ\\ell\. The target output for the QCNN is a realization of some state\|ψα⟩\|\\psi\_\{\\alpha\}\\rangle, i\.e\., there is always a QCNN that recognizes an input state\|ψα⟩\|\\psi\_\{\\alpha\}\\ranglewith some deterministic measurement outcome\.

The QCNN training procedure is supervised approximation process defined over a training set𝒯train\\mathcal\{T\}\_\{\\text\{train\}\}and a testing set𝒯test\\mathcal\{T\}\_\{\\text\{test\}\}, where a training samplex→∈𝒯train\\vec\{x\}\\in\\mathcal\{T\}\_\{\\text\{train\}\}is provided to the algorithm together with its true labelm​\(x→\)∈\{\+1,−1\}m\(\\vec\{x\}\)\\in\\\{\+1,\-1\\\}\. Here,m:𝒯train∪𝒯test→\{\+1,−1\}m:\\mathcal\{T\}\_\{\\text\{train\}\}\\cup\\mathcal\{T\}\_\{\\text\{test\}\}\\rightarrow\\left\\\{\+1,\-1\\right\\\}is the underlying true map\. The true labels of the test set𝒯test\\mathcal\{T\}\_\{\\text\{test\}\}are not given to the algorithm, however, is given labeled training data\{\(x→,m​\(x→\)\)\}x→∈𝒯train\\left\\\{\(\\vec\{x\},\\,m\(\\vec\{x\}\)\)\\right\\\}\_\{\\vec\{x\}\\in\\mathcal\{T\}\_\{\\text\{train\}\}\}\. The QCNN learns an approximate mapm~:𝒯train→\{\+1,−1\}\\tilde\{m\}:\\mathcal\{T\}\_\{\\text\{train\}\}\\rightarrow\\\{\+1,\-1\\\}by minimizing a loss function assigns a penalty to deviations ofm~​\(x→\)\\tilde\{m\}\(\\vec\{x\}\)fromm​\(x→\)m\(\\vec\{x\}\)across the training set\. Following training, the learned functionm~\\tilde\{m\}is evaluated on unseen samplesx→∈𝒯test\\vec\{x\}\\in\\mathcal\{T\}\_\{\\text\{test\}\}, where the goal is to minimize the generalization errorℰ=Prx→∈𝒯test⁡\[m~​\(x→\)≠m​\(x→\)\]\\mathcal\{E\}=\\Pr\_\{\\vec\{x\}\\in\\mathcal\{T\}\_\{\\text\{test\}\}\}\\left\[\\tilde\{m\}\(\\vec\{x\}\)\\neq m\(\\vec\{x\}\)\\right\]\. The training procedure will be discussed more in Section[III\-A](https://arxiv.org/html/2607.05724#S3.SS1)

### II\-BRecognizing a 1D SPT Phase

A QCNN takes advantage of the layered structure of the MERA to classify phase, where the input layer takes in an unknownNN\-input quantum state\. After the input layer, the hidden layers are the convolutional layer, pooling layer, and fully connected layer\[[7](https://arxiv.org/html/2607.05724#bib.bib7)\], similarly to the CNN\. The convolutional layer is given as a quasi\-local unitaryUℒU\_\{\\mathcal\{L\}\}and is applied in a translationally invariant manner for finite depth\. During pooling, a portion of qubits is measured, and the outcome determines the unitary rotationVℒV\_\{\\mathcal\{L\}\}applied to nearby qubits, allowing for nonlinearity through the reduction of degrees of freedom\. Unitary layers, such as unitaries in the convolutional and fully layers, apply quasi\-local operations to existing qubits, while isometry layers, like those in the pooling layer, remove qubits in a state before applying unitary transformations on nearby states\. The number of convolution and pooling layers remains fixed throughout, while the model learns the unitaries\. The convolution and pooling layers are appliedddtimes\. When the system is sufficiently small, a fully connected layer is applied on the remaining qubits as a unitaryFF\. The outcome of the circuit is measured and the prediction is obtained\.

The final operator measured by the circuit recognizes the SPT phase\. The set of phases considered consists of ground states\{\|ψGS⟩\}\\\{\|\\psi\_\{\\mathrm\{GS\}\}\\rangle\\\}corresponding to a family of Hamiltonians defined on anNN\-site spin\-12\\frac\{1\}\{2\}chain with open boundary conditions

H=−J​∑i=1N−2Zi​Xi\+1​Zi\+2−h1​∑i=1NXi−h2​∑i=1N−1Xi​Xi\+1,H=\-J\\sum\_\{i=1\}^\{N\-2\}Z\_\{i\}X\_\{i\+1\}Z\_\{i\+2\}\-h\_\{1\}\\sum\_\{i=1\}^\{N\}X\_\{i\}\-h\_\{2\}\\sum\_\{i=1\}^\{N\-1\}X\_\{i\}X\_\{i\+1\},\(2\)whereXiX\_\{i\}andZiZ\_\{i\}are Pauli operators acting on siteii,J\>0J\>0is the three\-body interaction strength,h1h\_\{1\}is the transverse field strength, andh2h\_\{2\}is the nearest\-neighbor coupling strength\. Theℤ2×ℤ2\\mathbb\{Z\}\_\{2\}\\times\\mathbb\{Z\}\_\{2\}symmetry protection of the Haldane chain is generated by globalπ\\pi\-rotations of every spin around theXXandYYaxes,

Rx=∏jei​π​σjxandRy=∏jei​π​σjy,R\_\{x\}=\\prod\_\{j\}e^\{i\\pi\\sigma\_\{j\}^\{x\}\}\\quad\\text\{and\}\\quad R\_\{y\}=\\prod\_\{j\}e^\{i\\pi\\sigma\_\{j\}^\{y\}\},\(3\)whereσjx\\sigma\_\{j\}^\{x\}andσjy\\sigma\_\{j\}^\{y\}are the spin\-12\\frac\{1\}\{2\}operators at sitejj\.

The one\-parameter family of Hamiltonians considered for the Haldane phase, which is a SPT phase, is defined on a11\-dimensional chain ofNNspin\-11particles with open boundary conditions

HHaldane=JH​∑j=1N𝐒j⋅𝐒j\+1\+ω​∑j=1N\(Sjz\)2,H\_\{\\text\{Haldane\}\}=J\_\{H\}\\sum\_\{j=1\}^\{N\}\\mathbf\{S\}\_\{j\}\\cdot\\mathbf\{S\}\_\{j\+1\}\+\\omega\\sum\_\{j=1\}^\{N\}\\left\(S\_\{j\}^\{z\}\\right\)^\{2\},\(4\)where𝐒j\\mathbf\{S\}\_\{j\}denotes the vector of spin\-11operators at lattice sitejj,JH\>0J\_\{H\}\>0represents the exchange coupling strength, andω∈ℝ\\omega\\in\\mathbb\{R\}is the single\-ion anisotropy parameter\.

The detection of the phase is as a measurement of the nonzero expected value⟨Zm⟩\\langle Z\_\{m\}\\ranglemiddle qubitmmof the chain in the Z basis\. Quantitatively, given a phase\|ψα⟩∈𝒫\|\\psi\_\{\\alpha\}\\rangle\\in\\mathcal\{P\}, where𝒫\\mathcal\{P\}denotes the set of ground states belonging to a given phase, the expected value is computed directly from the state vector as

⟨Zm⟩=∑ℓλm​\|⟨ℓ\|ψα⟩\|2,\\langle Z\_\{m\}\\rangle=\\sum\_\{\\ell\}\\lambda\_\{m\}\\,\|\\langle\\ell\|\\psi\_\{\\alpha\}\\rangle\|^\{2\},\(5\)wherem=⌊N\+12⌋m=\\left\\lfloor\\frac\{N\+1\}\{2\}\\right\\rfloor, and the sum runs over all computational basis states\|ℓ⟩\|\\ell\\rangle, andλm∈\{\+1,−1\}N\\lambda\_\{m\}\\in\\\{\+1,\-1\\\}^\{N\}is the eigenvalue of the Pauli\-ZZoperator on the middle qubit in the basis state\|ℓ⟩\|\\ell\\rangle\.

### II\-CStabilizer Codes

QEC is crucial for protecting quantum information from background noise and gate errors that occur during processing\. To achieve this, error correction protocols redundantly encode logical qubits into many physical qubits, some of which are ancilla qubits, to ensure errors can be detected and corrected\. The ancilla qubits in a stabilizer code are repeatedly measured by parity check operators so that the wavefunction of the state of the logical qubits is not collapsed\. The syndrome information is then used to detect and possibly correct errors that are believed to have occurred during processing\.

The number of errors that can be detected and corrected is dependent on the code\. An\[\[n,k,d\]\]\[\[n,k,d\]\]QEC code encodeskklogical qubits into a QEC codeword of lengthnnto protect the system during quantum processing, where the code distancedddetermines the minimum number of physical qubit errors needed to cause a logical error\. A code of distanceddcan correct up to⌊d−12⌋\\left\\lfloor\\frac\{d\-1\}\{2\}\\right\\rfloorerrors and detect up tod−1d\-1errors\.

The information protected by the stabilizer code is defined over a stabilizer group𝒮\\mathcal\{S\}, which is the Abelian subgroup of the Pauli groupP⊗nP^\{\\otimes n\}\. This group does not include−I⊗n\-I^\{\\otimes n\}whereIIis the2×22\\times 2identity matrix inPP\. The stabilizer group𝒮\\mathcal\{S\}hasn−kn\-kindependent generators𝒮i\\mathcal\{S\}\_\{i\},i=1,…,n−ki=1,\.\.\.,n\-k\. Each element in the stabilizer group𝒮=⟨𝒮1,𝒮2,…,𝒮n−k⟩\\mathcal\{S\}=\\left\\langle\\mathcal\{S\}\_\{1\},\\mathcal\{S\}\_\{2\},\\ldots,\\mathcal\{S\}\_\{n\-k\}\\right\\rangleis termed a stabilizer\. The stabilizer generators perform distinct operations on the data qubits, however all stabilizer generators leave the encoded quantum state unchanged\.

### II\-DBivariate Bicycle Codes

The BB code is a type of Calderbank\-Shor\-Steane \(CSS\) stabilizer code that leverages bivariate polynomials over the quotient ringR=𝔽2​\[x,y\]/\(xℓ−1,ym−1\)R=\\mathbb\{F\}\_\{2\}\[x,y\]/\(x^\{\\ell\}\-1,y^\{m\}\-1\), whereℓ\\ellandmmare positive integers, and𝔽2=\{0,1\}\\mathbb\{F\}\_\{2\}=\\\{0,1\\\}is the binary field\. An\[\[n,k,d\]\]\[\[n,k,d\]\]BB codeQC⁡\(A,B\)\\operatorname\{QC\}\(A,B\)is a CSS code defined by two polynomialsA,B∈RA,B\\in R, whereQC⁡\(A,B\)\\operatorname\{QC\}\(A,B\)denotes the associated quasi\-cyclic code generated by the pair\(A,B\)\(A,B\), and has parameters

n\\displaystyle n=\\displaystyle=2​ℓ​m,\\displaystyle 2\\ell m,\(6\)k\\displaystyle k=\\displaystyle=2⋅dim\(ker⁡\(𝐇X\)∩ker⁡\(𝐇Z\)\),\\displaystyle 2\\cdot\\dim\\left\(\\ker\(\\mathbf\{H\}^\{X\}\)\\cap\\ker\(\\mathbf\{H\}^\{Z\}\)\\right\),\(7\)d\\displaystyle d=\\displaystyle=min⁡\{\|v\|:v∈ker⁡\(𝐇X\)∖rs⁡\(𝐇Z\)\},\\displaystyle\\min\\left\\\{\|v\|:v\\in\\ker\(\\mathbf\{H\}^\{X\}\)\\setminus\\operatorname\{rs\}\(\\mathbf\{H\}^\{Z\}\)\\right\\\},\(8\)where𝐇X\\mathbf\{H\}^\{X\}and𝐇Z\\mathbf\{H\}^\{Z\}are the parity\-check matrices of the CSS code corresponding to theXX\-type andZZ\-type stabilizer checks, respectively, and𝐇\\mathbf\{H\}denotes a binary matrix\.ker⁡\(𝐇\)\\operatorname\{ker\}\(\\mathbf\{H\}\)denotes the set of vectors orthogonal to each row of𝐇\\mathbf\{H\},rs⁡\(𝐇\)\\operatorname\{rs\}\(\\mathbf\{H\}\)is the linear span of its rows, and\|v\|=∑i=1nvi\|v\|=\\sum\_\{i=1\}^\{n\}v\_\{i\}is the Hamming weight ofv∈𝔽2nv\\in\\mathbb\{F\}\_\{2\}^\{n\}\.

#### II\-D1Encoding Logical Qubits

The check matrices for a BB code denotedQC⁡\(A,B\)\\operatorname\{QC\}\(A,B\)with lengthn=2​ℓ​mn=2\\ell mis

𝐇X=\[A∣B\]and𝐇Z=\[BT∣AT\],\\mathbf\{H\}^\{X\}=\[A\\mid B\]\\quad\\text\{and\}\\quad\\mathbf\{H\}^\{Z\}=\[B^\{T\}\\mid A^\{T\}\],\(9\)where the vertical bar indicates stacking matrices horizontally, andTTdenotes matrix transpose\. The bivariate polynomials are a pair of matrices

A=A1\+A2\+A3andB=B1\+B2\+B3,A=A\_\{1\}\+A\_\{2\}\+A\_\{3\}\\quad\\text\{ and \}\\quad B=B\_\{1\}\+B\_\{2\}\+B\_\{3\},\(10\)such that each matrixAiA\_\{i\}andBjB\_\{j\}is a power ofxxoryy, whiledim⁡𝐇X=dim⁡𝐇Z=\(n/2\)×2\\operatorname\{dim\}\\mathbf\{H\}^\{X\}=\\operatorname\{dim\}\\mathbf\{H\}^\{Z\}=\(n/2\)\\times 2\. Each rowv∈𝔽2nv\\in\\mathbb\{F\}\_\{2\}^\{n\}in𝐇X\\mathbf\{H\}^\{X\}is aXX\-type operator, and similarly each rowv∈𝔽2nv\\in\\mathbb\{F\}\_\{2\}^\{n\}in𝐇Z\\mathbf\{H\}^\{Z\}is aZZ\-type operator, is defined as

X​\(v\)=∏j=1nXjvjandZ​\(v\)=∏j=1nZjvj,X\(v\)=\\prod\_\{j=1\}^\{n\}X\_\{j\}^\{v\_\{j\}\}\\quad\\text\{ and \}\\quad Z\(v\)=\\prod\_\{j=1\}^\{n\}Z\_\{j\}^\{v\_\{j\}\},\(11\)respectively\. The codes in Table 1 are described using the linear subspace associated with check matrices and polynomialsAAandBBfor high\-rate, shown in\[[5](https://arxiv.org/html/2607.05724#bib.bib18),[32](https://arxiv.org/html/2607.05724#bib.bib14)\]\.

#### II\-D2Syndrome Extraction

The code standardQC⁡\(A,B\)\\operatorname\{QC\}\(A,B\)has a syndrome measurement \(SM\) circuit that continuously measures the syndrome of each check operator, which requiresnndata qubits andnnancillary check qubits to record the measured syndromes, for a total of2​n2nphysical qubits\. The SM circuit only appliesCNOTs to pairs of qubits that are connected in the Tanner graph\.

The SM circuit starts and ends with a code dependent initialization and measurement cycle that determines the logical qubit initialization suited for the initial state and measuring logical qubits in a proper basis\. The rest of the SM circuit is comprised ofNcN\_\{c\}number of syndrome cycles \(SC\) repeated throughout the circuit periodically, where each SC measures syndromes for allnncheck operators of the code\.

#### II\-D3Decoding Algorithm

Consider a\[\[n,k,d\]\]\[\[n,k,d\]\]BB code and an SM circuit𝒰\\mathcal\{U\}constructed withNcN\_\{c\}syndrome cycles defined in the SM section\. Error correction is performed using a classical algorithm that takes a measured error syndrome as input and returns an estimate of the Pauli error that occurred during computation on the data qubits, accounting for all faults provided by the SM circuit\. Note that the syndrome circuit itself may contain measurement errors\. A proper guess occurs if the estimated Pauli error is a subset of the actual Pauli error up to a product computed via the check operators\.

The SM step is followed by belief propagation with an ordered statistics postprocessing step decoder \(BP\-OSD\)\[[5](https://arxiv.org/html/2607.05724#bib.bib18)\]\. The unknown error for the linearized noise model used in the BB code isξ\\xi,DDwhich denotes the decoding matrix, andσ=\[sU\|sF\]\\sigma=\\left\[s^\{U\}\|s^\{F\}\\right\]\(where the vertical bar indicates stacking matrices horizontally\) is the measured error syndrome\. Using information sets ranked according to their reliability, BP\-OSD finds an information setIIwith the largest reliability\. The output of BP\-OSD is the solution of the systemD​ξ=σD\\xi=\\sigmabased on the most reliable information setII\. This information set is used as the solution for the minimum weight errorξ∗=ξ∗​\(s\)∈\{0,1\}N\\xi^\{\*\}=\\xi^\{\*\}\(s\)\\in\\\{0,1\\\}^\{N\}optimization problem\. The guess for the unknown logical syndrome is given as

sL=DL​ξ∗,s^\{L\}=D^\{L\}\\xi^\{\*\},\(12\)whereLLindexes the logical qubit degree\-of\-freedom\. The BP\-OSD decoder is applied separately to the decoding matricesDxD\_\{x\}andDzD\_\{z\}, which are constructed from the parity\-check matrices𝐇X\\mathbf\{H\}^\{X\}and𝐇Z\\mathbf\{H\}^\{Z\}, respectively\. The result is a guessedXX\-type andZZ\-type errorsExE\_\{x\}andEzE\_\{z\}, where the guessed final error isE∗=Ex∗​Ez∗E^\{\*\}=E^\{\*\}\_\{x\}E^\{\*\}\_\{z\}\. This BP\-OSD computes the upper bound for the code distancedd\. A large number of candidate BB codes withn=𝒪​\(100\)n=\\mathcal\{O\}\(100\)qubits is then searched that satisfy the criteria for the BB code given in the original paper\[[5](https://arxiv.org/html/2607.05724#bib.bib18)\]\.

#### II\-D4Error Correction

After executing a quantum circuit the syndrome measurement outcomes are processed using BP\-OSD\. Using the final guessed X\-typeExE\_\{x\}and Z\-type errorEzE\_\{z\}are used to correct the data qubits without measurement occurring\.

## IIIProposed Bivariate\-Bicycle Coded QCNN

SC1\\mathrm\{SC\_\{1\}\}XchecksX\_\{\\text\{checks\}\}𝒞¯QCNN\\overline\{\\mathcal\{C\}\}\_\{\\text\{QCNN\}\}ZchecksZ\_\{\\text\{checks\}\}𝒩\\mathcal\{N\}XϕX\_\{\\phi\}𝒞¯QCNN−1\\overline\{\\mathcal\{C\}\}\_\{\\text\{QCNN\}\}^\{\-1\}ZϕZ\_\{\\phi\}SC2\\mathrm\{SC\_\{2\}\}\|0⟩\|0\\rangle\|0⟩\|0\\rangle\|0⟩\|0\\rangle\|ψl⟩\|\\psi\_\{l\}\\rangle\|0⟩\|0\\rangleflf\_\{l\}EncodingDecoding

Figure 2:QEC\-QCNN circuit\.In this section, we describe the parameters of the44\-distance BB code and the44\-qubit QCNN, followed by a description of how the QCNN is embedded within the BB code using transversal operations, ensuring that all QCNN operations remain within the protected code space\. Finally, the Tanner graph topology of the trained network is used to correct errors in the QCNN\.

### III\-ATraining the QCNN

We used the QCNN for quantum phase recognition \(QPR\) by applying it to a class of one\-dimensional many\-body systems\. Specifically, the class considered is aℤ2×ℤ2\\mathbb\{Z\}\_\{2\}\\times\\mathbb\{Z\}\_\{2\}phase\. The ground states were numerically obtained using an infinite\-size DMRG algorithm, following the method outlined in\[[25](https://arxiv.org/html/2607.05724#bib.bib9)\]with a maximum bond dimension of150150\. The symmetry operators generating theℤ2×ℤ2\\mathbb\{Z\}\_\{2\}\\times\\mathbb\{Z\}\_\{2\}symmetry are

Xeven=∏i∈evenXiandXodd=∏i∈oddXi,X\_\{\\text\{even\}\}=\\prod\_\{i\\in\\text\{even\}\}X\_\{i\}\\quad\\text\{and\}\\quad X\_\{\\text\{odd\}\}=\\prod\_\{i\\in\\text\{odd\}\}X\_\{i\},\(13\)whereXi\{X\_\{i\}\}is the Pauli\-XXoperator acting on siteii\. Each ground state energy density is obtained as a function ofh2h\_\{2\}for fixedh1h\_\{1\}\. We then compute its second\-order derivative to locate phase transitions\. This approach to locating phase transitions via the second\-order derivative of ground state energy densities follows\[[7](https://arxiv.org/html/2607.05724#bib.bib7)\]\.

As illustrated in Fig\.[3](https://arxiv.org/html/2607.05724#S4.F3), the QCNN circuit classifies whether the SPT phase exists as a\|S=1\|\|S=1\|Haldane chain or has transitioned to a\|S=1\|\|S=1\|paramagnetic phase or antiferromagnetic phase\. Whenω\\omegais zero or sufficiently small relative toJJ, the ground state, the phase belongs to the SPT phase\. The phase is antiferromagnetic when the staggered fieldh1h\_\{1\}is sufficiently large relative toJJ, breaking theℤ2×ℤ2\\mathbb\{Z\}\_\{2\}\\times\\mathbb\{Z\}\_\{2\}symmetry of the SPT phase\. The critical point is identified ash2/J=0\.423h\_\{2\}/J=0\.423using infinite size DMRG numerical simulations\.

During training, the untrained model is given a classified training set𝒯train=\{\(\|ψα⟩,yα\):α=1,…,M\}\\mathcal\{T\}\_\{\\text\{train\}\}=\\\{\(\|\\psi\_\{\\alpha\}\\rangle,y\_\{\\alpha\}\):\\alpha=1,\.\.\.,M\\\}, whereyαy\_\{\\alpha\}is the respective0or11binary classification for some input state\|ψα⟩\|\\psi\_\{\\alpha\}\\rangle\. The expected output of the QCNN is computed asf\{Ui,Vi,F\}​\(\|ψα⟩\)f\_\{\\left\\\{U\_\{i\},V\_\{i\},F\\right\\\}\}\\left\(\\left\|\\psi\_\{\\alpha\}\\right\\rangle\\right\)\(The expected output computation will be discussed more in Section[III\-C](https://arxiv.org/html/2607.05724#S3.SS3)\)\. The cost function of the model is computed using the mean\-squared error \(MSE\)

MSE=12​M​∑α=1M\(yi−f\{Ui,Vj,F\}​\(\|ψα⟩\)\)2\.\\texttt\{MSE\}=\\frac\{1\}\{2M\}\\sum\_\{\\alpha=1\}^\{M\}\\left\(y\_\{i\}\-f\_\{\\left\\\{U\_\{i\},V\_\{j\},F\\right\\\}\}\\left\(\\left\|\\psi\_\{\\alpha\}\\right\\rangle\\right\)\\right\)^\{2\}\.\(14\)The QCNN is constructed using isometric tensors in the pooling layer, where measuring a qubit lowers degrees of freedom for feature extraction\[[7](https://arxiv.org/html/2607.05724#bib.bib7)\]\. Since these measurement\-based pooling operations are non\-unitary and irreversible, standard quantum backpropagation cannot be applied through the full QCNN\. An analytical gradient parameter\-shift rule is used as an approximation scheme to compute the gradients, which provides unbiased gradient estimates for the variational unitary layers preceding the measurements\.

The model learns by iteratively optimizing all initially assigned unitaries until convergence via a finite\-difference scheme given the MSE cost function\. To calculate the gradient matrixMMusing the finite\-difference scheme, the parameterized weightsΘ\\Thetain the QCNN circuit are perturbed by shiftingθj∈Θ\\theta\_\{j\}\\in\\Thetaby±ϵ\\pm\\epsilonfor each parameter in the quantum circuit\. The weights are then updated via gradient descent as

Θ←Θ−η​M,\\Theta\\leftarrow\\Theta\-\\eta\\,M,\(15\)whereη\\etais the learning rate andMMis the gradient matrix of the MSE with respect toΘ\\Theta\.

We evaluate phase recognition performance using 40 spin\-12\\frac\{1\}\{2\}ground states sampled along theh2h\_\{2\}axis \(h1=0h\_\{1\}=0\), with QCNN weights trained from0\.6160\.616to0\.1740\.174after 100\-iterations under ideal, noiseless statevector simulation\. We will discuss the configuration of the simulation in Section[IV](https://arxiv.org/html/2607.05724#S4)\. The trained model produces expectation values⟨Zm⟩\\langle Z\_\{m\}\\ranglethat decrease monotonically from approximately0\.940\.94nearh2=0h\_\{2\}=0toward0\.280\.28at the far end of the paramagnetic region, yielding a smooth decision boundary consistent with the theoretical phase transition ath2/J=0\.423h\_\{2\}/J=0\.423\.

### III\-BOptimizing Quantum Error Correction

The QEC code is a BB code with parameters\[\[18,4,4\]\]\[\[18,4,4\]\], wherel=3l=3andm=3m=3, using\[a1,a2,a3\]=\[1,0,1\]\[a\_\{1\},a\_\{2\},a\_\{3\}\]=\[1,0,1\]and\[b1,b2,b3\]=\[2,0,2\]\[b\_\{1\},b\_\{2\},b\_\{3\}\]=\[2,0,2\]\. We derive the corresponding check matrices𝐇X=\[A\|B\]\\mathbf\{H\}^\{X\}=\[A\|B\]and𝐇Z=\[BT\|AT\]\\mathbf\{H\}^\{Z\}=\[B^\{T\}\|A^\{T\}\]and compute the code parametersk=4k=4andd=4d=4\. The BB code circuit is then constructed in accordance with the architecture described in Section[II\-D](https://arxiv.org/html/2607.05724#S2.SS4)\.

The feed\-forward correction layer is trained and evaluated on stochastic Pauli noise injected into the\[\[18,4,4\]\]\[\[18,4,4\]\]BB code\. Specifically, the four logical data qubits admit24=162^\{4\}=16distinct computational basis states, such that the training set is

ℱtrain=\{y∈\{0,1\}4\},\\mathcal\{\\mathcal\{F\}\_\{\\text\{train\}\}\}\\;=\\;\\bigl\\\{\\,y\\in\\\{0,1\\\}^\{4\}\\bigr\\\},\(16\)where each elementy=\(y1,y2,y3,y4\)y=\(y\_\{1\},y\_\{2\},y\_\{3\},y\_\{4\}\)encodes which of the four logical qubits has suffered a bit\-flip error\. At each SPSA iteration, every patterny∈ℱtrainy\\in\\mathcal\{\\mathcal\{F\}\_\{\\text\{train\}\}\}is prepared by applyingXXgates to the mapped physical qubits\. Concretely, the encoder usesζ\\zetalogical to map each basis state\|y⟩\\lvert y\\rangle, wherey∈\{0,1\}4y\\in\\\{0,1\\\}^\{4\}, to a codeword\|y¯⟩∈\(ℂ2\)⊗11\|\\overline\{y\}\\rangle\\in\(\\mathbb\{C\}^\{2\}\)^\{\\otimes 11\}via transversal gate operations\. Given a physical quantum circuit𝒞\\mathcal\{C\}acting on44logical qubits, we construct a physically encoded circuit𝒞enc\\mathcal\{C\}\_\{\\text\{enc\}\}via the transversal gate operations\. One classical register records the1111bits data measurement outcomes, respectively\. using propagated through the full syndrome\-extraction and feed\-forward correction circuit, and decoded to yield a predicted logical labely^\\hat\{y\}\.

Given the logical state space is finite and small the loss in Eq\. \([17](https://arxiv.org/html/2607.05724#S3.E17)\) is computed over allN=16N=16patterns at every iteration, making each update equivalent to a full\-batch gradient step over the complete input distribution\. This exhaustive evaluation guarantees that the trained parametersϕ\\boldsymbol\{\\phi\}are assessed on every failure mode of the code, and that a loss of0certifies perfect logical\-state recovery across all possible single\-layer error configurations\.

To correct errors, we augment the syndrome extraction circuit with a parametrized feed\-forward layer\. This layer introduces2​m2mtrainable rotation anglesϕi∈\[0,2​π\)\\phi\_\{i\}\\in\[0,2\\pi\), whereRy​\(ϕiX\)R\_\{y\}\(\\phi\_\{i\}^\{X\}\)andRy​\(ϕiZ\)R\_\{y\}\(\\phi\_\{i\}^\{Z\}\)denote single\-qubityy\-axis rotation gates parameterized by theiith trainable angle corresponding to theXX\-check andZZ\-check ancilla qubits, respectively\. These rotations are applied to the physical data qubits immediately following the syndrome vector, using controlled X and Z rotations\.

The objective used during training are the logical state or logical labels\. The set of true logical labels\{yi\}i=1N\\\{y\_\{i\}\\\}\_\{i=1\}^\{N\}and the corresponding predicted logical labels\{y^i\}i=1N\\\{\\hat\{y\}\_\{i\}\\\}\_\{i=1\}^\{N\}given inℱtrain\\mathcal\{F\}\_\{\\text\{train\}\}returned by the decoder after circuit execution, the loss is given as

Loss​\(ϕ\)=1N​∑i=1N𝟏​\[y^i≠yi\],\\texttt\{Loss\}\(\\boldsymbol\{\\phi\}\)\\;=\\;\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\\\!\\left\[\\hat\{y\}\_\{i\}\\neq y\_\{i\}\\right\],\(17\)
whereϕ\\boldsymbol\{\\phi\}denotes the trainable parameters of the QEC feed\-forward layer, distinct from the QCNN circuit parameters𝜽\\boldsymbol\{\\theta\}, and𝟏​\[⋅\]\\mathbf\{1\}\[\\cdot\]denotes the indicator function\. The loss therefore lies in\[0,1\]\[0,1\], with0indicating perfect correction and11indicating every frame is mislabeled\. This normalization gives the SPSA gradient estimator a stationary, bounded signal that is independent of the number of test cases and the code distance\.

The correction layer is embedded inside a quantum circuit evaluated by a noisy stochastic simulator\. Since the loss function is a non\-differentiable indicator function gradients ofLosswith respect to the trainable parametersϕ\\boldsymbol\{\\phi\}cannot be computed analytically\[[30](https://arxiv.org/html/2607.05724#bib.bib41)\]\. SPSA is therefore employed as the optimizer, such that at each iterationkk, the algorithm evaluates the a single syndrome cycle circuit33times, regardless of the number of parameterspp, making it practical for circuits with many syndrome checks\.

Each training iteration proceeds in four stages\.

1. 1\.Center evaluation\.The unperturbed circuit is evaluated to obtain the center\-point lossLoss0=Loss​\(ϕ\(k\)\)\\texttt\{Loss\}\_\{0\}=\\texttt\{Loss\}\(\\boldsymbol\{\\phi\}^\{\(k\)\}\), where,ϕ\(k\)\\boldsymbol\{\\phi\}^\{\(k\)\}is weight vector at thekkth iteration and is recorded for the learning\-rate adaptation rule described below\.
2. 2\.Positive perturbation\.A simultaneous perturbation vector𝚫∈\{−1,\+1\}p\\boldsymbol\{\\Delta\}\\in\\\{\-1,\+1\\\}^\{p\}is drawn independently and uniformly at random\. The parameters are shifted toϕ\(k\)\+c​𝚫\\boldsymbol\{\\phi\}^\{\(k\)\}\+c\\,\\boldsymbol\{\\Delta\}where,c=0\.2c=0\.2rad is the perturbation magnitude, and the circuit is evaluated to obtainLoss\+\\texttt\{Loss\}\_\{\+\}\.
3. 3\.Negative perturbation\.The parameters are shifted toϕk−c​𝚫\\boldsymbol\{\\phi\}^\{k\}\-c\\,\\boldsymbol\{\\Delta\}\. Using the same realization of𝚫\\boldsymbol\{\\Delta\}, and the circuit is evaluated to obtainLoss−\\texttt\{Loss\}\_\{\-\}\.
4. 4\.Gradient step\.The SPSA gradient estimate for parameteriiis g^i=Loss\+−Loss−2​c​Δi,\\hat\{g\}\_\{i\}\\;=\\;\\frac\{\\texttt\{Loss\}\_\{\+\}\-\\texttt\{Loss\}\_\{\-\}\}\{2\\,c\\,\\Delta\_\{i\}\},\(18\) and a gradient descent step is applied ϕik\+1=\(ϕi​k−αk​g^i\)mod2​π,\\phi\_\{i\}^\{k\+1\}\\;=\\;\\left\(\\phi\_\{i\}\{k\}\-\\alpha\_\{k\}\\,\\hat\{g\}\_\{i\}\\right\)\\\!\\\!\\mod 2\\pi,\(19\)whereαk\\alpha\_\{k\}is the current learning rate\. All parameters are updated simultaneously in a single step, with the base parameters restored from a snapshot taken before the perturbations were applied\. Angles are wrapped modulo2​π2\\pithroughout\.

The learning rateαk\\alpha\_\{k\}is adapted at every step by a bold\-driver rule\[[2](https://arxiv.org/html/2607.05724#bib.bib39)\]\. After the gradient step, the center\-point lossLoss0\\texttt\{Loss\}\_\{0\}of the current iteration is compared against that of the previous iteration\. If the loss has decreased, then the learning rate is increased by1\.051\.05, encouraging larger steps when progress is being made\. If the loss has increased or remained constant, the learning rate is halved, contracting the search radius to

αk\+1=\{1\.05​αkifLoss0\(k\)<Loss0\(k−1\),0\.5​αkotherwise\.\\alpha\_\{k\+1\}\\;=\\;\\begin\{cases\}1\.05\\,\\alpha\_\{k\}&\\text\{if \}\\texttt\{Loss\}\_\{0\}^\{\(k\)\}<\\texttt\{Loss\}\_\{0\}^\{\(k\-1\)\},\\\\\[4\.0pt\] 0\.5\\,\\alpha\_\{k\}&\\text\{otherwise\.\}\\end\{cases\}\(20\)
All trainable parameters are initialized to zero, corresponding to identity rotations on every check qubit\. Under this initialization the correction layer has no effect on the circuit output, so the very first center\-point evaluation measures the uncorrected logical error rate\. Subsequent iterations introduce progressively stronger corrections as SPSA explores the parameter space\. A seeded random number generator is used for both the initialization procedure and the Bernoulli draws that produce the𝚫\\boldsymbol\{\\Delta\}vectors, ensuring that training runs are exactly reproducible\.

The decoding offline stage uses BB code decoding matrices to decodes the linearized noise model using BP\-OSD\[[5](https://arxiv.org/html/2607.05724#bib.bib18)\]by taking a measured error syndrome as input and outputting an estimated final error that occurred on the data qubits during computation\.

### III\-CError Correction for QCNN

The QCNN suffers from the error\-prone nature of quantum operations throughout the entirety of the circuit, so it naturally motivates the need for QEC\. Our proposed QCNN architecture constructs the layers using transversal gate operations that conform to the Tanner graph structure of the BB code, denoted𝒞¯QCNN\\overline\{\\mathcal\{C\}\}\_\{\\text\{QCNN\}\}\. Thus, allowing gates in each layer to preserve the code’s locality and operations to occur in the protected subspace\. The variational parameters are continuous and therefore implement non\-Clifford rotations, which cannot be implemented transversally in the BB data block, allowing for errors to propagate through the circuit\.

To mitigate this, while avoiding the high overhead of magic state injection \(MSI\), we employ a FFNN as a syndrome\-based neural decoder interleaved between each convolutional, pooling, and fully connected layer of the QCNN and the next\. Thus, rather than correcting the non\-Clifford rotations themselves, the FFNN intercepts the X\-syndromes and Z\-syndromes produced by the BB code’s stabilizer checks after each noisy layer, and predicts per\-qubit Pauli\-X and Pauli\-Z correction decisions before the next layer executes\. This approach confines error propagation to within individual layers rather than allowing it to compound across the full circuit depth\.

Concretely, this correction occurs between each hidden layer, prior to the next variational layer, so that the parameter updates computed by SPSA reflect gradient estimates over a cleaner logical state rather than one corrupted by accumulated gate noise, as would occur if correction were deferred to end\-of\-circuit\.

This process circumvents the classical post\-processing bottleneck of BP\-OSD, where belief propagation would need to complete on the order of 1μ​s\\mu s\[[11](https://arxiv.org/html/2607.05724#bib.bib31),[6](https://arxiv.org/html/2607.05724#bib.bib32)\]for superconducting processors to keep pace with circuit execution\. Current state\-of\-the\-art decoders trade accuracy for latency to meet this tight decoding time budget\[[31](https://arxiv.org/html/2607.05724#bib.bib33)\], and BP\-OSD does not meet this requirement at scale\. Otherwise, physical qubits must sit idle between layers, during whichT1T\_\{1\}\(energy relaxation\) andT2T\_\{2\}\(dephasing\) processes accumulate additional errors on the quantum hardware\.

The FFNN’s learnable weights are one per edge of the BB code’s Tanner graph, such that the structure of the check and variable nodes respect the code’s locality, mirroring the same geometric constraints imposed on the transversal gates\. Its input is the X\-syndrome bit string extracted from stabilizer CX measurements on the\[\[18,4,4\]\]\[\[18,4,4\]\]BB code’s ancilla qubits, and its output is a binary correction vector indicating which physical data qubits require a corrective CX operation to return the logical state to the code space\. Where, the\[\[18,4,4\]\]\[\[18,4,4\]\]BB code’s stabilizer generators and logical operators are constructed following the procedure presented in\[[5](https://arxiv.org/html/2607.05724#bib.bib18)\]\(see section[III\-B](https://arxiv.org/html/2607.05724#S3.SS2)\)\.

The44\-spin SPT input state is encoded into an1818\-physical\-qubit state representing the logical state\|ψl⟩\|\\psi\_\{l\}\\rangleof the\[\[18,4,4\]\]\[\[18,4,4\]\]code\. The encoding is performed by mapping the four physical input spins to logical qubit degrees of freedom according to the sparse parity\-check structure of the code\. For each logical circuit𝒞\\mathcal\{C\}, this procedure traverses every layer of the QCNN, including the input layer and final layer\. The transversal operations then replaces each logical instruction with up to\|𝒫c\|×\|𝒫t\|\|\\mathcal\{P\}\_\{c\}\|\\times\|\\mathcal\{P\}\_\{t\}\|physical gates\. For a code with block sizennand a circuit of depthddcontainingggtwo\-qubit gates, the encoded circuit has depth at most𝒪​\(d\)\\mathcal\{O\}\(d\)and gate count at most𝒪​\(g⋅n2\)\\mathcal\{O\}\(g\\cdot n^\{2\}\)\. Specifically, each gateG∈𝒞G\\in\\mathcal\{C\}acting on logical qubitqiq\_\{i\}is replaced by the transversal operation

G⊗\|𝒫i\|=⨂j∈𝒫iGj,G^\{\\otimes\|\\mathcal\{P\}\_\{i\}\|\}=\\bigotimes\_\{j\\in\\mathcal\{P\}\_\{i\}\}G\_\{j\},\(21\)where𝒫i⊆\{0,1,…,n−1\}\\mathcal\{P\}\_\{i\}\\subseteq\\\{0,1,\\ldots,n\-1\\\}is the set of physical qubit indices corresponding to theiith logical qubit under the code’s qubit mappingζ:qi↦𝒫i\\zeta:q\_\{i\}\\mapsto\\mathcal\{P\}\_\{i\}\.

A logical CNOT between control qubitqcq\_\{c\}and target qubitqtq\_\{t\}is implemented transversally as

CNOT⊗=⨂j∈𝒫ck∈𝒫t,j≠kCNOTj→k,\\mathrm\{CNOT\}^\{\\otimes\}=\\bigotimes\_\{\\begin\{subarray\}\{c\}j\\in\\mathcal\{P\}\_\{c\}\\\\ k\\in\\mathcal\{P\}\_\{t\},\\;j\\neq k\\end\{subarray\}\}\\mathrm\{CNOT\}\_\{j\\to k\},\(22\)applying a physical CNOT from each control physical qubit to each target physical qubit, excluding self\-connections\. Gates involving stabilizer check qubits or reset operations are passed through without transversal expansion, as these are already defined at the physical level in the encoded circuit\.

The𝒞¯QCNN\\overline\{\\mathcal\{C\}\}\_\{\\text\{QCNN\}\}then acts on the logical state, which undergoes noise𝒩\\mathcal\{N\}as\.

fl=∑\|ψl⟩⁣∈\{\|±x,y,z⟩\}⟨ψl\|ℳq−1\(𝒩\(ℳq\(\|ψl⟩⟨ψl\|\)\)\)\|ψl⟩,f\_\{l\}=\\sum\_\{\\lvert\\psi\_\{l\}\\rangle\\in\\\{\\lvert\{\\pm x,y,z\}\\rangle\\\}\}\\langle\\psi\_\{l\}\\rvert\\,\\mathcal\{M\}\_\{q\}^\{\-1\}\\\!\\left\(\\mathcal\{N\}\\\!\\left\(\\mathcal\{M\}\_\{q\}\\\!\\left\(\\lvert\\psi\_\{l\}\\rangle\\langle\\psi\_\{l\}\\rvert\\right\)\\right\)\\right\)\\lvert\\psi\_\{l\}\\rangle,\(23\)whereℳq\\mathcal\{M\}\_\{q\}\(ℳq−1\)\(\\mathcal\{M\}\_\{q\}^\{\-1\}\)denotes the encoding \(decoding\) map of the BB code,\|±x,y,z⟩\\lvert\{\\pm x,y,z\}\\rangleare the±1\\pm 1eigenstates of the PauliXX,YY,ZZoperators, andflf\_\{l\}is the logical output of

𝒞¯QCNN,\{Ui,Vj,F\}\(\|ψl⟩\)\.\\overline\{\\mathcal\{C\}\}\_\{\\mathrm\{QCNN\},\\,\\\{U\_\{i\},\\,V\_\{j\},\\,F\\\}\}\\\!\\left\(\\lvert\\psi\_\{l\}\\rangle\\right\)\.\(24\)
The initial physical system two 9\-qubit data registers𝒟L\\mathcal\{D\}\_\{L\}and𝒟R\\mathcal\{D\}\_\{R\}, one 9\-qubitXX\-check ancilla register𝒜X\\mathcal\{A\}\_\{X\}, and one 9\-qubitZZ\-check ancilla register𝒜Z\\mathcal\{A\}\_\{Z\}\. However, the transversal qubit mapping has two 9\-qubit data registers𝒟L\\mathcal\{D\}\_\{L\}and𝒟R\\mathcal\{D\}\_\{R\}which are redundant\. The four logical qubits map onto only 11 distinct physical qubit indices\{0,…,10\}\\\{0,\\ldots,10\\\}, with qubits\{0,2,5,6\}\\\{0,2,5,6\\\}shared across multiple logical qubits and the remaining indices each carrying information unique to a single logical qubit\. Consequently, the two data registers collapse into a single 11\-qubit register𝒟\\mathcal\{D\}, reducing the full system from 36 to 29 qubits without loss of information\.

The encoder usesζ\\zetato map each basis state\|i⟩\\lvert i\\rangle, wherei∈\{0,1\}4i\\in\\\{0,1\\\}^\{4\}, to a codeword\|i¯⟩∈\(ℂ2\)⊗11\\lvert\\overline\{i\}\\rangle\\in\(\\mathbb\{C\}^\{2\}\)^\{\\otimes 11\}via transversal gate operations\. The evolved circuit is a Hilbert space of dimension2ntotal2^\{n\_\{\\text\{total\}\}\}, wherentotal=nleft\+ddata\+nright=29n\_\{\\text\{total\}\}=n\_\{\\text\{left\}\}\+d\_\{\\text\{data\}\}\+n\_\{\\text\{right\}\}=29\. Given99X\-check ancillas,1111data qubits, and99Z\-check ancillas\. Given the ancilla qubits are initialized to\|0⟩\|0\\rangleand, in the noiseless statevector simulation, remain unentangled with the data register, the data\-qubit marginal state can be extracted without a partial trace\. The reduced statevector is then decoded into the logical basis by iterating over all2ddata2^\{d\_\{\\text\{data\}\}\}computational basis states of the data register\. The corresponding logical state is then determined by

𝐯=\[\(LX​𝐛\)mod2\],\\mathbf\{v\}\\;=\\;\\left\[\\,\\bigl\(L\_\{X\}\\,\\mathbf\{b\}\\bigr\)\\bmod 2\\,\\right\],\(25\)The complex amplitudeαj\\alpha\_\{j\}is then accumulated into the entry of the logical statevector\|ψl⟩∈ℂ2N\|\\psi\_\{l\}\\rangle\\in\\mathbb\{C\}^\{2^\{N\}\}indexed byv=∑kvk​2N−1−kv=\\sum\_\{k\}v\_\{k\}\\,2^\{N\-1\-k\}, so that

\|ψl⟩=∑j=02dphys−1αj​\|𝐯​\(𝐛j\)⟩\.\|\\psi\_\{l\}\\rangle\\;=\\;\\sum\_\{j=0\}^\{2^\{d\_\{\\text\{phys\}\}\}\-1\}\\alpha\_\{j\}\\,\\bigl\|\\mathbf\{v\}\(\\mathbf\{b\}\_\{j\}\)\\bigr\\rangle\.\(26\)The resulting2N2^\{N\}\-dimensional statevector is passed to the prediction stage, where the expectation value of the middle\-qubit observable is evaluated to produce the final classification output\.

## IVExperimental Results and Discussion

### IV\-ASimulation Setup

We let each gate type is subject to an independent error probabilitypp, with errors injected as single\- or two\-qubit Pauli operators drawn uniformly from\{I,X,Y,Z\}\\\{I,X,Y,Z\\\}\. The two\-qubit CNOT gates are assigned errors from the full 15\-element two\-qubit Pauli group\{I,X,Y,Z\}⊗2∖\{I​I\}\\\{I,X,Y,Z\\\}^\{\\otimes 2\}\\setminus\\\{II\\\}, with each error type drawn uniformly given that an error occurs with probabilitypp\. Idle qubit errors are inserted to maintain a consistent time boundary across all data qubits\. Idle qubits are modeled per layer, such that any qubit not involved in an active gate receives a uniformly random Pauli error\{I,X,Y,Z\}\\\{I,X,Y,Z\\\}with probabilitypp, independently per qubit\. Measurement operations are subject to a pre\-measurement bit\-flip, phase\-flip, orYY\-error at the same rate\. Further, we implement a stochastic Pauli noise model applied at the circuit level\[[5](https://arxiv.org/html/2607.05724#bib.bib18)\]\. The considered noise model configuration is used to sweep over physical error ratesp∈\{0\.0001,0\.001,0\.003\}p\\in\\\{0\.0001,0\.001,0\.003\\\}, which span the near\-threshold regime relevant to NISQ\-era demonstrations\. This noise model approximates a symmetric depolarizing channel on each gate type and serves as a baseline for evaluating the robustness of our BB\-code\-protected QCNN against gate\-level noise\. All noise simulations used the same seeded random number generator for reproducibility\.

### IV\-BTraining and Evaluation

The 4\-qubit QCNN was first trained under ideal statevector simulation using a dataset of 40 labeled ground states evenly spaced along the lineh2=0h\_\{2\}=0, where the Hamiltonian is exactly solvable via the Jordan\-Wigner transformation and labeled±1\\pm 1\[[29](https://arxiv.org/html/2607.05724#bib.bib11)\]\. Training ran for 100 iterations using the finite\-difference gradient scheme described in Section II\-B\.

Fig\.[3](https://arxiv.org/html/2607.05724#S4.F3)plots and compares statevector training training loss for44\-qubit QCNN and 11\-qubit transversal QCNN under error ratesp∈\{p\\in\\\{0\.0%\\%, 0\.01%\\%, 0\.1%\\%, 0\.3%\}\\%\\\}\.

![Refer to caption](https://arxiv.org/html/2607.05724v1/qcnn_noisy_sweep.png)Figure 3:Statevector training loss vs iteration for44\-qubit QCNN and 11\-qubit transversal QCNN\.In Fig\.[3](https://arxiv.org/html/2607.05724#S4.F3)a, under noiseless \(i\.e\.,p=0p=0\) statevector training, the standard QCNN converges smoothly and monotonically from0\.6160\.616at to0\.1740\.174and reached a stable plateau by iteration 60\. In comparison the transverse QCNN started at1\.4421\.442and exhibits substantial loss variance during the first 55 iterations, with a quick peak exceeding 2\.0 and recurring peaks near iterations 20 and 45\. After iteration 55, the transversal circuit stabilizes and plateaus at0\.6160\.616, with the lowest loss at iteration 19 of0\.190\.19and oscillating around losses of about0\.250\.25from iterations 20 to 60\. This variability and performance gap of the transversal QCNN under noiseless conditions suggest structural constraints due to the transversal gate set\. We believe that the application of identical unitary operation fault\-tolerantly across all physical reduces, the effective expressibility of the ansatz\. Updates made to the weights are magnified by the repetition of the gates across many physical qubits\. While this property is designed to minimize physical error rates on a logical qubit, it also causes large shifts for the finite\-difference scheme when traversing the gradient\. This can be seen by large steps taken during early\-training, before plateauing after the learning loss was lowered, due to fine tuning from the bold driver\.

In Fig\.[3](https://arxiv.org/html/2607.05724#S4.F3)b, forp=0\.001p=0\.001, the standard QCNN fails to learn and converge entirely\. The loss starts at0\.6320\.632and after 100 iterations is at0\.6520\.652\. The QCNN reaches its lowest loss of0\.420\.42at iteration 2, after which it never drops below a loss of0\.60\.6\. The transversal QCNN starts at0\.8020\.802oscillates early on, however, start learning and descending around iteration 45, ending at0\.3480\.348\. The lowest loss recorded for the transversal QCNN was0\.2240\.224at iteration 46\. The transversal QCNN is noise\-resilient by construction\. Given the transversal QCNN learned better than the standard QCNN without noise, we believe the error corrected QCNN will train better, however, with limited hardware, we leave simulating 29 qubits on classical hardware for future research\.

In Fig\.[3](https://arxiv.org/html/2607.05724#S4.F3)c, forp=0\.003p=0\.003, the highest noise level tested for training\. The standard QCNN begins at a loss of0\.7140\.714at iteration 1 and never improves in any sustained amount\. We assume this is because the bold\-driver schedule triggers repeated halving events from as early as iteration 3, collapsing the learning rate from10510^\{5\}to below10−810^\{\-8\}by iteration 100, at which point the loss reads1\.5651\.565to a loss higher than the starting loss\.

In Fig\.[3](https://arxiv.org/html/2607.05724#S4.F3)d, forp=0\.003p=0\.003, the transversal QCNN follows a similar loss volatility but exhibits markedly higher loss magnitudes throughout\. The initial loss started at0\.7140\.714in iteration 1, the model oscillates chaotically across the full 100 iterations, with losses regularly exceeding1\.01\.0and peaking above1\.61\.6on multiple occasions at iterations 8 and 38, respectively\. The learning rate undergoes the same rapid collapse as the standard variant, falling below10−810^\{\-8\}before iteration 100, and the final loss of1\.5651\.565provides no improvement over initialization\. Neither architecture recovers a downward trend after the learning rate saturates near zero, and neither approaches the0\.3480\.348final loss achieved by the transversal QCNN atp=0\.001p=0\.001\. This shows that an error rate ofp=0\.003p=0\.003is too high for trainability of both the standard and transversal QCNN under the present finite\-difference scheme configuration\. The noise\-resilience advantage of the transversal design observed atp=0\.001p=0\.001does not extend to this error rate\.

![Refer to caption](https://arxiv.org/html/2607.05724v1/qcnn_v_trans_qcnn_min.png)Figure 4:Training loss versus iteration \(a\) the 4\-qubit QCNN and \(b\) the 11\-qubit transversal QCNN\.Fig\.[4](https://arxiv.org/html/2607.05724#S4.F4)plots the best\-so\-far training loss over 100 iterations for the standard QCNN \(a\) and transversal QCNN \(b\) under different NISQ noise levelsp∈\{0,0\.0001,0\.001,0\.003\}p\\in\\\{0,0\.0001,0\.001,0\.003\\\}\. Each plotted value reflects the lowest loss achieved up to that iteration rather than the raw per\-iteration loss\. Under noiseless conditions, the standard QCNN converges smoothly and monotonically to a final loss of approximately0\.1740\.174, whereas all noisy variants stall early\. Atp=0\.001p=0\.001andp=0\.003p=0\.003plateauing immediately near0\.620\.62and0\.570\.57respectively, andp=0\.0001p=0\.0001halting around0\.430\.43after a few iterations\. All four noise levels for the transversal QCNN descend rapidly within the first 10 iterations and continue improving, with the noiseless curve reaching approximately0\.210\.21,p=0\.0001p=0\.0001settling near0\.220\.22, and even the noisiest conditions \(p=0\.001p=0\.001andp=0\.003p=0\.003\) converging to losses of roughly0\.430\.43and0\.500\.50respectively\. This contrast demonstrates that the transversal architecture retains trainability at noise levels that completely suppress learning in the standard QCNN, confirming its resilience at NISQ\-era error rates\.

![Refer to caption](https://arxiv.org/html/2607.05724v1/error_effect_on_precision.png)Figure 5:The violin plots of QCNN predictions \(blue\) vs Transversal QCNN \(orange\)\. Values are normalized to the noise\-free median\.Fig\.[5](https://arxiv.org/html/2607.05724#S4.F5)shows violin plots of qubit prediction values for the standard QCNN \(in blue\) and logical prediction values for the transversal QCNN \(in orange\) across Pauli noise levels normalized to the noise\-free statevector predictions\. Each noise level hasn=50n=50predictions over 25 random seeds\. Atp=0\.0001p=0\.0001, both architectures remain tightly concentrated near zero, indicating minimal deviation from the noiseless baseline, as noise increases top\>=0\.001p\>=0\.001, both distributions shift downward and broaden, reflecting degraded and increasingly variable predictions\. The standard QCNN consistently produces a more compact and symmetric distribution with its median tracking closer to zero at lower noise levels, reaching approximately−1\.05\-1\.05atp=0\.005p=0\.005and−1\.2\-1\.2atp=0\.01p=0\.01\. The transversal QCNN, by contrast, exhibits wider distributions with heavier tails extending well below−2\-2at higher noise levels, and while its median remains slightly higher than the standard QCNN at−0\.7\-0\.7and−0\.6\-0\.6forp=0\.005p=0\.005andp=0\.01p=0\.01respectively, this comes with substantially greater variance and outliers\. It can be concluded that the transversal gate structure preserves higher median prediction quality at elevated noise levels at the cost of greater variance, and that this preservation of median predictions may allow the optimizer to identify clearer local minima during noisy training\.

![Refer to caption](https://arxiv.org/html/2607.05724v1/qec_qcnn_precision.png)Figure 6:44\-qubit QCNN prediction of an antiferromagnetic and paramagnetism phase with the trained model\. The QCNN \(blue\) and the QEC\-QCNN \(orange\) are shown\.The QEC\-QCNN[6](https://arxiv.org/html/2607.05724#S4.F6)shows that the FFNN trained on Clifford gates over corrects for errors, and Non\-Clifford Gates are not apart of the training data, and appear to the FFNN soft decoder as errors\. This causes bit and phase flips to the entangled data, causing predictions to entropy to predictions of 1\. To improve results, the FFNN soft decoder should be trained with the QCNN model, for noise\-aware error correction, or trained on a pre\-trained model\. Thus, errors in application match errors seen during training\.

### IV\-CQubit Resource Analysis

The toric surface code encodes logical qubits into a physical qubit lattice with overhead scaling asn∝d2n\\propto d^\{2\}, achieving an error threshold of approximately0\.67%0\.67\\%\-0\.81%0\.81\\%under circuit\-level depolarizing noise using gauge\-fixing techniques\[[14](https://arxiv.org/html/2607.05724#bib.bib26)\], below which increasingddexponentially suppresses logical errors\. However, encodingk=4k=4logical qubits at distanced=4d=4requires roughly6464physical qubits plus400400–1,0001\{,\}000additional qubits per magic state distillation factory\[[23](https://arxiv.org/html/2607.05724#bib.bib20),[8](https://arxiv.org/html/2607.05724#bib.bib38)\]to supply the non\-Clifford gates essential to QCNN variational layers at current gate error rates\[[18](https://arxiv.org/html/2607.05724#bib.bib34)\], placing total physical qubit requirements between10410^\{4\}and10610^\{6\}\[[10](https://arxiv.org/html/2607.05724#bib.bib36),[22](https://arxiv.org/html/2607.05724#bib.bib37),[1](https://arxiv.org/html/2607.05724#bib.bib35)\]\. This is a prohibitive overhead for current hardware\. By contrast, the\[\[18,4,4\]\]\[\[18,4,4\]\]BB code encodes the same44logical qubits into2929physical qubits via a shared\-index transversal mapping[III\-C](https://arxiv.org/html/2607.05724#S3.SS3), with non\-Clifford corrections handled by the interleaved FFNN decoder, eliminating factory overhead at the cost of a lower error threshold \(∼0\.3%\\sim 0\.3\\%\)\. This makes it far better suited to near\-term operation\.

## VConclusion and Future Directions

In this study, we considered the problem of training instability of QCNN observed under NISQ hardware noise\. We proposed the integration of BB code to overcome this issue\. The\[\[18,4,4\]\]\[\[18,4,4\]\]BB code provides a mapping for physical to logical qubits and spins via the sparse parity check matrices to transversally encode a 4\-qubit QCNN into an 11\-qubit QCNN logical structure\. The transversal QCNN, by contrast, retains meaningful learning progress atp=0\.001p=0\.001, and converges to a final loss of 0\.348, demonstrating that fault\-tolerant encoding confers measurable noise resilience even without a fully trained decoder\. At higher noise rates, such asp=0\.003p=0\.003, neither architecture recovers\. This indicates that the gradient landscape is too difficult for current QCNN models without meaningful error correction\. These findings necessitate the joint error mitigation and error correction between the FFNN soft decoder and the transversal QCNN\. When higher qubit counts become available, qLDPC codes with larger distances and more logical qubits can be applied to deeper QCNN models\. Further, by co\-training the decoder alongside the QCNN variational parameters, or by initializing from a pre\-trained QCNN model, the correction layer can be exposed to the same error distribution encountered during inference, aligning the noise seen at training time with that seen in application\.

## Acknowledgment

The authors would like to thank Jorma Kilpi and Olli Apilo of VTT Technical Research Centre of Finland for their insightful feedback to this work\.

## References

- \[1\]R\. Babbush, J\. R\. McClean, M\. Newman, C\. Gidney, S\. Boixo, and H\. Neven\(2021\)Focus beyond quadratic speedups for error\-corrected quantum advantage\.PRX Quantum2,pp\. 010103\.External Links:[Document](https://dx.doi.org/10.1103/PRXQuantum.2.010103),2011\.04149Cited by:[§IV\-C](https://arxiv.org/html/2607.05724#S4.SS3.p1.15)\.
- \[2\]R\. Battiti\(1992\-03\)First\- and second\-order methods for learning: between steepest descent and Newton’s method\.Neural Computation4\(2\),pp\. 141–166\.Cited by:[§III\-B](https://arxiv.org/html/2607.05724#S3.SS2.p10.3)\.
- \[3\]P\. Bermejoet al\.\(2026\-04\)Quantum convolutional neural networks are effectively classically simulable\.PRX Quantum7\(2\),pp\. 020304\.External Links:[Document](https://dx.doi.org/10.1103/8qt9-72ts)Cited by:[§I\-A](https://arxiv.org/html/2607.05724#S1.SS1.p3.1)\.
- \[4\]J\. Biamonteet al\.\(2017\)Quantum machine learning\.Nature549,pp\. 195–202\.Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[5\]S\. Bravyi, A\. W\. Cross, J\. M\. Gambetta,et al\.\(2024\)High\-threshold and low\-overhead fault\-tolerant quantum memory\.Nature627,pp\. 778–782\.Cited by:[§I\-A](https://arxiv.org/html/2607.05724#S1.SS1.p2.5),[§I\-B](https://arxiv.org/html/2607.05724#S1.SS2.p1.5),[§I](https://arxiv.org/html/2607.05724#S1.p2.5),[§II\-D1](https://arxiv.org/html/2607.05724#S2.SS4.SSS1.p1.16),[§II\-D3](https://arxiv.org/html/2607.05724#S2.SS4.SSS3.p2.19),[§II\-D3](https://arxiv.org/html/2607.05724#S2.SS4.SSS3.p2.7),[§III\-B](https://arxiv.org/html/2607.05724#S3.SS2.p13.1),[§III\-C](https://arxiv.org/html/2607.05724#S3.SS3.p5.2),[§IV\-A](https://arxiv.org/html/2607.05724#S4.SS1.p1.8)\.
- \[6\]L\. Cauneet al\.\(2024\)Demonstrating real\-time and low\-latency quantum error correction with superconducting qubits\.arXiv preprint arXiv:2410\.05202\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.05202)Cited by:[§III\-C](https://arxiv.org/html/2607.05724#S3.SS3.p4.3)\.
- \[7\]I\. Cong, S\. Choi, and M\. D\. Lukin\(2019\-12\)Quantum convolutional neural networks\.Nature Physics15\(12\),pp\. 1273–1278\.Cited by:[§I\-A](https://arxiv.org/html/2607.05724#S1.SS1.p1.3),[§I\-B](https://arxiv.org/html/2607.05724#S1.SS2.p1.5),[§II\-A](https://arxiv.org/html/2607.05724#S2.SS1.p1.1),[§II\-A](https://arxiv.org/html/2607.05724#S2.SS1.p3.3),[§II\-B](https://arxiv.org/html/2607.05724#S2.SS2.p1.5),[§III\-A](https://arxiv.org/html/2607.05724#S3.SS1.p1.8),[§III\-A](https://arxiv.org/html/2607.05724#S3.SS1.p3.7)\.
- \[8\]B\. Eastin and E\. Knill\(2009\)Restrictions on transversal encoded quantum gate sets\.Physical Review Letters102\(11\),pp\. 110502\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevLett.102.110502)Cited by:[§IV\-C](https://arxiv.org/html/2607.05724#S4.SS3.p1.15)\.
- \[9\]J\. M\. Gambetta, J\. M\. Chow, and M\. Steffen\(2017\)Building logical qubits in a superconducting quantum computing system\.npj Quantum Information3,pp\. 1–7\.External Links:[Document](https://dx.doi.org/10.1038/s41534-016-0004-0)Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[10\]C\. Gidney and M\. Ekerå\(2021\)How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits\.Quantum5,pp\. 433\.External Links:[Document](https://dx.doi.org/10.22331/q-2021-04-15-433),1905\.09749Cited by:[§IV\-C](https://arxiv.org/html/2607.05724#S4.SS3.p1.15)\.
- \[11\]Google Quantum AI and Collaborators\(2025\)Quantum error correction below the surface code threshold\.Nature638,pp\. 920–926\.Cited by:[§III\-C](https://arxiv.org/html/2607.05724#S3.SS3.p4.3)\.
- \[12\]D\. Gottesman\(2014\)Fault\-tolerant quantum computation with constant overhead\.Quantum Information and Computation14\(15–16\)\.External Links:[Document](https://dx.doi.org/10.26421/QIC14.15-16-7)Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p2.5)\.
- \[13\]V\. Havlíčeket al\.\(2019\)Supervised learning with quantum\-enhanced feature spaces\.Nature567,pp\. 209–212\.Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[14\]O\. Higgott and N\. P\. Breuckmann\(2021\)Subsystem codes with high thresholds by gauge fixing and reduced qubit overhead\.Physical Review X11\(3\)\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevX.11.031039)Cited by:[§IV\-C](https://arxiv.org/html/2607.05724#S4.SS3.p1.15)\.
- \[15\]M\. Hirvensalo\(2001\)Quantum computing\.2nd edition,Springer,Berlin, Germany\.External Links:[Document](https://dx.doi.org/10.1007/978-3-662-04461-2)Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[16\]S\. Hochreiter\(1998\)The vanishing gradient problem during learning recurrent neural nets and problem solutions\.International Journal of Uncertainty, Fuzziness and Knowledge\-Based Systems6\(2\),pp\. 107–116\.Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[17\]H\. Huang, R\. Kueng, and J\. Preskill\(2021\)Information\-theoretic bounds on quantum advantage in machine learning\.Physical Review Letters126\(19\),pp\. 190505\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevLett.126.190505)Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[18\]IBM Quantum\(2024\-11\)IBM quantum delivers on 2022 100×100 performance challenge\.Note:IBM Quantum BlogAccessed: 2024External Links:[Link](https://www.ibm.com/quantum/blog/qdc-2024)Cited by:[§IV\-C](https://arxiv.org/html/2607.05724#S4.SS3.p1.15)\.
- \[19\]A\. Yu\. Kitaev\(2003\)Fault\-tolerant quantum computation by anyons\.Annals of Physics303\(1\),pp\. 2–30\.External Links:[Document](https://dx.doi.org/10.1016/S0003-4916%2802%2900018-0)Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p2.5)\.
- \[20\]Y\. LeCun, Y\. Bengio, and G\. Hinton\(2015\-05\)Deep learning\.Nature521\(7553\),pp\. 436–444\.External Links:[Document](https://dx.doi.org/10.1038/nature14539)Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1),[§II\-A](https://arxiv.org/html/2607.05724#S2.SS1.p1.1)\.
- \[21\]Y\. LeCun and Y\. Bengio\(1995\)Convolutional networks for images, speech, and time\-series\.InThe Handbook of Brain Theory and Neural Networks,Cited by:[§II\-A](https://arxiv.org/html/2607.05724#S2.SS1.p1.1)\.
- \[22\]J\. Leeet al\.\(2021\)Even more efficient quantum computations of chemistry through tensor hypercontraction\.PRX Quantum2,pp\. 030305\.External Links:[Document](https://dx.doi.org/10.1103/PRXQuantum.2.030305),2011\.03494Cited by:[§IV\-C](https://arxiv.org/html/2607.05724#S4.SS3.p1.15)\.
- \[23\]D\. Litinski\(2019\-12\)Magic state distillation: Not as costly as you think\.Quantum3,pp\. 205\.External Links:[Document](https://dx.doi.org/10.22331/q-2019-12-02-205),1905\.06903Cited by:[§IV\-C](https://arxiv.org/html/2607.05724#S4.SS3.p1.15)\.
- \[24\]J\. R\. McCleanet al\.\(2018\)Barren plateaus in quantum neural network training landscapes\.Nature Communications9,pp\. 4812\.Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[25\]I\. P\. McCulloch\(2008\)Infinite size density matrix renormalization group, revisited\.arXiv preprint arXiv:0804\.2509\.Cited by:[§III\-A](https://arxiv.org/html/2607.05724#S3.SS1.p1.3)\.
- \[26\]M\. A\. Nielsen and I\. L\. Chuang\(2011\)Quantum computation and quantum information: 10th anniversary edition\.Cambridge University Press,Cambridge, U\.K\.\.Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[27\]R\. Pascanu, T\. Mikolov, and Y\. Bengio\(2013\)On the difficulty of training recurrent neural networks\.InProceedings of the 30th International Conference on Machine Learning,Cited by:[§I](https://arxiv.org/html/2607.05724#S1.p1.1)\.
- \[28\]D\. Pomarico\(2023\)Multiscale entanglement renormalization ansatz: causality and error correction\.Dynamics3\(3\),pp\. 622–635\.External Links:[Document](https://dx.doi.org/10.3390/dynamics3030033)Cited by:[§II\-A](https://arxiv.org/html/2607.05724#S2.SS1.p2.18)\.
- \[29\]S\. Sachdev\(2011\)Quantum phase transitions\.2nd edition,Cambridge University Press,Cambridge, U\.K\.\.Cited by:[§IV\-B](https://arxiv.org/html/2607.05724#S4.SS2.p1.2)\.
- \[30\]J\. C\. Spall\(1992\)Multivariate stochastic approximation using a simultaneous perturbation gradient approximation\.IEEE Transactions on Automatic Control37\(3\),pp\. 332–341\.External Links:[Document](https://dx.doi.org/10.1109/9.119632)Cited by:[§III\-B](https://arxiv.org/html/2607.05724#S3.SS2.p8.5)\.
- \[31\]J\. Viszlaiet al\.\(2025\)Low latency GNN accelerator for quantum error correction\.arXiv preprint arXiv:2603\.22149\.Cited by:[§III\-C](https://arxiv.org/html/2607.05724#S3.SS3.p4.3)\.
- \[32\]K\. Wang, Z\. Lu, C\. Zhang,et al\.\(2026\)Demonstration of low\-overhead quantum error correction codes\.Nature Physics22,pp\. 308–314\.Cited by:[§I\-A](https://arxiv.org/html/2607.05724#S1.SS1.p2.5),[§I\-B](https://arxiv.org/html/2607.05724#S1.SS2.p1.5),[§II\-D1](https://arxiv.org/html/2607.05724#S2.SS4.SSS1.p1.16)\.

Similar Articles

Decoherence as Defence and the Magnitude of Noise Regularisation: A Rigorous N -Qubit Theory of Stochastic Quantum Neural Networks for Adversarially Robust Network Intrusion Detection

arXiv cs.CL

This paper presents a rigorous N-qubit theory of stochastic quantum neural networks (SQNNs) for adversarially robust network intrusion detection, proving a decoherence-contraction theorem and showing that depolarising noise provides robustness against adversarial attacks, with experiments on the NSL-KDD dataset.