Surface Code Runs Below Threshold on IBM Heavy-Hex – But Only in One Direction at a Time

September 15, 2026 – Researchers at the University of Southern California, quantum software company Quantum Elements, and the Instituto de Física Teórica (UAM-CSIC) in Madrid demonstrated subthreshold surface-code scaling on IBM Heron processors whose heavy-hex connectivity does not natively match the code, according to a paper published July 29 in Nature Communications.

The paper was co-authored by Arian Vezvaee and Cesar Benito as equal first authors, alongside Mario Morford-Oberst, Alejandro Bermudez, and Daniel Lidar. Bermudez and Lidar supervised the project, and the fold-unfold embedding builds on a 2025 theoretical study by Benito, Bermudez, and colleagues. Lidar directs USC’s Center for Quantum Information Science & Technology and serves as chief scientific officer at Quantum Elements; Vezvaee is a quantum research scientist at the company.

The team embedded the surface code on 156-qubit IBM Heron processors using a depth-minimizing “fold-unfold” SWAP strategy with bridge ancillas. IBM’s heavy-hex architecture places qubits on the sites and links of a honeycomb lattice – a lower-connectivity layout than the square grid the surface code requires. To work around the resulting routing delays and idle gaps, the researchers paired the SWAP embedding with robust dynamical decoupling (DD) – the same techniques Quantum Elements later productized as its Orbit Qiskit Function – suppressing coherent ZZ crosstalk and non-Markovian dephasing that accumulate during idle periods.

The experiment scaled from a 37-qubit distance-3 surface code to 65-qubit anisotropic configurations at distances (3, 5) and (5, 3), running up to 10 full quantum error correction cycles at circuit depths exceeding 140, with a total of roughly 2,200 entangling gates.

In the direction of increased distance, the larger codes showed lower logical error rates, with suppression factors between 1.23 and 1.46. Measured against the best-performing distance-3 patch instead of the average of three patches, the factors fell to 1.01–1.10. Averaged over all four logical input states, they were 0.67 and 0.77, below 1.

The paper also introduced a SPAM-aware entanglement fidelity (EF) metric, which the authors argued provides a more rigorous benchmark for below-threshold performance than the widely used single-parameter suppression factor. Using the EF metric, the team did not find global, state-independent subthreshold scaling for the full logical channel on ibm_aachen, the top-performing processor in the study. On ibm_marrakesh, a second Heron processor, the team observed what they described as the onset of subthreshold scaling at later QEC cycles, reaching 95% statistical confidence at the ninth cycle. Calibrated noise simulations indicated that a roughly 30% reduction in current error rates on a slightly larger processor would achieve full isotropic scaling.

The researchers warned that subthreshold scaling claims require careful consideration of dynamical decoupling effects, showing that without DD, logical error rates can decrease with code distance in a way that mimics genuine subthreshold behavior but reverses once DD is properly applied. They described this as “spurious subthreshold scaling.”

The work was funded by IARPA’s Entangled Logical Qubits program, the U.S. Army Research Office, and DARPA, and extends a collaboration between USC and Quantum Elements that also produced a quantum Monte Carlo method for noisy circuit simulation, published in Physical Review Letters earlier in 2026.


Part 2: My Analysis

The headline result – subthreshold surface-code scaling on IBM’s heavy-hex architecture – is technically correct and genuinely useful. The authors qualify the claim carefully, and the subthreshold scaling they report is anisotropic.

The team could grow the code distance in one direction at a time, suppressing either X-type or Z-type logical errors, but not both simultaneously. When they combined all four basis states using their entanglement fidelity metric, the larger codes did not outperform the smaller ones across the board on ibm_aachen. On ibm_marrakesh, the larger code gradually overtook the smaller one from the sixth cycle onward – what the authors call the onset of subthreshold scaling, though the statistical confidence reached 95% only at the ninth cycle. Each QEC cycle takes roughly 8 μs on heavy-hex versus 1.1 μs on Google’s native square grid. The paper does not break that gap down. Part of it follows from the layout: heavy-hex can measure only half the stabilizers in parallel, so every full cycle needs two rounds of syndrome extraction, and the extra depth and idle time compound physical errors in the unprotected direction. As I noted when analyzing Google’s Willow result in December 2024, the square-grid Sycamore/Willow processors were purpose-built for the surface code. The IBM Heron processors were not, and the gap shows.

What it is: the first demonstration that surface-code error suppression works on a superconducting processor whose qubit connectivity was designed for engineering reasons other than surface-code performance. IBM chose heavy-hex to reduce frequency crowding, lower crosstalk, and ease fabrication – practical manufacturing tradeoffs that most QPU architects will face. The result means that a chip designed around those tradeoffs is not locked out of surface-code error correction. That is a useful data point for anyone planning a quantum computing facility where the QPU vendor and the QEC strategy may not come from the same roadmap.

The Suppression Factor in Context

The directional suppression factors of 1.23 to 1.46 are lower than both Google’s Willow suppression factor of 2.14 and the 1.56 to 2.15 range reported across native-connectivity surface-code demonstrations, from Google’s dynamic surface codes and Willow chip to Harvard/QuEra’s neutral-atom runs. That gap was expected.

The more interesting number is the 30% noise reduction threshold the authors calculated would deliver genuine isotropic (5, 5)-versus-(3, 3) scaling on a heavy-hex chip. That is a hardware improvement target an engineering team can plan around. Current Heron processors are 6 qubits short of hosting the (5, 5) code at all, and the noise rates need to drop by less than a third. Neither constraint looks like a multi-year obstacle.

The EF Metric and the Measurement Gap

The paper’s secondary contribution may prove more durable than the experimental result itself. The entanglement fidelity metric the team introduced exposes a genuine vulnerability in the standard benchmarking approach. The single-parameter suppression factor assumes stationary noise, no SPAM errors, and purely unital (Pauli-only) logical channels, and the IBM processor data deviated from all three. The three-parameter model the authors validated through the Akaike Information Criterion showed that logical SPAM and non-unital noise were both present – and that switching between one-, two-, and three-parameter fits moved the suppression factor by up to 14.93% for the (3, 5) code and 17.22% for the (5, 3) code.

That variation matters when the field is comparing results across hardware platforms. A suppression factor of 1.5 that absorbs SPAM into the per-cycle error and a suppression factor of 1.5 that separates SPAM are reporting different things. Vezvaee et al. argue for fitting-model-free comparisons that account for SPAM, non-unitality, and cycle-dependent noise, and the broader error correction community will need to adopt them as groups publish results from more processors, codes, and modalities.

The DD Warning

The finding on spurious subthreshold scaling deserves attention from anyone evaluating quantum error correction claims. The team showed that without dynamical decoupling, the (5, 3) code appeared to outperform the (3, 3) code under the EF metric – but only because the smaller code was hurt more by unmitigated idle noise than the larger code. Once each code was independently optimized with DD, the apparent scaling advantage disappeared on ibm_aachen and narrowed to a late-cycle onset on ibm_marrakesh.

The practical implication: any QEC result that does not optimize error-suppression strategies per-code-size before comparing across sizes is potentially reporting an artifact. The warning applies beyond IBM hardware. Google’s Willow experiments applied DD only to idling data qubits during readout and reset, but the integration was simpler because the surface-code circuits on Willow’s native square grid had no idle gaps to fill. As research groups move to non-native geometries, multi-code comparisons, and denser QEC cycles, they will each need to control for DD effects as a source of systematic error.

What This Means for the Path to a CRQC

In my CRQC Quantum Capability Framework, this result applies to two capabilities, B.3 (Below-Threshold Operation & Scaling) and B.4 (Qubit Connectivity & Routing Efficiency). It extends below-threshold evidence to a non-native architecture, and it validates that connectivity-aware embedding can partially compensate for reduced hardware connectivity – a tradeoff every modality will confront at larger scale.

It does not change the timeline to a CRQC. The result is anisotropic scaling at small code distances (3 to 5 in one direction) on a processor of 156 qubits, four to five orders of magnitude fewer than the physical qubit counts a cryptographically relevant machine would require. What it does is broaden the hardware base that can credibly participate in the error correction race. Until this paper, below-threshold surface-code demonstrations on superconducting hardware had been achieved only on processors whose connectivity was purpose-built for the code – Google’s Sycamore and Willow processors and USTC’s Zuchongzhi 3.2. IBM’s commercial heavy-hex architecture, designed for other engineering priorities, had not entered that list.

For CISOs and security architects tracking the pace of quantum computing, the relevant takeaway is not this particular suppression factor. It is the pattern: error correction techniques are becoming increasingly hardware-agnostic. The surface code now shows directional error suppression on heavy-hex, where no qubit has more than three neighbors, and Google’s dynamic surface-code experiments ran it with circuits that use three couplers per qubit instead of four. The number of engineering paths to fault tolerance is growing, and that growth compresses the timeline uncertainty around when any of them reaches scale.

The date by which organizations must migrate does not depend on whether heavy-hex processors reach isotropic scaling next quarter or next year. Regulatory deadlines for post-quantum migration are already set. The engineering base on which a future fault-tolerant machine could be built keeps widening, and every increment in that base reduces the probability that an unforeseen technical obstacle blocks all paths simultaneously.

The post appeared first on PostQuantum - Quantum Computing, Quantum Security, PQC.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论