Research analysis · Compute

A 10,000-node annealer, measured four nodes at a time

A preprint introduces Apollo, a claimed 10,000-node room-temperature quantum annealer in 16 nm CMOS drawing about 0.5 W, and supports the claim with careful bench data from a 350 nm test chip running four stochastic analog nodes. Both halves of that sentence matter. The measured part is legitimate low-power analog-compute work; the flagship part is a performance model, and the distance between them is a case study in how hardware claims should be audited.

Source: Quantum-Driven Neuromorphic Computing for Million-Qubit-Scale Workloads, arXiv preprint (quant-ph), submitted 11 June 2026. Primary source. Read: full HTML version including the architecture sections, all experimental tables, the spin-glass benchmark section and the figure captions.

What the work claims

Ivanov, Rahmeh, Nascimento and Herrmann, affiliated with a Liechtenstein holding establishment and one university institute, claim a neuromorphic processor called Apollo built from p-qubits: bistable analog units whose switching probability follows a sigmoidal function of a local field, with stochasticity injected by on-chip entropy circuits said to derive true randomness from quantum-mechanical processes such as electron tunnelling.1 The headline system is 10,000 p-qubits in 16 nm mixed-signal CMOS, running clockless continuous-time dynamics at roughly 12.5 picoseconds per transition and no more than 10 femtojoules per transition, inside a 0.5 W analog power envelope. Through the Suzuki-Trotter correspondence, the authors claim this classical stochastic network reproduces the annealing dynamics of a transverse-field quantum annealer. On a three-dimensional spin-glass benchmark previously run on D-Wave hardware, they report matching the quantum-critical scaling seen in the superconducting device and reaching lower ground-state energies, at 10^3 nanoseconds runtime against the published 10^5 nanoseconds.2

The paper is therefore two documents fused together: an experimental section describing a real but small chip, and a system section projecting a chip that was not measured. The claims that make it spectacular all belong to the second document.

How it works

The measured device, Apollo-RC1, is a 350 nm reconfigurable analog CMOS chip in the style of a field-programmable analog array: tunable operational transconductance amplifiers, floating-gate elements for non-volatile analog weight storage, switched capacitors and a routing fabric. Each p-qubit sums weighted currents from its neighbors in the analog domain (the vector-matrix multiplication), adds noise from its entropy unit, and passes the sum through an amplifier whose sigmoid-shaped response sets the instantaneous switching probability. With no clock, the network relaxes continuously, and annealing is performed by ramping the effective noise down while couplings dominate, analogous to lowering the transverse field in a quantum annealer. The bench results cover four-node systems: clean sigmoid transfer curves, an entropy source whose bit bias, serial correlation and NIST SP 800-90B min-entropy statistics are statistically indistinguishable from a commercial ID Quantique Quantis quantum random number generator across 1,000 test segments, Gibbs-distribution sampling with Kullback-Leibler divergence under 1 percent, stable operation from 0 to 85 °C, and an FPGA control unit driving the chip hardware-in-the-loop.

The 10,000-node Apollo exists as theory plus a performance model. The 16 nm transition rate, the 0.5 W envelope, the energy-per-flip ranking table and the aggregate flip-rate figure of about 10^17 flips per second in a ten-by-ten multi-package assembly are all extrapolations from the 350 nm device physics and scaling laws, not measurements of 16 nm silicon. This is a legitimate way to design a chip; it is not a legitimate way to claim a result, and the paper alternates between the two registers.

Where a skeptic should push

The single most load-bearing assumption is that dynamics measured on four time-multiplexed analog nodes transfer quantitatively to 10,000 parallel nodes in a new process node. Everything headline-grabbing rests on it, and nothing in the paper tests it. On the measured chip itself, the physics is plausible and the bench work is careful; the entropy statistics tables are genuinely impressive until you notice that the IQEU and Quantis means agree to three decimal places in three separate tables (7.673 bits of min-entropy in both, 7.823 in both standard tables), a degree of agreement between two physically independent noise sources that deserves an explanation the paper does not give.

There are also internal consistency problems that a careful reader should not skate past. The abstract claims at most 10 femtojoules per transition, while the energy table reports 0.63 femtojoules as this work's value, a factor of sixteen apart, with the table itself labeled as estimated. The claimed aggregate rate of roughly 10^17 flips per second exceeds what 10,000 nodes at the stated 12.5 picoseconds could deliver by about two orders of magnitude, and only closes if the ten-by-ten package assembly, which is not described as having been built, exists. The spin-glass comparison reproduces the D-Wave study's construction faithfully, but the Apollo curve is generated on the modeled system and compared against digitized literature data from a different device with different anneal scheduling, at two orders of magnitude shorter runtime; agreement in power-law slope is suggestive, not confirmation. Finally, the figure captions describe Hadamard gates, X gates and plus-state superpositions on an analog stochastic chip that has no gates, vocabulary borrowed from gate-model quantum computing that signals how loosely the quantum framing is being applied. The entropy source is never shown to be quantum rather than thermal; Johnson noise passes the same statistical suites. None of this proves the work wrong. It establishes that the demonstrated fraction is four analog nodes and the asserted fraction is nearly everything else.

What modeled silicon means for MEA-side compute

The direct relevance to microelectrode array hardware is not annealing; it is the audit template. The MEA field is being offered analog and neuromorphic front-ends, on-chip spike sorters and closed-loop controllers whose datasheets follow exactly this structure: a small measured array, a large extrapolated system, an energy figure that counts only the core. Here the claimed 0.5 W analog core sits beside a control FPGA measured at 1.815 W, before any host, converters or I/O enter the budget, and the per-flip ranking quietly excludes everything around the array. When the same rhetorical pattern appears in a neural-interface compute brochure, the questions are the same ones this paper fails to pre-empt: how many nodes are on the die you actually measured, what does the energy number exclude, and is the benchmark against live hardware or against published curves.

The genuine opportunity underneath the hype deserves stating plainly, because the physics is real. Continuous-time, clockless analog dynamics with femtojoule-class updates is the right compute model for always-on work at the array edge: spike detection, artifact rejection, Bayesian decoding of stochastic neural data, and closed-loop triggering, where a digital DSP chain costs milliwatts per channel and analog summation with floating-gate weights costs orders of magnitude less. A four-node proof that analog stochastic units sample correct Gibbs distributions is a meaningful step toward probabilistic front-ends that treat neural noise as the native signal rather than the enemy.

The threat is twofold. Practically, analog parameters drift: this design leans on tunable gain and floating-gate storage to hold its operating point, and in a closed-loop neural stimulator, silent analog drift is a safety problem before it is an accuracy problem. Culturally, quantum-class claims attached to extrapolated silicon set a hype precedent that will be quoted back at the neural-interface field the next time an analog front-end benchmark circulates without its denominator. The correction this paper needs, replication of the scaling claims on fabricated 16 nm silicon with a live head-to-head benchmark, is exactly the correction MEA compute claims will need as well.

The bottom line

Established: a 350 nm reconfigurable analog chip can implement stochastic p-bit dynamics that sample correct Boltzmann statistics at four nodes, with an entropy source that passes NIST randomness suites indistinguishably from a certified quantum random number generator. Not established: the 16 nm, 10,000-node system, the quantum origin of the entropy, the femtojoule and picosecond figures, and the quantum-advantage reproduction, all of which rest on modeling and literature comparison. What would confirm the claims is taped-out 16 nm silicon measured with the same rigor as RC1, a head-to-head spin-glass run against a live annealer, and a physical discriminator showing the entropy source is quantum rather than thermal. What would break them is failure of the scaled silicon to approach the modeled dynamics, or disclosure that the benchmark curves are simulated. Until then the honest reading is a solid small analog-compute result wearing the press release of a revolution.

Frequently asked questions

What is a p-qubit?

A probabilistic bit: an analog unit with two preferred output states whose probability of flipping depends sigmoidally on a weighted sum of neighbor states plus injected noise. It is classical stochastic hardware, not a quantum two-level system, whatever the name suggests.

What did the authors actually measure?

A 350 nm reconfigurable analog test chip running four physical p-qubits: sigmoid transfer curves, entropy statistics against a commercial quantum random number generator, four-node Gibbs sampling with under 1 percent divergence from theory, and hardware-in-the-loop operation driven by an FPGA control unit.

Why do the energy numbers matter?

Because they are the headline and they do not reconcile. The abstract allows at most 10 femtojoules per transition while the comparison table lists 0.63 femtojoules as this work's figure, marked as estimated, and neither figure includes the 1.815 W FPGA controller. Energy-per-operation claims are only as honest as the system boundary they draw.

Can a classical room-temperature device be quantum-equivalent?

The Suzuki-Trotter correspondence maps a d-dimensional transverse-field quantum system onto a d-plus-one-dimensional classical stochastic network, so classical hardware can in principle reproduce the equilibrium statistics. Equivalence in scaling exponents on one benchmark is evidence about dynamics, not proof of quantumness, and the paper offers no test that distinguishes quantum entropy from thermal noise.

Why does a microelectrode array site cover an annealing chip?

Because the claim architecture transfers directly. Analog compute blocks marketed for on-array neural decoding use the same pattern of small measured demos, large extrapolated systems, and core-only energy figures. This preprint is a worked example of which questions to ask before believing the datasheet.

What would make the claims credible?

Silicon: a fabricated and measured 16 nm array rather than a scaling model. Benchmarks: a live head-to-head against an operating annealer instead of comparison to published curves. Physics: a discriminator experiment showing the entropy source is genuinely quantum rather than thermal noise passing the same statistical tests.

References

  1. A. Ivanov, S. Rahmeh, E. G. S. Nascimento, D. Herrmann. Quantum-Driven Neuromorphic Computing for Million-Qubit-Scale Workloads. arXiv:2606.12968. 2026. https://arxiv.org/abs/2606.12968. Accessed 2026-09-03.
  2. A. D. King et al. Quantum critical dynamics in a 5,000-qubit programmable spin glass. Nature 617, 61-66. 2023. https://doi.org/10.1038/s41586-023-05867-2. Accessed 2026-09-03.