Byte-identical edge inference, and what it buys the closed-loop rig
The same neural network, the same input, different outputs depending on which kernel the runtime silently selected: this fragmentation, previously documented on x86, is now characterized on ARM, the substrate most embedded acquisition nodes actually run. The surprise is the fix. INT8 post-training quantization does not merely speed inference up; its quantized-graph structure makes outputs bit-exact across Cortex-A53, A72, and A76, while FP32 diverges on every image tested.
Source: INT8 Quantization Makes ARM Edge Inference Dispatch-Invariant, arXiv:2607.23227v1 [cs.ET], 25 July 2026. Primary source. Read the full arXiv HTML version, including the per-layer remaining-precision tables and the H1+H2 mechanism sections.
What the work claims
Cruz Romero and Maldonado Guerra, both at Capicu Technologies, claim two things. First, a negative: on ARM edge hardware under default CPU dispatch, the microarchitecture is not observable in a fixed network's outputs, a 12-layer CNN trained on CIFAR-10 produced byte-identical outputs on every one of 1,000 test images across four Raspberry Pi nodes spanning Cortex-A53, A72, and A76, on two SoC families and two ONNX Runtime builds. But the execution provider is observable: the same FP32 model on the same Cortex-A76 chip diverged from its own outputs on 1,000 of 1,000 images when XNNPACK rather than the CPU provider executed it, losing on average 8.03 of 23 mantissa bits per output.1
Second, a positive: INT8 QDQ post-training quantization collapses both axes of variation to a single equivalence class. All five tested conditions, three microarchitectures, two SoC vendors, and both execution providers, produced byte-identical outputs for the small CNN on all 1,000 images, and the result extended to MobileNetV2 and ResNet50V2 under TensorFlow Lite with confirmed different microkernel dispatch: every one of 500 ImageNet images per model was classified identically on Cortex-A76 (using the SDOT dot-product instruction) and Cortex-A72 (using NEON multiply-accumulate), with all 109 intermediate tensors of MobileNetV2 and all 156 of ResNet50V2 byte-identical on a 50-image subset.1
This is a primary empirical result with a mechanistic explanation, not a position piece. Its weight is limited by being an arXiv preprint from a two-author company team, with the dispatch positive measured on a single small model, but the hardware protocol is unusually careful: inputs were SHA-256-verified byte-identical across nodes, single-threaded inference, and every claimed microkernel difference is corroborated by timing signatures rather than assumed.
How it works
The background is the equivalence-class framework of Schlögl et al., who showed that on x86, configurations of hardware, runtime, and execution provider fragment into classes whose members produce byte-identical outputs: one TensorFlow binary running ResNet-18 on 75 x86 platforms split into as many as 26 equivalence classes, aligned with SSE versus AVX versus AVX-512 kernel boundaries. The cause is floating-point non-associativity: SIMD width changes the order of reduction sums, and rounding makes order matter. Remaining precision, the count of leading identical mantissa bits (capped at 23 for IEEE binary32), measures the drift.1
The authors localize the ARM FP32 divergence to the first convolution: layer-wise extraction of prefix subgraphs shows Conv0 diverging with a mean remaining precision of only 2.56 bits between providers, partially attenuating to 15.32 bits at the output as ReLU clipping and MaxPool selection smooth the low-order noise. That profile kills an error-accumulation story and identifies the kernel, not the depth, as the source.1
The INT8 mechanism has two halves they call H1 and H2. H1, discrete-grid absorption: if a convolution's inputs lie exactly on the integer quantization grid, its FP32 output is byte-identical across providers, the authors verify this by inserting a single QuantizeLinear/DequantizeLinear pair at the input, which makes Conv0 bit-exact. H2, the requantization ratchet: a full QDQ graph restores grid membership at every layer boundary, so the precondition survives the whole network. Inserting the input pair alone fails immediately: Conv1 reverts to divergence because Conv0's continuous outputs left the grid. Under the hood, the runtime's optimizer routes quantized convolutions to integer kernels that accumulate into INT32, and integer addition is associative, so reduction order, SIMD width, and microkernel choice become invisible. The authors prove a bound: any channel dimension up to 131,072 accumulates exactly in INT32 without overflow, since each INT8 product is at most 16,256 in magnitude; every standard edge CNN sits orders of magnitude inside it.1
They also nail down why x86 behaves differently: the PMADDUBSW instruction saturates pairwise INT16 intermediates (200 by 90 plus 210 by 90 gives 36,900, clipped to 32,767), and different SSE and AVX implementations partition the reduction tree differently, so which pairs saturate depends on the kernel. The divergence on x86 is integer saturation, not floating-point rounding, and ARM has no analogue instruction. Intel's VNNI extension, which accumulates directly into INT32, mirrors the ARM design and resolves it.1
Two hardware details deserve quoting because instrument builders will recognize them. The dispatch difference on the production CNNs is evidenced indirectly but convincingly by timing: INT8 inference on the A76 is 5.04 times faster than FP32 on ResNet50V2, above the 4x memory-bandwidth reduction, consistent with SDOT doing four multiply-accumulates per cycle; on the A72 the INT8 speedup is about 1.0x (137.80 ms versus approximately 138 ms), exactly what the absence of SDOT predicts. And the Pi 5 ran at 88 to 91 degrees Celsius during sustained ResNet50V2 inference, adding latency variability that forced the authors to treat timing evidence as strong but indirect, no perf-stat instruction counts were taken.1
Where a skeptic should push
The most load-bearing assumption is that the timing signatures actually prove different microkernels executed. The authors are honest that they did not: direct instruction-level profiling is named as future work, the working XNNPACK execution provider on ARM was available only in one Debian package on one node, and an earlier attempt had to be discarded entirely when pip ONNX Runtime silently accepted an XNNPACK provider request but resolved every run to the CPU provider. That silent fallback is itself a finding, a textbook case of what the ML-systems literature calls configuration debt, and the authors' remedy, read the resolved provider from per-run telemetry, never from the session object, is advice every acquisition engineer should recognize from a different domain: trust the telemetry, not the front panel.1
Second, the headline reproducibility gain is bought with quantization error, and the trade is quantified rather than hidden. On the small CNN, INT8 changed 27 of 1,000 predicted labels relative to FP32 (2.7 percent), concentrated in the low-confidence tail: mean FP32 confidence on the disagreement set was 0.42 versus 0.83 overall, and a 0.5 confidence threshold recovered full label consistency. On the production CNNs, INT8 versus FP32 top-1 disagreement was 8.0 percent on MobileNetV2 and 6.8 percent on ResNet50V2. Every one of those disagreements was identical on both platforms, so the hardware axis contributes zero label drift, but a rig that is reproducibly different from its FP32 reference is still different. For a closed-loop system whose stimulation policy was tuned against FP32 confidences, switching precision regime is a retuning event, not a silent upgrade.1
Third, scope. FP32 cross-hardware divergence on the production CNNs was not studied, so the comforting "microarchitecture is not observable" claim is demonstrated only for the small model's operator set. The mechanism is stated as an empirical property of specific operators under specific runtime builds, not a theorem. Cortex-A55, A78, Cortex-X, and Qualcomm Oryon are untested, and the x86 result means the invariance does not transfer off ARM.1
What dispatch-invariance means for the MEA acquisition chain
An MEA acquisition chain increasingly ends in a classifier. Organoid screening platforms score network bursts and rhythmicity states on the acquisition host; closed-loop rigs trigger stimulation from a model's output; chronic and wireless systems push inference onto the embedded node precisely because raw multichannel streams do not fit down a telemetry link. Every one of those deployments inherits the phenomenon this paper measures. If the same burst-detection network runs FP32 on a lab PC upgraded over five years, or across a fleet of mixed-vendor ARM controllers, the continuous outputs, confidences, anomaly scores, and any threshold sitting on them drift by kernel dispatch alone. The authors measured 100 percent output-hash divergence with zero label flips for their model, which is exactly the failure mode that escapes accuracy-based validation while corrupting anything downstream that consumes continuous values. A stimulation threshold placed on a model confidence lives in that blind spot.
The non-obvious implication is that INT8 is not primarily an efficiency decision for this field; it is a verification decision. Byte-identity across microarchitectures gives the instrumentation world something it has never had at the compute layer: a cheap commissioning test for closed-loop reproducibility. Hash the output tensor on two rigs and compare; if both run QDQ-quantized graphs, agreement is structural, not statistical. That turns "did the model change when we moved the rig" from a forensic question into a pass-fail check, and it applies equally to vendor qualification: two nominally identical acquisition workstations can be proven decision-identical for a fixed network and input set.
The threat is the mirror image. Most published closed-loop MEA work has been validated on whatever hardware the lab owned, in whatever precision the framework defaulted to, with no record of which kernels executed. If FP32 dispatch can shift confidences by eight mantissa bits, then a closed-loop result replicated on a different machine is not a replication of the same controller, it is a nearby controller. As regulators and funders start asking acquisition-system builders for reproducibility evidence, rigs whose decision layer cannot demonstrate byte-identity will carry an unquantified variance term in every claim. There is also a subtler threat to the hype side of edge AI in instruments: the same paper shows hardware fingerprinting from model outputs, a supply-chain audit technique demonstrated on x86, provably cannot work on ARM INT8 because there is no divergence to fingerprint. Audit and attestation tooling for heterogeneous instrument fleets must target other layers, which is a real constraint on procurement designs that assumed it.
One practical detail belongs in every builder's notebook: the silent execution-provider fallback. A runtime that accepts a provider request and quietly runs another is the software twin of a headstage that accepts a sampling-rate command and streams at another, and the remedy is the same in both cases. Instrument software that logs configuration intent instead of resolved configuration is writing fiction, and this paper's discarded experiment shows the failure occurs in production-grade runtimes, not in student code.
The bottom line
Established, on ARM and within its stated scope: FP32 outputs of a fixed CNN are byte-identical across microarchitectures under CPU dispatch but diverge on every image when the execution provider changes; INT8 QDQ outputs are bit-exact across Cortex-A53, A72, and A76, across providers, and across confirmably different microkernels, for networks up to ResNet50V2; and the x86 exception is pinned to a single saturating instruction. Not established: instruction-level confirmation of dispatch, FP32 behavior on production CNNs, any non-Raspberry-Pi silicon, and any neural workload beyond image classification. For MEA instrumentation the takeaway is precise: treat INT8 QDQ as the reproducible mode of edge inference, budget its 2.7 to 8 percent label drift against FP32 as a retuning event, and add output-hash commissioning to closed-loop rig qualification. What would confirm the broader claim is perf-stat verification on additional cores; what would break it is a mainstream ARM INT8 kernel with a saturating intermediate, which is precisely the x86 disease this paper shows ARM avoided.
Frequently asked questions
What is an equivalence class in this context?
A set of hardware, runtime, and execution-provider configurations whose members produce byte-identical outputs for the same model and input. Schlögl et al. showed one TensorFlow binary running ResNet-18 on 75 x86 platforms splits into up to 26 such classes along SSE, AVX, and AVX-512 kernel boundaries. This paper shows INT8 QDQ inference on ARM collapses all tested configurations into a single class.
What is remaining precision?
The number of leading mantissa bits two floating-point outputs share, capped at 23 for IEEE binary32. A value of 23 means byte-identical; lower values measure drift. On the small CNN, FP32 outputs from XNNPACK versus the CPU provider on the same Cortex-A76 chip had a mean remaining precision of 14.97 bits, meaning about 8 bits lost on average, even though no predicted label changed.
Why does INT8 quantization make outputs identical?
Two structural properties the authors call H1 and H2. H1: when a convolution's inputs lie exactly on the integer quantization grid, its output is dispatch-deterministic. H2: the QDQ graph re-quantizes every activation at every layer boundary, restoring grid membership throughout the network. Quantized convolutions then execute in integer kernels accumulating into INT32, and integer addition is associative, so reduction order and SIMD width no longer matter. They prove exactness holds for channel dimensions up to 131,072.
Why does x86 not show the same invariance?
A single instruction, PMADDUBSW, saturates pairwise INT16 intermediates in the x86 INT8 GEMM path: two products whose sum exceeds 32,767 are clipped, and different SSE and AVX kernels group operands differently, so which pairs saturate depends on the kernel. ARM's SDOT and NEON paths accumulate into INT32 with no saturation step, which is why the divergence mechanism has no ARM analogue. Intel's VNNI extension removes the saturation and behaves like ARM.
Does quantization change what the model decides?
Yes, and the paper quantifies it honestly. On the small CNN, INT8 changed 27 of 1,000 labels versus FP32, all in the low-confidence region where mean FP32 confidence was 0.42 versus 0.83 overall. On MobileNetV2 and ResNet50V2, INT8 versus FP32 top-1 disagreement was 8.0 and 6.8 percent. Crucially, every disagreement was identical on both ARM platforms, so quantization shifts decisions but hardware does not.
What should an MEA rig builder do differently?
Three things. Deploy decision-making networks as INT8 QDQ graphs if cross-hardware reproducibility matters, and treat any precision switch as a retuning event. Add an output-hash commissioning test: run a fixed input set on two rigs and compare hashes, which is only meaningful because QDQ makes agreement structural. And log the resolved execution provider from per-run telemetry, because this study caught a production runtime silently substituting providers, which would have silently invalidated any experiment that treated the provider as a controlled factor.
References
- S. A. Cruz Romero, S. E. Maldonado Guerra. INT8 Quantization Makes ARM Edge Inference Dispatch-Invariant. arXiv:2607.23227v1 [cs.ET], 25 July 2026. https://arxiv.org/abs/2607.23227. Accessed 2026-09-28.