One curve for every substrate: separability scaling and the array readout
A team spanning FEMTO-ST in Besancon, Tampere, Lodz and TU Berlin has found that physical neural networks obey a task-determined power law: classification loss falls as a power of a separability statistic that can be computed from the system's raw responses without training anything. The law holds across femtosecond-pulse propagation in optical fibre, transverse-mode dynamics in a large-area vertical-cavity laser, and simulated networks of coupled nonlinear oscillators, with data from physically distinct systems collapsing onto a single curve per benchmark. For microelectrode arrays, the result is a way to price the recording chain into the cost of any computation that consumes it, and a discipline on the claim that exotic substrates compute their way past silicon.
Source: Power law scaling for classification accuracy in physical neural networks, arXiv (cs.ET), 2026-06-30 (v1). Primary source. Read in full (arXiv HTML of v1, including the Hotelling trace derivation, all three system descriptions, and the layer-resolved training analysis).
What the work claims
This is a primary result combining experiment and simulation: two physical systems measured in the lab, one generic oscillator network simulated, all evaluated on standard image-classification benchmarks.1 A physical neural network (PNN) is a real physical system, here an optical fibre or a laser cavity, whose natural dynamics transform an injected input into a high-dimensional response from which a simple linear readout extracts an answer; the physical system is the hidden layer, and its appeal is that it may execute transformations at speeds and energies a digital simulation of them cannot afford. The field's chronic problem is evaluative: performance can only be known after a full training procedure, which for physical hardware is slow, and which gives no transferable figure of merit for comparing substrates.
The paper's claim is that the Hotelling Trace Criterion (HTC), a classical multivariate statistic normalized as the trace of the inverse within-class scatter matrix times the between-class scatter matrix, is that missing figure of merit. HTC is computed from the system's response to labeled inputs alone, with no training. Across all three substrates and both benchmark tasks, classification loss follows a power law in HTC, the data collapsing onto one curve per task rather than one per system: Pearson correlation coefficients of -0.991 for MNIST and -0.969 for fashion-MNIST, with fitted exponents of -0.47 and -0.29 respectively. Once the exponent is established from a few trained calibration systems, subsequent performance predictions require only HTC measurements. A second claim is diagnostic: tracking HTC layer by layer through training of a three-hidden-layer oscillator network shows the first layer saturating and then declining, deviating from the power law that the second and third layers follow, revealing an underutilized layer invisible to the global loss; the corresponding accuracy gain from adding the third layer is marginal, 96.3% to 96.8%.
How it works
HTC measures class separability of a representation: between-class scatter over within-class scatter, summed over discriminants. Intuitively it asks how far apart the class centroids sit in the system's response space relative to how spread out each class is around its own center. A PNN whose dynamics fold all inputs onto nearly the same state has tiny between-class scatter and is useless regardless of how many physical degrees of freedom it nominally possesses, and because HTC is a property of the measured responses, it can be evaluated for any physical system that can be probed with labeled data, which is what makes it substrate-agnostic.
The experimental platforms are deliberately heterogeneous. The first injects encoded images as weak adjustments to femtosecond optical pulses propagating in highly nonlinear fibre, where spectral broadening and nonlinear mixing generate the hidden-layer response. The second injects an image-bearing optical field into a large-area vertical-cavity surface-emitting laser, whose transverse lasing modes interact nonlinearly in a single optical transit. The third is a simulated network of 50 coupled nonlinear oscillators per layer, three hidden layers deep, trained by error backpropagation on the full 70,000-image MNIST set, serving as a generic substrate-independent reference. Sweeping a physical control parameter, input pulse power in the fibre case, moves both HTC and classification loss along the same power-law line, and the fitted exponent agrees across experimental and simulated realizations of the same task even where absolute experimental-simulation agreement was only partial, supporting the authors' interpretation that the exponent is a property of the task, not the substrate.
Where a skeptic should push
The load-bearing assumption is that separability in the measured response predicts separability after the readout, which for the linear readouts studied here is close to definitional. HTC is computed on the hidden-layer responses, and the trained component in these systems is precisely a linear map from those responses to class labels; for linear readouts, high response-space separability is almost the thing being scored. The bolder claim, that the power law extrapolates to trained nonlinear readouts or to other task classes, such as regression or temporal prediction, is asserted with a citation to related work, not demonstrated. Anyone applying HTC to a system with a nonlinear decoder should treat the exponent as recalibration-required.
Second, the benchmarks are MNIST and fashion-MNIST, image tasks routed into optical and oscillator substrates. The collapse of three systems onto one line is striking, but it is collapse across variations within a benchmark family, not across task types, and the two tasks differ enough to change the exponent from -0.47 to -0.29. Task-dependence of the exponent is the finding and the limitation at once: the law tells you less than it appears to, since the exponent itself must be established per task, and the paper offers no way to predict it from the task, only to fit it from calibration systems.
Third, the correlation numbers, while excellent, come from log-log fits over parameter sweeps whose points are not independent: repeated measurements of one physical device at many settings dominate the fibre and laser datasets, so the effective sample size behind a Pearson coefficient of -0.991 is smaller than the point count suggests. Accuracy is a thresholded quantity and noisier than the mean-squared-error loss used for fitting, so headline accuracy predictions inherit that noise. None of this overturns the mechanism, but it calibrates how much weight a single fitted exponent should carry.
Finally, the layer-wise analysis rests on one trained network of one architecture. The underutilized-first-layer result is suggestive and mechanistically plausible, matching known diminishing returns of depth on MNIST, but a 96.3% to 96.8% gap is small, and generalizing the diagnostic would need its own evidence.
Pricing the readout into physical computation
The source never mentions electrodes, organoids, or electrophysiology; what follows is this analysis's extrapolation, grounded in the paper's mathematics. A microelectrode array recording from living neural tissue, with evoked responses collected as spike features or field potentials, is structurally a PNN readout problem: the tissue is the physical system, the array's acquisition chain is the probe, and any decoder built on the recorded features is a readout trained on a transformed, noisy version of the physical state. The paper's framework then prices the recording chain explicitly, because within-class scatter is exactly where front-end noise lands. Input-referred amplifier noise, electrode drift, sampling jitter, and spike-sorting errors all widen the cloud of repeated responses to the same input class, raising the within-class scatter term that sits in HTC's denominator. The power law converts that into a quantitative claim an instrumentation engineer can use: every microvolt of input-referred noise is not merely a signal-integrity statistic but a tax on any downstream computation consuming the recording, moving the operating point down a task-determined curve whose exponent can be measured with a handful of trained calibration runs.
That yields two concrete uses. The first is benchmarking without torture. Training a decoder on electrophysiology is expensive in a way the paper's optical calibration is not: recordings are hours long, biological state drifts across sessions, and repeated training runs consume preparation lifetime. A no-training separability metric computed on recorded responses would let an array platform compare stimulation encodings, electrode subsets, or amplifier configurations before committing to full decoder training, and would give stimulation hardware a design objective, choose the encoding that maximizes separability of evoked states, that is measurable in closed loop. The second use is hype correction with a mechanism attached: physically distinct substrates collapse onto one task-determined curve, so elegance of substrate does not move you off the curve, only along it. Applied to claims about computing on living neural tissue, the discipline is sharp. An organoid-plus-array system with thousands of electrodes but a response state that collapses onto a low-dimensional manifold offers no more classification capacity than the effective dimensionality of that manifold, whatever the channel count says, and a vendor or lab reporting accuracy anecdotes without a fitted scaling exponent is reporting a point, not a scaling law.
The genuine threat cuts in the opposite direction as well. If separability, not channel count, is the currency, the push toward ever denser arrays is exposed to a specific failure mode this paper names directly: nominal dimensionality high, effective dimensionality collapsed, capacity no better than a far smaller system. An acquisition chain that pours bandwidth into channels whose responses are redundant spends power and data movement on nothing the computation can use, and the HTC framing gives a measurement, not a slogan, for saying so. The dual-use note is short: a separability-optimal stimulation encoding is also a separability-optimal perturbation of the tissue, so closed-loop encoding search needs the same duty-cycle and charge-density limits as any stimulation protocol, because the metric it optimizes contains no biological cost term. The practical caveat is calibration: the exponent must be fitted per task with trained calibration systems, and on biological preparations that calibration is exactly the expensive step the framework is meant to amortize, so the saving is real only for platforms that benchmark many hardware variants against a stable task.
The bottom line
Established: across two measured and one simulated physical neural network, classification loss follows a power law in the Hotelling trace criterion, with per-task exponents, correlations of -0.991 and -0.969 on MNIST and fashion-MNIST, and experimental-simulation slope agreement supporting the substrate-independence of the exponent. Hypothesis: that the law extends to nonlinear readouts, other task classes, and far-from-optical substrates, including electrophysiological ones; the paper demonstrates none of these. For array instrumentation, the durable content is a measurement philosophy: separability of the recorded state, computable without training, is the figure of merit that ties front-end noise and encoding design to downstream compute accuracy, and channel count is not. What would confirm the transfer: an HTC-versus-accuracy power law fitted on array-recorded evoked responses with a linear decoder, plus a demonstration that the fitted exponent predicts accuracy for a held-out hardware configuration. What would break it: failure of separability to predict readout accuracy once the decoder becomes nonlinear, or a strong dependence of the exponent on the recording hardware rather than the task.
Frequently asked questions
What is a physical neural network?
A real physical system whose natural dynamics transform an injected input into a high-dimensional response, used as the hidden layer of a neural computation. A simple readout, typically linear, is trained on the responses. The physical system does the heavy transformation; the appeal is speed and energy efficiency.
What is the Hotelling Trace Criterion?
A classical multivariate statistic: the trace of the inverse within-class scatter matrix multiplied by the between-class scatter matrix. It measures how far apart class centroids sit in a representation relative to the spread of each class around its own center, and here it is computed from raw system responses with no training.
Which physical systems were tested?
Three: femtosecond optical pulse propagation in highly nonlinear fibre, image-bearing optical injection into a large-area vertical-cavity surface-emitting laser, and a simulated network of 50 coupled nonlinear oscillators per layer with three hidden layers, trained by backpropagation on the 70,000-image MNIST set.
How strong is the claimed scaling law?
Classification loss follows a power law in HTC with fitted exponents of -0.47 for MNIST and -0.29 for fashion-MNIST, and Pearson correlations of -0.991 and -0.969. Data from the physically distinct systems collapse onto one curve per task, and experimental and simulated points share the same exponent per task.
What does this mean for microelectrode arrays?
Within-class scatter, the denominator of HTC, is where recording-chain noise lands. Front-end noise, drift, and sampling error therefore tax any computation consuming the recording, and a no-training separability metric can compare amplifier settings, electrode subsets, and stimulation encodings before committing to full decoder training, and can expose when dense arrays deliver high nominal but low effective dimensionality.
What are the limits of applying it to biological substrates?
The exponent must be fitted per task using trained calibration systems, and calibration on biological preparations is itself expensive. The law is demonstrated for linear readouts on image benchmarks; nonlinear decoders, other task types, and electrophysiological substrates are extrapolations the paper does not test.
References
- Ermolaev, Hary, Skalli, Genty, Gebski, Czyszanowski, Reitzenstein, Lott, Dudley, Brunner. Power law scaling for classification accuracy in physical neural networks. arXiv:2606.31588 (cs.ET), 2026. https://arxiv.org/abs/2606.31588. Accessed 2026-09-18.