Pruned analog splines put nonlinear math before the ADC
Researchers at KIT and the University of Albany have built analog Kolmogorov-Arnold networks on a flexible thin-film transistor process and shown that hardware-aware training plus coefficient-level pruning can cut a spline block's area by 55 percent and its power by 50 percent, with approximation accuracy that improves rather than degrades. Everything is simulated; but the recipe, train the algorithm against a circuit-level error model of the imperfect hardware, is the part the mixed-signal world should borrow.
Source: Co-Optimization of Analog Kolmogorov-Arnold Networks for Low-Power Function Approximation in Flexible Electronics, IEEE Journal on Emerging and Selected Topics in Circuits and Systems, September 2026 (arXiv:2606.27892). Primary source. Read the full arXiv HTML version, including Table II and the PVT sweep.
What the work claims
Duarte, Zervakis, Tahoori, and Nassif claim a framework, Analog Kolmogorov-Arnold Networks (AKANs), for approximating nonlinear multivariate functions directly in the analog domain on flexible electronics. The target class of functions is exactly the awkward, transcendental workload that sensor front ends carry: neural network activation functions, sensor calibration curves that map raw readings to physical units, and preprocessing such as logarithmic compression and power operations for feature extraction. Doing these in analog avoids the energy of converting everything to digital first and then computing there.1
The method has two halves. During training, the loss sees a circuit-level error model derived from transient SPICE simulations of unipolar IGZO (indium gallium zinc oxide) thin-film transistor circuits, so the learned spline weights are compensated for the hardware's systematic non-idealities in advance. After training, pruning operates on the polynomial coefficients of each quadratic spline, and each removed coefficient corresponds to the physical removal of a circuit block, a multiplier, a squarer, or an adder, from the IGZO implementation. Across four benchmark datasets the authors report average area and power reductions of 29.8 and 30.5 percent, best-case 55 and 50.3 percent, and a normalized mean percentage error (NMPE) that improves under pruning because sparsifying spline coefficients regularizes the fit.1
How it works
The Kolmogorov-Arnold representation theorem says any continuous multivariate function can be written as a composition of additions and single-variable functions. A KAN is that construction turned into a network: instead of fixed activation functions on weighted sums, the learnable elements are univariate spline functions on the edges. Here each spline is quadratic, parameterized by three coefficients, k0, k1, and k2, which map one-to-one onto analog hardware: k1 multiplies, k2 squares, k0 adds. That mapping is why pruning is hardware-meaningful in a way it rarely is: zeroing k2 does not merely sparsify a tensor, it deletes a squarer from the schematic.1
The error model is the engineering core. The team ran 1,000 transient SPICE simulations of the unipolar IGZO circuits across the design space, ranked configurations by NMPE, and derived the representative error curve from the top 30. The 5th to 95th percentile envelope of that calibrated subset spans at most 189.82 mV across the input range, and the curve is stable to the selection threshold, deviating less than 27 mV from the top-30 model whether the top 10, 50, or 100 simulations are used. In the top-30 set, NMPE values span minus 13 to plus 13 percent, with more than half between minus 5 and plus 5. This error curve is baked into training, so the network learns weights whose errors cancel the circuit's systematic deviations rather than suffering them.1
The hardware numbers are modest but honest. One unpruned spline block occupies 45,570 square micrometers and draws 241.8 microwatts at 1.0 V, 27 degrees C, at the typical-typical corner. Aggressive pruning of the k1 or k2 coefficients brings the block to 20,536 square micrometers and 120 microwatts, reductions of 55 and 50.4 percent; the intermediate k1-only prune lands at 45.1 and 49.6 percent. Against a 27-corner PVT sweep the approximation holds: within the wearable temperature range of 27 to 40 degrees C the NMPE deviates at most 0.74 percentage points from nominal, and even at 0 degrees C the worst-case NMPE is 2.50 percent. Validation runs on four datasets chosen to mimic flexible-sensor workloads: PPG-DaLiA photoplethysmography, ECG5000 electrocardiograms, household electric power consumption, and the Iris set, through a small 1-to-3-to-1 network.1
Where a skeptic should push
The single most load-bearing assumption is that a representative error model can stand in for the actual device population. The authors are admirably explicit about the gap: a full Monte Carlo mismatch and yield analysis is named as future work, because local device mismatch, random spatial variation in threshold voltage and mobility, perturbs each spline circuit independently in ways the systematic PVT model does not capture. Everything else in the paper, the pruning gains, the accuracy improvements, the temperature robustness, is computed inside the model's own loop. When a result shows that pruning improves accuracy, one should ask how much of that improvement is a regularization effect of the training pipeline itself, since the same pipeline defines both the accuracy metric and the hardware error it claims to beat. The finding is best read as internal consistency of a co-design methodology, not as a measured device property.
Two quantitative cautions follow. First, the ±13 percent NMPE tail of the calibrated configuration set is the operational boundary, and the paper itself notes that downstream effects near that edge remain unexamined; only the ±5 percent region has validated negligible impact on the chosen tasks. Second, the headline savings are block-level. The paper is careful to say the system-level impact depends on what fraction of the chip the splines occupy, and the absolute figure, 241.8 microwatts per unpruned block at 1.0 V, is not small at array scale: a naive per-channel deployment across a few hundred channels would land in the tens of milliwatts before the ADCs and amplifiers are counted. The wins are real for sparse preprocessing duty cycles, not for continuously running one block per channel. There is also no aging or bias-stress drift data; the PVT sweep covers corners, not months of operation, which is the timescale that matters for implanted or long-duration recording hardware.1
What pre-ADC analog splines ask of the MEA chain
For microelectrode array hardware, the paper sits in the middle of the argument that decides how far arrays can scale. A high-density MEA acquisition chain is, per channel, an electrode, a low-noise amplifier, a filter, an ADC, and a digital backend, and the channel count is bounded by the aggregate ADC bandwidth and data rate, not by electrode pitch. Any nonlinear preprocessing that happens before the converter, log compression to tame the spike dynamic range, power and energy features, per-electrode calibration that maps the node voltage to a physical quantity, reduces what must be digitized and streamed. That is precisely the function class this paper implements in analog, and its significance is not the 241.8 microwatt block; it is a demonstrated methodology for making an analog preprocessor whose accuracy is specified against the imperfections of the actual process it will be fabricated in.
The non-obvious implication is that device mismatch stops being purely a yield problem and becomes a training problem. Mixed-signal MEA front ends have always paid for mismatch twice: once in trimming and calibration circuitry, once in channels discarded at test. The AKAN recipe inverts the logic: characterize the process's error distribution in simulation, then train the algorithm so that its weights absorb the systematic part. Pruning that deletes physical blocks then doubles as a per-device calibration knob, since a device whose measured error curve differs from nominal can be re-pruned to a different coefficient set. If that retraining loop is cheap enough, the front end ships uncalibrated and calibrates itself in software against a measured error curve. No one has demonstrated this loop end to end; the paper's framework is the first half of it, and the missing Monte Carlo analysis is exactly where the second half would fail or succeed.
The genuine opportunity, then, is for conformable arrays. Flexible thin-film front ends are the plausible route to electrodes that wrap three-dimensional organoid surfaces rather than plate them, and this paper shows that the same process can host meaningful preprocessing, not just routing and amplification. The genuine threat is quieter. If analog preprocessing absorbs the feature extraction and compression layer, the value in the acquisition chain migrates away from ADC and DSP vendors toward whoever owns the co-design toolchain and the calibration loop, and array makers who treat the analog front end as a dumb pipe risk commoditization. The second threat is epistemic: acquisition chains whose accuracy claims rest on simulation-internal error models will fail in exactly the way this paper concedes it has not yet tested, when local mismatch and long-term drift sit on top of microvolt-class neural signals. For an instrument whose entire claim is measurement fidelity, an unverified ±13 percent tail in the preprocessor is not a rounding error; it is a systematic error source upstream of everything downstream.
The bottom line
Established, by simulation: a hardware-software co-optimization flow in which IGZO TFT circuit errors derived from 1,000 SPICE runs are modeled during KAN training; coefficient pruning that physically removes multiplier, squarer, and adder blocks; block-level savings up to 55 percent area and 50.4 percent power, with 29.8 and 30.5 percent averages across four sensor datasets; NMPE that improves under pruning; and PVT robustness within 0.74 percentage points across the 27 to 40 degrees C wearable range. Not established: any measured silicon, any Monte Carlo mismatch or yield result, any aging behavior, or any system-level energy figure for an array of channels. For MEA instrumentation the paper is a portable recipe with an honest list of its own missing verification steps. What would confirm it is a fabricated flexible spline block whose measured error curve lands inside the simulated envelope and whose mismatch spread matches the Monte Carlo analysis the authors defer. What would break it is local device mismatch dominating the systematic PVT error the training absorbs, in which case per-device retraining costs would eat the area and power the pruning saved.
Frequently asked questions
What is a Kolmogorov-Arnold network?
A network architecture based on the Kolmogorov-Arnold representation theorem, which states that any continuous multivariate function can be decomposed into additions and univariate functions. In a KAN the learnable elements are those univariate functions, implemented as splines, rather than fixed nonlinearities applied to weighted sums. In this work each spline is quadratic with three coefficients that map directly onto analog circuit blocks: a multiplier, a squarer, and an adder.
What does the pruning actually remove?
Pruning operates on the spline coefficients k0, k1, and k2, and because each coefficient corresponds to a physical circuit block, removing a coefficient means deleting that block from the IGZO schematic. The most aggressive configuration, pruning k1 and k2, cuts a block from 45,570 square micrometers and 241.8 microwatts to 20,536 square micrometers and 120 microwatts, savings of 55 and 50.4 percent. Averaged across the four benchmark datasets, pruning saves 29.8 percent of area and 30.5 percent of power.
Why does pruning improve accuracy instead of hurting it?
The authors attribute it to regularization: zeroing small spline coefficients constrains the spline shapes and prevents overfitting, so the normalized mean percentage error drops relative to the unpruned baseline, by about 1.1 percentage points on average across pruning configurations. A skeptic should note that the accuracy metric and the hardware error model come from the same simulation pipeline, so part of this improvement may be a property of the training loop rather than of any physical device.
How is the hardware error modeled during training?
From 1,000 transient SPICE simulations of the unipolar IGZO thin-film transistor circuits. Configurations are ranked by approximation error, the top 30 by NMPE define a representative error curve, and the 5th to 95th percentile envelope of that subset spans at most 189.82 mV across the input range. The curve is stable to subset size, deviating less than 27 mV from the top-30 model whether the top 10, 50, or 100 simulations are used, and it is folded into training so learned weights compensate systematic circuit error in advance.
What does this have to do with microelectrode arrays?
Channel count in a high-density MEA is bounded by aggregate ADC bandwidth and data rate, so any nonlinear preprocessing done before the converter, such as log compression, power features, or calibration curves, shrinks what must be digitized. The paper supplies a method for building such analog preprocessors with accuracy specified against the fabrication process's own errors, on flexible substrates that could one day wrap three-dimensional organoid surfaces instead of plating them.
What is the biggest unresolved risk?
Local device mismatch. The authors state that a full Monte Carlo mismatch and yield analysis is future work, and mismatch, random variation in threshold voltage and mobility, perturbs each circuit independently in ways the systematic PVT error model does not capture. For microvolt-class neural recording, an unverified approximation error tail of up to ±13 percent sitting upstream of the ADC is a potential systematic error source, and long-term bias-stress drift of the thin-film transistors is not characterized at all.
References
- P. Lozano Duarte, G. Zervakis, M. B. Tahoori, S. R. Nassif. Co-Optimization of Analog Kolmogorov-Arnold Networks for Low-Power Function Approximation in Flexible Electronics. IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 2026 (arXiv:2606.27892). https://arxiv.org/abs/2606.27892. Accessed 2026-09-25.