An LLM evolving its own prompts designs leaner multipliers
Tomasovic and Sekanina have built a co-evolutionary loop in which an off-the-shelf large language model designs approximate 8-bit multiplier circuits while a second loop evolves the prompt templates that steer it, no circuit-specific training required. The system found 70 worst-case-error-versus-area trade-offs absent from the 24,912-circuit EvoApproxLib library, 24 of them on the global Pareto front, using one to two orders of magnitude fewer circuit evaluations than classic genetic programming. The multiplier is a small digital block, but it is the most-frequently-executed operation in the neural-signal DSP sitting behind every large-scale recording array, and the method's hidden constraint, exactness when an operand is zero, is suspiciously well matched to what neural data looks like.
Source: Multi-Objective Coevolution of Prompts and Templates for Circuit Approximation, arXiv:2606.13089, to appear at PPSN 2026, Trento. Primary source. Read the full arXiv HTML version, including the method, setup, results, and timing sections.
What the work claims
Approximate computing is the deliberate design of arithmetic circuits that compute inexactly in exchange for area, power, and latency; it pays off in error-tolerant workloads such as neural networks, where multiplications dominate both training and inference. The authors' claim is that a common, untrained LLM can be made competitive with the heavily optimized circuits in EvoApproxLib, a library produced by years of Cartesian genetic programming search, if the prompting strategy itself is evolved rather than hand-written.1
The system coevolves two populations. One is circuits, each an 8-bit unsigned multiplier expressed as a few hundred lines of C-like Boolean expressions; new candidates are generated when Qwen3-Coder-480B, a 480-billion-parameter mixture-of-experts coding model with about 35 billion active parameters per inference, edits an existing circuit according to a template. The other population is the templates themselves, short prompt blueprints with placeholders for target error, target area, and a parent circuit; templates are mutated and crossed over by a smaller model, GPT-OSS-120B, and are scored by how many valid circuits they produce and how close those circuits land to the target specification. Selection in both loops is NSGA-II Pareto ranking via the pymoo library, on the two objectives of approximation error and estimated chip area. In long runs across 32 target error-area pairs, the system produced 70 trade-offs for worst-case error and area that EvoApproxLib does not contain, 24 of them non-dominated against the entire library, and 63 unseen trade-offs for the mean-squared-error variant, 23 on the global Pareto front.1
How it works
The circuit representation is chosen to sit inside the LLM's training distribution: a netlist of readable logic equations rather than Verilog, which is underrepresented in code corpora. Error is measured exhaustively over all 65,536 input vectors by a C program with OpenMP, so no statistical sampling hides a bad corner. Area is estimated, not synthesized, by summing per-gate sizes for a 45 nm technology, an inverter at 1.40 square micrometers, a two-input XOR at 4.69. One fitness constraint does heavy lifting: a candidate that fails to compute the exact product whenever either operand is zero is assigned infinite error. This is a known requirement for neural-network accelerators, where zero activations, as after ReLU clipping, are common, and it silently shapes the entire search.1
The coevolution is what makes an untrained LLM useful. The authors report that both one-shot prompting and a simple hill-climb over mutation templates failed to reach competitive circuits; evolving the templates, at population 8 with mutation and crossover each at probability 0.9, alongside the circuit population of 30 with three LLM calls per template per generation, did. Under a fixed budget of 450 large-model calls per run, coevolution clearly beat hill-climbing on Pareto-front hypervolume, and a deep-generation, few-calls-per-template configuration matched a shallow, many-calls configuration with no statistically significant difference under a Mann-Whitney U test at alpha 0.05. Wall-clock was roughly 207 to 223 minutes per run, because a single Qwen call averages 10.0 seconds against 0.5 seconds for an exhaustive circuit evaluation, so LLM inference, not circuit checking, dominates cost. The comparison point is stark: classical CGP search for a comparable single target can run to a million generations, 7.8 to 147.7 minutes per target on a 2.4 GHz Xeon, and the LLM loop reaches competitive designs with one to two orders of magnitude fewer circuit evaluations.1
Where a skeptic should push
The most load-bearing assumption is that error metrics computed over uniformly distributed inputs predict behavior on real workloads. Worst-case error and mean-squared error over all 2^16 input pairs treat every operand combination as equally likely. Neural data is nothing like that: extracellular recordings are dominated by low-amplitude background with brief, sparse, high-amplitude events, so the input distribution to any downstream fixed-point kernel is heavy-tailed and correlated in time. An approximate multiplier tuned for uniform inputs may spend its error budget exactly where the data is densest, or, worse, may corrupt the quiet baseline where downstream algorithms, spike detectors, covariance estimators, drift trackers, are actually listening. The paper offers no application-level validation, and the authors list a statistical analysis of long-run behavior as future work.
Second, area is an estimate from a gate-size table. There is no synthesis, no place-and-route, no measured power, and no delay numbers at all; the Pareto fronts are drawn in the space of error and estimated area only, so the claimed trade-offs could compress or vanish after physical design, where wirelength and glitching, not gate count, decide multiplier energy. Third, reproducibility: the design operator is a stochastic LLM at temperature 0.4 accessed through shared infrastructure, so the same run cannot be re-executed bit-for-bit, which matters more than it might seem, because for an instrument that must be qualified and documented, a design flow whose intermediate steps are non-deterministic is a regulatory headache even when its outputs are excellent. The honest summary: compelling search efficiency, weak evidence about the metrics a hardware engineer would sign off on.1
Approximate arithmetic and the array back end
For microelectrode array hardware the digital back end is where channel count goes to die. A dense CMOS array can put tens of thousands of electrodes on one headstage, and the moment spike detection, feature extraction, or compression moves on-chip to keep the data rate survivable, every channel owns a stream of multiply-accumulates, and the multiplier becomes the largest repeated digital block in the per-channel budget. Shrinking it is therefore not an abstract optimization: at array scale, multiplier area and energy per channel set how many channels fit under a power cap. This paper matters because it collapses the cost of customizing that block. EvoApproxLib-quality approximate multipliers previously required specialist evolutionary-search infrastructure; here, an untrained LLM plus template coevolution reaches and extends that frontier in a few hundred to a few thousand model calls, which puts approximate-DSP design inside the reach of a small instrumentation team that has an error budget but no circuit-evolution expert.
The non-obvious implication is the zero-operand constraint, and it deserves a second look from anyone recording neural signals. The requirement that the multiplier be exact whenever an operand is zero is, in effect, a requirement that the circuit be bit-perfect in silence. Extracellular recordings spend most of their time near baseline; a DSP back end built from these multipliers would multiply exact zeros exactly, most of the time, and spend its approximation error only during active epochs, where a trained spike classifier is most tolerant of perturbation. That is a genuinely elegant match between a design constraint and the statistics of the signal, and it suggests a broader design philosophy for acquisition hardware: define exactness where the biology is quiet, tolerate error where the biology is loud. Nobody has demonstrated that pairing; the paper hands over both halves.1
The threats are equally concrete. The first is silent corruption of the signal the field cares about most: sub-threshold and local-field-potential structure lives precisely in the low-amplitude regime, and if approximate error concentrates on small operands, a back end tuned on uniform-input error metrics would degrade the quiet dynamics while reporting excellent average error. Any adoption needs workload-aware error evaluation on recorded traces, which this paper does not attempt. The second threat is governance of the design flow itself: stochastic, non-deterministic generation complicates verification, traceability, and regulatory documentation, and a medical-adjacent instrument vendor cannot ship a multiplier whose provenance is a prompt template that no longer exists. For array makers, the opportunity is a cheap, custom, per-product approximate DSP block; the risk is discovering, after tape-out or in the field, that the error metric optimized was the wrong one for tissue.1
The bottom line
Established: a co-evolutionary loop that evolves circuits and LLM prompt templates together, with an untrained 480B-parameter model as the mutation operator, discovers 8-bit approximate multipliers with 70 worst-case-error-area and 63 mean-squared-error-area trade-offs absent from a 24,912-circuit reference library, 24 and 23 of them globally non-dominated, using one to two orders of magnitude fewer circuit evaluations than classical Cartesian genetic programming. Not established: synthesized or measured area, power, or delay; behavior on any real, structured workload; reproducibility of the design flow; or statistical rigor on the long runs, which the authors defer. For MEA instrumentation the takeaway is twofold: approximate multipliers are a legitimate lever on the per-channel DSP power wall, and the exact-on-zero constraint is a template for designing error out of the quiet parts of a neural signal. What would confirm the thesis is an approximate-multiplier DSP pipeline evaluated end to end on recorded extracellular traces, showing detection and decoding metrics survive the approximation. What would break it is error concentrated at small operands, which would quietly poison exactly the sub-threshold information high-density arrays exist to capture.
Frequently asked questions
What exactly is being coevolved?
Two populations in interlocked loops. One is a population of 30 candidate 8-bit approximate multiplier circuits, each written as a list of C-like Boolean expressions, where new circuits are generated by having Qwen3-Coder-480B edit an existing circuit according to a prompt template, with three model calls per template each generation. The other is a population of 8 prompt templates, short blueprints with placeholders for target error, area, and parent circuit, which a smaller model, GPT-OSS-120B, mutates and recombines. Templates are scored by the fraction of valid circuits they produce and how closely those circuits hit the target specification, so the system evolves not just circuits but the instructions for making circuits.
How do the results compare to EvoApproxLib?
EvoApproxLib contains 24,912 unsigned 8-bit approximate multipliers produced by extensive Cartesian genetic programming search. Across 32 target error-area specifications, the co-evolutionary system produced 70 worst-case-error-versus-area trade-offs not present in that library, 24 of them on the global Pareto front, and 63 unseen trade-offs under the mean-squared-error metric, 23 globally non-dominated. The authors attribute the gain to coevolving templates, since one-shot prompting and plain hill-climbing over a fixed template failed to reach competitive results.
What is the zero-operand constraint?
A candidate circuit that fails to compute the exact product whenever either operand is zero is assigned infinite error and discarded. This is a known deployment requirement for neural-network accelerators, because zero activations are common, for example after ReLU clipping. In effect the constraint makes the multiplier exact in silence and allows approximation only when both operands are active, which happens to mirror the statistics of extracellular recordings, mostly quiet baseline with sparse high-amplitude events.
Is the claimed efficiency gain real?
Circuit evaluation is cheap, about 0.5 seconds per exhaustive check over all 65,536 inputs, while each Qwen call averages 10.0 seconds, so LLM inference dominates the roughly 207 to 223 minute runtimes. The efficiency claim is about search: one to two orders of magnitude fewer circuit evaluations than classical CGP, which can consume up to a million generations per target. Whether the total wall-clock or dollar cost is lower depends on LLM inference pricing, which the authors note sat outside their control.
What does this have to do with microelectrode arrays?
In a high-density array, on-chip DSP per channel, spike detection, feature extraction, compression, is a stream of multiply-accumulates, and the multiplier is the largest repeated digital block in the per-channel power and area budget. A method that lets a small team custom-design approximate multipliers against a stated error budget attacks that budget directly, and the exact-on-zero constraint aligns naturally with neural data that is quiet most of the time and active in bursts.
What is the biggest unresolved risk?
The error metrics assume uniformly distributed inputs, so they may not predict behavior on structured, heavy-tailed neural data, and approximation error concentrated on small operands would corrupt the sub-threshold regime that downstream analysis relies on. On top of that, area is estimated from a gate-size table rather than synthesized or measured, there are no power or delay figures, and the stochastic LLM design operator undermines bit-for-bit reproducibility, which matters for verification and regulatory documentation.
References
- M. Tomasovic, L. Sekanina. Multi-Objective Coevolution of Prompts and Templates for Circuit Approximation. To appear at Parallel Problem Solving from Nature (PPSN), Trento, 2026; arXiv:2606.13089. https://arxiv.org/abs/2606.13089. Accessed 2026-09-26.