Research analysis · Signal integrity

The averaging step buried a real burst phenotype on the array

A knockout that looked electrophysiologically silent in the standard per-well summary turned out to have a clear, right-shifted burst-duration distribution once the raw events were pooled. The finding is a biology result, but the sharpest reading for hardware is that the data-reduction stage of a microelectrode-array pipeline is not a neutral convenience: it decides what reaches the analyst.

Source: ST3GAL3 loss-of-function disrupts synaptic integrity and excitatory/inhibitory cortical dynamics, bioRxiv preprint, posted 2026-06-25. Primary source. Read in full, including methods, MEA results, supplementary statistics and discussion.

What the work claims

This is a primary result. The authors made an isogenic pair of human induced pluripotent stem cell (iPSC) lines, identical except that one carries a CRISPR/Cas9 knockout of ST3GAL3, a sialyltransferase (an enzyme that attaches sialic acid sugars to proteins) whose loss is linked to intellectual disability and infantile epilepsy. They differentiated both lines into cortical neurons by two independent routes, a directed protocol read at day in vitro (DIV) 30 and an induced Ngn2/Ascl1 protocol read at DIV 49, then characterised the networks with microelectrode-array (MEA) recording and RNA sequencing.1

The headline is that losing ST3GAL3 disrupts the excitatory/inhibitory balance of the network, and that the functional signature of this is aberrant bursting, specifically prolonged and more variable bursts, backed by broad downregulation of glutamatergic and GABAergic synaptic genes. What makes the paper interesting from an instrumentation standpoint is not the headline itself but the road the authors had to travel to reach it, because the effect was absent in the first, most conventional way they looked at their own MEA data.

How the measurement was made, and where it split

The recordings used a Multi Channel Systems MEA2100 with six-well chips carrying titanium electrodes on a 200 micrometre pitch, roughly ten electrodes per well, sampled at 10 kHz for 20 minutes of spontaneous activity at 37 degrees Celsius. Spikes were detected offline and reduced to the usual family of network metrics: spike rate, percentage of random (non-burst) spikes, mean burst duration, mean spikes per burst, mean inter-burst interval and mean burst frequency.

Analysed the standard way, by averaging each metric over time and electrodes to one number per well and comparing genotypes, the knockout and wild-type were indistinguishable. Principal component analysis showed no separation, and a PERMANOVA (a permutation test for multivariate group differences) returned R squared of 0.031 with p of 0.682, with comparable within-group dispersion. On the per-well means the experiment was, to a first approximation, negative.

The signal appeared only when the authors stopped averaging. Pooling individual burst events across the dataset, the knockout produced a visibly broader and right-shifted burst-duration distribution. Stratified into percentile bins, medium-length bursts fell from 67 percent of events in wild-type to 45 percent in knockout, while long bursts rose from 7 percent to 17 percent and very long bursts from 2 percent to 13 percent. A related count, the fraction of electrodes that detected any bursting, was 35.4 percent for knockout versus 17.4 percent for wild-type. The same data reduced two different ways gave two different answers.

Where a skeptic should push

The load-bearing question is whether the event-level difference is a genotype effect or an artefact of how events were pooled. Pooling every burst into one distribution is pseudoreplication: bursts from the same well are not independent, and the knockout arm carried two extra replicate wells, which alone inflates its event count and its tail. The authors are aware of this and did the right thing, fitting a Bayesian multinomial regression with a random intercept per well, which is the correct way to stop well-level structure from masquerading as an effect. But the eye-catching raw proportions (17 percent versus 2 percent very long bursts) are the pseudoreplicated numbers, not the modelled effect sizes, and should not be read as the magnitude of anything.

Two internal signals argue for caution. First, the authors report a near-perfect correlation between spikes-per-burst and burst duration (r of about 0.906) and correctly call it an artefact of the burst-detection settings, since a longer detected burst mechanically contains more spikes. That is a warning that the burst metrics are strongly shaped by detector parameters, so a distributional shift in burst duration can partly reflect how the detector interacts with a slightly different firing regime rather than a distinct class of biological burst. Second, the burst-detecting-electrode fraction, 35 versus 17 percent, is only half a biological quantity; the other half is coverage and coupling, that is, how many of the sparse electrodes happened to sit under active tissue, which is a property of plating and the array, not of the genotype. The authors themselves label the hyperexcitability read inconclusive and note that per-well analyses stayed non-significant, and they call, rightly, for more sophisticated analytics than mean comparisons. I would go one step further: a reporting inconsistency in the correlation interval they print (a confidence interval whose lower bound is negative for a positive r near 0.906) suggests the statistics deserve a careful second pass before anyone builds on the exact figures.

So the honest state of the evidence is: the event-level distribution differs, the per-well means do not, and whether the former survives a balanced, well-as-unit analysis is not yet established. That is a narrower claim than the abstract's, and it is the correct one.

What sparse grids and mean values hide

For anyone building or buying acquisition hardware for living tissue, the useful lesson is that the instrument does not end at the amplifier. Here the limiting stage was neither the electrodes nor the front-end noise; it was the reduction of a 20-minute, ten-channel recording to one mean per well. That step is usually performed inside vendor software and treated as a given, yet it was the element that turned a real distributional difference into a null result. A screening campaign that trusts the per-well dashboard, which is exactly how throughput MEA work is run, would have filed this genotype as no effect. The threat is concrete and directional: false negatives in drug and disease-model screens, produced not by a weak effect but by an aggregation that discards the shape of the distribution the effect lives in.

The second lesson is about electrode yield. On a grid of roughly ten electrodes per well at 200 micrometre pitch, only 17 to 35 percent of electrodes saw bursts, so most channels contributed silence, and the one metric that did separate the groups was itself partly a coverage statistic. This is where the hardware roadmap and the analysis roadmap meet. Higher-density CMOS arrays, with thousands of electrodes at tens of micrometres, do two things at once: they raise the probability that some electrode sits under active tissue, decoupling the yield confound from biology, and they supply enough independent observations to model well-level and electrode-level structure honestly rather than pooling it away. Retaining raw waveforms rather than only on-device summaries is the cheaper half of the same fix, because it lets an analyst recompute event-level statistics the vendor pipeline never exposed.

For organoid work the warning is stronger, not weaker. A three-dimensional organoid couples to a planar array only across a small basal contact zone, so the effective electrode yield is lower and more variable than in the monolayer cultures studied here, and the temptation to summarise a noisy, sparsely sampled recording into a single tidy number is greater. The same study run on organoids over a low-channel planar MEA, read through per-well means, would be even more likely to report nothing while a genuine phenotype sat in the tails. The opportunity is the mirror image: high-density arrays plus distributional, event-level analytics plus raw-waveform retention are precisely the combination that would let an organoid platform detect the kind of effect this paper nearly missed.

The bottom line

Established: in these cultures the per-well mean network metrics do not differ by genotype, the pooled event-level burst-duration distribution is broader and right-shifted in the knockout, and the transcriptomic data show synaptic-gene downregulation. Hypothesis, and labelled as such by the authors: that the knockout network is hyperexcitable. What would confirm it is a pre-registered analysis with balanced replicates that treats the well as the unit, uses a mixed-effects model, and ideally a high-density array with matched electrode yield; if the event-level shift survives that, it is real, and if it collapses, it was pseudoreplication and coverage. For hardware the conclusion is independent of how the biology resolves: the acquisition chain includes its data-reduction stage, and on this evidence that stage, not the analog front-end, was the part that decided what could be seen.

Frequently asked questions

Did the microelectrode array fail to detect the effect?

No. The electrodes and amplifiers captured the activity; the effect was lost at the analysis stage, when 20 minutes of multi-channel data were averaged to a single value per well. The same recordings revealed the difference once individual burst events were examined instead of well means.

Why do per-well averages behave so differently from pooled events?

Averaging collapses the distribution to its centre and discards its shape. The genotype difference here lived in the tails, more long and very long bursts, so a comparison of means was blind to it while a comparison of full distributions was not.

Is the pooled-event result trustworthy on its own?

Only with care. Pooling bursts across wells is pseudoreplication, and the knockout arm had extra replicate wells that inflate its event count. The authors mitigated this with a per-well random intercept in a Bayesian model, but the raw pooled proportions overstate the effect and should not be read as effect sizes.

What does electrode yield have to do with the biological claim?

The most separating metric, the fraction of electrodes detecting bursts, is partly a coverage statistic: it depends on how many sparse electrodes happen to sit under active tissue. On a roughly ten-electrode grid that confound is large, which is why a higher-density array would help isolate biology from coupling.

What is the practical takeaway for MEA screening platforms?

Do not trust per-well summary metrics as the primary readout, retain raw waveforms so event-level statistics can be recomputed, and prefer higher channel counts that supply enough independent observations to model structure rather than average it away. These steps directly reduce the risk of false negatives in disease-model and drug screens.

References

  1. Diouf D, Tsounis DL, Pishva E, Vanmierlo T, et al. ST3GAL3 loss-of-function disrupts synaptic integrity and excitatory/inhibitory cortical dynamics. bioRxiv. 2026. https://www.biorxiv.org/content/10.64898/2026.06.21.733355. Accessed 2026-07-25.