Research analysis · Acquisition chain

Two hundred thousand fixations held still by engineering, not luck

A new MEG dataset follows five participants through ten sessions each and over 200,000 self-guided fixations, and its authors get away with it because they treat head position, eye-tracker drift and ocular artefact rejection as measurement problems with reported error distributions. That discipline is exactly what chronic microelectrode array recording lacks, and most of it transfers.

Source: Active Visual Semantics: A large-scale MEG and eye-tracking dataset for understanding visual intelligence in action, Sulewski, Amme, König, Hebart and Kietzmann, arXiv:2609.01055, 2026. Primary source. Read: the full 24-page preprint PDF, including the acquisition methods, preprocessing pipeline and the head-position precision analysis.

What the work claims

This is a data-resource paper: a primary dataset with quality-control analyses, not a hypothesis test. The authors record magnetoencephalography and eye movements while five participants freely explore 4,080 natural scenes over ten sessions each, producing 203,356 fixation epochs during scene viewing plus 61,624 more during a verbal captioning task, with no enforced central fixation and no stimulus repetitions to average over.1

The claim that matters technically is that active, self-timed viewing, which destroys the two props conventional MEG analyses rely on, randomised onsets and repeated stimuli, can still yield single-trial signal, provided the acquisition chain is engineered to hold the sensor-to-source geometry constant and to tie every artefact decision to an independent measurement. Using encoding models fed by ResNet50 image features, they show a single fixation's target is decodable from the gradiometer signal, peaking at 114 milliseconds with a mean correlation of 0.045 across subjects, and rising to about 0.2 at posterior sensors. Using representational similarity analysis, they show category-level structure that is highly reliable across participants, with an inter-subject noise ceiling of 0.838 and alignment of 0.620 with mid-level network features at the same 114 millisecond peak, in both sensor and source space.1

How it works

The instrumentation is a 306-channel whole-head MEG system with 102 magnetometers and 204 planar gradiometers, sampled at 1 kHz with an online bandpass of 0.1 to 330 Hz, recorded alongside an EyeLink 1000 eye tracker also at 1 kHz. Five head-position indicator coils track the head continuously, and the head shape is digitised before each session with a Polhemus FASTRAK system. The engineering centerpiece is mechanical: each participant's head is immobilised by a foam cast milled from a 3D head scan to fill the gap between head and helmet, allowing natural eye movement while suppressing head movement. Measured from the coil recordings, between-session repositioning error is 1.92 mm, 2.07 mm and 2.87 mm along the three axes, and within-session error across blocks is 1.20, 1.11 and 1.75 mm. The authors benchmark this against the 4 to 5 mm typical of conventional surface- or fiducial-based co-registration, and against the movement thresholds that trigger re-acquisition in compensation protocols.1

The artefact strategy is equally instrumented rather than judgmental. The eye tracker is calibrated on a nine-point grid at the start of each session and after every break, achieving a mean calibration error of 0.268 degrees, and every scene presentation is gated on a successful drift correction against the fixation cross, with a median correction magnitude of 0.440 degrees and 92.9 percent of corrections below one degree. Ocular and cardiac components are then rejected per session by independent component analysis, but the rejection criterion is correlation with the measured gaze channels, not component morphology: components in the top five percent of absolute correlation with horizontal or vertical gaze are removed, which discards an average of 7.8 ocular components per session. The neural preprocessing runs temporal signal-space separation with movement compensation, a 0.2 to 200 Hz bandpass, downsampling to 500 Hz, and epochs of minus 500 to plus 800 milliseconds around each fixation onset, with no baseline correction.1

The analysis chain converts this hygiene into statistics. Per-channel ridge regressions map ResNet50 embeddings of each fixation target onto the fixation-locked gradiometer response; representational similarity analysis over 171 object categories, benchmarked against an inter-subject noise ceiling rather than against chance alone, localises the reliable structure with a 20 mm geodesic searchlight in source space reconstructed with individual boundary-element forward models. The posterior-to-anterior gradient in the results, from a noise ceiling of 0.66 in early visual cortex down to 0.50 in frontal eye field and dorsolateral prefrontal cortex, tracks where visual category information actually lives, and the model alignment of 0.33 to 0.37 in posterior regions against 0.13 to 0.14 in frontal regions is only interpretable because the cross-session geometry was stable enough for source-level comparison.1

Where a skeptic should push

The most load-bearing number in the paper is also its weakest: five participants. The design is an intensive within-subject one that trades population inference for per-fixation power, following the same logic as other densely sampled resources, but it means every group-level claim, including the tidy reliability gradient across cortical regions, rests on n = 5 with a narrow age range. Effect sizes here describe these five brains, and the authors say as much in their limitations.1

Second, the headline single-trial encoding correlation of 0.045 is small. The authors' defence is that it is a lower bound, from a single linear model per channel on un-averaged, never-repeated events, and that conventional paradigms roughly double encoding performance by averaging repetitions. That is fair, but a skeptic should still treat the single-epoch result as evidence that information survives, not as evidence that it is strong; and the elevation of representational alignment already at fixation onset, about 0.408 for the best model layer, warns that temporal smear between adjacent fixations is real and partially contaminates the category structure being celebrated.1

Third, MEG is a volume-conduction measurement: the sensors are centimetres from the sources, the signals are field projections, and ocular artefacts are large, stereotyped and peri-ocular. The gaze-anchored ICA strategy is well suited to exactly this physics. Transferring its logic to other modalities is not automatic, because what counts as an independent reference channel depends on where the artefacts actually couple in.

The array cannot cast its tissue in foam

The microelectrode array faces the same problem as this MEG study, one level harder. Every claim that spans more than one recording day, which is to say every chronic claim an array makes, depends on knowing where each electrode sits relative to each source neuron across time. The AVS solution was to spend its engineering budget before any data existed: immobilise the head mechanically, sense its position continuously with coils, digitise the geometry, gate every trial on a measured drift correction, and validate artefact rejection against the one instrument that independently watches the artefact source. Between-session repositioning error under 3 mm is not a preprocessing achievement; it is a mechanical and metrological one that preprocessing then inherits.1

An array cannot do the mechanical half of this, and that is the non-obvious point. Tissue in chronic implants and organoid cultures grows, softens, migrates and flows; unlike a skull in a helmet, it cannot be cast in foam, and its motion is sometimes the signal, as when organoid development or electrode migration changes which neurons a channel sees. So the MEG pattern inverts: where MEG froze the source and tracked the sensor, an array must freeze neither and instead treat relative geometry as a continuously measured covariate. The closest existing instrument to the AVS head coils is impedance tomography at the electrode array, which can in principle watch the interface drift; what the field mostly does instead is re-cluster spikes after the fact and call the result the same units, which is the array equivalent of assuming the head never moved and blaming the analysis.1

The artefact lesson transfers more directly. The AVS pipeline rejected components by their measured correlation with the gaze channels, at a fixed, pre-registered cutoff, and reported the rejected components' topographies as validation, including cardiovascular ones. Arrays are full of artefacts with known physical causes: stimulation recharge transients, reference drift, pump and incubator cycling, temperature steps. Tying rejection decisions to a measured reference channel, a thermistor trace, a stimulation current log, an optical motion readout, is what turns artefact handling from morphology-based judgment, where every lab quietly diverges, into a reportable measurement with an error distribution, like the 0.268 degree calibration error and the 92.9 percent sub-degree drift corrections reported here.1

There is also a subtler acquisition-chain point in the numbers. The payoff of all this discipline is that single, never-repeated events carry usable signal: one fixation, seen once, decoded at a 114 millisecond peak. The field habitually buys single-trial capability with more channels or more averaging; this paper bought it with geometry stability and artefact hygiene on a sensor whose physics are far less kind than an electrode's. If a centimetre-scale, volume-conducted, ocular-artefact-ridden measurement can be held to sub-3 mm cross-session geometry, the near-field electrode interface, microns from its sources, has no excuse for treating day-to-day drift as unmeasured noise. The honest baseline for the next generation of chronic array papers is to report their equivalent of the repositioning-error table: a measured, session-by-session number for electrode-to-tissue geometry, not an assumption of it.1

The bottom line

Established: with individualised mechanical stabilisation, continuous position sensing, gated drift correction and gaze-anchored artefact rejection, cross-session MEG source analysis holds up under fully naturalistic active viewing, and the dataset's quality-control numbers are reported with the honesty of an instrumentation paper. Asserted, not yet established: that the same discipline, inverted for a moving source, would recover unit identity in chronic array recordings; nobody has yet built the electrode-side equivalent of the HPI coil. What would confirm it: chronic array studies that report measured electrode-to-tissue geometry per session and show their cross-session statistics improve when that covariate is included. What would break the transfer: evidence that interface drift in tissue is so fast or so nonlinear that no continuous electrical or optical proxy can track it, in which case the foam-cast era of array analysis is over before it began and only within-session claims are defensible.1

Frequently asked questions

What is the Active Visual Semantics dataset?

It is a simultaneously recorded MEG and eye-tracking dataset in which five participants freely explored 4,080 natural scenes across ten sessions each, yielding over 200,000 self-guided fixation epochs, per-fixation object labels, a semantic captioning task on 25 percent of trials, and structural MRI scans for source reconstruction.

How did the authors keep head position stable across ten sessions?

Each participant wore an individualised foam cast milled from a 3D head scan that filled the space between head and MEG helmet. Continuous head-position indicator coils then measured repositioning errors of 1.92 to 2.87 mm between sessions and 1.11 to 1.75 mm within sessions, well below the 4 to 5 mm of conventional co-registration.

Why does artefact rejection use the eye tracker?

Independent component analysis finds signal components but cannot say which are artefacts. The authors rejected the components most correlated with the measured gaze channels, an average of 7.8 per session, so every rejection decision is anchored to an independent physical measurement rather than to component shape, which is a judgment call.

What does single-trial encoding mean here?

A linear model predicts the MEG response to an individual fixation from deep-network features of the fixated image patch, with no averaging across repetitions because none exist. The mean correlation peaks at 114 milliseconds and is modest, about 0.045 across subjects and up to about 0.2 at posterior sensors, which the authors frame as a lower bound on what the dataset contains.

What does this have to do with microelectrode arrays?

Chronic array recordings face the same cross-session identity problem: which electrode sees which neuron today versus next week. The MEG solution combines mechanical immobilisation, continuous position sensing, drift-gated acquisition and reference-anchored artefact rejection. An array cannot immobilise living tissue, so it must invert the pattern and measure electrode-to-tissue geometry continuously instead of assuming it.

What are the limits of this paper?

Five participants with an intensive within-subject design limit population-level inference; the single-trial effects are small; and some representational structure is already present at fixation onset, indicating temporal overlap between neighbouring fixations. The dataset is a resource for active-vision research, not a benchmark of encoding-model quality.

References

  1. P. Sulewski, C. Amme, P. König, M. N. Hebart, T. C. Kietzmann. Active Visual Semantics: A large-scale MEG and eye-tracking dataset for understanding visual intelligence in action. arXiv:2609.01055. 2026. https://arxiv.org/abs/2609.01055. Accessed 2026-09-10.