Research analysis · Signal acquisition

Spike timing deserves better than a bin: log-time encoding at the MEA front end

An event-vision paper from Linkoping University shows that most of what a recurrent network learns about time can be baked into the front-end event representation itself, as overlapping logarithmic temporal basis functions updated in constant work per event. The sensing problem it solves is structurally identical to the microelectrode array readout problem, and the engineering consequences transfer.

Source: Efficient Multi-Timescale Event Representations for Feed-Forward Object Detection, NeVi workshop at ECCV 2026 (arXiv preprint), September 2026. Primary source. Read the full arXiv PDF text, including all result tables.

What the work claims

Event cameras output asynchronous, timestamped brightness-change events at microsecond resolution, and object detectors built on them usually handle time with recurrent hidden state that accumulates events over successive windows. Lundell, Forssén, Wadenbäck and Lundmark instead encode temporal history directly into the input representation: each event is projected onto four overlapping logarithmically spaced B-spline functions of its age, multiplied by a local spatial confidence, and fed to a purely feed-forward detector. Their claim is that carefully designed representations capture much of the temporal information that recurrent models learn, without any sequential state.1

This is a methods and benchmarking paper, not a biological result, and its metrics are detection benchmarks (PEDRo and Gen1 automotive datasets), not neural signals. The claim is bounded: they do not say recurrence is obsolete. They show representation design recovers a substantial, measurable share of what recurrence buys, on fixed feed-forward hardware.

How it works

The baseline they beat is CSTR, a compact three-channel surface of recent events. Their representation assigns each incoming event a temporal age and evaluates that age against four second-order B-splines spanning 1.2 to 35 ms. The logarithmic spacing is the point: resolution near zero age (recent events) is fine, while old history is compressed into few channels. Positive and negative event polarities get separate channels, and a normalized local confidence, computed with a 3x3 binomial kernel over temporally weighted event counts, marks how well supported each temporal response is.1

The hardware-relevant trick is the last stage. Analytical B-splines are expensive to recompute per event, so the authors fit, by least squares, a learned polynomial mapping from a small set of recursively decaying exponential states onto the spline responses. A second-order fit uses 9 coefficients per spline at an approximation error of 1.44 x 10^-5; a fourth-order fit uses 34 coefficients and reaches 1.03 x 10^-6. Because the exponentials can be updated recursively, each event costs constant work, and the whole representation is maintained event by event without ever recomputing a window.1

The numbers: on PEDRo (40 ms windows), spline encoding with local confidence lifts AP50:95 from 0.566 (CSTR) to 0.630, and against exponential decay it raises small-object AP from 0.015 to 0.062 and small-object recall from 0.142 to 0.219. On Gen1, the spline representation reaches 0.370 AP50:95 versus 0.355 for CSTR, and a window study pushes that to 0.406 with 150 ms windows and finer near-term resolution. The recursive polynomial replacement costs almost nothing: 0.626 AP50:95 versus 0.636 for the analytical spline. Against recurrent ReYOLO variants, the feed-forward detector posts the best AP50 on PEDRo (0.918) but the recurrent models still win AP50:95, especially on Gen1 where they exploit 550 ms of context versus 150 ms.1

Where a skeptic should push

The single most load-bearing assumption is that benchmark gains in automotive event vision transfer to other asynchronous event domains. That is an analogy, not a result. The datasets are driving and pedestrian scenes; nothing in the paper touches neural spike trains, whose statistics (burst structure, refractory dead time, drifting firing rates over days of recording) differ from photon-driven event generation. Treat every number above as evidence about representations in general, not as a measured property of neural signals.

Second, the margins are uneven. The Gen1 gains are small (0.355 to 0.370 AP50:95), local confidence helps on PEDRo and not on Gen1, and small-object AP in absolute terms remains terrible (0.062 at best). Much of the improvement comes from simply lengthening the temporal window, which any pipeline can do. Third, the recurrent baselines still win where long context matters, and the authors say so plainly. And the detector itself is a 25M-parameter convolutional network evaluated on GPU, so nothing here has been demonstrated on the kind of edge silicon an MEA front end would actually use.

Log-time encoding at the MEA front end

Here is why this matters to anyone building microelectrode array instrumentation. An event-camera stream and a thresholded multiunit spike stream are the same kind of object: asynchronous, sparse, timestamped, polarity-bearing, with useful information at microsecond to second timescales. Yet most MEA analysis pipelines quantize spikes into fixed raster bins of 10, 20, or 50 ms before any learning happens. That binning is the frame-camera move: it throws away exactly the fine timing this paper shows carries outsized discriminative weight. Their small-object AP moved more than fourfold when the finest temporal scale was properly placed, and their window study showed the shortest timescale parameter, not the window length, was the sensitive knob. If a detector on vision events is that sensitive to temporal placement, a decoder on spike timing, where information is known to live at millisecond precision, has no excuse for blind binning.

The opportunity is architectural. The recursive exponential-polynomial update is constant work per event with no window recompute, which maps cleanly onto per-channel digital logic or analog leaky integrators placed at the array edge. Feed-forward decoding then removes the sequential hidden-state dependency that makes recurrent pipelines awkward to parallelize and hard to bound in latency. For closed-loop stimulation, where the whole point is a decision within a known millisecond budget, a representation that carries its own temporal context and never touches a window buffer is a genuine gift: latency becomes a throughput question, not a recurrence-depth question.

The threat is quieter. A logarithmic basis spends its channels on recent history and compresses the tail, which is where slow biological change lives. Electrode drift, seal degradation, gradual shifts in firing rate as a culture matures are all slow-moving signals that a log-time front end represents at very coarse resolution. The paper's own confidence channels were designed for spatial support of recent events, not for flagging slow degradation, so a pipeline built this way could decode today's spikes beautifully while staying blind to the instrument failing over days. The second threat is calibration: their best temporal parameters did not transfer between PEDRo and Gen1, so a representation tuned on one preparation or one recording regime should be assumed wrong for another until revalidated.

The bottom line

Established: on event-vision benchmarks, multi-timescale log-time basis representations with per-event recursive updates recover a substantial share of recurrent temporal modeling, at near-zero approximation cost, in a fully feed-forward pipeline. Hypothesis, untested here: the same construction preserves spike-timing information that binning destroys, at MEA front ends. What would confirm it is a direct comparison on real array data: binned rasters versus log-time basis features fed to matched decoders, with detection or decoding latency and accuracy both reported. What would break it is evidence that neural temporal statistics punish log-compression of the tail, where adaptation and drift live, more than vision statistics do. Either way, the bin should now have to justify its existence.

Frequently asked questions

What is an event camera, and why is it comparable to an MEA?

An event camera reports per-pixel brightness changes asynchronously with microsecond timestamps and polarity, instead of frames. A thresholded MEA channel reports per-electrode threshold crossings asynchronously with microsecond timestamps. Both are sparse event streams, which is why front-end representations designed for one are worth testing on the other.

What does the logarithmic B-spline encoding actually do?

Each event's age is evaluated against four overlapping B-spline curves spaced logarithmically in time. Recent events fall on finely spaced curves where timing differences are preserved; older events land on coarsely spaced curves where only coarse history survives. The result is a fixed set of feature channels that contain multi-timescale temporal context without any recurrent state.

How large is the improvement over the compact baseline?

On PEDRo, AP50:95 rises from 0.566 for the three-channel CSTR baseline to 0.630 for the spline representation with local confidence, with small-object recall more than tripling from 0.072 to 0.219. On Gen1 the gain is smaller, from 0.355 to 0.370, and the recurrent baselines still lead where long temporal context matters.

Does this make recurrent networks unnecessary?

No. The authors show representation design recovers much of what recurrence buys, and their feed-forward detector posts the best AP50 on PEDRo, but recurrent ReYOLO variants retain the lead on AP50:95, particularly on Gen1 where they use several hundred milliseconds of propagated context. The honest claim is a trade-off, not a replacement.

Why is the recursive polynomial approximation important for hardware?

Evaluating analytical B-splines per event is costly. The polynomial mapping lets the representation be reconstructed from a handful of recursively updated exponential states, so each event costs constant work with no window recompute. That is the property that makes the scheme plausible for per-channel logic or analog integrators at the array edge.

What would it take to validate this on MEA data?

A controlled comparison on real array recordings: identical decoders fed either conventionally binned rasters or log-time basis features, with accuracy and closed-loop latency both reported, plus a test of whether slow drift and degradation signals survive the compressed temporal tail. Until that exists, the MEA application is a well-motivated hypothesis, not a finding.

References

  1. F. Lundell, P.-E. Forssén, M. Wadenbäck, A. Lundmark. Efficient Multi-Timescale Event Representations for Feed-Forward Object Detection. NeVi workshop, ECCV 2026 (arXiv:2609.05049). 2026. https://arxiv.org/abs/2609.05049. Accessed 2026-09-22.