Feedback Signals Are Broader and More Bidirectional Than Stimulus Signals Across the Mouse Brain
1 Original Question
Are feedback signals (reward vs. error) more widespread or more focal than stimulus signals?
2 Introduction
Understanding how outcomes are encoded across the brain is a central question in systems neuroscience. When a mouse performs a decision-making task, two broad classes of events shape neural activity: sensory input (the appearance of a visual stimulus) and outcome (reward or punishment following a correct or incorrect choice). These events carry fundamentally different types of information — sensory signals describe the state of the world, while feedback signals instruct the organism about the value of its actions and must be broadcast to update future behaviour.
The International Brain Laboratory (IBL) Brain Wide Map (BWM) provides a unique opportunity to address this question at scale. With 699 Neuropixels probe insertions across 279 brain areas in 139 mice from 12 laboratories, it is the largest simultaneous-multiregion dataset currently available (Steinmetz et al. 2021; International Brain Laboratory 2025).
Prior work has shown that both sensory and reward signals are widely distributed across the brain (Musall et al. 2019; Stringer et al. 2019), but a direct, brain-wide comparison of how broadly each signal is encoded — and crucially, whether the nature of modulation (excitation vs. suppression) differs — has not been systematically performed on a dataset of this scale.
3 Question Refinement
3.1 From qualitative to quantitative
The original question contained two ambiguities that required explication before analysis could proceed:
“Feedback signals” — The stored event-response features capture all feedback trials pooled (correct and incorrect). Separating reward from error would require recomputing from raw spike shards. We focus first on the pooled question: does feedback, as an event class, differ from stimulus in its brain-wide reach?
“More widespread” — Breadth was operationalised as the fraction of good-quality neurons per brain region that show significant modulation, using a precomputed modulation index (MI) with a threshold.
3.2 Modulation index definition
The IBL bwm_ephys dataset (v1.1.0) provides a stored modulation_index for every unit × event combination:
\[\text{MI} = \frac{\text{peak FR} - \text{baseline FR}}{\text{peak FR} + \text{baseline FR}}\]
where peak FR is the maximum firing rate in the 0–200 ms post-event window (5 ms bins, 20 ms Gaussian smoothing) and baseline FR is computed over the 200 ms pre-event window. MI ranges from −1 (complete suppression) to +1 (complete excitation). A unit is classified as modulated when |MI| > 0.1, excited when MI > 0.1, and suppressed when MI < −0.1.
3.3 Excited vs. suppressed units
A key finding from metric validation was that the two events differ not just in how many neurons they modulate, but in the direction of modulation.
This finding led to three distinct, testable hypotheses:
- H1 (Suppression asymmetry): Feedback recruits more suppressed neurons per region than stimOn.
- H2 (Excitation asymmetry): StimulusOn recruits more excited neurons per region than feedback.
- H3 (Total breadth): Feedback modulates a greater fraction of neurons per region overall.
4 Locked Analysis Plan
4.1 Data
- Dataset: IBL Brain Wide Map,
bwm_ephysv1.1.0. - 75,395 good-quality units, 698 probe insertions, 266 Beryl-atlas regions, 12 laboratories.
- Events compared:
stimOnvs.feedback(all trials pooled). - Regions included: ≥ 10 good units present in both events (198 qualifying regions).
4.2 Data split
To prevent p-hacking, the dataset was split before any analysis into:
- Exploration set: 66 subjects (35,714 units, 248 regions), randomly selected with stratification by lab (seed = 42).
- Confirmation set: 73 subjects (39,681 units, 237 regions), held out until hypothesis lock.
4.3 Statistical tests (locked after exploratory analysis)
All three hypotheses were tested using the Wilcoxon signed-rank test on per-region paired differences (198 region pairs, one-sided, α = 0.05):
| Hypothesis | Comparison |
|---|---|
| H1 | feedback frac_suppressed > stimOn frac_suppressed |
| H2 | stimOn frac_excited > feedback frac_excited |
| H3 | feedback frac_modulated > stimOn frac_modulated |
5 Exploratory Results
5.1 Per-region responsive fractions
At |MI| > 0.1, across 198 qualifying regions:
| Metric | stimOn | feedback |
|---|---|---|
| Mean fraction modulated | 0.581 | 0.617 |
| Mean fraction excited | 0.559 | 0.478 |
| Mean fraction suppressed | 0.022 | 0.139 |
| Regions with >50% modulated | 144 | 174 |
5.2 Brain maps
5.3 Notable regions
Regions most strongly recruited by stimulus (stimOn excited fraction): medial habenula (MH, 93%), nucleus incertus (NI, 88%), claustrum (CL, 86%), anterior ventral nucleus (AV, 84%), ventral posterolateral thalamus (VPL, 81%).
Regions most strongly excited by feedback: auditory cortex (AUDv, 88%), parafascicular nucleus (PCN, 78%), ventral posteromedial thalamus (VPM, 67%), anterior cerebellar lobule (ANcr2, 65%).
Regions most strongly suppressed by feedback: hypoglossal nucleus (XII, 50%), lateral reticular nucleus (LRN, 38%), spinal trigeminal nucleus (SPVI, 30%), thalamic reticular nucleus (TRN, 28%), dorsal raphe (DR, 31%).
Regions with largest feedback excess over stimulus: dorsal raphe (DR, +62%), triangular septal nucleus (TTd, +45%), medial septal nucleus (MS, +42%), lateral preoptic area (LPO, +41%), visceral cortex (VISC, +35%).
6 Confirmatory Results
All three hypotheses were confirmed on the held-out 73-subject confirmation set (198 paired regions, Wilcoxon signed-rank, one-sided, α = 0.05):
| Hypothesis | Description | Mean (stimOn) | Mean (feedback) | p-value | Significance |
|---|---|---|---|---|---|
| H1 | frac_suppressed: feedback > stimOn | 0.018 | 0.131 | 2.5 × 10⁻³² | *** |
| H2 | frac_excited: stimOn > feedback | 0.551 | 0.486 | 2.7 × 10⁻⁷ | *** |
| H3 | frac_modulated: feedback > stimOn | 0.569 | 0.616 | 9.0 × 10⁻⁵ | *** |
7 Discussion
7.1 Summary
Across 198 brain regions and 73 held-out mice, we find that feedback signals are both more widespread and qualitatively different in character compared to visual stimulus signals. The most striking distinction is not breadth but directionality: stimulus onset drives nearly pure excitation (55% of neurons excited, 2% suppressed), while feedback drives a substantial suppression component (49% excited, 13% suppressed) that is visible throughout the brain.
7.2 Relation to prior work
The widespread distribution of outcome signals is consistent with the view that the brain broadcasts feedback information broadly to update representations across many circuits (Schultz et al. 1997; Rangel et al. 2008). The dominance of suppression in feedback responses — particularly in the brainstem, reticular nucleus, and neuromodulatory regions — is less commonly highlighted and may reflect inhibitory gating of irrelevant sensory or motor processing during outcome evaluation.
The strong feedback modulation of the dorsal raphe (+62% excess over stimulus) is consistent with the established role of serotonergic neurons in encoding outcomes and waiting costs (Nakamura 2008; Liu et al. 2014). The thalamic reticular nucleus (TRN) suppression is also consistent with its role in gating thalamocortical transmission — outcome signals may suppress sensory relay to redirect processing toward outcome evaluation networks.
The high feedback modulation in the medial septum and triangular septal nucleus (TTd) is consistent with hippocampal-septal circuits processing spatial and temporal context around reward delivery.
7.3 Caveats
Pooled feedback: All feedback trials (correct and incorrect) are pooled in the stored features. Reward and error signals may have distinct spatial signatures; this should be examined in follow-up work using spike-level recomputation.
Modulation index as a proxy: The stored MI uses the peak of the smoothed PSTH in a 0–200 ms window. Units with slow-onset responses (>200 ms) may be misclassified as unmodulated. The MI is also symmetric with respect to the sign of rate change only through the absolute-value threshold; different shapes (e.g. transient excitation followed by suppression) can yield the same scalar.
Threshold choice: Results were stable across |MI| thresholds from 0.05 to 0.30, but the relative ranking of feedback vs. stimOn total breadth (H3) reverses at thresholds above ~0.15. The asymmetric suppression finding (H1) is robust across all thresholds tested.
Statistical unit: The Wilcoxon test treats brain regions as exchangeable units. Regions vary substantially in unit count (10 to 566), and results from small-n regions should be interpreted cautiously.
Confounds: Stimulus and feedback are separated in time within trials but may co-occur with movement and decision signals. The stored MI does not control for trial-type identity; a formal comparison would require GLM-based analyses.
7.4 Lessons for future AI-assisted analyses
The most scientifically interesting finding (suppression asymmetry) was not part of the original question and emerged from the metric validation step. Plotting excited vs. suppressed separately — rather than only |MI| — is a productive early diagnostic for any modulation comparison and should be added as a default step.
The stored
event_response_featuresare sufficient for broad brain-wide screening questions, but separate reward/error analysis requires spike-level recomputation. The agent should recommend this follow-up explicitly when pooled feedback is used.
8 Suggested Instruction File Updates
The following suggestions are proposed for improving future agent performance. No files have been edited; these are suggestions for user review.
8.1 Addition to skills/ibl-analyze/references/bwm_analysis_patterns.md
Modulation direction diagnostic (default for any event-response comparison) When comparing neural modulation across events or conditions, always plot excited and suppressed fractions separately before reporting total |MI| breadth. Suppression asymmetries can be the dominant signal but are invisible in total-modulation summaries.
8.2 Addition to skills/ibl-analyze/references/scientific_context_and_metric_semantics.md
Feedback pooling note: The stored
event_response_featuresfeedback rows pool correct and incorrect trials. For analyses where reward vs. error distinction is scientifically relevant, flag this as a proxy and plan spike-level recomputation stratified byfeedbackTypefrom the trials table.
9 Methods
Data: IBL Brain Wide Map, bwm_ephys v1.1.0. 75,395 good-quality neurons from 698 Neuropixels probe insertions, 139 mice, 12 laboratories.
Modulation index: Precomputed stored feature. MI = (peak_FR − baseline_FR) / (peak_FR + baseline_FR). Baseline: −200 to 0 ms; peak window: 0–200 ms post-event; 5 ms bins, 20 ms Gaussian smoothing; all trials pooled per event.
Threshold: |MI| > 0.1. Units with MI > 0.1 classified as excited; MI < −0.1 as suppressed. Regions with < 10 good units excluded (198 of 266 Beryl regions retained).
Data split: 66 subjects for exploration (stratified by lab, seed = 42), 73 for confirmation. All confirmatory tests run exactly once on the held-out set.
Statistics: Wilcoxon signed-rank test on per-region paired differences (198 pairs), one-sided, α = 0.05. No multiple-comparison correction applied across the three pre-specified hypotheses.
Atlas: Beryl mapping (IBL atlas). Brain maps generated with iblatlas.plots.plot_scalar_on_slice.
Software: Python 3.14, ibllib, iblatlas, one-api, scipy, pandas, matplotlib, ibl-ai-agent v0.1.0.