Ephys Atlas region classifier (2026_W32, Cosmos)
Predicts the brain region of each Neuropixels recording channel from electrophysiological
features alone -- no histology required. Trained by the
International Brain Laboratory on the Ephys Atlas
feature release 2026_W32.
What you can and cannot do without IBL access. The model runs for anyone. Computing the input features from raw Neuropixels data needs
ibllib/ibl-neuropixel, and the raw data itself is IBL-hosted. To try the model immediately, use the bundled sample underexample/-- no account, no raw data, no S3.
Quickstart
import pandas as pd
from ephysatlas import load_pretrained
model = load_pretrained("int-brain-lab/ea-decoder-channel-xgboost", revision="2026_W32")
df = pd.read_parquet("example/features_sample.parquet") # or your own features
out = model.predict(df)
print(out[["predicted_acronym", "prediction_probability", "fold_agreement"]].head())
load_pretrained is the entry point for every ephysatlas model, whatever its family — it reads
ephysatlas_model.json and returns the right wrapper. Use it rather than importing a concrete
class, so your code keeps working as the package evolves.
predict returns one row per input channel, indexed identically to the input:
predicted_acronym, its Allen predicted_atlas_id, the fold-averaged
prediction_probability, a fold_agreement column (fraction of the 5 folds voting
for the winner -- the natural uncertainty signal), and a p_<acronym> column per class.
The prediction columns are namespaced so that df.join(out) works: the feature table already
carries histology-derived acronym / atlas_id columns, and predictions must not shadow them.
Inputs
- 50 features, listed in
ephysatlas_model.jsonunderinputs.features. Every one must be present;predictraises and names anything missing. - Indexed by
(pid, channel), one row per recording channel. - Must be the denoised aggregated features of vintage
2026_W32-- that is, theraw_ephys_features_denoised.pqttable produced by the Ephys Atlas aggregation pipeline, as loaded byephysatlas.data.read_features_from_disk. Units are baked into that table by the pipeline (RMS features in dB,spike_countin log2), so feeding raw features, or features from a vintage whose units differ, produces confident nonsense. Runmodel.selftest()to confirm your install reproduces the shipped output before trusting it.
Performance
Pooled out-of-fold accuracy: 0.5962 over 13 Cosmos regions,
765 insertions. Splits are by insertion (pid), so no channel from a test
insertion appears in training. See confusion_matrix.png.
Limitations
- Trained on IBL Neuropixels 1.0 recordings in mouse. Transfer to NP2, other species or other rigs is untested.
- Coverage follows IBL brain-wide-map targeting; rare regions are under-represented.
Cosmosis a coarse parcellation. Predictions are per-channel and spatially unregularised -- neighbouring channels can disagree.- Known-misaligned insertions were excluded from training.
Reproducibility
Pin the revision. revision="2026_W32" is an immutable tag. Omitting revision resolves
to main, which tracks whichever model is currently recommended and will change when a new
feature vintage is published — fine for a first look, not for anything you publish or re-run.
ephysatlas_model.json records the training-time environment (xgboost, scikit-learn, numpy,
ephysatlas, python) and random_seed. Verify your install reproduces the shipped output:
model.selftest()
Note scikit-learn<1.9 is required (1.9 broke OneToOneFeatureMixin.get_feature_names_out,
which the feature transformer relies on).
Citation
Please cite the International Brain Laboratory Ephys Atlas. Model id 2026_W32_Cosmos_guiltless-orange-mallard,
feature vintage 2026_W32.
- Downloads last month
- 4