nm000110 NEMAR-native dataset

CHB-MIT

The CHB-MIT Scalp EEG Database is a collection of continuous EEG recordings from 24 pediatric subjects with intractable epilepsy, acquired at Children's Hospital Boston. The dataset comprises 686 EEG scans containing 198 annotated seizure events, recorded at 256 Hz using the International 10-20 electrode system with bipolar montages. Subjects (5 males and 18 females, ages 1.5-22 years) were monitored for several days following anti-seizure medication withdrawal to characterize seizure patterns and assess surgical intervention candidacy. The original EEG data in .edf format from PhysioNet has been converted to BIDS format with standardized channel naming and preserved non-EEG channels. Note: Surrogate dates were used to replace protected health information, which may result in inaccurate age calculations in automated reports.

AI-generated description, may include mistakes
EEG Zarr fidelity issue
Issues GitHub

Download this dataset

Pick a method. Large datasets skip the zip and use the streaming methods below — all resumable. Full download guide →

  1. Download archive (.zip) — 25.4 GB

    A single zip of the published version. Best for small/medium datasets.

    Download zip

  2. NEMAR CLI recommended

    Pulls the pinned version + annexed data and resumes cleanly. Install nemar-cli →

    nemar dataset download nm000110
  3. DataLad

    Clone the dataset repo and fetch file content on demand. Docs →

    datalad clone https://github.com/nemarDatasets/nm000110 nm000110
    cd nm000110 && datalad get .
  4. git-annex

    Plain git + git-annex against the dataset repo. Docs →

    git clone https://github.com/nemarDatasets/nm000110 nm000110
    cd nm000110 && git annex get .
  5. Direct files (wget / curl / rclone)

    Every file with a stable, range-resumable URL from the manifest. Needs curl, jq, wget (or rclone/aria2c). Docs →

    curl -s https://data.nemar.org/nm000110/v1.0.1/manifest.json | jq -r '.[].bytes_url' > urls.txt
    wget -xc -i urls.txt

Compute on this dataset

Two routes today, with a third (in-browser one-click submission) landing soon.

  1. NeuroScience Gateway (NSG) portal.

    NSG runs EEGLAB / Brainstorm / MNE pipelines on supercomputing time donated by SDSC. Create an account, point a job at this dataset's S3 prefix (s3://nemar/nm000110), and submit.
    nsgportal.org →

  2. Local processing with nemar-cli.

    Pull the dataset to your machine and run any toolbox locally. Honors the published version pinning.

    npm install -g nemar-cli
    nemar dataset clone nm000110
    cd nm000110 && nemar dataset get
  3. Just the files.

    rclone, aria2c, or any HTTPS client works against data.nemar.org/nm000110/ — the manifest carries presigned S3 URLs.

Direct compute access is coming soon. One-click NSG submission from this page is scoped for a follow-up phase. Tracked on nemarOrg/website#6.

Citations

    Loading demographics…

    Files

    Loading file index…

    Signal viewer

    How to use the data (for agentic research) license, citation, download commands

    What it is

    Modalities
    EEG
    Participants
    24
    Size
    42.6 GB
    Tasks
    rest

    License and terms

    License
    ODC-By-1.0
    Recommended citation
    Connolly, J., Edwards, H., Bourgeois, B., Treves, S. T., Shoeb, A., & Guttag, J. (2026). CHB-MIT (Version v1.0.1) [Data set]. NEMAR. https://doi.org/10.82901/nemar.nm000110

    Where the bytes are

    Latest version (always current)
    https://data.nemar.org/nm000110/latest/

    How to download

    The dataset
    nemar dataset download nm000110 Clones and fetches in one step. Content under stimuli/ and derivatives/ is skipped by default because those trees can be large; add --stimuli --derivatives for the whole thing.
    A subset, one step
    nemar dataset download nm000110 --subjects sub-01,02 Also filters by --sessions, --tasks, --runs, --datatypes, --include and --exclude.
    A subset, step 1
    nemar dataset clone nm000110 Clones git-annex pointers only; fetches no file content. Creates ./nm000110.
    A subset, step 2
    cd nm000110 The get command below reads the clone's annex, so it only works from inside the clone.
    A subset, step 3
    nemar dataset get <files> Pulls the files you actually need. Skips stimuli/ and derivatives/ unless the path you ask for is under one of them.
    One small file
    https://data.nemar.org/nm000110/v1.0.1/participants.tsv A direct HTTPS fetch works for any single file.