nm000153 NEMAR-native dataset

Langer et al. 2017 — Multimodal Resource for Studying Information Processing in the Developing Brain (MIPDB)

MIPDB is a multimodal neuroimaging resource comprising high-density EEG (128-channel, 500 Hz) and eye-tracking data collected from 111 participants spanning childhood to adulthood. Participants completed a standardized battery of cognitive and perceptual tasks including resting state, surround suppression, naturalistic viewing, contrast-change detection, sequence learning, and symbol search. This dataset enables investigation of information-processing maturation across human development.

AI-generated description, may include mistakes
EEG Zarr unverifiable
Issues GitHub

Download this dataset

dataset 110.3 GB exceeds 100.0 GB archive limit; use direct download. Use one of the streaming methods below — all resumable. Full download guide →

  1. NEMAR CLI recommended

    Pulls the pinned version + annexed data and resumes cleanly. Install nemar-cli →

    nemar dataset download nm000153
  2. DataLad

    Clone the dataset repo and fetch file content on demand. Docs →

    datalad clone https://github.com/nemarDatasets/nm000153 nm000153
    cd nm000153 && datalad get .
  3. git-annex

    Plain git + git-annex against the dataset repo. Docs →

    git clone https://github.com/nemarDatasets/nm000153 nm000153
    cd nm000153 && git annex get .
  4. Direct files (wget / curl / rclone)

    Every file with a stable, range-resumable URL from the manifest. Needs curl, jq, wget (or rclone/aria2c). Docs →

    curl -s https://data.nemar.org/nm000153/v1.0.0/manifest.json | jq -r '.[].bytes_url' > urls.txt
    wget -xc -i urls.txt

Compute on this dataset

Two routes today, with a third (in-browser one-click submission) landing soon.

  1. NeuroScience Gateway (NSG) portal.

    NSG runs EEGLAB / Brainstorm / MNE pipelines on supercomputing time donated by SDSC. Create an account, point a job at this dataset's S3 prefix (s3://nemar/nm000153), and submit.
    nsgportal.org →

  2. Local processing with nemar-cli.

    Pull the dataset to your machine and run any toolbox locally. Honors the published version pinning.

    npm install -g nemar-cli
    nemar dataset clone nm000153
    cd nm000153 && nemar dataset get
  3. Just the files.

    rclone, aria2c, or any HTTPS client works against data.nemar.org/nm000153/ — the manifest carries presigned S3 URLs.

Direct compute access is coming soon. One-click NSG submission from this page is scoped for a follow-up phase. Tracked on nemarOrg/website#6.

Citations

    Loading demographics…

    Files

    Loading file index…

    Signal viewer

    How to use the data (for agentic research) license, citation, download commands, Zarr access

    What it is

    Modalities
    EEG
    Participants
    111
    Size
    110 GB
    Tasks
    block01, block02, block03, block04, block05, block06, block07, block08, block09, block10, block11, block12, block13, block14, block15, block16, block17, block18, block19, block398

    License and terms

    License
    CC-BY-NC-SA-3.0
    Note
    Non-commercial use only (CC-BY-NC-SA-3.0).
    Recommended citation
    Langer, N., Ho, E. J., Alexander, L. M., Xu, H. Y., Jozanovic, R. K., Henin, S., Petroni, A., Cohen, S., Marcelle, E. T., Parra, L. C., Milham, M. P., & Kelly, S. P. (2026). Langer et al. 2017 — Multimodal Resource for Studying Information Processing in the Developing Brain (MIPDB) (Version v1.0.0) [Data set]. NEMAR. https://doi.org/10.82901/nemar.nm000153

    Where the bytes are

    Latest version (always current)
    https://data.nemar.org/nm000153/latest/

    How to download

    The dataset
    nemar dataset download nm000153 Clones and fetches in one step. Content under stimuli/ and derivatives/ is skipped by default because those trees can be large; add --stimuli --derivatives for the whole thing.
    A subset, one step
    nemar dataset download nm000153 --subjects sub-01,02 Also filters by --sessions, --tasks, --runs, --datatypes, --include and --exclude.
    A subset, step 1
    nemar dataset clone nm000153 Clones git-annex pointers only; fetches no file content. Creates ./nm000153.
    A subset, step 2
    cd nm000153 The get command below reads the clone's annex, so it only works from inside the clone.
    A subset, step 3
    nemar dataset get <files> Pulls the files you actually need. Skips stimuli/ and derivatives/ unless the path you ask for is under one of them.
    One small file
    https://data.nemar.org/nm000153/v1.0.0/participants.tsv A direct HTTPS fetch works for any single file.

    Working with the Zarr copy

    1. Start at the index
    https://zarr.nemar.org/nm000153/zarr/index.json The mandatory entry point. Never hardcode a bucket path.
    2. Pick a store entry
    stores[].zarr, stores[].groups[].name These two fields exist in every index format version, so a recipe that keys on them works against the whole catalog while the back conversion is still in flight.
    3. Build the store URI
    s3://nemar/nm000153/zarr/{store.zarr} Derivable from the store entry alone. An index at format_version 3 or later also publishes contract_base, data_base and s3_uri; use them when they are there, never require them.
    4. Open the store anonymously
    zarr.open_group(store=..., mode="r", zarr_format=3) Anonymous FsspecStore.from_url in region us-east-2, no credentials. zarr_format=3 is required: without it zarr-python probes for Zarr v2 sidecars, and because anonymous ListBucket is denied, S3 answers a missing key with 403 rather than 404 and the open raises.
    5. Read the level-0 array
    root[store.groups[0].name]["0"] Level 0 is the full-rate signal. Never read a view/ array for inference; those exist for display.
    6. Dequantize the samples
    physical = digital * scale + offset scale and offset are attributes of the level-0 array, one entry per channel; the unit is on the group's channels attribute.
    7. Slice, don't download
    signal[0:4, 0:500] Stream a window of channels and samples; download only when you will touch most of the array.
    8. Know the HTTP contract
    index.json Only index.json is always proxied and edge-cached. A plain GET for a store object, manifest.json or events.parquet 302s to the public S3 object for non-browser clients, so follow redirects, and HEAD is never redirected.
    9. Read the attribution before reuse
    root.attrs["nemar"] The store carries its own dataset id, DOI, license, citation and source commit.
    10. Filter for pipelines
    has_zarr=1 This is the converted filter. has_zarr_verified is the stricter one, and its result set can be empty until the daily fidelity sweep reaches a dataset; verification is reported, never a precondition for serving (nemar-cli ADR 0005).