nm000275 NEMAR-native dataset

Multi-channel EEG recordings during a sustained-attention driving task

This dataset comprises multi-channel EEG recordings from 27 healthy adults performing a sustained-attention driving task in a virtual-reality simulator across 62 sessions. The task involved event-related lane-departure paradigms where participants maintained vehicle position on a simulated highway, with recordings capturing 32-channel EEG (30 scalp + 2 mastoid references) sampled at 500 Hz alongside vehicle position data. The dataset includes approximately 82 hours of raw, unfiltered EEG data with over 27,000 lane-departure trials and associated behavioral markers, providing a resource for investigating fatigue, drowsiness, and sustained attention mechanisms.

AI-generated description, may include mistakes
Issues GitHub

Download this dataset

Pick a method. Large datasets skip the zip and use the streaming methods below — all resumable. Full download guide →

  1. Download archive (.zip) — 19.4 GB

    A single zip of the published version. Best for small/medium datasets.

    Download zip

  2. NEMAR CLI recommended

    Pulls the pinned version + annexed data and resumes cleanly. Install nemar-cli →

    nemar dataset download nm000275
  3. DataLad

    Clone the dataset repo and fetch file content on demand. Docs →

    datalad clone https://github.com/nemarDatasets/nm000275 nm000275
    cd nm000275 && datalad get .
  4. git-annex

    Plain git + git-annex against the dataset repo. Docs →

    git clone https://github.com/nemarDatasets/nm000275 nm000275
    cd nm000275 && git annex get .
  5. Direct files (wget / curl / rclone)

    Every file with a stable, range-resumable URL from the manifest. Needs curl, jq, wget (or rclone/aria2c). Docs →

    curl -s https://data.nemar.org/nm000275/v1.0.0/manifest.json | jq -r '.[].bytes_url' > urls.txt
    wget -xc -i urls.txt

Compute on this dataset

Two routes today, with a third (in-browser one-click submission) landing soon.

  1. NeuroScience Gateway (NSG) portal.

    NSG runs EEGLAB / Brainstorm / MNE pipelines on supercomputing time donated by SDSC. Create an account, point a job at this dataset's S3 prefix (s3://nemar/nm000275), and submit.
    nsgportal.org →

  2. Local processing with nemar-cli.

    Pull the dataset to your machine and run any toolbox locally. Honors the published version pinning.

    npm install -g nemar-cli
    nemar dataset clone nm000275
    cd nm000275 && nemar dataset get
  3. Just the files.

    rclone, aria2c, or any HTTPS client works against data.nemar.org/nm000275/ — the manifest carries presigned S3 URLs.

Direct compute access is coming soon. One-click NSG submission from this page is scoped for a follow-up phase. Tracked on nemarOrg/website#6.

Citations

    Use this data

    What it is

    Modalities
    EEG
    Participants
    27
    Size
    36.4 GB
    Tasks
    driving
    HED version
    8.2.0

    License and terms

    License
    CC-BY-4.0
    Recommended citation
    Cao, Z., Chuang, C., King, J., & Lin, C. (2026). Multi-channel EEG recordings during a sustained-attention driving task (Version v1.0.0) [Data set]. NEMAR. https://doi.org/10.82901/nemar.nm000275

    Where the bytes are

    Latest version (always current)
    https://data.nemar.org/nm000275/latest/

    How to download

    The dataset
    nemar dataset download nm000275 Clones and fetches in one step. Content under stimuli/ and derivatives/ is skipped by default because those trees can be large; add --stimuli --derivatives for the whole thing.
    A subset, one step
    nemar dataset download nm000275 --subjects sub-01,02 Also filters by --sessions, --tasks, --runs, --datatypes, --include and --exclude.
    A subset, step 1
    nemar dataset clone nm000275 Clones git-annex pointers only; fetches no file content. Creates ./nm000275.
    A subset, step 2
    cd nm000275 The get command below reads the clone's annex, so it only works from inside the clone.
    A subset, step 3
    nemar dataset get <files> Pulls the files you actually need. Skips stimuli/ and derivatives/ unless the path you ask for is under one of them.
    One small file
    https://data.nemar.org/nm000275/v1.0.0/participants.tsv A direct HTTPS fetch works for any single file.

    Working with the Zarr copy

    1. Start at the index
    https://zarr.nemar.org/nm000275/zarr/index.json The mandatory entry point. Never hardcode a bucket path.
    2. Pick a store entry
    stores[].zarr, stores[].groups[].name These two fields exist in every index format version, so a recipe that keys on them works against the whole catalog while the back conversion is still in flight.
    3. Build the store URI
    s3://nemar/nm000275/zarr/{store.zarr} Derivable from the store entry alone. An index at format_version 3 or later also publishes contract_base, data_base and s3_uri; use them when they are there, never require them.
    4. Open the store anonymously
    zarr.open_group(store=..., mode="r", zarr_format=3) Anonymous FsspecStore.from_url in region us-east-2, no credentials. zarr_format=3 is required: without it zarr-python probes for Zarr v2 sidecars, and because anonymous ListBucket is denied, S3 answers a missing key with 403 rather than 404 and the open raises.
    5. Read the level-0 array
    root[store.groups[0].name]["0"] Level 0 is the full-rate signal. Never read a view/ array for inference; those exist for display.
    6. Dequantize the samples
    physical = digital * scale + offset scale and offset are attributes of the level-0 array, one entry per channel; the unit is on the group's channels attribute.
    7. Slice, don't download
    signal[0:4, 0:500] Stream a window of channels and samples; download only when you will touch most of the array.
    8. Know the HTTP contract
    index.json Only index.json is always proxied and edge-cached. A plain GET for a store object, manifest.json or events.parquet 302s to the public S3 object for non-browser clients, so follow redirects, and HEAD is never redirected.
    9. Read the attribution before reuse
    root.attrs["nemar"] The store carries its own dataset id, DOI, license, citation and source commit.
    10. Filter for pipelines
    has_zarr=1 This is the converted filter. has_zarr_verified is the stricter one, and its result set can be empty until the daily fidelity sweep reaches a dataset; verification is reported, never a precondition for serving (nemar-cli ADR 0005).

    Loading demographics…

    Files

    Loading file index…

    Signal viewer