nm000323 NEMAR-native dataset
Lee et al. 2019 (ERP) — EEG dataset and OpenBMI toolbox for three BCI paradigms: an investigation into BCI illiteracy
This dataset comprises EEG recordings from 54 healthy participants performing a P300-based brain-computer interface speller task, designed to investigate BCI illiteracy. The study includes 62 EEG channels and 4 EMG channels sampled at 1000 Hz, with participants completing offline training and online test phases using a 36-symbol row-column speller paradigm. The dataset achieved 96.7% accuracy with an 11.1% BCI illiteracy rate, providing a comprehensive resource for evaluating P300-based BCI performance and individual differences in BCI competence.
AI-generated description, may include mistakesLoading demographics…
Coming soon. Per-file data-quality summaries are precomputed by the NEMAR processing pipeline. The static aggregate is on the way — tracked at nemar-cli#511.
Files
How to use the data (for agentic research) license, citation, download commands, Zarr access
What it is
- Modalities
- EEG
- Participants
- 54
- Size
- 132 GB
- Tasks
- p300
- HED version
- 8.4.0
License and terms
- License
- GPL-3.0
- Recommended citation
- Lee, M., Kwon, O., Kim, Y., Kim, H., Lee, Y., Williamson, J., Fazli, S., & Lee, S. (2026). Lee et al. 2019 (ERP) — EEG dataset and OpenBMI toolbox for three BCI paradigms: an investigation into BCI illiteracy (Version v1.0.4) [Data set]. NEMAR. https://doi.org/10.82901/nemar.nm000323
Where the bytes are
- Latest version (always current)
- https://data.nemar.org/nm000323/latest/
- This version (v1.0.4)
- https://data.nemar.org/nm000323/v1.0.4/
How to download
- The dataset
-
nemar dataset download nm000323Clones and fetches in one step. Content under stimuli/ and derivatives/ is skipped by default because those trees can be large; add --stimuli --derivatives for the whole thing. - A subset, one step
-
nemar dataset download nm000323 --subjects sub-01,02Also filters by --sessions, --tasks, --runs, --datatypes, --include and --exclude. - A subset, step 1
-
nemar dataset clone nm000323Clones git-annex pointers only; fetches no file content. Creates ./nm000323. - A subset, step 2
-
cd nm000323The get command below reads the clone's annex, so it only works from inside the clone. - A subset, step 3
-
nemar dataset get <files>Pulls the files you actually need. Skips stimuli/ and derivatives/ unless the path you ask for is under one of them. - One small file
- https://data.nemar.org/nm000323/v1.0.4/participants.tsv A direct HTTPS fetch works for any single file.
Assess fit without downloading
- Participants table
- https://data.nemar.org/nm000323/v1.0.4/participants.tsv
- Dataset description
- https://data.nemar.org/nm000323/v1.0.4/dataset_description.json
- Directory listing
- https://data.nemar.org/nm000323/v1.0.4/?format=json
- Catalog record
- https://api.nemar.org/datasets/nm000323
Working with the Zarr copy
- 1. Start at the index
- https://zarr.nemar.org/nm000323/zarr/index.json The mandatory entry point. Never hardcode a bucket path.
- 2. Pick a store entry
-
stores[].zarr, stores[].groups[].nameThese two fields exist in every index format version, so a recipe that keys on them works against the whole catalog while the back conversion is still in flight. - 3. Build the store URI
-
s3://nemar/nm000323/zarr/{store.zarr}Derivable from the store entry alone. An index at format_version 3 or later also publishes contract_base, data_base and s3_uri; use them when they are there, never require them. - 4. Open the store anonymously
-
zarr.open_group(store=..., mode="r", zarr_format=3)Anonymous FsspecStore.from_url in region us-east-2, no credentials. zarr_format=3 is required: without it zarr-python probes for Zarr v2 sidecars, and because anonymous ListBucket is denied, S3 answers a missing key with 403 rather than 404 and the open raises. - 5. Read the level-0 array
-
root[store.groups[0].name]["0"]Level 0 is the full-rate signal. Never read a view/ array for inference; those exist for display. - 6. Dequantize the samples
-
physical = digital * scale + offsetscale and offset are attributes of the level-0 array, one entry per channel; the unit is on the group's channels attribute. - 7. Slice, don't download
-
signal[0:4, 0:500]Stream a window of channels and samples; download only when you will touch most of the array. - 8. Know the HTTP contract
-
index.jsonOnly index.json is always proxied and edge-cached. A plain GET for a store object, manifest.json or events.parquet 302s to the public S3 object for non-browser clients, so follow redirects, and HEAD is never redirected. - 9. Read the attribution before reuse
-
root.attrs["nemar"]The store carries its own dataset id, DOI, license, citation and source commit. - 10. Filter for pipelines
-
has_zarr=1This is the converted filter. has_zarr_verified is the stricter one, and its result set can be empty until the daily fidelity sweep reaches a dataset; verification is reported, never a precondition for serving (nemar-cli ADR 0005).