Tutorial: streaming NWB data from DANDI
This tutorial opens neurophysiology data from the DANDI Archive without downloading entire files, first from Python and then through a FUSE mount. It follows the DANDI tutorial Streaming and interacting with NWB data from DANDI, but gets the data through git-annex instead of the DANDI API.
The data are position tracking and spike times recorded from the medial entorhinal cortex of rats (Dandiset 000582; Sargolini et al., Science, 2006), stored in NWB files.
Setup
Besides datalad-fuse and git-annex (see Installation), install
the packages used to read and plot the data:
$ python3 -m pip install datalad-fuse h5py pynwb matplotlib
The FUSE part of the tutorial also uses h5ls from the HDF5 command-line
tools (sudo apt-get install hdf5-tools on Debian/Ubuntu, or conda install
-c conda-forge hdf5); any other program that reads files would do as well.
Run all commands and Python code from the same directory: the one in which
you clone the dandiset (so that it contains 000582/).
Get the dandiset
Public dandisets are mirrored as git-annex repositories on GitHub, under
https://github.com/dandisets. Each asset is an annexed file, for which
git-annex knows the URLs of its content on the archive’s S3 bucket and on
the DANDI API. Clone the dandiset with DataLad (or plain git clone):
$ datalad clone https://github.com/dandisets/000582
...
install(ok): /home/me/000582 (dataset)
$ du -sh 000582
2.2M 000582
DataLad prints a few [INFO] messages while cloning, including a
suggestion to enable siblings; datalad-fuse does not need any of that.
The clone holds all file names but no annexed content, which would be 1.86 GB. Annexed files are symlinks that point to content which is not there:
$ ls -l 000582/sub-10073/
... sub-10073_ses-17010302_behavior+ecephys.nwb -> ../.git/annex/objects/vq/Z5/SHA256E-s15657857--43b3...
git annex whereis shows where the content of a file can be found:
$ git -C 000582 annex whereis sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb
whereis sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb (1 copy)
00000000-0000-0000-0000-000000000001 -- web
...
web: https://api.dandiarchive.org/api/assets/2b9e441b-.../download/
web: https://dandiarchive.s3.amazonaws.com/blobs/26a/22c/26a22c31-...?versionId=...
ok
The web: lines are the URLs that datalad-fuse will read the content
from, trying them in this order.
A quick check that the content can be reached: datalad fsspec-head
prints the first bytes of the file, here the signature of an HDF5 file:
$ datalad fsspec-head -d 000582 -c 8 sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb | od -c
0000000 211 H D F \r \n 032 \n
0000010
Read the data from Python
DatasetAdapter gives access to the
files of one dataset, addressed by paths relative to the dataset’s top
directory. Its open() method returns a file object that h5py, and hence
PyNWB, can read from:
from contextlib import closing
import h5py
import matplotlib.pyplot as plt
import numpy as np
import pynwb
from datalad_fuse.fsspec import DatasetAdapter
nwb_path = "sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb"
# caching=False: keep fetched data in memory only (see below)
with closing(DatasetAdapter("000582", caching=False)) as dsa:
with dsa.open(nwb_path) as f, h5py.File(f, "r") as h5:
with pynwb.NWBHDF5IO(file=h5) as io:
nwbfile = io.read()
print(nwbfile.subject)
# Position of the animal, from the first LED on its head
position = nwbfile.processing["behavior"]["Position"]["SpatialSeriesLED1"]
ts = position.timestamps[:]
x = position.data[:, 0]
# Sorted units, with the spike times of each unit
units_df = nwbfile.units.to_dataframe()
plt.plot(ts, x)
plt.xlabel("Time (seconds)")
plt.ylabel("X coordinate of the subject")
plt.show()
fig, ax = plt.subplots()
for i, spike_times in enumerate(units_df["spike_times"]):
ax.plot(spike_times, np.ones_like(spike_times) + i, "|", markersize=20)
ax.set(xlabel="Time (s)", ylabel="Unit number")
plt.show()
Some things to note:
PyNWB and h5py read data lazily:
position.datais not loaded until it is sliced, and only the parts of the file needed for that slice are fetched. So read what you need inside thewithblocks;units_df,tsandxare in-memory copies that remain usable afterwards. For interactive work, e.g. in Jupyter, see Working interactively.No DANDI-specific code was needed: the file was found by its path in the dataset, and its URL came from git-annex. The same code works for any dataset whose annexed content is reachable over HTTP(S).
From here on, the analysis part of the DANDI tutorial (e.g. computing tuning
curves with pynapple) applies to nwbfile
unchanged, as long as it runs while the file is open.
Read the data through a FUSE mount
A FUSE mount gives the same access to programs that expect file names rather than Python file objects. First create an empty directory to mount the dataset on (the mount point), next to the dataset rather than inside it:
$ mkdir mnt
datalad fusefs keeps running for as long as the dataset is mounted, so
start it in a terminal of its own:
$ datalad fusefs -d 000582 --foreground mnt
Continue in a second terminal, in the same directory. In the mount, annexed files look like regular files, and their content is fetched when it is read:
$ ls -l mnt/sub-10073/
-rw-r--r-- 1 me me 15657857 May 28 09:22 sub-10073_ses-17010302_behavior+ecephys.nwb
$ h5ls mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb
acquisition Group
analysis Group
file_create_date Dataset {1}
general Group
identifier Dataset {SCALAR}
processing Group
session_description Dataset {SCALAR}
session_start_time Dataset {SCALAR}
specifications Group
stimulus Group
timestamps_reference_time Dataset {SCALAR}
units Group
Python code can now use plain file names:
import pynwb
with pynwb.NWBHDF5IO("mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb") as io:
print(io.read().units.to_dataframe())
When done, unmount the dataset by pressing Ctrl-C in the terminal
running datalad fusefs, or with fusermount -u mnt (see Command-line usage).
Keep fetched data in a cache
By default, fetched data are kept in memory only while a file is open, so opening the file again fetches the data again. With caching enabled, they are stored on disk, inside the dataset, and reused in later sessions too:
with closing(DatasetAdapter("000582", caching=True)) as dsa:
...
or, for a mount:
$ datalad fusefs -d 000582 --foreground --caching ondisk mnt
Remove the cache with datalad fsspec-cache-clear -d 000582 when it is no
longer needed. See Caching for details.
Pin the version of the data
The dandiset repository records every change of the dandiset, and each
published version of the dandiset is a git tag. Checking out a tag gives you
exactly the files of that version, and datalad-fuse then reads the content
of those versions of the files:
$ git -C 000582 tag
0.251111.2151
$ git -C 000582 checkout 0.251111.2151
Git then reports a “detached HEAD”: you are looking at a past version rather
than at a branch. To return to the latest state of the dandiset, check out
its main branch, which is called draft for dandisets:
$ git -C 000582 checkout draft
After checking out another version, create a new adapter or remount the
dataset, as they do not notice such changes. Recording the commit (git -C
000582 rev-parse HEAD) along with your results tells exactly which data they
were computed from.
Next steps
How it works explains where
datalad-fuselooks for content, and what that means for performance and caching.Command-line usage and Python usage describe the command line and Python interfaces in detail.
Troubleshooting helps when something does not work.