Tutorial: streaming NWB data from DANDI

This tutorial opens neurophysiology data from the DANDI Archive without downloading entire files, first from Python and then through a FUSE mount. It follows the DANDI tutorial Streaming and interacting with NWB data from DANDI, but gets the data through git-annex instead of the DANDI API.

The data are position tracking and spike times recorded from the medial entorhinal cortex of rats (Dandiset 000582; Sargolini et al., Science, 2006), stored in NWB files.

Setup

Besides datalad-fuse and git-annex (see Installation), install the packages used to read and plot the data:

$ python3 -m pip install datalad-fuse h5py pynwb matplotlib

The FUSE part of the tutorial also uses h5ls from the HDF5 command-line tools (sudo apt-get install hdf5-tools on Debian/Ubuntu, or conda install -c conda-forge hdf5); any other program that reads files would do as well.

Run all commands and Python code from the same directory: the one in which you clone the dandiset (so that it contains 000582/).

Get the dandiset

Public dandisets are mirrored as git-annex repositories on GitHub, under https://github.com/dandisets. Each asset is an annexed file, for which git-annex knows the URLs of its content on the archive’s S3 bucket and on the DANDI API. Clone the dandiset with DataLad (or plain git clone):

$ datalad clone https://github.com/dandisets/000582
...
install(ok): /home/me/000582 (dataset)
$ du -sh 000582
2.2M 000582

DataLad prints a few [INFO] messages while cloning, including a suggestion to enable siblings; datalad-fuse does not need any of that.

The clone holds all file names but no annexed content, which would be 1.86 GB. Annexed files are symlinks that point to content which is not there:

$ ls -l 000582/sub-10073/
... sub-10073_ses-17010302_behavior+ecephys.nwb -> ../.git/annex/objects/vq/Z5/SHA256E-s15657857--43b3...

git annex whereis shows where the content of a file can be found:

$ git -C 000582 annex whereis sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb
whereis sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb (1 copy)
     00000000-0000-0000-0000-000000000001 -- web
...
  web: https://api.dandiarchive.org/api/assets/2b9e441b-.../download/
  web: https://dandiarchive.s3.amazonaws.com/blobs/26a/22c/26a22c31-...?versionId=...
ok

The web: lines are the URLs that datalad-fuse will read the content from, trying them in this order.

A quick check that the content can be reached: datalad fsspec-head prints the first bytes of the file, here the signature of an HDF5 file:

$ datalad fsspec-head -d 000582 -c 8 sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb | od -c
0000000 211   H   D   F  \r  \n 032  \n
0000010

Read the data from Python

DatasetAdapter gives access to the files of one dataset, addressed by paths relative to the dataset’s top directory. Its open() method returns a file object that h5py, and hence PyNWB, can read from:

from contextlib import closing

import h5py
import matplotlib.pyplot as plt
import numpy as np
import pynwb

from datalad_fuse.fsspec import DatasetAdapter

nwb_path = "sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb"

# caching=False: keep fetched data in memory only (see below)
with closing(DatasetAdapter("000582", caching=False)) as dsa:
    with dsa.open(nwb_path) as f, h5py.File(f, "r") as h5:
        with pynwb.NWBHDF5IO(file=h5) as io:
            nwbfile = io.read()
            print(nwbfile.subject)

            # Position of the animal, from the first LED on its head
            position = nwbfile.processing["behavior"]["Position"]["SpatialSeriesLED1"]
            ts = position.timestamps[:]
            x = position.data[:, 0]

            # Sorted units, with the spike times of each unit
            units_df = nwbfile.units.to_dataframe()

plt.plot(ts, x)
plt.xlabel("Time (seconds)")
plt.ylabel("X coordinate of the subject")
plt.show()

fig, ax = plt.subplots()
for i, spike_times in enumerate(units_df["spike_times"]):
    ax.plot(spike_times, np.ones_like(spike_times) + i, "|", markersize=20)
ax.set(xlabel="Time (s)", ylabel="Unit number")
plt.show()

Some things to note:

  • PyNWB and h5py read data lazily: position.data is not loaded until it is sliced, and only the parts of the file needed for that slice are fetched. So read what you need inside the with blocks; units_df, ts and x are in-memory copies that remain usable afterwards. For interactive work, e.g. in Jupyter, see Working interactively.

  • No DANDI-specific code was needed: the file was found by its path in the dataset, and its URL came from git-annex. The same code works for any dataset whose annexed content is reachable over HTTP(S).

From here on, the analysis part of the DANDI tutorial (e.g. computing tuning curves with pynapple) applies to nwbfile unchanged, as long as it runs while the file is open.

Read the data through a FUSE mount

A FUSE mount gives the same access to programs that expect file names rather than Python file objects. First create an empty directory to mount the dataset on (the mount point), next to the dataset rather than inside it:

$ mkdir mnt

datalad fusefs keeps running for as long as the dataset is mounted, so start it in a terminal of its own:

$ datalad fusefs -d 000582 --foreground mnt

Continue in a second terminal, in the same directory. In the mount, annexed files look like regular files, and their content is fetched when it is read:

$ ls -l mnt/sub-10073/
-rw-r--r-- 1 me me 15657857 May 28 09:22 sub-10073_ses-17010302_behavior+ecephys.nwb
$ h5ls mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb
acquisition              Group
analysis                 Group
file_create_date         Dataset {1}
general                  Group
identifier               Dataset {SCALAR}
processing               Group
session_description      Dataset {SCALAR}
session_start_time       Dataset {SCALAR}
specifications           Group
stimulus                 Group
timestamps_reference_time Dataset {SCALAR}
units                    Group

Python code can now use plain file names:

import pynwb

with pynwb.NWBHDF5IO("mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb") as io:
    print(io.read().units.to_dataframe())

When done, unmount the dataset by pressing Ctrl-C in the terminal running datalad fusefs, or with fusermount -u mnt (see Command-line usage).

Keep fetched data in a cache

By default, fetched data are kept in memory only while a file is open, so opening the file again fetches the data again. With caching enabled, they are stored on disk, inside the dataset, and reused in later sessions too:

with closing(DatasetAdapter("000582", caching=True)) as dsa:
    ...

or, for a mount:

$ datalad fusefs -d 000582 --foreground --caching ondisk mnt

Remove the cache with datalad fsspec-cache-clear -d 000582 when it is no longer needed. See Caching for details.

Pin the version of the data

The dandiset repository records every change of the dandiset, and each published version of the dandiset is a git tag. Checking out a tag gives you exactly the files of that version, and datalad-fuse then reads the content of those versions of the files:

$ git -C 000582 tag
0.251111.2151
$ git -C 000582 checkout 0.251111.2151

Git then reports a “detached HEAD”: you are looking at a past version rather than at a branch. To return to the latest state of the dandiset, check out its main branch, which is called draft for dandisets:

$ git -C 000582 checkout draft

After checking out another version, create a new adapter or remount the dataset, as they do not notice such changes. Recording the commit (git -C 000582 rev-parse HEAD) along with your results tells exactly which data they were computed from.

Next steps