Data formats & splits¶
Every dataset — whether fetched via mirdata or recorded by hand — is read
through one on-disk format: ../musicality_db/<name>/tracks/ (audio) and
../musicality_db/<name>/annotations/ (.beats + .meta.json), a
plain directory layout centralized in musicality.dataformats
(musicality/dataformats/dataformat.yaml, data_dir currently
../musicality_db — a sibling git+dvc repo, cloned by
tools/setup_remote.sh). TempoDataset/BeatDataset, the desktop
annotator, and the mobile companion all read exclusively through this
format via musicality.dataformats.track_io — none of them import
mirdata.
mirdata is acquisition-only: tools/download_dataset.py fetches a
dataset (ballroom, brid, hainsworth, rwc_classical, rwc_jazz, rwc_popular,
groove_midi, guitarset, …) into its own raw layout. A migration tool then
bridges it into this project’s format before anything else touches it:
tools/migrate_mirdata_dataset.py— for mirdata datasets. Writesannotations/*.beats/*.meta.jsonfrom mirdata’s own beat annotations, and moves (not copies) each track’s audio intotracks/— mirdata is never read again for a migrated track. Handles datasets whose on-disk audio layout doesn’t match mirdata’s own index (seetools/mirdata_audio.py).tools/migrate_rwc_genre.py— for CSV-annotated datasets (e.g. hand-corrected RWC annotations) that already have audio undertracks/.tools/migrate_gtzan.py— for GTZAN. Its audio is downloaded separately, via the data dir’s owndl_gtzan.py(HuggingFace’s marsyas/gtzan mirror — mirdata’sgtzan_genreaudio host is a dead link), and its beat annotations come from CPJKU’s beat_this_annotations project, already dropped in as this project’s own.beatsformat. The two sources number tracks with a one-off mismatch; this tool resolves that offset, then moves audio intotracks/same as the other tools.
A dataset with no tracks/ folder is treated as not-yet-migrated: the
annotator won’t list it, and tools.annotator.data raises an error
naming the migration command to run, rather than silently falling back to
mirdata.
Data format¶
.beats files are <time> <position> per line — seconds, then the
1-indexed bar/count position, when annotated:
10.949773 1
11.247052 2
11.653333 3
Count across whatever you consider one bar — the count does not have to match
the group_size a model is trained at. A bar counted 1..8 against a
4-beat model is folded to 1,2,3,4,1,2,3,4 when the dataset is loaded, so
beat 5 correctly trains as a downbeat (see
musicality.loaders.beat_dataset.fold_positions()). A count that is not
a whole multiple of group_size — 6 against 4, say — has no consistent
folding; those tracks still contribute their beats, but their bar positions are
masked out of the loss rather than folded wrongly.
.meta.json carries fields with no place in .beats — all optional,
filled in incrementally as an annotation is worked on, and forward-compatible
(a file missing a newer field just falls back to that field’s default on
load):
Field |
Meaning |
|---|---|
|
Where the track was recorded/found |
|
Recording device (e.g. phone model, hostname) |
|
Free-text song structure notes |
|
Audio duration, in seconds |
|
Tempo statistics derived from the beat annotation itself (see
|
|
Who made this annotation — |
|
Whether the first tapped beat is the true start of a section
( |
|
Metadata schema version (currently |
Multiple people can annotate the same track independently:
annotator_id: null is the original, unsuffixed slot (every file saved
before multi-annotator support existed still resolves here — no migration
needed), and each named annotator gets their own subdirectory holding a
parallel .beats/.meta.json pair for the same track_id.
TempoDataset/BeatDataset read only the default (null) slot per
track — picking a specific annotator’s take for training is separate future
work.
Splits¶
tools/create_splits.py builds a train/val split for a migrated dataset
and saves it under splits_dir/<name>/{train,val}.txt — one
<dataset_name>/<track_id> line per track, not a positional index, so a
split’s contents can be read back, concatenated, or merged with another
dataset’s split independently of any one dataset instance’s ordering (see
musicality.splits.splitter.Splitter). One split per dataset serves
every task — <name>, or <name>-binary with --binary-only. Tempo
and beat runs read the same file: a track’s tempo label is derived from its
beat annotation, so the two tasks can never disagree about which tracks are
usable, and binary_only is the only flag that changes membership (see
musicality.splits.splitter.split_name).
Splitter.load_refs/.save_refs read and write this format directly;
TempoDataset/BeatDataset accept it straight via their refs=
argument, or Splitter(...).run() for the Subset-returning form used
during training.
Splitting a subset of a dataset: --contains¶
Some datasets encode a category in the track id — gtzan’s tracks are
blues_00001, jazz_00042, rock_00007, and so on. To work with one
of those groups alone, pass a substring:
uv run python tools/create_splits.py --datasets gtzan --contains blues
That keeps only the tracks whose id contains blues (matched
case-insensitively) and writes them to their own split name,
splits_dir/gtzan-blues/ — or gtzan-blues-binary under
--binary-only. The dataset’s full split, if it has one, is left
untouched.
The filter applies at creation time only, and deliberately so: what it
produces is an ordinary split file, indistinguishable downstream from any
other. Train on it with data.input=gtzan-blues, evaluate it by that
name, and combine groups with tools/merge_datasets.py --datasets
gtzan-blues gtzan-jazz --output gtzan-blues_jazz — no consumer needs to
know the split was ever narrowed, and nothing has to re-derive the filter
to reproduce it.
The filtering itself is musicality.dataformats.track_io.list_track_refs’
contains= argument, applied before either dataset is built, so a
narrowed run doesn’t pay to resolve the tracks it’s about to drop. Substring
matching is intentionally all it does — for anything finer, assemble a
folder of train.txt/val.txt by hand and point data.input at the
path (see below).
Flagged tracks are excluded from splits¶
A track whose metadata has warning (the annotator’s ⚠ Flag — “this
annotation looks wrong”) or needs_review (❓ To Review — “take another
look”) set to True is skipped by the splitter, and so never reaches
training or validation. The filter runs on both sides: create()
excludes flagged tracks from the pool before splitting, so they’re written
into neither train.txt nor val.txt, and every read path
(run(), load_refs, load_refs_from_dir) drops them again on
load. That second half is what makes flagging usable day to day — a track
flagged in the annotator after a split was generated falls out of the
next training run on its own, with no need to regenerate the split file or
re-run create_splits.py. Only the default annotation slot’s metadata is
consulted, matching which slot the loaders read. Clearing the flag in the
annotator puts the track straight back in.
A split’s tracks must be on disk¶
Reading a split checks that every track it lists actually resolves to files
— tracks/<id>.wav and annotations/<id>.beats — and raises
musicality.splits.splitter.MissingTrackDataError naming the corpora that
are short if any doesn’t. The check runs in _read_refs, so it covers
every read path (run(), load_refs, load_refs_from_dir) and
therefore every trainer and tools/eval_beat.py. Flagged tracks are
dropped before it, so a track that is both flagged and absent is simply out
of the split rather than an error.
This exists because the failure it catches is invisible otherwise. Tracks
whose audio couldn’t be resolved used to be dropped one by one with a single
printed line, so an incomplete dvc pull on a fresh remote instance — a
whole corpus that never arrived — produced a run that trained on a strictly
smaller dataset than its split describes and reported its metrics as if
nothing had happened. A run must fail on that, not shrink.
When the split is genuinely the thing that’s stale (tracks deleted from the
data repo, say), regenerate it with tools/create_splits.py rather than
loosening the check. The one caller that opts out is the annotator, which
reads splits only to badge tracks train/val in its tree
(Splitter.load_refs_from_dir(..., verify=False)) — browsing the data you
do have shouldn’t require having all of it.
Merging datasets¶
tools/merge_datasets.py --datasets ballroom brid --output ballroom_brid
combines several datasets’ existing splits into one merged split, keeping
train and val separate: each source dataset’s train tracks go into the
merged train split, and its val tracks into the merged val split — so a
track held out for one dataset stays held out in the merge. It writes
nothing under the source datasets’ own directories, and no merged dataset
directory of any kind — only
splits_dir/ballroom_brid/{train,val}.txt, in the same format
create_splits.py produces. There’s no --val-split argument: the
merge inherits whatever ratio each source was already split at, and fails
fast, before writing anything, if a requested source doesn’t have a split
yet (run tools/create_splits.py first).
A merged split name is then usable anywhere a real one is — e.g.
data.input: ballroom_brid in a training config — since
build_dataloaders/build_beat_dataloaders load a split’s refs and
construct TempoDataset/BeatDataset directly via refs=, with no
distinction between a plain dataset’s split and a merged one.
Sources that hold the same track¶
Two source names can cover the same tracks: a dataset and its -binary
variant (the same pool whenever the meter filter drops nothing — e.g.
swing), a dataset and a --contains subset of it (gtzan and
gtzan-blues), or simply the same name given twice. The merge handles
that in two steps, both before anything is written:
Repeats collapse. A track reached through several sources is written once, and the run prints how many repeats it collapsed. Without this the track is loaded twice per epoch — weighted double in training, counted twice in every validation metric.
Disagreements abort.
<name>and<name>-binaryare drawn independently (seesplit_name), so they partition the same pool differently: merging both would put roughly a fifth of the tracks into the merged train and val splits. There is no right side to pick, so the merge fails and names the tracks. Merge one split per pool of tracks.
Telling training which split to use: data.input¶
Training configs (configs/train_tempo.yaml, train_phase_beat.yaml,
train_beat_only.yaml) have one field, data.input, for naming the
split to train on. It’s read two ways, told apart by whether it contains a
/:
A bare name — e.g.
data.input=ballroomordata.input=ballroom_brid— looked up under the canonicalsplits_dir:splits_dir/<input>/{train,val}.txt. This is the normal case: everythingcreate_splits.pyandmerge_datasets.pyproduce lands undersplits_dirautomatically, so a name is all you need.A path — e.g.
data.input=../musicality_db/splits/ballroomordata.input=/anywhere/my_split— used directly as the folder to readtrain.txt/val.txtfrom, bypassingsplits_direntirely. Use this for a split that isn’t registered undersplits_dir— one assembled by hand, or one living outside this repo’s data directory.
See musicality.trainers.common.resolve_split_refs, shared by both the
tempo and beat trainers, for the exact dispatch.
API reference¶
Loads dataformat.yaml and exposes hardcoded directory names as a typed object. |
|