splitter¶
Functions
|
Return refs without the tracks |
|
Return whether ref's annotation is flagged as not fit for training. |
|
Return the files ref needs but doesn't have on disk. |
|
Split directory name for name: |
|
Raise unless every ref in refs resolves to files on disk. |
Classes
|
Loads persistent train/val splits for a dataset. |
Exceptions
A split lists tracks whose files are not on this machine. |
- exception MissingTrackDataError[source]¶
Bases:
FileNotFoundErrorA split lists tracks whose files are not on this machine.
Raised by
verify_refs_present(), and therefore by every path that reads atrain.txt/val.txt(see_read_refs()). The failure it exists to stop is a partial data pull — e.g.dvc pullfetching some corpora but not others on a fresh remote instance. Before this check, a corpus missing from disk was dropped track by track with nothing but a printed line, so a run could train on a strictly smaller dataset than its split describes and report metrics as if nothing had happened.
- class Splitter(dataset, splits_dir, name, val_split)[source]¶
Bases:
objectLoads persistent train/val splits for a dataset.
Splits are stored under
splits_dir/<name>/train.txtandval.txt, onedataset_name/track_idline per track (seeTrackRef) rather than positional indices — so a split’s contents can be resolved, concatenated with another dataset’s split, or loaded standalone (seeload_refs()) independently of any one dataset instance’s ordering. This is what letstools/merge_datasets.pycombine several datasets’ existing splits into one.run()only ever reads these files — it never generates a split, so the same files (typically version-controlled via DVC) produce the same split on every machine. Usecreate()(ortools/create_splits.py) to generate the files in the first place.Every read path also checks that the tracks a split lists are actually on disk, and raises
MissingTrackDataErrorif they aren’t (seeverify_refs_present()) — a split is a claim about which tracks a run uses, and a partial data pull must stop the run rather than quietly shrink it.Tracks their annotator flagged (
warningorneeds_reviewin the track metadata — seeis_flagged()) are skipped on both sides:create()never writes them into a new split, and every read path (run(),load_refs(),load_refs_from_dir()) drops them from splits saved before the flag was set. So flagging a track in the annotator takes it out of training and validation immediately, with no need to regenerate any split file.- Parameters:
dataset (Dataset) – The dataset to split. Must expose
.refs— alist[TrackRef]parallel to its samples (bothTempoDatasetandBeatDatasetdo).splits_dir (Path) – Root directory where split subfolders are stored.
name (str) – Dataset name, used as the subfolder under
splits_dir.val_split (float) – Fraction of the dataset to use for validation.
- create()[source]¶
Generate a new split and persist it to disk, overwriting any existing one.
Flagged tracks (see
is_flagged()) are excluded from the pool before splitting, so they land in neither side andval_splitis taken as a fraction of the eligible tracks only.- Returns:
Tuple of (train_ds, val_ds). Their
.indicesindex intoself.dataset, not into the filtered pool.- Return type:
tuple[Subset, Subset]
- static load_refs(splits_dir, name, verify=True)[source]¶
Return a split’s
(train_refs, val_refs)directly — no parent dataset needed. The read counterpart tosave_refs(), and what letsTempoDataset/BeatDatasetbe built straight from a split via theirrefs=argument, for a plain or a merged name alike.- Parameters:
verify (bool) – See
load_refs_from_dir().splits_dir (Path)
name (str)
- Raises:
FileNotFoundError – If no split has been generated for
nameyet.MissingTrackDataError – If
verifyand any listed track is not on disk.
- Return type:
- static load_refs_from_dir(split_path, verify=True)[source]¶
Return
(train_refs, val_refs)read straight from split_path’strain.txt/val.txt— the same formatload_refs()reads, but for a folder anywhere on disk rather than one registered under a canonicalsplits_dirby name. Lets a training config point directly at a folder of lists (data.input, when it contains a/) without going throughsplits_dirlookup at all.- Parameters:
verify (bool) – Fail if a listed track is not on disk — see
_read_refs(). Leave it on for anything that will load the audio; passFalseonly to read a split as a list of names.split_path (Path)
- Raises:
FileNotFoundError – If split_path has no
train.txt/val.txt.MissingTrackDataError – If
verifyand any listed track is not on disk.
- Return type:
- run()[source]¶
Return (train_ds, val_ds) loaded from disk.
- Raises:
FileNotFoundError – If no split has been generated for
nameyet.MissingTrackDataError – If the split lists tracks that are not on disk.
- Returns:
Tuple of (train_ds, val_ds).
- Return type:
tuple[Subset, Subset]
- static save_refs(splits_dir, name, train_refs, val_refs)[source]¶
Persist
(train_refs, val_refs)tosplits_dir/<name>/. The write counterpart toload_refs(); also whatcreate()uses internally.
- drop_flagged_refs(refs, context='')[source]¶
Return refs without the tracks
is_flagged()rejects, printing a one-line count when anything was dropped.
- is_flagged(ref)[source]¶
Return whether ref’s annotation is flagged as not fit for training.
Reads the track’s default-slot metadata (the same slot
TempoDataset/BeatDatasetread) and reports True if either annotator flag is set:warning— “this annotation looks wrong” (the annotator’s ⚠ Flag).needs_review— “take another look” (the annotator’s ❓ To Review).
A track with no metadata file at all is not flagged.
- Parameters:
ref (TrackRef)
- Return type:
bool
- missing_files(ref)[source]¶
Return the files ref needs but doesn’t have on disk.
A track is loadable when both halves of this project’s format are present: its audio under
tracks/and its default-slot annotation underannotations/(see docs/source/data.rst). Either one absent makes the track unusable, so both are reported.- Returns:
The absent paths — empty when the track is fully present.
- Parameters:
ref (TrackRef)
- Return type:
list[Path]
- split_name(name, binary_only=False)[source]¶
Split directory name for name:
<name>, or<name>-binary.One split per dataset serves every task. A track’s tempo label is derived from its beat annotation (
TempoDatasetreadsbpm_median, which the migration tools compute from the beats file), so a track is usable for tempo exactly when it is usable for beats. The two pools cannot diverge, and a task-specific split would only be the same tracks drawn into a different train/val partition.binary_onlyis the one flag that does change membership — it drops tracks whose annotated cycle isn’t a multiple of two (seeBeatDataset) — so it gets its own split rather than silently reusing the other variant’s held-out tracks.Nothing else is namespaced: not the task, not
group_size(an unrepresentable meter keeps the track and switches its position supervision off rather than dropping it). So a beat-phase run, a beat-only run and a tempo run over one dataset all hold out the exact same tracks, which is what keeps their numbers comparable.- Parameters:
name (str)
binary_only (bool)
- Return type:
str
- verify_refs_present(refs, context='')[source]¶
Raise unless every ref in refs resolves to files on disk.
- Parameters:
refs (list[TrackRef]) – Refs to check, typically one side of a split.
context (str) – Where the refs came from (a split file), named in the error so a failure points at the list to fix.
- Raises:
MissingTrackDataError – If any ref is missing its audio or its annotation, with a per-corpus breakdown — a partial pull usually takes out whole corpora, and the count per corpus is what tells that apart from a handful of individually broken tracks.
- Return type:
None