splitter

Functions

drop_flagged_refs(refs[, context])

Return refs without the tracks is_flagged() rejects, printing a one-line count when anything was dropped.

is_flagged(ref)

Return whether ref's annotation is flagged as not fit for training.

missing_files(ref)

Return the files ref needs but doesn't have on disk.

split_name(name[, binary_only])

Split directory name for name: <name>, or <name>-binary.

verify_refs_present(refs[, context])

Raise unless every ref in refs resolves to files on disk.

Classes

Splitter(dataset, splits_dir, name, val_split)

Loads persistent train/val splits for a dataset.

Exceptions

MissingTrackDataError

A split lists tracks whose files are not on this machine.

exception MissingTrackDataError[source]

Bases: FileNotFoundError

A split lists tracks whose files are not on this machine.

Raised by verify_refs_present(), and therefore by every path that reads a train.txt/val.txt (see _read_refs()). The failure it exists to stop is a partial data pull — e.g. dvc pull fetching some corpora but not others on a fresh remote instance. Before this check, a corpus missing from disk was dropped track by track with nothing but a printed line, so a run could train on a strictly smaller dataset than its split describes and report metrics as if nothing had happened.

class Splitter(dataset, splits_dir, name, val_split)[source]

Bases: object

Loads persistent train/val splits for a dataset.

Splits are stored under splits_dir/<name>/train.txt and val.txt, one dataset_name/track_id line per track (see TrackRef) rather than positional indices — so a split’s contents can be resolved, concatenated with another dataset’s split, or loaded standalone (see load_refs()) independently of any one dataset instance’s ordering. This is what lets tools/merge_datasets.py combine several datasets’ existing splits into one.

run() only ever reads these files — it never generates a split, so the same files (typically version-controlled via DVC) produce the same split on every machine. Use create() (or tools/create_splits.py) to generate the files in the first place.

Every read path also checks that the tracks a split lists are actually on disk, and raises MissingTrackDataError if they aren’t (see verify_refs_present()) — a split is a claim about which tracks a run uses, and a partial data pull must stop the run rather than quietly shrink it.

Tracks their annotator flagged (warning or needs_review in the track metadata — see is_flagged()) are skipped on both sides: create() never writes them into a new split, and every read path (run(), load_refs(), load_refs_from_dir()) drops them from splits saved before the flag was set. So flagging a track in the annotator takes it out of training and validation immediately, with no need to regenerate any split file.

Parameters:
  • dataset (Dataset) – The dataset to split. Must expose .refs — a list[TrackRef] parallel to its samples (both TempoDataset and BeatDataset do).

  • splits_dir (Path) – Root directory where split subfolders are stored.

  • name (str) – Dataset name, used as the subfolder under splits_dir.

  • val_split (float) – Fraction of the dataset to use for validation.

create()[source]

Generate a new split and persist it to disk, overwriting any existing one.

Flagged tracks (see is_flagged()) are excluded from the pool before splitting, so they land in neither side and val_split is taken as a fraction of the eligible tracks only.

Returns:

Tuple of (train_ds, val_ds). Their .indices index into self.dataset, not into the filtered pool.

Return type:

tuple[Subset, Subset]

static load_refs(splits_dir, name, verify=True)[source]

Return a split’s (train_refs, val_refs) directly — no parent dataset needed. The read counterpart to save_refs(), and what lets TempoDataset/BeatDataset be built straight from a split via their refs= argument, for a plain or a merged name alike.

Parameters:
Raises:
  • FileNotFoundError – If no split has been generated for name yet.

  • MissingTrackDataError – If verify and any listed track is not on disk.

Return type:

tuple[list[TrackRef], list[TrackRef]]

static load_refs_from_dir(split_path, verify=True)[source]

Return (train_refs, val_refs) read straight from split_path’s train.txt/val.txt — the same format load_refs() reads, but for a folder anywhere on disk rather than one registered under a canonical splits_dir by name. Lets a training config point directly at a folder of lists (data.input, when it contains a /) without going through splits_dir lookup at all.

Parameters:
  • verify (bool) – Fail if a listed track is not on disk — see _read_refs(). Leave it on for anything that will load the audio; pass False only to read a split as a list of names.

  • split_path (Path)

Raises:
  • FileNotFoundError – If split_path has no train.txt/val.txt.

  • MissingTrackDataError – If verify and any listed track is not on disk.

Return type:

tuple[list[TrackRef], list[TrackRef]]

run()[source]

Return (train_ds, val_ds) loaded from disk.

Raises:
  • FileNotFoundError – If no split has been generated for name yet.

  • MissingTrackDataError – If the split lists tracks that are not on disk.

Returns:

Tuple of (train_ds, val_ds).

Return type:

tuple[Subset, Subset]

static save_refs(splits_dir, name, train_refs, val_refs)[source]

Persist (train_refs, val_refs) to splits_dir/<name>/. The write counterpart to load_refs(); also what create() uses internally.

Parameters:
  • splits_dir (Path)

  • name (str)

  • train_refs (list[TrackRef])

  • val_refs (list[TrackRef])

Return type:

None

drop_flagged_refs(refs, context='')[source]

Return refs without the tracks is_flagged() rejects, printing a one-line count when anything was dropped.

Parameters:
  • refs (list[TrackRef]) – Refs to filter.

  • context (str) – Short label for the printed message (e.g. the split file the refs came from), so several calls in one run stay tellable apart.

Return type:

list[TrackRef]

is_flagged(ref)[source]

Return whether ref’s annotation is flagged as not fit for training.

Reads the track’s default-slot metadata (the same slot TempoDataset/BeatDataset read) and reports True if either annotator flag is set:

  • warning — “this annotation looks wrong” (the annotator’s ⚠ Flag).

  • needs_review — “take another look” (the annotator’s ❓ To Review).

A track with no metadata file at all is not flagged.

Parameters:

ref (TrackRef)

Return type:

bool

missing_files(ref)[source]

Return the files ref needs but doesn’t have on disk.

A track is loadable when both halves of this project’s format are present: its audio under tracks/ and its default-slot annotation under annotations/ (see docs/source/data.rst). Either one absent makes the track unusable, so both are reported.

Returns:

The absent paths — empty when the track is fully present.

Parameters:

ref (TrackRef)

Return type:

list[Path]

split_name(name, binary_only=False)[source]

Split directory name for name: <name>, or <name>-binary.

One split per dataset serves every task. A track’s tempo label is derived from its beat annotation (TempoDataset reads bpm_median, which the migration tools compute from the beats file), so a track is usable for tempo exactly when it is usable for beats. The two pools cannot diverge, and a task-specific split would only be the same tracks drawn into a different train/val partition.

binary_only is the one flag that does change membership — it drops tracks whose annotated cycle isn’t a multiple of two (see BeatDataset) — so it gets its own split rather than silently reusing the other variant’s held-out tracks.

Nothing else is namespaced: not the task, not group_size (an unrepresentable meter keeps the track and switches its position supervision off rather than dropping it). So a beat-phase run, a beat-only run and a tempo run over one dataset all hold out the exact same tracks, which is what keeps their numbers comparable.

Parameters:
  • name (str)

  • binary_only (bool)

Return type:

str

verify_refs_present(refs, context='')[source]

Raise unless every ref in refs resolves to files on disk.

Parameters:
  • refs (list[TrackRef]) – Refs to check, typically one side of a split.

  • context (str) – Where the refs came from (a split file), named in the error so a failure points at the list to fix.

Raises:

MissingTrackDataError – If any ref is missing its audio or its annotation, with a per-corpus breakdown — a partial pull usually takes out whole corpora, and the count per corpus is what tells that apart from a handful of individually broken tracks.

Return type:

None