augmentations

Waveform augmentations for tempo estimation training.

All augmentations operate on (1, T) float32 tensors. The composed TempoAugmenter is the public entry point; use build_augmenter() to construct one from a Hydra config section.

Functions

build_augmenter(cfg)

Build a TempoAugmenter from a Hydra augmentations config section.

build_beat_phase_augmenter(cfg)

Build a BeatPhaseAugmenter from a Hydra augmentations config section.

Classes

AddNoise([std])

Add white Gaussian noise at a fixed standard deviation.

AugmentedBeatDataset(subset, augmenter, ...)

Wraps a BeatDataset and applies BeatPhaseAugmenter at item-access time.

AugmentedDataset(subset, augmenter, ...)

Wraps a dataset and applies TempoAugmenter at item-access time.

BeatPhaseAugmenter([time_stretch, gain, noise])

Composes waveform+target augmentations for beat-phase training.

FrameTimeStretch([min_rate, max_rate])

Randomly stretch or compress audio and resample the frame-level target to match.

RandomGain([min_db, max_db])

Scale amplitude by a random gain drawn uniformly from a dB range.

TempoAugmenter([time_stretch, gain, noise])

Composes waveform augmentations and returns an updated (wav, tempo) pair.

TimeStretch([min_rate, max_rate])

Randomly stretch or compress audio and scale the tempo label accordingly.

class AddNoise(std=0.005)[source]

Bases: object

Add white Gaussian noise at a fixed standard deviation.

Parameters:

std (float) – Noise standard deviation (relative to full-scale ±1 audio).

class AugmentedBeatDataset(subset, augmenter, sample_rate, n_samples, n_frames)[source]

Bases: Dataset

Wraps a BeatDataset and applies BeatPhaseAugmenter at item-access time.

Use this to augment only the training split while leaving the validation split unchanged — the two are independent dataset instances (see build_beat_dataloaders()), so wrapping one has no effect on the other.

Parameters:
  • subset – A BeatDataset (or torch.utils.data.Subset of one).

  • augmenter (BeatPhaseAugmenter) – The augmenter to apply.

  • sample_rate (int) – Sample rate of the audio (passed to the augmenter).

  • n_samples (int) – Fixed clip length in samples (passed to the augmenter).

  • n_frames (int) – Fixed target length in frames (passed to the augmenter).

class AugmentedDataset(subset, augmenter, sample_rate, n_samples)[source]

Bases: Dataset

Wraps a dataset and applies TempoAugmenter at item-access time.

Use this to augment only the training split while leaving the validation split unchanged — the two are independent dataset instances (see build_dataloaders()), so wrapping one has no effect on the other.

Parameters:
  • subset – A TempoDataset (or torch.utils.data.Subset of one).

  • augmenter (TempoAugmenter) – The augmenter to apply.

  • sample_rate (int) – Sample rate of the audio (passed to the augmenter).

  • n_samples (int) – Fixed clip length in samples (passed to the augmenter).

class BeatPhaseAugmenter(time_stretch=None, gain=None, noise=None)[source]

Bases: object

Composes waveform+target augmentations for beat-phase training.

Applied in order:

  1. FrameTimeStretch — changes length; re-crops/pads both the waveform and the target back to their expected fixed lengths.

  2. RandomGain

  3. AddNoise

Parameters:
class FrameTimeStretch(min_rate=0.85, max_rate=1.15)[source]

Bases: object

Randomly stretch or compress audio and resample the frame-level target to match.

Same resampling mechanism as TimeStretch (treating the waveform as recorded at sr * rate and resampling back to sr), but instead of rescaling a scalar label, resamples the target’s time axis by the same rate so beat/one/last/mask events stay aligned with the stretched audio’s frames. Frame count and sample count scale identically under a constant hop length, so using one shared rate for both is what keeps them in sync.

Parameters:
  • min_rate (float) – Minimum speed multiplier (< 1 = slower, lower tempo).

  • max_rate (float) – Maximum speed multiplier (> 1 = faster, higher tempo).

class RandomGain(min_db=-6.0, max_db=6.0)[source]

Bases: object

Scale amplitude by a random gain drawn uniformly from a dB range.

Parameters:
  • min_db (float) – Lower bound of the gain range in dB (negative = quieter).

  • max_db (float) – Upper bound of the gain range in dB (positive = louder).

class TempoAugmenter(time_stretch=None, gain=None, noise=None)[source]

Bases: object

Composes waveform augmentations and returns an updated (wav, tempo) pair.

Applied in order: 1. TimeStretch — changes length; re-crops/pads back to n_samples. 2. RandomGain 3. AddNoise

Parameters:
class TimeStretch(min_rate=0.85, max_rate=1.15)[source]

Bases: object

Randomly stretch or compress audio and scale the tempo label accordingly.

Implemented via resampling: treating the waveform as if it were recorded at sr * rate Hz and resampling back to sr changes its duration by 1 / rate, which is equivalent to speeding up (rate > 1) or slowing down (rate < 1). Pitch also shifts as a side-effect, which is acceptable for tempo estimation.

Parameters:
  • min_rate (float) – Minimum speed multiplier (< 1 = slower, lower tempo).

  • max_rate (float) – Maximum speed multiplier (> 1 = faster, higher tempo).

build_augmenter(cfg)[source]

Build a TempoAugmenter from a Hydra augmentations config section.

Returns None if cfg.enabled is false or no individual augmentation is enabled, so callers can skip wrapping altogether.

Parameters:

cfg (DictConfig)

Return type:

TempoAugmenter | None

build_beat_phase_augmenter(cfg)[source]

Build a BeatPhaseAugmenter from a Hydra augmentations config section.

Mirrors build_augmenter(), using the frame-target-aware FrameTimeStretch in place of TimeStretch.

Returns None if cfg.enabled is false or no individual augmentation is enabled, so callers can skip wrapping altogether.

Parameters:

cfg (DictConfig)

Return type:

BeatPhaseAugmenter | None