augmentations¶
Waveform augmentations for tempo estimation training.
All augmentations operate on (1, T) float32 tensors. The composed
TempoAugmenter is the public entry point; use
build_augmenter() to construct one from a Hydra config section.
Functions
|
Build a |
Build a |
Classes
|
Add white Gaussian noise at a fixed standard deviation. |
|
Wraps a |
|
Wraps a dataset and applies |
|
Composes waveform+target augmentations for beat-phase training. |
|
Randomly stretch or compress audio and resample the frame-level target to match. |
|
Scale amplitude by a random gain drawn uniformly from a dB range. |
|
Composes waveform augmentations and returns an updated (wav, tempo) pair. |
|
Randomly stretch or compress audio and scale the tempo label accordingly. |
- class AddNoise(std=0.005)[source]¶
Bases:
objectAdd white Gaussian noise at a fixed standard deviation.
- Parameters:
std (float) – Noise standard deviation (relative to full-scale ±1 audio).
- class AugmentedBeatDataset(subset, augmenter, sample_rate, n_samples, n_frames)[source]¶
Bases:
DatasetWraps a
BeatDatasetand appliesBeatPhaseAugmenterat item-access time.Use this to augment only the training split while leaving the validation split unchanged — the two are independent dataset instances (see
build_beat_dataloaders()), so wrapping one has no effect on the other.- Parameters:
subset – A
BeatDataset(ortorch.utils.data.Subsetof one).augmenter (BeatPhaseAugmenter) – The augmenter to apply.
sample_rate (int) – Sample rate of the audio (passed to the augmenter).
n_samples (int) – Fixed clip length in samples (passed to the augmenter).
n_frames (int) – Fixed target length in frames (passed to the augmenter).
- class AugmentedDataset(subset, augmenter, sample_rate, n_samples)[source]¶
Bases:
DatasetWraps a dataset and applies
TempoAugmenterat item-access time.Use this to augment only the training split while leaving the validation split unchanged — the two are independent dataset instances (see
build_dataloaders()), so wrapping one has no effect on the other.- Parameters:
subset – A
TempoDataset(ortorch.utils.data.Subsetof one).augmenter (TempoAugmenter) – The augmenter to apply.
sample_rate (int) – Sample rate of the audio (passed to the augmenter).
n_samples (int) – Fixed clip length in samples (passed to the augmenter).
- class BeatPhaseAugmenter(time_stretch=None, gain=None, noise=None)[source]¶
Bases:
objectComposes waveform+target augmentations for beat-phase training.
Applied in order:
FrameTimeStretch— changes length; re-crops/pads both the waveform and the target back to their expected fixed lengths.
- Parameters:
time_stretch (FrameTimeStretch | None)
gain (RandomGain | None)
noise (AddNoise | None)
- class FrameTimeStretch(min_rate=0.85, max_rate=1.15)[source]¶
Bases:
objectRandomly stretch or compress audio and resample the frame-level target to match.
Same resampling mechanism as
TimeStretch(treating the waveform as recorded atsr * rateand resampling back tosr), but instead of rescaling a scalar label, resamples the target’s time axis by the samerateso beat/one/last/mask events stay aligned with the stretched audio’s frames. Frame count and sample count scale identically under a constant hop length, so using one sharedratefor both is what keeps them in sync.- Parameters:
min_rate (float) – Minimum speed multiplier (< 1 = slower, lower tempo).
max_rate (float) – Maximum speed multiplier (> 1 = faster, higher tempo).
- class RandomGain(min_db=-6.0, max_db=6.0)[source]¶
Bases:
objectScale amplitude by a random gain drawn uniformly from a dB range.
- Parameters:
min_db (float) – Lower bound of the gain range in dB (negative = quieter).
max_db (float) – Upper bound of the gain range in dB (positive = louder).
- class TempoAugmenter(time_stretch=None, gain=None, noise=None)[source]¶
Bases:
objectComposes waveform augmentations and returns an updated (wav, tempo) pair.
Applied in order: 1.
TimeStretch— changes length; re-crops/pads back ton_samples. 2.RandomGain3.AddNoise- Parameters:
time_stretch (TimeStretch | None)
gain (RandomGain | None)
noise (AddNoise | None)
- class TimeStretch(min_rate=0.85, max_rate=1.15)[source]¶
Bases:
objectRandomly stretch or compress audio and scale the tempo label accordingly.
Implemented via resampling: treating the waveform as if it were recorded at
sr * rateHz and resampling back tosrchanges its duration by1 / rate, which is equivalent to speeding up (rate > 1) or slowing down (rate < 1). Pitch also shifts as a side-effect, which is acceptable for tempo estimation.- Parameters:
min_rate (float) – Minimum speed multiplier (< 1 = slower, lower tempo).
max_rate (float) – Maximum speed multiplier (> 1 = faster, higher tempo).
- build_augmenter(cfg)[source]¶
Build a
TempoAugmenterfrom a Hydra augmentations config section.Returns
Noneifcfg.enabledis false or no individual augmentation is enabled, so callers can skip wrapping altogether.- Parameters:
cfg (DictConfig)
- Return type:
TempoAugmenter | None
- build_beat_phase_augmenter(cfg)[source]¶
Build a
BeatPhaseAugmenterfrom a Hydra augmentations config section.Mirrors
build_augmenter(), using the frame-target-awareFrameTimeStretchin place ofTimeStretch.Returns
Noneifcfg.enabledis false or no individual augmentation is enabled, so callers can skip wrapping altogether.- Parameters:
cfg (DictConfig)
- Return type:
BeatPhaseAugmenter | None