tcn¶
Classes
|
Local time-frequency processing in front of the 1D trunk. |
|
Additive sinusoidal positional encoding (Vaswani et al., 2017). |
|
One transformer-encoder-style block: self-attention sublayer, then a feedforward sublayer, each wrapped in its own residual connection and LayerNorm. |
|
Dilated TCN for tempo regression (Davies & Böck, 2019), or per-frame beat-phase detection. |
- class Conv2dStem(n_mels, channels=16, n_layers=3, freq_pool=3)[source]¶
Bases:
ModuleLocal time-frequency processing in front of the 1D trunk.
TCNTempoNet’s first operation used to beConv1d(n_mels, channels, kernel_size=1): at every frame, a fixed linear mixture of all mel bands, after which the band axis is gone and every remaining layer convolves over time only. Nothing in the model ever saw a time-frequency neighbourhood.That is the wrong first operation for beat tracking, for two reasons (
plans/08_rethinking_the_approach.md§2.1 has the long version):A 1x1 mixer is frequency-absolute; onsets are frequency-relative. It learns one weight per band, applied identically at every frame — so it can learn “these bands matter”, but not “energy rose in whichever band it rose in”. A kick drum and a walking bass note are the same event shape at different absolute frequencies, and a mixer must spend separate output channels on each register to detect the same thing twice. A
Conv2dshares one kernel across frequency and gets that equivariance for free, which is the reason to put audio on a log-frequency axis at all.Mixing before differencing lets onsets cancel. Onset strength is a difference across time within a band. With the bands summed first, a band rising and another falling by the same weighted amount produces a flat mixture, and the event is gone before any layer could see it. That is exactly a harmonic change with no percussive attack — the dominant downbeat cue wherever no drum marks the bar.
Every published tracker does local spectro-temporal processing first (madmom’s TCN opens with 3x3 convolutions and frequency max-pooling; Beat This! uses frequency-wise partial attention), and both reach their numbers with a weak decoder or none at all, which is what points at the front end.
Shape, with the defaults and
n_mels=128:(B, 128, T) log-mel, as the trunk used to receive it (B, 1, 128, T) band axis promoted to a spatial axis (B, 16, 128, T) Conv2d 3x3 + BN + GELU (B, 16, 42, T) MaxPool2d((3, 1)) — frequency only (B, 16, 42, T) Conv2d 3x3 + BN + GELU (B, 16, 14, T) MaxPool2d((3, 1)) (B, 16, 14, T) Conv2d 3x3 + BN + GELU (B, 224, T) frequency folded into channels
Time is never pooled. The frame rate is the output resolution, and the trunk’s dilations are what buy context — pooling time here would spend precision the 70 ms evaluation tolerance cannot afford. The 3x3 kernels do widen the receptive field by
2 * n_layersframes, which is negligible beside the trunk’s1 + 2 * sum(dilations).- Parameters:
n_mels (int) – Number of input mel bands.
channels (int) – Feature maps per 2D layer. 16 is madmom-scale; the stem is meant to be cheap next to the trunk, not to hold capacity.
n_layers (int) – Number of
Conv2dblocks. Frequency is pooled after every block except the last, son_layers=3pools twice.freq_pool (int) – Frequency pooling factor per pool.
- Raises:
ValueError – If n_layers is below 1, or if the pooling schedule would leave fewer than one frequency bin.
- forward(x)[source]¶
- Parameters:
x (Tensor) – Normalised log-mel, shape
(B, n_mels, T).- Returns:
(B, out_channels, T)— T unchanged.- Return type:
Tensor
- n_freq¶
Frequency bins surviving the pooling schedule.
- out_channels¶
Channel count the trunk’s
input_projmust expect.
- class PositionalEncoding(channels)[source]¶
Bases:
ModuleAdditive sinusoidal positional encoding (Vaswani et al., 2017).
Self-attention has no built-in notion of frame order (unlike convolution or recurrence), so this injects one: each position gets a fixed sin/cos pattern that varies by position and by channel pair, added directly to the input.
- Parameters:
channels (int) – Channel width. Must be even — sin/cos are paired per two channels.
- class SelfAttentionBlock(channels, n_heads)[source]¶
Bases:
ModuleOne transformer-encoder-style block: self-attention sublayer, then a feedforward sublayer, each wrapped in its own residual connection and LayerNorm. Lets every frame’s representation draw on every other frame in the sequence, unlike the TCN trunk’s fixed dilated-conv receptive field (see docs/beat_phase_context_ideas.md).
- Parameters:
channels (int) – Channel width (attention embedding dim).
n_heads (int) – Number of attention heads.
- class TCNTempoNet(n_mels=128, sample_rate=22050, hop_length=512, channels=32, n_layers=8, dropout=0.3, n_outputs=1, frame_level=False, use_self_attention=False, n_attn_layers=1, n_attn_heads=4, conv2d_stem=False, stem_channels=16, stem_layers=3, stem_freq_pool=3, fixed_norm=False)[source]¶
Bases:
ModuleDilated TCN for tempo regression (Davies & Böck, 2019), or per-frame beat-phase detection.
Applies a log-mel transform, projects to the TCN channel width, then runs a stack of dilated 1D residual convolutions with exponentially growing dilation (1, 2, 4, …, 2^(n_layers-1)).
Two output modes, controlled by
frame_level:frame_level=False(default): globally pools over time, then a small FC head produces scalar/bin regression or classification logits. Input: (B, 1, T) → Output: (B,) or (B, n_outputs).frame_level=True: skips the pool; a 1x1 conv head produces per-frame logits instead (e.g. beat/one/last for beat-phase detection). Input: (B, 1, T) → Output: (B, n_outputs, T’) or (B, T’) if n_outputs == 1, where T’ is the mel transform’s frame count. Sigmoid is not applied — pair withBCEWithLogitsLossdownstream, matching the classification mode’s convention of returning raw logits.
Receptive field is
1 + (kernel_size - 1) * sum(dilations)frames — with the default schedule1, 2, ..., 2^(n_layers-1)that is2^(n_layers+1) - 1, so 511 frames ≈ 11.9 s atn_layers=8,hop_length=512. (This docstring used to quotekernel_size × (2^n_layers − 1), a loose upper bound ~1.5x the truth; seedocs/beat_phase_context_ideas.mdandplans/08§1.2/§3.1.) The same trunk is shared between both modes, so this is unaffected byframe_level. Aconv2d_stemadds2 × stem_layersframes to that, which is noise beside it.The receptive field is only real if the input is at least that long. Every trunk conv uses
padding=dilation, so a layer whose dilation exceeds the input length has both off-centre taps in zero padding at every frame and collapses to a 1x1 conv. On a 16 s clip athop_length=512(689 frames) that is any layer past the ninth. Seeplans/08§3.1.conv2d_stemis the one structural option here. Off, the first operation is a 1x1 mix over mel bands and no layer ever sees a time-frequency neighbourhood; on, a smallConv2dStemruns first. See that class for why it exists.- Parameters:
n_mels (int) – Number of mel filterbanks.
sample_rate (int) – Audio sample rate used to build the mel transform.
hop_length (int) – Hop length for the mel transform. Controls temporal resolution (smaller = more frames per second). Defaults to 512 (≈43 fps at 22050 Hz).
channels (int) – Channel width for the TCN.
n_layers (int) – Number of dilated layers; dilation doubles per layer. Keep the dilation of the deepest layer (
2^(n_layers-1)frames) below the input sequence length, or that layer only ever convolves padding.dropout (float) – Dropout probability applied right before each head’s final
Conv1d/Linear— the pooled regression head’s lastLinear, the frame head’s 1x1 conv, or (whenuse_self_attention=True) thebeat_head/phase_head1x1 convs.n_outputs (int) – Output dimension. In pooled mode,
1for scalar regression, > 1 for classification over tempo bins. In frame-level mode, the number of per-frame target channels (e.g. 3 for beat/one/last).frame_level (bool) – If
True, produce per-frame outputs instead of pooling over time.use_self_attention (bool) – Frame-level mode only. If
True, splits the frame head in two:beat_headreads straight off the TCN trunk (unchanged, already accurate), whilephase_headroutes the remainingn_outputs - 1channels through a positional encoding + a stack ofSelfAttentionBlock, giving them context beyond the trunk’s fixed dilated-conv receptive field. Output channel order is alwaysbeatfirst, then the phase channels —(beat, one, last)formusicality.losses.beat_phase.beat_phase_loss(), or(beat, pos_1, ..., pos_G)formusicality.losses.beat_position.beat_position_loss(). See docs/beat_phase_context_ideas.md.n_attn_layers (int) – Number of stacked
SelfAttentionBlockinphase_head. Only used whenuse_self_attention=True.n_attn_heads (int) – Attention heads per
SelfAttentionBlock. Only used whenuse_self_attention=True.conv2d_stem (bool) – Run a
Conv2dStembetween the mel and the trunk instead of projecting the raw bands. Defaults toFalsehere so that checkpoints predating the stem reconstruct into identical parameter shapes from their own saved hyperparameters; the shipped frame-level backbones turn it on. Measured cost at theirchannels=32: 32,805 parameters to 40,773 — the stem itself is 4.9 k, the rest isinput_projwidening from 128 to 224 inputs.stem_channels (int) – Feature maps per stem layer.
conv2d_stemonly.stem_layers (int) –
Conv2dblocks in the stem; frequency is pooled after all but the last.conv2d_stemonly.stem_freq_pool (int) – Frequency pooling factor per stem pool.
conv2d_stemonly.fixed_norm (bool) – Normalise the log-mel with frozen per-band statistics (
norm_mean/norm_std, filled byfit_input_stats()at the start of training and saved in the checkpoint) instead of statistics taken over the input tensor. Off reproduces what every checkpoint up to v6 was trained with; on removes a train/inference mismatch, since the per-tensor statistics depend on whether the input is a 16 s crop or a whole track.plans/08§2.1 item 4 has the measurement.