DrumTranscriber v2
The first of our two networks: it listens to a drum recording and writes out what it hears, bar by bar.
DrumTranscriber v2 is the first of the two neural networks in the CedarGrooveNET pipeline. Its job: listen to a drum recording and write out what it hears as MIDI-style drum events. It's an encoder-decoder Transformer with ~12.6M parameters, split into separate ONNX encoder and decoder graphs (~54 MB total) so the encoder can run once per clip while the decoder generates tokens autoregressively.
How it works
- Audio preprocessing: The uploaded file is decoded to mono float32, resampled to 22,050 Hz, and normalized to −23 LUFS using EBU R128 loudness measurement — the same target loudness the model was trained on. Without this step the transcriber misses quiet kicks and over-triggers on loud cymbals.
- Mel spectrogram: A 128-bin log-mel spectrogram is computed (2048-point FFT, 256-sample hop, Hann window, reflect padding, Slaney-normalized librosa filterbank). This gives the Transformer a time-frequency image of the audio with 256 time frames.
- Bar windowing: The model works best on focused, short context, so audio is split into one-bar windows based on the detected BPM (or 120 if no BPM hint is provided). Each bar gets its own forward pass and its events are offset by the window start time.
- Encoder pass: The encoder ingests the mel spectrogram and produces a memory tensor (one per window).
- Autoregressive decoding: Starting from a BOS token, the decoder predicts the next token one at a time. Tokens are grouped in triplets of
(time, instrument, velocity), drawn from a 302-token vocabulary (256 time tokens at 10 ms resolution, 11 instrument tokens, 32 velocity bins, plus BOS/EOS/PAD). Decoding stops at EOS or after 128 tokens. - Quantization: Events are snapped to the 16th-note grid based on the target BPM and routed to the 11 drum classes shared with the generator. The last event of each bar is dropped to avoid the model predicting the next bar's downbeat before it has heard it.
What it's good & not so good at
- Works well on: clean drum-only stems, drum loops under 45 seconds, common kit layouts (kick, snare, hi-hats, toms, cymbals, claps).
- Struggles with: mixed bass + drums, heavy distortion, uncommon percussion, very fast fills (sub-32nd-note rolls).
- Limitation: feeding a full mix gets partial results at best — use the stem splitter to isolate the drums first.
Why transcribe to MIDI at all? Because MIDI is editable. Once your loop is on the grid you can nudge ghost notes, swap kicks, lock the hi-hats, regenerate individual rows — and export the result to your DAW as MIDI that triggers your own samples. The audio you dropped is just the seed.