How CedarGroove works
Two models of our own, two we build on, and a server we run ourselves. What leaves your machine, and for how long, is spelled out below.
CedarGrooveNET is an AI-powered drum pattern tool that takes either a MIDI file or an audio drum loop (WAV, MP3, FLAC, AIFF) and extends it into a new pattern you can drag straight into your DAW. Rather than producing random beats or relying on simple algorithmic rules, it pairs two neural networks trained on thousands of real drum recordings.
If you start from audio, the DrumTranscriber v2 — a Transformer encoder-decoder that emits MIDI-style tokens — listens to the loop and writes out what each drum is playing, step by step. If you start from MIDI, it skips straight to the second stage. Either way, you end up with a 16th-note grid that's ready to edit, extend, or regenerate.
The generation stage is an encoder-decoder LSTM (Long Short-Term Memory) network. The encoder reads your input pattern and compresses it into a latent representation that captures the rhythmic essence — the groove, the feel, the dynamics. The decoder then unfolds that representation into new steps that continue where your pattern left off, producing output that sounds like a natural extension of what you started with.
The model supports 11 drum classes: Kick, Snare, Closed Hi-Hat, Open Hi-Hat, Low Tom, Mid Tom, High Tom, Crash Cymbal, Ride Cymbal, Clap, and Rimshot. This comprehensive instrument set covers the full palette of modern drum programming, from four-on-the-floor house kicks to intricate jungle breakbeats.
What makes CedarGrooveNET different from a simple random pattern generator is that the model understands musical structure. It has learned that kicks anchor the downbeat, snares land on the backbeat, hi-hats provide rhythmic texture and motion, and toms and crashes are used for fills and transitions. It knows that a sparse, open pattern calls for a different continuation than a dense, driving groove. It understands velocity dynamics — that ghost notes on the snare create a different feel than full-velocity rimshots.
The generation is not just random — the model has learned the statistical relationships between drum instruments across thousands of patterns. When it sees a two-step kick pattern, it knows that certain snare placements are more likely than others. When it detects a syncopated hi-hat rhythm, it adjusts the generated continuation to maintain that syncopation rather than defaulting to straight 16ths.
- Drum transcription model: DrumTranscriber v2 — encoder-decoder Transformer (~12.6M params, ~54 MB split ONNX).
- Generation model: Encoder-decoder 2-layer LSTM (~6.3 MB ONNX), executed via ONNX Runtime for fast, portable inference.
- Audio DSP (backend): NumPy + SciPy (resampling), pyloudnorm (EBU R128), librosa-compatible mel filterbank.
- Preview engine (browser): Web Audio API with 11 synthesized drum voices generated entirely in the browser — no sample libraries required.
- Backend: FastAPI + ONNX Runtime, single-threaded sessions for consistent latency.
- No external dependencies: The macOS app runs entirely locally, with no cloud services and no account required. The browser version sends your file to our own server, never a third-party cloud.
Read on: the drum transcriber, the generation model, the nine reference genres, and the research it stands on.
What we build on
- Stem separation: Demucs (
htdemucs) runs out of process on our server. A job and its stems live for one hour so you can download them, then they are deleted. - Pitched audio to MIDI: Basic Pitch for the notes and SwiftF0 for the pitch track, one instrument per file. The result also comes back as MusicXML and an engraved score.
- Key and BPM: estimated from the file itself in a few seconds; nothing is stored.
The server, and what it keeps
FastAPI and ONNX Runtime on a machine we run ourselves; nothing goes to a third-party cloud. Drum transcription keeps your audio in memory only and discards it once the result is returned. A stem split or an audio-to-MIDI job is kept on the server for one hour so you can download it, then deleted; the MIDI from an audio-to-MIDI transcription stays in your account for twelve months, or until you delete it.
Questions
How does the AI generate drum continuations?
An encoder-decoder 2-layer LSTM network reads your input pattern, compresses it into a latent representation that captures the groove, then unfolds that representation into new 16th-note steps. Post-processing applies temperature sampling, input-class filtering, genre constraints, and humanize for a musical result.
How does CedarGrooveNET compare to AnthemScore?
AnthemScore is a paid commercial transcription tool focused on pitched instruments and chord detection — it does not have built-in drum transcription. CedarGrooveNET is free, drum-specific, and adds a generative continuation step AnthemScore lacks. If you need to transcribe a drum loop to MIDI, CedarGrooveNET is purpose-built for that; AnthemScore is for transcribing melodies and chords.
How does CedarGrooveNET compare to Demucs?
Demucs and CedarGrooveNET solve different problems and are complementary. Demucs is the source-separation model behind the free stem splitter on this site at cedargroove.com/stem-splitter/, which pulls the vocals, drums, bass and other stems out of a full mix. CedarGrooveNET takes an isolated drum stem (or a clean drum recording) and converts it to MIDI, then optionally extends the pattern. The two run back to back here: split the song, then press Transcribe these drums on the drum stem.
What can CedarGrooveNET not do?
The drum transcriber does drums only; bass, piano, guitar and vocals go through the audio-to-MIDI converter, which also writes MusicXML and a score. Neither edits pitch or tuning the way Melodyne does, and neither is a DAW or a plugin: they export files for use elsewhere. The transcriber is not a stem separator on its own, but the stem splitter pulls drums, vocals, bass and other out of a full mix, and its drum stem feeds straight into the transcriber. Clips longer than 6 minutes are rejected.