The generation model
An encoder-decoder LSTM that reads your pattern and unfolds what comes next, then a post-processing chain that keeps it musical.
Architecture
Once your pattern is on the grid (either drawn, MIDI-loaded, or transcribed from audio), the generator takes over. It uses an encoder-decoder architecture built with 2-layer LSTM (Long Short-Term Memory) networks. LSTMs are a type of recurrent neural network specifically designed to learn long-range dependencies in sequential data — making them ideal for understanding rhythmic patterns that unfold over time.
- The encoder processes your input pattern (up to 32 steps) and compresses it into a hidden state — a compact numerical representation that captures the rhythmic essence of what you played. This hidden state encodes information about which drums are active, their velocity relationships, the density and syncopation of the pattern, and its overall feel.
- The decoder takes this hidden state and generates new steps (up to 16 at a time), unfolding the compressed representation back into a full drum pattern that continues naturally from your input.
- Each step is represented as 24 features: 11 hit presence flags (one per drum class), 11 velocity values (normalized 0-1), and 2 positional encodings (bar position and beat index) that help the model understand where it is in the musical measure.
- The output is 22 values per step: 11 hit probabilities (the likelihood of each drum triggering) and 11 velocity predictions (how hard each drum should be hit).
Training
The model was trained on thousands of drum patterns spanning multiple genres and styles. During training, it learned the statistical relationships between instruments — how kick patterns relate to snare placement, how hi-hat density affects the overall groove feel, how toms and crashes are used to create fills and signal transitions between sections.
The model uses the ONNX (Open Neural Network Exchange) format for efficient cross-platform inference. At only 6.3 MB, the model is small enough for real-time generation with near-instantaneous response times.
Post-Processing pipeline
The raw model output goes through a sophisticated multi-stage post-processing pipeline before becoming the final pattern you see on the grid:
- Temperature Sampling: Each drum class has its own temperature preset tuned for musical results. For example, kicks and snares are more constrained (0.75) to maintain solid rhythmic foundations, while hi-hats are more free (0.90) to allow textural variation. The global temperature knob scales all of these proportionally.
- Density Scaling: Hit probabilities are scaled by the density parameter, giving you direct control over how busy or sparse the generated pattern is.
- Anchor Strength: The model reinforces drums that were present in your input pattern and suppresses those that were not. This prevents the model from "hallucinating" instruments you did not include — if your input only uses kick, snare, and hi-hat, the continuation will respect that palette.
- Input Class Filtering: Any drum class not present in your input is completely filtered out of the generation, ensuring clean, focused output.
- Genre Constraints: When a genre is selected, hand-authored step-weight grids modulate the hit probabilities. These arrays encode where kicks and snares typically fall in that genre's patterns.
- Style preset blending: Selected style presets inject characteristic pattern tendencies and velocity ranges from specific sub-styles, guiding the generation toward a particular sound.
- Humanize: Random timing and velocity jitter is applied to the final output for a more natural, human feel. The amount is controllable from subtle to pronounced.
Inference
Inference runs on ONNX Runtime, the industry-standard engine for deploying machine learning models. The session is configured for single-threaded execution for consistent, predictable performance. Tensor shapes are discovered dynamically from the model at load time, meaning the system automatically adapts to future model upgrades without code changes.
Generation is near-instantaneous — typically completing in under 100 milliseconds. This makes it practical to experiment freely, generating and regenerating patterns until you find the perfect groove.