MulTTiPop: A Multitrack Transcription Dataset for Pop Music

2026-07-09Sound

SoundMachine Learning
AI summary

The authors created MulTTiPop, a collection of pop music clips paired with detailed MIDI files showing the notes played. This collection covers various genres and decades from the 1930s to the 2000s and totals 3.5 hours of music. They carefully matched MIDI files to audio by aligning their beats and tempos. When testing current music transcription computer programs on this dataset, the authors found these programs still struggle, with the best one only correctly identifying about 38% of note starts. The dataset and examples are available online for others to use and improve transcription models.

automatic music transcriptionMIDIpop musicbeat trackingtempo warpingLakh MIDI datasetTheoryTab datasetOnset F1music segmentationmultitrack recordings
Authors
Nathan Pruyne, Benjamin Stoler, William Chen, Chien-yu Huang, Shinji Watanabe, Chris Donahue
Abstract
We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evaluate state-of-the-art automatic music transcription models on MulTTiPop and find substantial room for improvement, with the best model achieving 38% Onset F1. More details and sound examples of MulTTiPop are available at https://gclef-cmu.org/multtipop.