Abstract
Comment: 8 pages, 4 figures. Associated web preview available at https://gclef-cmu.org/multtipop
MulTTiPop is a multitrack transcription dataset of 572 pop-music segments (3.5 hours) spanning 1930s–2000s. The authors matched Lakh MIDI and TheoryTab metadata, manually located anchor beats, then applied beat-tracking and tempo-warping to align MIDI with audio. Evaluations show SOTA transcription models score only 38% Onset F1, exposing substantial gaps for future work.
MulTTiPop contains 572 pop-music segments totaling 3.5 hours of audio, with songs from the 1930s through the 2000s.
Abstract
Comment: 8 pages, 4 figures. Associated web preview available at https://gclef-cmu.org/multtipop