ArXiv

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

Authors
Nathan Pruyne, Benjamin Stoler, William Chen...
Categories
cs.SD, cs.LG
arXiv
https://arxiv.org/abs/2607.08756v1
PDF
https://arxiv.org/pdf/2607.08756v1

Brief

MulTTiPop is a multitrack transcription dataset of 572 pop-music segments (3.5 hours) spanning 1930s–2000s. The authors matched Lakh MIDI and TheoryTab metadata, manually located anchor beats, then applied beat-tracking and tempo-warping to align MIDI with audio. Evaluations show SOTA transcription models score only 38% Onset F1, exposing substantial gaps for future work.

Why it matters

MulTTiPop contains 572 pop-music segments totaling 3.5 hours of audio, with songs from the 1930s through the 2000s.

Key details

  • Dataset construction used metadata matching between the Lakh MIDI and TheoryTab collections, manual identification of an anchor beat, audio beat-tracking, and tempo/time warping to align MIDI with audio.
  • Evaluation of state-of-the-art automatic music transcription models on MulTTiPop finds the best model achieves 38% Onset F1, indicating substantial room for improvement; dataset preview available at https://gclef-cmu.org/multtipop.
Cleaned source text

Abstract

Comment: 8 pages, 4 figures. Associated web preview available at https://gclef-cmu.org/multtipop