This is the accompanying demo page to our datasets Beethoven Symphony Excerpt Dataset (BSED) and Beethoven Symphony Dataset (BSD). The data can be found on Zenodo.
If you use BSED/BSD in your research, please cite the following paper:
@article{BerendesSMAM26_BSED,
author = {Berendes, Hans-Ulrich and Saha, Abhirup and Maman, Ben and Arifi-Müller, Vlora and Müller, Meinard},
title = {Beethoven Symphony Excerpt Dataset ({BSED}): An Evaluation Dataset for Orchestral Music Transcription},
journal = {Transactions of the International Society for Music Information Retrieval},
volume = {9},
number = {1},
pages = {405--422},
year = {2026},
doi = {10.5334/tismir.343},
}
Orchestral music poses significant challenges for Automatic Music Transcription (AMT) due to its dense polyphony, diverse instrument timbres, and varied acoustic conditions. While AMT research for solo instruments and chamber music has advanced considerably, orchestral transcription remains underexplored, primarily because of the lack of high-quality, time-aligned datasets. In this work, we make three main contributions. First, we introduce the Beethoven Symphony Excerpt Dataset (BSED), a carefully curated, publicly available benchmark for evaluating orchestral AMT systems. BSED contains 20 short excerpts from Beethoven's nine symphonies, each provided in five audio versions: four distinct concert recordings and one synthetic rendition, with an approximate total duration of 37 minutes. For each excerpt, we supply symbolic scores in multiple formats (PDF, MusicXML, Sibelius, MIDI, CSV) alongside high-quality note-level annotations obtained through robust, manually verified score-audio alignment, refined with transcription-based onset features. Second, to facilitate model development, we release the Beethoven Symphonies Dataset (BSD), a large-scale orchestral training set comprising 62 hours of public-domain recordings with time-aligned annotations derived from digital scores and structurally verified. While less controlled than BSED, BSD offers rich diversity across performances, conductors, and acoustic environments. Third, we establish a baseline for orchestral AMT by training an instrument-agnostic note-level transcription model on BSD and evaluating it on both BSED and the independent PHENICX dataset. Together, BSED and BSD address a critical gap in orchestral MIR resources and provide a foundation for advancing AMT research toward more robust and generalizable systems.
The following table gives an overview of the 20 excerpts. For simplicity we only show one example audio version per excerpt, however, in the dataset, we provide 5 audio versions per excerpt. The example audio shown here is the version from the Blomstedt recordings (splitID 4).
The BSED-IDs are clickable and lead to a more detailed page with score and audio for each excerpt.
| BSED-ID | MovementID | Measures | # Notes | Avg. Dur. [s] | Ex. Audio Version |
|---|---|---|---|---|---|
| BSED-01 | Op021-01 | 1-4 | 159 | 26.8 | |
| BSED-02 | Op021-02 | 6-16 | 256 | 21.6 | |
| BSED-03 | Op036-01 | 0-3 | 73 | 17.4 | |
| BSED-04 | Op036-01 | 35-48 | 556 | 19.5 | |
| BSED-05 | Op036-03 | 1-26 | 396 | 17.0 | |
| BSED-06 | Op055-01 | 1-17 | 318 | 21.5 | |
| BSED-07 | Op055-02 | 143-148 | 347 | 23.1 | |
| BSED-08 | Op055-03 | 231-264 | 234 | 21.8 | |
| BSED-09 | Op060-01 | 1-6 | 72 | 28.5 | |
| BSED-10 | Op060-02 | 98-101 | 279 | 21.6 | |
| BSED-11 | Op067-01 | 1-21 | 247 | 21.2 | |
| BSED-12 | Op067-04 | 110-120 | 367 | 16.7 | |
| BSED-13 | Op068-01 | 1-26 | 300 | 27.7 | |
| BSED-14 | Op068-05 | 194-201 | 500 | 16.9 | |
| BSED-15 | Op092-02 | 1-18 | 171 | 16.9 | |
| BSED-16 | Op092-04 | 31-49 | 525 | 16.5 | |
| BSED-17 | Op093-01 | 1-20 | 769 | 22.2 | |
| BSED-18 | Op093-02 | 30-35 | 322 | 17.8 | |
| BSED-19 | Op125-01 | 374-385 | 376 | 21.7 | |
| BSED-20 | Op125-03 | 3-7 | 102 | 27.3 |
You can find the same score-audio player view for each excerpt by clicking on the BSED-ID in the table above.
The data provided on this website is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
If you publish results obtained using this data, please cite the paper mentioned above.
This work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Grant No. 500643750 (MU 2686/15-1). The International Audio Laboratories Erlangen are a joint institution of the Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU) and Fraunhofer Institute for Integrated Circuits IIS.