A perceptual measure for evaluating the resynthesis of automatic music transcriptions

Screenshot of the listening test

Federico Simonetta, S. Ntalampiras, F. Avanzini

Published in Multimedia Tools and Applications 2022

Written with GPT-5.4

The key question is simple: if you change the instrument and the acoustics but keep the same MIDI, are listeners still hearing the same interpretation?

This paper is built around a very concrete musical doubt. When an automatic transcription system turns a recording into MIDI, what survives of the performance? Notes and timing, yes. But what about the sense that a player is shaping a phrase in a particular room, on a particular instrument, for a particular sound?

The listening test is the center of the paper, and the screenshot above gives the right feel for it. Participants did not judge a score or a waveform. They listened to short excerpts and decided whether the interpretation still felt the same after resynthesis. That makes the work more interesting than a standard transcription benchmark, because it shifts the question from correctness on paper to credibility in the ear.

The main result is not flattering to MIDI-only evaluation. If the acoustic context changes, the same MIDI can be heard as a different interpretation. In other words, a performance is not just a list of note events. Part of what listeners recognize as expressive intent lives in the interaction between playing, instrument, and space.

That is the eye-catching idea here. The paper separates “performance” from “interpretation”: the performed gesture is tied to a concrete context, while the intended musical shape is what one hopes to preserve. Once that distinction is made, a plain note-based metric starts to look too thin. A transcription can be objectively close to the reference and still miss what listeners care about.

The proposed perceptual measure is valuable for that reason. It does not claim to solve musical expression in general. It simply tracks listener judgments better than common MIDI-based scores. That is a modest claim, but a useful one, especially for restoration, resynthesis, and any system that wants to sound convincing rather than merely correct.