
Federico Simonetta
Published in the Transactions of the International Society for Music Information Retrieval (TISMIR) 2025
Written with GPT-5.4
Many papers in this area promise to identify composers from symbolic scores. This survey asks a better question: when should anyone trust that result?
That is the real interest of the paper. It reviews 58 studies, but it does not treat them as a parade of methods. It checks what they actually measure. A system can reach a high score for the wrong reason. It may learn to separate broad historical periods, genres, or national traditions rather than the fine details that matter in authorship attribution.
The image above says something important on its own. The literature keeps circling around a familiar canon: Bach, Mozart, Haydn, Beethoven. That makes sense historically, but it also creates a comfort zone. Results can look solid when the musical distances are large and the repertories are already well structured.
The paper’s strongest point is its criticism of evaluation habits. Too many studies rely on plain accuracy, even when the datasets are unbalanced. Others use validation protocols that are too weak for the claims they make. For a benchmark paper this is already a problem. For a real attribution case, it is much worse.
So the eye-catching perspective here is not “can a machine spot Bach?” It is that the main bottleneck is often not the model but the test. The survey argues for stricter cross-validation, better metrics such as balanced accuracy, cleaner datasets, and clearer separation between generic composer classification and genuine authorship attribution.
That keeps the paper grounded. It does not deny that computational methods can help musicology. It says that they help only when the experiment is built to rule out easy shortcuts. In this field, caution is not a limitation. It is part of the result.