Paper dossier

U-MusT: A Unified Framework for Cross-Modal Translation of Score Images, Symbolic Music, and Performance Audio

Detail viewSimilarity handoff

Review source metadata, abstract, authors, topics, and local similarity context before moving into explanation and ranking views.

Paper year

2025

Citations

0

Authors

0

Topic labels

0

Paper ID: W4417094688edge sliceunknown source slug

Source readout

Source and corpus status

Venue

Unknown venue

Source slug

unknown

Corpus placement

Controlled edge slice

Similarity rows

Not available yet

Ranking readout

Where this paper lands in the current run

Ranking details could not be loaded (API 503).

Abstract

Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between such modalities are established as core tasks of music information retrieval, such as automatic music transcription (audio-to-MIDI) and optical music recognition (score image to symbolic score). However, most past work on multimodal translation utilizes specialized models trained for each translation task. In this paper, we propose a unified framework based on a common tokenization strategy. We use dedicated separate models for the Image-to-Audio and Audio-to-Image directions, sharing an identical encoder-decoder architecture to handle each task within a coherent framework. Two key factors make this unified approach viable: a new large-scale dataset, and the tokenization of each modality. Firstly, we propose a new dataset that consists of more than 1,300 hours of paired audio-score image data collected from YouTube videos, which is an order of magnitude larger than any existing music modal translation datasets. Secondly, our unified tokenization framework discretizes score images, audio, MIDI, and MusicXML into a sequence of tokens, enabling standard encoder-decoder Transformers to tackle multiple crossmodal translation as one coherent sequence-to-sequence task. Experimental results confirm that our unified framework improves upon single-task baselines in several key areas, notably reducing the symbol error rate for optical music recognition from 24.58% to a state-of-the-art 13.67%, while also seeing substantial improvements across the other translation tasks. Notably, our approach achieves the first musically-coherent score-image-conditioned audio generation, marking a significant breakthrough in cross-modal music generation.

Authors

No authors available.

Neighborhood labels

Topics

0 labels

Topic labels are imported metadata and can be noisy; use them as coarse navigation hints, not authoritative classifications.

Neighbor surface

Similar papers

Similar papers use a separately configured neighbor embedding; it may differ from the embedding version used by the current ranked run.

No embedding-backed neighbors available for this paper/version yet.

Next handoff

Best next moves from here

01

Check recommendation families

Use Recommended to see whether this paper behaves like an emerging or undercited signal in the current ranked feed, or how it appears on the bridge preview / diagnostics view.

02

Inspect nearby topics

Use Trends to understand whether its attached labels are heating up or cooling down inside the curated corpus.

03

Cross-check evaluation baselines

Use Evaluation to compare the dossier readout against citation and recency baselines for the same resolved family run.