Paper year
2025
Detect emerging, bridge-candidate, and undercited papers inside a curated audio-ML corpus, then expose the signals behind every recommendation.
Paper dossier
Review source metadata, abstract, authors, topics, and local similarity context before moving into explanation and ranking views.
Paper year
2025
Citations
0
Authors
0
Topic labels
0
Source readout
Unknown venue
unknown
Controlled edge slice
Not available yet
Ranking readout
Ranking details could not be loaded (API 503).
Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between such modalities are established as core tasks of music information retrieval, such as automatic music transcription (audio-to-MIDI) and optical music recognition (score image to symbolic score). However, most past work on multimodal translation utilizes specialized models trained for each translation task. In this paper, we propose a unified framework based on a common tokenization strategy. We use dedicated separate models for the Image-to-Audio and Audio-to-Image directions, sharing an identical encoder-decoder architecture to handle each task within a coherent framework. Two key factors make this unified approach viable: a new large-scale dataset, and the tokenization of each modality. Firstly, we propose a new dataset that consists of more than 1,300 hours of paired audio-score image data collected from YouTube videos, which is an order of magnitude larger than any existing music modal translation datasets. Secondly, our unified tokenization framework discretizes score images, audio, MIDI, and MusicXML into a sequence of tokens, enabling standard encoder-decoder Transformers to tackle multiple crossmodal translation as one coherent sequence-to-sequence task. Experimental results confirm that our unified framework improves upon single-task baselines in several key areas, notably reducing the symbol error rate for optical music recognition from 24.58% to a state-of-the-art 13.67%, while also seeing substantial improvements across the other translation tasks. Notably, our approach achieves the first musically-coherent score-image-conditioned audio generation, marking a significant breakthrough in cross-modal music generation.
No authors available.
Neighborhood labels
Topic labels are imported metadata and can be noisy; use them as coarse navigation hints, not authoritative classifications.
Neighbor surface
Similar papers use a separately configured neighbor embedding; it may differ from the embedding version used by the current ranked run.
No embedding-backed neighbors available for this paper/version yet.
Next handoff
01
Use Recommended to see whether this paper behaves like an emerging or undercited signal in the current ranked feed, or how it appears on the bridge preview / diagnostics view.
02
Use Trends to understand whether its attached labels are heating up or cooling down inside the curated corpus.
03
Use Evaluation to compare the dossier readout against citation and recency baselines for the same resolved family run.