Paper year
2025
Detect emerging, bridge-candidate, and undercited papers inside a curated audio-ML corpus, then expose the signals behind every recommendation.
Paper dossier
Review source metadata, abstract, authors, topics, and local similarity context before moving into explanation and ranking views.
Paper year
2025
Citations
0
Authors
0
Topic labels
0
Source readout
Unknown venue
unknown
Controlled edge slice
Not available yet
Ranking readout
Ranking details could not be loaded (API 503).
Citizen science engages volunteers to contribute data to scientific projects, often through visual annotation tasks. Hearing based activities are rare and less well understood. Having high quality annotations of performed music structures is essential for reliable algorithmic analysis of recorded music with applications ranging from music information retrieval to music therapy. Music annotations typically begin with an aural input combined with a variety of visual representations, but the impact of the visuals and aural inputs on the annotations are not known. Here, we present a study where participants annotate music segmentation boundaries of variable strengths given only visuals (audio waveform or piano roll) or only audio or both visuals and audio simultaneously. Participants were presented with the set of 33 contrasting theme and variations extracted from a through-recorded performance of Beethoven's 32 Variations in C minor, WoO 80, under differing audiovisual conditions. Their segmentation boundaries were visualized using boundary credence profiles and compared using the unbalanced optimal transport distance, which tracks boundary weights and penalizes boundary removal, and compared to the F-measure. Compared to annotations derived from audio/visual (cross-modal) input (considered as the gold standard for our study), boundary annotations derived from visual (unimodal) input were closer than those derived from audio (unimodal) input. The presence of visuals led to larger peaks in boundary credence profiles, marking clearer global segmentations, while audio helped resolve discrepancies and capture subtle segmentation cues. We conclude that audio and visual inputs can be used as cognitive scaffolding to enhance results in large-scale citizen science annotation of music media and to support data analysis and interpretation. In summary, visuals provide cues for big structures, but complex structural nuances are better discerned by ear.
No authors available.
Neighborhood labels
Topic labels are imported metadata and can be noisy; use them as coarse navigation hints, not authoritative classifications.
Neighbor surface
Similar papers use a separately configured neighbor embedding; it may differ from the embedding version used by the current ranked run.
No embedding-backed neighbors available for this paper/version yet.
Next handoff
01
Use Recommended to see whether this paper behaves like an emerging or undercited signal in the current ranked feed, or how it appears on the bridge preview / diagnostics view.
02
Use Trends to understand whether its attached labels are heating up or cooling down inside the curated corpus.
03
Use Evaluation to compare the dossier readout against citation and recency baselines for the same resolved family run.