Paper dossier

Predicting Perceived Semantic Expression of Functional Sounds Using Unsupervised Feature Extraction and Ensemble Learning

Detail viewSimilarity handoff

Review source metadata, abstract, authors, topics, and local similarity context before moving into explanation and ranking views.

Paper year

2026

Citations

0

Authors

5

Topic labels

3

Source readout

Source and corpus status

Venue

Transactions of the International Society for Music Information Retrieval

Source slug

tismir

Corpus placement

Core corpus

Similarity rows

6

Ranking readout

Where this paper lands in the current run

Ranking details could not be loaded (API 503).

Abstract

Functional sounds-typically brief, nonverbal audio cues used in the interfaces of electronic devices-play a critical role in human-machine interaction but remain largely unexplored within music information retrieval (MIR). This study proposes a data-driven framework that uses musically informed audio features to predict the perceived semantic expression of functional sounds. Our three-stage pipeline first uses unsupervised feature extraction to transform 805 functional sounds into high-level topic distributions for timbre, chroma, and loudness using Gaussian mixture models and latent Dirichlet allocation. Second, these features train multi-output regression models to predict 19 perceptual dimensions from the FBMUX framework, with a random forest regressor achieving the best performance. Finally, a listening experiment assesses how well the model predictions align with user perceptions. Interpretability analyses further reveal how individual features contribute to model predictions. This work contributes to MIR by expanding its scope to the domain of functional, non-musical audio. It presents a novel application of MIR techniques, demonstrating that structured, musically informed descriptors can support perceptual modeling in domains with limited data and high subjective variance. It contributes a transferable approach and highlights the potential of MIR to inform human-machine interaction and sound design.

Authors

  • Annika Frommholz
  • Steffen Lepa
  • Tom Virkus
  • Stefan Weinzierl
  • Johannes Helberger

Neighborhood labels

Topics

3 labels

Topic labels are imported metadata and can be noisy; use them as coarse navigation hints, not authoritative classifications.

Music and Audio ProcessingEmotion and Mood RecognitionMusic Technology and Sound Studies

Neighbor surface

Similar papers

6 total neighborsEmbedding v1-title-abstract-1536-cleantext-r3

Similar papers use a separately configured neighbor embedding; it may differ from the embedding version used by the current ranked run.

Next handoff

Best next moves from here

01

Check recommendation families

Use Recommended to see whether this paper behaves like an emerging or undercited signal in the current ranked feed, or how it appears on the bridge preview / diagnostics view.

02

Inspect nearby topics

Use Trends to understand whether its attached labels are heating up or cooling down inside the curated corpus.

03

Cross-check evaluation baselines

Use Evaluation to compare the dossier readout against citation and recency baselines for the same resolved family run.