Paper year
2025
Detect emerging, bridge-candidate, and undercited papers inside a curated audio-ML corpus, then expose the signals behind every recommendation.
Paper dossier
Review source metadata, abstract, authors, topics, and local similarity context before moving into explanation and ranking views.
Paper year
2025
Citations
0
Authors
0
Topic labels
0
Source readout
Unknown venue
unknown
Controlled edge slice
Not available yet
Ranking readout
Ranking details could not be loaded (API 503).
We propose DeepASA, a multi-purpose model for auditory scene analysis that performs multi-input multi-output (MIMO) source separation, dereverberation, sound event detection (SED), audio classification, and direction-of-arrival estimation (DoAE) within a unified framework. DeepASA is designed for complex auditory scenes where multiple, often similar, sound sources overlap in time and move dynamically in space. To achieve robust and consistent inference across tasks, we introduce an object-oriented processing (OOP) strategy. This approach encapsulates diverse auditory features into object-centric representations and refines them through a chain-of-inference (CoI) mechanism. The pipeline comprises a dynamic temporal kernel-based feature extractor, a transformer-based aggregator, and an object separator that yields per-object features. These features feed into multiple task-specific decoders. Our object-centric representations naturally resolve the parameter association ambiguity inherent in traditional track-wise processing. However, early-stage object separation can lead to failure in downstream ASA tasks. To address this, we implement temporal coherence matching (TCM) within the chain-of-inference, enabling multi-task fusion and iterative refinement of object features using estimated auditory parameters. We evaluate DeepASA on representative spatial audio benchmark datasets, including ASA2, MC-FUSS, and STARSS23. Experimental results show that our model achieves state-of-the-art performance across all evaluated tasks, demonstrating its effectiveness in both source separation and auditory parameter estimation under diverse spatial auditory scenes.
No authors available.
Neighborhood labels
Topic labels are imported metadata and can be noisy; use them as coarse navigation hints, not authoritative classifications.
Neighbor surface
Similar papers use a separately configured neighbor embedding; it may differ from the embedding version used by the current ranked run.
No embedding-backed neighbors available for this paper/version yet.
Next handoff
01
Use Recommended to see whether this paper behaves like an emerging or undercited signal in the current ranked feed, or how it appears on the bridge preview / diagnostics view.
02
Use Trends to understand whether its attached labels are heating up or cooling down inside the curated corpus.
03
Use Evaluation to compare the dossier readout against citation and recency baselines for the same resolved family run.