Paper dossier

An Enhanced CNN Architecture for Music Genre Classification on the GTZAN Dataset

Detail viewSimilarity handoff

Review source metadata, abstract, authors, topics, and local similarity context before moving into explanation and ranking views.

Paper year

2025

Citations

1

Authors

0

Topic labels

0

Paper ID: W4415034087edge sliceunknown source slug

Source readout

Source and corpus status

Venue

Unknown venue

Source slug

unknown

Corpus placement

Controlled edge slice

Similarity rows

Not available yet

Ranking readout

Where this paper lands in the current run

Ranking details could not be loaded (API 503).

Abstract

A principal objective within contemporary Music Information Retrieval (MIR) research is the development of automated systems for genre classification, especially due to the exponential proliferation of digital audio content on platforms such as streaming services, online radio, and algorithmically generated playlists. Manual annotation is no longer viable, thereby necessitating scalable and intelligent classification solutions. Music Genre Classification Using Convolutional Neural Networks presents a comprehensive examination of automatic music genre classification using deep learning frameworks, augmented by signal processing and traditional machine learning methodologies. The GTZAN genre collection, comprising 1,000 audio tracks each with a duration of 30 seconds, serves as the primary dataset. This benchmark includes ten balanced musical genres: blues, classical, country, disco, hip-hop, jazz, metal, pop, reggae, and rock. Feature extraction is performed using both time-domain and frequencydomain techniques. To address the challenges inherent in modeling complex, high-dimensional audio data, Music Genre Classification Using Convolutional Neural Networks proposes a specialized CNN architecture that utilizes log-mel spectrogram representations of the audio signal as two-dimensional input. Data augmentation techniques such as noise injection, pitch shifting, and time stretching are employed to improve model robustness and generalization across diverse musical content. The CNN model achieves an average classification accuracy of <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{9 5. 2 \%}$</tex>, demonstrating strong capability in learning genrespecific acoustic patterns. Analysis of the confusion matrix reveals classification challenges in genres with overlapping sonic characteristics, such as classical and jazz or rock and metal. Nevertheless, the high precision and recall across most categories affirm the effectiveness of CNN-based methods for music genre recognition and their applicability to large-scale music retrieval and recommendation systems.

Authors

No authors available.

Neighborhood labels

Topics

0 labels

Topic labels are imported metadata and can be noisy; use them as coarse navigation hints, not authoritative classifications.

Neighbor surface

Similar papers

Similar papers use a separately configured neighbor embedding; it may differ from the embedding version used by the current ranked run.

No embedding-backed neighbors available for this paper/version yet.

Next handoff

Best next moves from here

01

Check recommendation families

Use Recommended to see whether this paper behaves like an emerging or undercited signal in the current ranked feed, or how it appears on the bridge preview / diagnostics view.

02

Inspect nearby topics

Use Trends to understand whether its attached labels are heating up or cooling down inside the curated corpus.

03

Cross-check evaluation baselines

Use Evaluation to compare the dossier readout against citation and recency baselines for the same resolved family run.