Menu

🎧 Expanding the GTZAN Dataset: A Journey from YouTube to Mel Spectrograms
📰
0

🎧 Expanding the GTZAN Dataset: A Journey from YouTube to Mel Spectrograms

DEV Community: datascience·Alejandro Tacoronte González·5 months ago
#sIKFPcFo
#dev#strong#project#lofi#article#video
Reading 0:00
15s threshold

I am currently finishing my specialization in Artificial Intelligence and Big Data, and I’ve decided to document the progress of my final project. This work integrates everything I’ve learned in the modules of AI Models (Modelos de Inteligencia Artificial) and Machine Learning Systems (Sistemas de Aprendizaje Automático). The goal? A robust Music Genre Classifier. But as any data scientist will tell you, the model is only as good as the data. Today, I focused on building the "Data Kitchen": the pipeline that fetches, cleans, and prepares audio for training. 1. The Challenge: Expanding the GTZAN Dataset While the GTZAN dataset is the industry standard, it lacks modern genres. To make my project unique, I wanted to include Lofi and Rap another others. I used yt-dlp to source high-quality audio from YouTube. However, I ran into a classic "Junior vs. Environment" boss fight: FFmpeg. Technical Tip: Even if you install FFmpeg via Conda, Windows sometimes hides the binaries from your Python subprocesses.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More