TL;DR
my first project: train a keyword spotting model under 1M parameters on Google Speech Commands, built from scratch instead of starting from a pretrained speech model, and measure its accuracy, model size, and inference latency.
the task
given a 1-second audio clip, predict which word was spoken.
"yes" → YES
"no" → NO
"stop" → STOP
"go" → GO
what I want to be able to explain
- why 16 kHz?
- what is a waveform?
- what does the FFT do?
- why Mel spectrograms?
- what does the CNN learn?
- what is cross-entropy?
- why did accuracy improve?
- why is the model 400K parameters?
- where does the model fail?
the plan
about 2 to 3 focused hours a day.
| day | what I’ll do |
|---|---|
| 2 | waveform, sample rate, spectrogram, Mel spectrogram |
| 3 | explore Speech Commands; build the dataset loader and preprocessing |
| 4 | understand a small CNN and build the first architecture |
| 5 | train it; loss, optimizer, batches, epochs |
| 6 | evaluate: accuracy, confusion matrix, the mistakes |
| 7 | improve: augmentation, architecture, preprocessing |
| 8 | measure parameters, model size, and inference latency |
| 9 | make it smaller, from 1M toward 500K parameters |
| 10 | clean the training and inference code; reproduce the results |
| 11 | release irangareddy/kws-* on Hugging Face and write it up |
after this
project 2 is harder: take a small pretrained speech model and fine-tune it for ASR. going from a tiny model I understand end to end to a modern pretrained one is the point.