TL;DR

my first project: train a keyword spotting model under 1M parameters on Google Speech Commands, built from scratch instead of starting from a pretrained speech model, and measure its accuracy, model size, and inference latency.

the task

given a 1-second audio clip, predict which word was spoken.

"yes"  → YES
"no"   → NO
"stop" → STOP
"go"   → GO

what I want to be able to explain

  • why 16 kHz?
  • what is a waveform?
  • what does the FFT do?
  • why Mel spectrograms?
  • what does the CNN learn?
  • what is cross-entropy?
  • why did accuracy improve?
  • why is the model 400K parameters?
  • where does the model fail?

the plan

about 2 to 3 focused hours a day.

day what I’ll do
2 waveform, sample rate, spectrogram, Mel spectrogram
3 explore Speech Commands; build the dataset loader and preprocessing
4 understand a small CNN and build the first architecture
5 train it; loss, optimizer, batches, epochs
6 evaluate: accuracy, confusion matrix, the mistakes
7 improve: augmentation, architecture, preprocessing
8 measure parameters, model size, and inference latency
9 make it smaller, from 1M toward 500K parameters
10 clean the training and inference code; reproduce the results
11 release irangareddy/kws-* on Hugging Face and write it up

after this

project 2 is harder: take a small pretrained speech model and fine-tune it for ASR. going from a tiny model I understand end to end to a modern pretrained one is the point.