2026 · completed
Transcribing every call
An in-house speech-to-text pipeline that transcribes all of a call centre’s calls at a fraction of the vendor’s cost.
ML research internship, AI Lab, Selectra, Madrid, June–September 2026.
- word error rate, raw model → final pipeline (preprocessing + fine-tuning)
- 29.4 → 17.1%
- lower running cost than the best vendor ($1.5 vs $222 per 1,000 h of audio)
- 150×
- of calls quality-checked, previously a random sample
- 100%
How it works
Calls are pulled nightly, loudness-normalised, cut to speech by a voice-activity detector, filtered for noise, and transcribed by an open-weight model (NVIDIA Parakeet, 600M parameters) fine-tuned on real call recordings; each fine-tune costs one to two dollars and fifteen minutes; the batch runs on spot GPU instances with cost caps and automated tests.
Results
The fine-tuned model edges past the best vendor (17.1% against 17.5% word error rate) for $1.5 per thousand hours of audio instead of $222: about $800 a year instead of $120,000, on more than 550,000 hours. The whole volume is now transcribed and quality-checked, not a sample.
Technical details
Model
NVIDIA's Parakeet: a small open-weight speech model of 600 million parameters, a FastConformer encoder followed by a token-and-duration transducer, cheap enough to run on one GPU. Out of the box it produced sentences that were never said, wrote text over inaudible segments and switched to English in French or Spanish calls.
- Base model
- NVIDIA Parakeet (600M parameters, open weights)
- Architecture
- FastConformer encoder, TDT transducer decoder
- Why a small model
- domain vocabulary (contracts, meter readings) beats size, and it is cheap to run
Preprocessing
Recording levels varied a lot from call to call. Loudness normalisation (EBU R128) followed by a voice-activity detector (Silero VAD), which removes silences, hold tones and music to keep only speech, cut the word error rate by six points without touching the model.
- Loudness
- EBU R128 normalisation
- Voice activity
- Silero VAD
- Gain
- ≈ −6 WER points, model unchanged
Fine-tuning
There was no annotated data at the start, so I transcribed ninety minutes of calls by hand to get a reference, and every version of the system was scored on that same test set. Synthetic voices did not match real agents, and clips that were too short lost the context an attention model needs; both were dropped. Over the summer I ran many fine-tunes on real call recordings, and the last one passed the vendor in early August.
- Test reference
- 90 minutes of calls transcribed by hand
- One fine-tuning run
- about 15 minutes, $1–2
- Result
- 29.4% → 17.1% word error rate; best vendor 17.5%
Serving
The pipeline went to production with help from the DevOps team and its automated tests. Spot GPU instances, rented at a discount because the provider can reclaim them at any time, process the previous day's calls every night and then shut down. A spending cap is set and a cost report arrives every morning.
- Schedule
- nightly batch on the previous day's calls
- Compute
- AWS spot GPU instances
- Guards
- spending cap, morning cost report, automated tests
- Running cost
- $1.5 per 1,000 h of audio, ≈ $800 a year