2026 · completed

Transcribing every call

An in-house speech-to-text pipeline that transcribes all of a call centre’s calls at a fraction of the vendor’s cost.

ML research internship, AI Lab, Selectra, Madrid, June–September 2026.

word error rate, raw model → final pipeline (preprocessing + fine-tuning)
29.4 → 17.1%
lower running cost than the best vendor ($1.5 vs $222 per 1,000 h of audio)
150×
of calls quality-checked, previously a random sample
100%

How it works

Every call of the day before is transcribed overnight. The fine-tuned open-weight model replaced the external vendor.yesterday’s callsnightly batchvolume, speech, noisenormalise · VAD · filterspeech-to-textParakeet, fine-tunedtranscriptsevery call quality-checkedAWS · GPU spot instances · cost cap · automated tests
Fig. 1 Every call of the day before is transcribed overnight. The fine-tuned open-weight model replaced the external vendor.

Calls are pulled nightly, loudness-normalised, cut to speech by a voice-activity detector, filtered for noise, and transcribed by an open-weight model (NVIDIA Parakeet, 600M parameters) fine-tuned on real call recordings; each fine-tune costs one to two dollars and fifteen minutes; the batch runs on spot GPU instances with cost caps and automated tests.

Results

word error rate$ per 1,000 h of audioParakeet, as released29.4%not deployedMistral Voxtral Mini (batch)22.5%$90Gladia Solaria 317.5%$222fine-tuned, in-house17.1%$1.5
Fig. 2 Word error rate on the same held-out, hand-annotated calls (left, lower is better) and running cost per thousand hours of audio (right, log scale). The in-house model matches the best vendor at a hundred-and-fiftieth of its price.

The fine-tuned model edges past the best vendor (17.1% against 17.5% word error rate) for $1.5 per thousand hours of audio instead of $222: about $800 a year instead of $120,000, on more than 550,000 hours. The whole volume is now transcribed and quality-checked, not a sample.

Technical details

Model

NVIDIA's Parakeet: a small open-weight speech model of 600 million parameters, a FastConformer encoder followed by a token-and-duration transducer, cheap enough to run on one GPU. Out of the box it produced sentences that were never said, wrote text over inaudible segments and switched to English in French or Spanish calls.

Base model
NVIDIA Parakeet (600M parameters, open weights)
Architecture
FastConformer encoder, TDT transducer decoder
Why a small model
domain vocabulary (contracts, meter readings) beats size, and it is cheap to run

Preprocessing

Recording levels varied a lot from call to call. Loudness normalisation (EBU R128) followed by a voice-activity detector (Silero VAD), which removes silences, hold tones and music to keep only speech, cut the word error rate by six points without touching the model.

Loudness
EBU R128 normalisation
Voice activity
Silero VAD
Gain
≈ −6 WER points, model unchanged

Fine-tuning

There was no annotated data at the start, so I transcribed ninety minutes of calls by hand to get a reference, and every version of the system was scored on that same test set. Synthetic voices did not match real agents, and clips that were too short lost the context an attention model needs; both were dropped. Over the summer I ran many fine-tunes on real call recordings, and the last one passed the vendor in early August.

Test reference
90 minutes of calls transcribed by hand
One fine-tuning run
about 15 minutes, $1–2
Result
29.4% → 17.1% word error rate; best vendor 17.5%

Serving

The pipeline went to production with help from the DevOps team and its automated tests. Spot GPU instances, rented at a discount because the provider can reclaim them at any time, process the previous day's calls every night and then shut down. A spending cap is set and a cost report arrives every morning.

Schedule
nightly batch on the previous day's calls
Compute
AWS spot GPU instances
Guards
spending cap, morning cost report, automated tests
Running cost
$1.5 per 1,000 h of audio, ≈ $800 a year

← All projects