Temporal Taxation Compounds Under Post-Training Compression of Whisper Models
Automatic speech recognition models are typically audited for demographic fairness at full precision, but deployed models are often quantized, pruned, or distilled. This paper asks whether post-training weight compression, which changes model weights rather than the audio signal or its feature representation, redistributes error burden across demographic groups.
The study evaluates the Whisper family on Fair-Speech, Common Voice 25, and AfriSpeech-200. It tests 50% Wanda pruning, INT4 HQQ quantization, and distillation, and formalizes the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric.
On Fair-Speech, 50% Wanda pruning of Whisper-large-v3 sharply widens the Black/AA-vs-Asian temporal-taxation differential: the absolute word-error-rate gap between the worst- and best-served groups more than doubles. At an assumed five seconds of correction effort per transcription error, this is a rise from 30 to 64 seconds of correction time per minute of speech. The +111% relative increase is invariant to the assumed per-error cost, survives an audio-quality control, and is only partly mitigated by beam-search decoding, which still leaves an +86% increase. At edge model size, INT4 HQQ quantization compounds catastrophic transcript loops on West African accents by factors of five to seven. Distillation, by contrast, narrows demographic gaps in 21 of 27 evaluated settings, defined by teacher-student pair, precision, and dataset, with the exceptions concentrated on a single model pair.
The authors conclude that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.