1
Collect representative audio
Choose clips with the accents, background noise, vocabulary, and speaker changes you actually handle. Keep the audio and requested output format identical for both tests; otherwise a faster-looking result may reflect an easier input.
2
Time the complete path
For AWS Transcribe, include upload or stream setup, job completion, and retrieval. For self-hosted Whisper, include file transfer, queueing, model startup where relevant, inference, and any alignment or speaker-labeling pass.
3
Stop the clock after review
Have the same reviewer correct names, missing phrases, and speaker turns using the same acceptance rules. Record both machine turnaround and editing time: the first result to arrive is not necessarily the first one ready to publish.