There’s a particular kind of frustration reserved for the moment a finished voice-over gets kicked back by a broadcaster’s quality control department. The read was approved. The client signed off. And then a QC technician rejects it, not because it sounds wrong to a human ear, but because it falls outside a technical delivery specification.
This happens because professional audio delivery comes with technical requirements covering loudness, peak levels, file formats, sample rates, and audio quality. Many production and broadcast workflows use automated QC tools to measure these parameters alongside human review.
Here’s what broadcast, radio, and OTT workflows may check for, and how to give your files the best chance of clearing technical QC the first time.
Choosing the right voice-over services means considering not only the performance, but also the recording quality and technical requirements of the final production.
First: What “Dry” Actually Means (and Why It’s Delivered That Way)
A dry voice-over file is an isolated vocal performance, typically without music, sound effects, reverb, or other creative mix processing baked into the file, delivered as a clean, unmixed asset. This gives the post-production or dubbing team control over how the voice is mixed against picture, music, and sound design.
That flexibility also makes clean voice stems useful for creating different versions of the same production. A voice track can be incorporated into different mixes according to the requirements of the broadcaster, territory, or platform.
Because the voice is delivered as an isolated asset, technical and recording-quality issues can be particularly noticeable. Clicks, pops, unwanted background noise, clipping, excessive room sound, and other recording problems can all affect whether a file is accepted.
The Loudness Standards, Explained Without the Jargon
There is no single loudness specification that applies to every broadcaster, radio station, streaming service, or OTT platform. Many professional audio standards use measurement methods based on the ITU-R BS.1770 family of recommendations, but the required target depends on the delivery specification.
EBU R128 (Europe)
EBU R128 establishes -23 LUFS as the reference programme loudness level for European broadcasting. The recommendation also specifies a maximum permitted true-peak level of –1 dBTP for programme material.
EBU R128 uses Loudness Range (LRA) as a measurement of programme loudness variation, but it does not establish a universal 20 LU maximum that every programme must meet. Specific broadcasters and delivery specifications may impose additional requirements.
For UK productions, always check the current delivery specification provided by the broadcaster, distributor, or production partner rather than assuming that EBU R128 alone defines every technical requirement.
ATSC A/85 (US)
ATSC A/85 provides recommended practices for managing television loudness in the United States and uses -24 LKFS as the reference loudness level.
The U.S. regulatory framework also includes the CALM Act and related FCC rules governing the loudness of television commercials relative to surrounding programming. The exact technical requirements for a particular delivery should always be checked against the broadcaster or network’s current specification.
Important: LUFS and LKFS refer to essentially the same type of loudness measurement, although the terminology differs between standards and regions.
Streaming and OTT Platforms
Streaming platforms do not all use the same loudness target. Some publish platform-specific delivery specifications, while others apply loudness normalization during playback.
For example, Netflix’s published audio specifications use dialogue-gated loudness measurements for applicable programme mixes rather than simply applying a universal “-27 LUFS” rule to every piece of audio. Spotify uses loudness normalization around a -14 LUFS reference level for playback, but that should not be treated as a universal submission requirement. YouTube also performs its own loudness processing and should not be described as requiring every upload to be mastered to exactly -14 LUFS.
Podcast specifications vary by platform and publisher. Apple Podcasts, for example, recommends approximately -16 dB LKFS for stereo podcast audio, with a ±1 dB tolerance, alongside a maximum true-peak recommendation of -1 dBFS.
Always confirm the specific platform’s current technical specification before delivery. The practical takeaway is simple: “sounds about right” is not enough for technical compliance. A proper loudness meter is essential when a delivery specification includes loudness requirements.
What a QC Technician Is Actually Checking
Beyond loudness measurements, professional QC can include both technical measurements and listening tests. Common issues include:
- Editing artifacts – clicks, pops, abrupt cuts, or other audible problems caused by editing
- Unwanted room sound – excessive room tone, echo, or reverberation
- Clipping or distortion – peaks that exceed the recording system’s limits or otherwise create audible distortion
- Mouth noise and excessive breaths – depending on the project’s editing requirements
- Plosives – excessive bursts of low-frequency energy from sounds such as “P” and “B”
- Sibilance – excessively harsh “S” and “SH” sounds
- Hiss, hum, and electrical noise – unwanted noise in the recording
- Background noise – traffic, HVAC systems, computer fans, handling noise, or other unwanted sounds
- Over-processing – excessive noise reduction, EQ, compression, or other processing that negatively affects the recording
Not every QC workflow uses the same checklist. The exact acceptance criteria depend on the client, broadcaster, production company, marketplace, and final delivery specification.
File Format: The Part That’s Easy to Get Wrong
Professional voice-over and post-production workflows commonly use uncompressed PCM audio in a WAV container rather than MP3. The exact requirements vary, but 48 kHz and 24-bit WAV are common choices for video production.
MP3 is a lossy format that removes some audio information during encoding. It can be appropriate for certain consumer or preview applications, but it is generally not the preferred format when a production requires an uncompressed master or voice stem.
Sample rate and bit depth should be determined by the project’s delivery requirements. For video production, 48 kHz is commonly used, while 24-bit is widely preferred for recording and production because it provides greater available dynamic range and lower quantization noise than 16-bit.
Do not confuse sample-rate conversion with improving audio quality. If a recording was captured at 44.1 kHz and the final delivery requires 48 kHz, a properly performed sample-rate conversion can create a valid 48-kHz deliverable. Upsampling does not recover information that was not captured during the original recording, but it does not automatically mean the voice-over must be recorded again.
Always confirm the target format before recording whenever possible.
A Practical Pre-Delivery Checklist
Before a dry voice-over file goes to a broadcaster, production company, or platform, check:
- Loudness: Measure against the exact loudness target specified by the client or broadcaster rather than applying a universal number
- True peak: Confirm that the file meets the delivery specification’s maximum peak requirement
- Noise: Listen carefully to the recording’s noise floor at the beginning, end, and during pauses
- Editing: Check for clicks, pops, abrupt edits, excessive breaths, mouth noise, and other artifacts
- Recording quality: Check for plosives, sibilance, distortion, clipping, room echo, and background noise
- Format: Confirm WAV/PCM or the required delivery format, sample rate, bit depth, and channel configuration
- Final listen: Listen on accurate headphones or monitors in addition to checking technical measurements
Wrap Up – Getting It Right Starts Before the Session, Not After
The files that pass QC consistently usually start with a clean recording environment, appropriate microphone technique, controlled recording levels, and careful editing. Post-production can correct some problems, but excessive room noise, clipping, distortion, and poor recording technique cannot always be repaired without affecting the quality of the voice.
Bunny Studio, for example, identifies issues such as room echo, clipping, incorrect microphone distance, editing problems, and excessive mouth noise among the quality issues that can affect voice-over submissions. These criteria reflect Bunny Studio’s own quality-control requirements and should not be treated as universal requirements for every broadcaster or platform.
This is also why a polished demo should not be treated as proof of the recording environment behind every future session. The Voice Finder only accepts talent with professionally produced demos, and then quality checks the studios or recording spaces where those artists record. This adds an additional layer of confidence that the talent’s working environment is suitable for production requirements, not just that the demo sounds impressive.
The Voice Finder connects you with professional voice artists who understand the technical requirements involved in broadcast, radio, and OTT voice-over delivery. Browse experienced voice talent for projects that require clean, production-ready recordings.
No surprises. No unnecessary technical rejections. Just professional voice talent ready for your next production.
FAQ
What audio format is commonly used for professional voice-over delivery?
Uncompressed PCM audio in a WAV container is commonly used for professional production. 48 kHz, 24-bit WAV is a frequent choice for video projects, but the required format may vary.
Why might a technically correct voice-over still fail QC?
A file may meet loudness and format requirements but still fail because of clicks, pops, clipping, distortion, excessive room echo, background noise, mouth noise, plosives, or over-processing.