Understanding Television Better

Sound Mixing for a Film

By the Faller Editorial Team | | 6 min read
Audio mixer

When you watch a movie, you hear a soundtrack that’s been put together from many individual parts. Dialogue, sound effects, ambient sounds, and music are recorded separately and aren’t combined until the very end. Much of it sounds different in the finished film than it actually did on set.

This process is called sound mixing and takes place during post-production. Sound mixing plays a key role in determining how well the dialogue will be understood later on. This is because speech competes with everything else that goes onto the audio track.

Sound on Set

It all starts with the recording during filming. The sound crew records what the actors say on set. This recording is, of course, very important, but it often doesn't end up in the finished film as is.

Photo taken during filming

The microphone is usually suspended from a boom arm above the actors' heads or attached to their clothing as a lavalier microphone. The boom arm provides a more natural sound, while the lavalier microphone offers better protection against background noise. Both options are often used simultaneously so that there are options to choose from later. The recording captures not only the dialogue but everything that can be heard on set.

Recording is usually done on multiple tracks simultaneously so that you can choose between microphones during post-production. All tracks are time-stamped, which allows them to be synchronized with the video later on.

Limitations of the Original Audio

But that is precisely where the problem lies. A filming location is rarely quiet and peaceful. The scenery, a street, the wind, or even the crew itself end up in the shot. Furthermore, some scenes are shot in places where it’s simply impossible to record with minimal background noise.

In addition to technical reasons, there are linguistic ones as well. Sometimes a sentence doesn't sound right, sometimes a word gets lost, or sometimes the emphasis no longer fits the scene after the edit. In such cases, the original audio is replaced later.

Dubbing of the dialogues

If the original audio cannot be used optimally, the actors re-record their lines in the studio. This process is called ADR, which is usually written out as Automated Dialogue Replacement. In German, it is also referred to as “Nachvertonung” or “Dialogersatz.”

ADR in the Studio

The actor watches the scene and repeats the line until it matches the lip movements. The recording is made in a quiet room, without any background noise. That’s why it sounds extremely clear, but also different from the original audio.

This clarity must then be removed. A voice speaking in a hall, for example, is given artificial reverb. Without this additional effect, the dialogue would not fit into the scene.

Dubbed versions in other languages

In a German dubbed version, the same process happens all over again, only with different voice actors and a translated script. Here, too, the voice is recorded in the studio and then added to the existing mix.

Music and sound effects are often provided as a separate mix without dialogue. This version is called M&E, short for Music and Effects. It serves as the basis for adding new dialogue in other languages. Where sound effects were present beneath the original dialogue, they must, of course, be added back in.

Sounds, Atmosphere, and Music

In addition to the dialogue, there are three other elements that are recorded or composed separately. They run on their own tracks until the end. Only after they are put together do they form what you hear as the final film soundtrack.

Element Origins Relation to Speech Intelligibility
Dialog Filming on location or in the studio the part that is meant to be understood
Noises specially recorded or created on the computer mostly uncritical, because they are brief and selective
Atmosphere Recording Ambient Sound often runs in the background and can help bridge the gap with language
Music composed, recorded, or licensed often a competing factor, because it operates simultaneously and frequency ranges may overlap

All four tracks run on separate tracks until they are mixed together. The balance between speech and music is particularly crucial for the intelligibility of dialogue.

Foley and the Sound Effects Artists

Footsteps, doors, the rustling of clothing, or the sound of a glass being set down are recorded separately while the scene is playing. The technical term for this is “Foley”; in German, it’s called “Geräuschemacher.”

This work takes place in a studio with various floor coverings and a collection of objects. What you hear in the film—for example, the sound of footsteps on pebbles—is created there, not on location. On set, these sounds would be recorded along with everything else and could not be separated from the rest of the audio later on.

Sound Design for Everything Else

Not every sound can be recorded. A spaceship or a monster doesn't have a real sound. Sounds like these are created on a computer by combining sounds from various sources. Funnily enough, these sources often have nothing to do with the scene itself.

This step is called sound design and is distinct from sound recording, even though both end up on the same track group.

Background atmosphere

The atmosphere, also known as "atmo," is the continuous background sound of a scene—traffic outside the window, birdsong, or other ambient sounds. It ensures that a room actually sounds like a room and not like a studio recording.

Music and Language

The film score is usually created last, because the composer needs to see the edited film. It follows the plot and begins at specific points that are determined in advance in collaboration with the director. These points are recorded using a timestamp (also known as a timecode) so that the music and visuals will sync up later.

Music is the most challenging factor when it comes to the intelligibility of the spoken word. It can overlap with the frequency range of the human voice and often plays at the exact same time that someone is speaking.

Mixing in the recording studio

In the end, all the tracks come together. During mixing, a mixing engineer determines how loud each element is in relation to the others. This is where it’s decided whether the dialogue stands out or gets drowned out.

Separate Lanes

The individual elements remain separated into groups called "stems." Three stems are typical: one for dialogue, one for music, and one for sound effects. This separation is useful because it allows you to replace individual groups or adjust their levels later on.

For dubbed versions, this is a fundamental requirement. Newer techniques for improving speech intelligibility also focus on this very aspect by increasing the volume of the dialogue relative to the rest of the audio.

The Relationship Between Language and Everything Else

The degree to which dialogue stands out against the background is a creative decision. Some directors aim for a realistic effect and are willing to accept that the dialogue may be drowned out in the scene. Others bring the dialogue clearly to the forefront.

The Audio Engineering Society considers this ratio to be the key factor in the intelligibility of dialogue. According to research by the EBU, too small a difference between speech and background noise is one of the most common causes of complaints about television sound.

Volume Standards

For television, there are established guidelines that specify the volume levels at which programs should be mixed. This ensures, for example, that programs and commercials are comparable with one another. However, this standard regulates only the overall volume level of a program; it does not specify the relative levels of speech and music within that program.

Impact on TV Audio

A movie is mixed for a large, quiet theater with many speakers. On a television, that mix then encounters a living room, two small speakers on the TV itself, and background noise. What works as quiet dialogue in the theater can often get lost in your own living room.

Feature films, TV series, and movies are often mixed with too much dynamic range for broadcast and streaming, because the difference between the overall volume and the dialogue volume is too great. For this reason, separate versions with a lower dynamic range are frequently produced for television and streaming services.

In object-based formats, individual components of the audio can be transmitted as separate objects or groups with metadata. If the dialogue and background audio are separated, the dialogue portion can generally be handled differently or adapted to the listener.

Frequently asked questions

During mixing, all audio elements are combined, and their relative volumes are adjusted. Until that point, dialogue, sound effects, ambient sounds, and music are recorded on separate tracks. A sound mixer then decides how loud each element should be in relation to the others. This step is part of post-production and takes place after editing.

Because the original audio from the set is often not clear enough. Street noise, wind, air conditioning, or even the film crew itself end up on the recording. Even after editing, the emphasis might still be off. The actors then re-record their lines in the studio; this process is called ADR.

He records everyday sounds specifically for the film while the scene is being shot. These include footsteps, doors, the rustling of clothing, or a glass being set down. The studio has various types of flooring and a large collection of props available for this purpose. The technical term for this work is "Foley."

Music can overlap with the frequency range of the human voice and often plays at the exact moment someone is speaking. The degree of separation between the two is a creative decision made during mixing. Some films deliberately aim for a realistic feel and are willing to accept that dialogue may be less distinct as a result.

That's often the case. A cinema mix is designed for a large, quiet auditorium with many speakers and uses a wide dynamic range between quiet and loud passages. For television and streaming, separate versions with a narrower dynamic range are therefore often created. Whether such a version is broadcast depends on the broadcaster or provider.