Understand TV Better

Speech Frequency Range

By faller audio editorial team | | 6 min. read
Close-up Loudspeaker

The speech frequency range extends from approximately 100 to 8,000 Hertz. The fundamental tone of the voice is in the lower part. Vowels follow above that. Many consonants reach up into the high frequencies. Most of the sound energy is in the low and mid-range.

Many consonants have their most important components between about 1,000 and 4,000 Hertz and are significantly quieter than vowels. Nevertheless, they are more important for understanding because they help identify which word is meant. When watching TV, consonants are easily lost, especially if music and noises are in the same frequency range or if the speakers do not direct the sound towards the sofa.

What is the Speech Frequency Range

Frequency describes how often a sound wave vibrates per second. The unit for this is Hertz. Low tones have a low frequency, high tones have a high frequency. The human voice does not consist of a single tone, but of the fundamental tone, its overtones, and the formants. Together, they form the speech frequency range.

Fundamental Tone of the Voice

When speaking, the vocal cords in the larynx vibrate and produce the fundamental tone. It determines how high or low a voice sounds. For men, it averages around 120 Hertz, and for women, around 210 Hertz. For children, it is even higher.

However, the fundamental tone does not remain constant. Within a sentence, it rises and falls, for example, at the end of a question or during emphasis. This speech melody aids understanding. However, the fundamental tone is less important for distinguishing individual words.

Overtones and Timbre

Overtones are produced along with the fundamental tone. These are multiples of the fundamental frequency that extend far upward. Their strength differs from person to person and gives each voice its unique timbre. Timbre allows you to recognize, for example, familiar voices on TV, even if the actor or presenter is not currently on screen.

Formants and Vowels

On its way out, sound passes through the throat and mouth, and for some sounds, also through the nose. These spaces amplify individual frequency ranges, which are called formants. Depending on the position of the tongue, jaw, and lips, the formants shift. This is how an A, an I, or a U is produced.

To distinguish most vowels, the first two to three formants are sufficient. These formants lie between a few hundred Hertz and about 3,000 Hertz. For open vowels like A, the first formant is higher; for closed vowels like I or U, it is lower.

Consonants in the High Frequencies

Many consonants are not produced by the vocal cords, but by air flowing past a constriction in the mouth or being briefly stopped. This creates a hissing sound that is higher than the vowels. These sounds are quieter than vowels but carry a large part of the information that distinguishes one word from another.

Speech Component Typical Range Importance for Understanding
Fundamental Tone Men around 120 Hz, Women around 210 Hz Pitch and Speech Melody
Most Sound Energy approx. 250 to 500 Hz Fullness and Loudness of Voice
Vowel Formants a few hundred to approx. 3,000 Hz Distinction of Vowels
Sch (sh) around 2,500 to 3,000 Hz Distinction of S and Sch (sh)
S around 4,000 to 5,000 Hz strong hiss, distinction of S and Sch (sh)
F no clear peak, broadly distributed quiet, easily masked

Most energy is in the low range, while the components important for understanding are higher.

Sibilants, Fricatives, and Plosives

Sounds like S, Sch (sh), and F are called fricatives because air is forced through a narrow opening. For S, the strongest component is around 4,000 to 5,000 Hertz, but the hiss extends even higher. For Sch (sh), it is slightly lower, around 2,500 to 3,000 Hertz. This difference helps you hear, for example, whether a film refers to a “Tasse” (cup) or a “Tasche” (bag).

The F sound is significantly quieter and distributes its energy more evenly over a broad range. Therefore, it is particularly easily lost when other noises are heard simultaneously. The same applies to sounds like T, K, and P. They are called plosives because the air is briefly stopped and then released. This produces only a short sound.

Voiced and Voiceless Sounds

Many consonants exist in two variants. With W, the vocal cords vibrate; with F, they do not. The same applies to the soft S in “reisen” (to travel) and the sharp S in “reißen” (to tear). For voiced sounds, therefore, low components in the range of the fundamental tone are added, usually below 500 Hertz.

This difference in the low range helps distinguish words like “Wein” (wine) and “fein” (fine). However, a large part of the information here also lies in the hiss at the constriction and thus in the higher frequencies.

Quiet Components with Great Impact

Most of the speech sound energy lies between approximately 250 and 500 Hertz, i.e., in the range of the fundamental tone and first formants. Around 3,000 Hertz, speech is significantly quieter on average. Nevertheless, individual words are understood primarily through these upper frequencies.

If the sound lacks high-frequency components, a voice may still sound loud enough, but the words become blurred. “Tisch” (table) and “Fisch” (fish) are then harder to distinguish. Conversely, speech remains understandable even if the low frequencies are weak, for example, with a small speaker.

Speech in TV Sound

On television, voices are rarely heard alone. Music, noises, and ambient sound play simultaneously and often lie in the same frequency ranges as the consonants. Additionally, how the TV emits sound and how far away the sofa is also play a role.

Music and Noises in the Same Range

A loud sound masks quieter sounds with similar frequencies. If film music or noises are in the range between 1,000 and 4,000 Hertz, they precisely mask the consonants that are important for understanding. The vowels often remain audible, which is why the voice sounds loud but indistinct. In news and talk shows, there is usually no music under the voices. The consonants are then not masked, which is why speech is often easier to understand there than in a film.

With the TV’s equalizer, you can slightly increase the volume of the consonant range. However, this also makes music and noises in this range louder. With a TV speech amplifier like OSKAR, an algorithm extracts the voices from the rest of the TV sound before they are enhanced. Because the device can be placed next to the sofa, even the quiet consonants arrive without a long detour through the room.

Speakers and Sound Dispersion

Speakers emit high frequencies more directionally than low frequencies. Therefore, high tones primarily arrive where the speaker is pointing. Many flat-screen TVs have speakers on the bottom or back of the device. It is precisely the consonants that then first hit the wall or furniture.

Low-frequency components, such as the fundamental tone, on the other hand, spread more evenly in all directions. On the sofa, the voice often still sounds loud enough, even though the high-frequency components are missing.

A soundbar in front of the TV usually directs sound forward, i.e., towards the sofa. However, if it is placed on a shelf or behind a panel of the TV furniture, some high tones can still be lost. Therefore, consonants are better received if nothing obstructs the path between the speaker and the sofa.

Distance and Reverberation

The further the sofa is from the TV, the more sound arrives indirectly via walls, ceiling, and floor. These reflections arrive slightly later than the direct sound. They blur short sounds like T and K, in particular, because a loud vowel still reverberates in the room when the quiet consonant follows.

If the sofa is closer to the TV, more direct sound arrives. The consonants are then clearer to hear because there is less reverberation overlaying them.

Frequently Asked Questions

Spoken language uses frequencies from about 100 to 8,000 Hertz. At the lower end are low-frequency components like the fundamental tone of a male voice; at the upper end are sibilants like the S. In between, you find the overtones and formants that distinguish vowels.

When whispering, the vocal cords do not vibrate, so the fundamental tone is absent. However, the mouth and throat still form the formants of the vowels. Voiceless consonants like S, F, or T are produced as in normal speech. Because these components are more important for understanding than the fundamental tone, whispered speech remains intelligible as long as the room is quiet enough.

Singing essentially uses the same ranges, but often extends higher and lower in pitch than speech. When singing, vowels are also held longer than when speaking. Consonants therefore recede somewhat, which is why song lyrics are sometimes harder to understand than spoken sentences.

The human voice does not have a single frequency. The fundamental tone is perceived as pitch, which is almost an octave higher for women than for men. When speaking, it constantly moves up and down. At the same time, overtones and formants, which extend up to several thousand Hertz, also resonate.

For speech, the range between approximately 1,000 and 4,000 Hertz is particularly important. A speaker should reproduce it cleanly and without dips. Low frequencies below 100 Hertz are hardly needed for voices; they primarily give volume to music and noises. However, the frequency specification in the data sheet does not reveal how evenly a device reproduces these frequencies or in which direction it emits the sound.