Formants, Spectrum, and Where “Brightness” and “Weight” Come From

The same pitch can sound completely different. What matters is not only how fast the vocal folds vibrate, but also how resonance and spectral shaping transform the sound.

First the Source, Then the Filter

A useful approximation is the source–filter model: periodic vocal-fold vibration provides the fundamental frequency and harmonics, while the pharyngeal, oral, and nasal cavities strengthen or weaken different frequencies. Strongly enhanced bands form formants. Even when the fundamental frequency stays the same, changing tongue position and mouth shape changes vowels and timbre.

This explains why “lower pitch” does not automatically mean “heavier voice,” and why changing mouth shape or resonance placement can noticeably change how a voice sounds.

What Spectral Statistics Describe

FeatureIntuitive MeaningCommon Misconception
Spectral CentroidThe “center of mass” of spectral energy; higher values are often associated with a brighter or sharper sound.It is not a fixed formant and can also be pushed upward by noise.
Spectral Roll-OffThe frequency at which cumulative energy reaches a specified proportion, reflecting how far energy extends into higher frequencies.Encoding, microphone frequency response, and sibilance can all change it.
Spectral FlatnessWhether the spectrum looks more like clear harmonics or uniform noise.High spectral flatness does not mean “poor sound quality”; breathy voice and ambient sounds naturally have different spectral characteristics.
FormantsLocal spectral peaks created by vocal-tract filtering.Automatic estimates are affected by fundamental frequency, vowel, sample rate, and parameter settings.

The Microphone Also Shapes the Sound

Close recording can boost low frequencies, off-axis recording can reduce some highs, room reflections can create new peaks and dips, and phone noise reduction and automatic gain can alter the spectrum over time. So what the algorithm measures first is a file jointly produced by “speaker + content + room + device + processing chain.”

CVoice combines multiple spectral statistics into perceptual dimensions. Reborn also uses transients, event density, and dynamic range to characterize sound type. Neither translates a single spectral number directly into a person’s identity.

Resonance Provides Clues, Not Proof of Identity

Vocal-tract structure, articulation, language, training, emotion, and devices can all affect resonance. Statistical trends may exist between groups, but distributions from individual recordings overlap substantially. Using formants to describe “this recording sounds relatively bright” is descriptive; using them to confirm actual sex, age, or anatomy goes beyond what the evidence supports.

How to Read the LabelsTreat “warm,” “clear/bright,” and “weighty” as relative positions of the current recording along designed dimensions, not as immutable traits of the speaker.

References