What is the algorithm actually observing in a voice recording?

A voice is not a fixed label. It is the result of vocal production, wording, emotion, equipment, and environment together. Understanding those factors is more useful than remembering only a test label.

From air vibration to a computable signal

A microphone converts air vibrations into a continuous stream of digital samples. The algorithm does not "know you like a person would"; it measures periodicity, energy distribution, rate of change, and noise over short windows, then organizes multiple signals into readable descriptions.

The same person will not produce identical recorded spectra across different phones, rooms, and distances; even on the same device, speaking softly, excitement, fatigue, or deliberately lowering the voice can change the result. CVoice therefore describesthe voice presented in this recording

Pitch: a primary clue to sound periodicity

For voiced sounds with stable periodicity, the fundamental frequency—commonly called F0—can be estimated. Higher or lower F0 clearly affects perception, but a person's pitch continually moves with intonation, so an average alone misses a great deal of information.

CVoice also considers pitch range, short-term variation, and trackability. FrostNote's real-time detection places more emphasis on stability around the current note. Background music, whistling, strong breathiness, and multiple people speaking at once can all make fundamental-frequency tracking difficult.

How voice-quality dimensions affect perception

Brightness

Usually related to energy in higher-frequency regions and the spectral centroid. Being too close to the microphone, device noise reduction, or harsh noise can all change it.

Weight

A combined impression from low-to-mid-frequency energy, glottal excitation, and resonance structure. It does not simply mean loudness and is not exclusive to any gender.

Softness and clarity

Breathiness, consonant edges, noise, and articulation all contribute. Soft does not necessarily mean muffled, and clear does not necessarily mean sharp.

Stability and activity

Observe short-term changes in pitch, energy, and rhythm. Steady reading and emotionally expressive natural speech produce different trajectories.

What formants can tell us

The vocal tract behaves like a filter whose shape constantly changes. Tongue position, mouth opening, and vocal-tract length create several formants that are important to vowels and timbre. Formants can explain differences that pitch alone cannot, but they are also affected by spoken content, recording bandwidth, and tracking errors.

Important boundariesFormants and pitch can support probabilistic descriptions of "voice tendencies," but they cannot by themselves confirm biological sex, actual age, or identity. The CVoice family consistently presents these results as perceptual impressions and acoustic behavior in the current recording.

How to record more stable test audio

  1. Use your natural conversational voice and speak one or two complete sentences continuously rather than reading isolated words.
  2. Record in a quiet room when possible, with background music, voice changers, strong reverb, and other people's speech turned off.
  3. Keep roughly a palm's distance from the microphone. Too close can cause clipping or plosives; too far away can cause noise reduction to remove useful detail.
  4. Keep your volume comfortable; there is no need to deliberately raise, lower, or constrict your voice to obtain a particular result.
  5. When comparing two results, use the same device, similar phrases, and a similar microphone distance whenever possible.

If you upload AMR, M4A, or other mobile recordings, the service may need to convert the format first. Conversion can make the audio readable by the algorithm, but it cannot restore frequency detail already lost in the original recording.

Why whistles and non-speech sounds are handled separately

A whistle can have a clear pitch while lacking the vocal-tract resonance and continuous phoneme structure of natural speech. If forced into standard voice labels, the algorithm can still produce numbers but create a false sense of certainty. CVoice therefore routes clearer non-speech vocalizations to special-sound results and excludes them from gender/age impressions and standard rankings.