Measurement, Impression, and Identity: Three Levels of Conclusions That Must Not Be Mixed

Voice does carry information, but "a measurable difference exists" does not mean "we can confirm who a person is." A responsible report must state the strength of evidence behind each inference.

What are the three levels of conclusions?

LevelExampleHow it can be stated
Acoustic measurementThe recording's median F0, energy range, spectral centroid, and pause ratio.When quality is sufficient and the method is clear, values and uncertainty can be reported.
Perceptual impressionAn impression of brighter, softer, tenser, more active, or more masculine/feminine expression.It should be described as the current recording's relative impression on a specific dimension.
Identity and state factsActual age, gender identity, personality, disease, honesty, or ability.A single ordinary recording usually cannot confirm these reliably and should not be presented as a detection result.

Why group trends cannot be applied directly to individuals

Even when an acoustic feature has different group averages, the distributions may overlap substantially. Individuals are also affected by language, accent, training, hormones, physical state, context, equipment, and intentional expression. Turning a statistical association into a personal fact from one short recording magnifies both error and bias.

"Age impression" and "gender impression" are better treated as coordinates of auditory style: they describe how an expression is perceived at that moment, not what an ID document or the body says. Users can intentionally change these expressions, and that change itself need not be treated as deception.

Emotion, stress, and fatigue require particular caution

Speaking rate, pauses, pitch variation, and voice quality may relate to state, but the same signal can have many causes. Low energy may come from fatigue, calmness, a distant microphone, or deliberately soft speech; pauses may come from thinking, language unfamiliarity, or the way a text is arranged.

Accordingly, Insight can only say things such as "more voice cues associated with tension are present" or "the current expression is relatively calm"; it cannot diagnose anxiety, depression, or voice disorders. Arcana's voice-state cues can only support questioning, not read minds.

How to word a responsible report

More appropriate

"In this recording," "may be related to...," "the current sample shows...," or "recording quality limits confidence."

Avoid

"You are...," "The algorithm proved...," "100% accurate," or "Your voice reveals your true personality."

Algorithmic transparency is more than listing feature names. It also means explaining the source of training or rules, applicable inputs, failure conditions, retention periods, and which decisions the conclusions should not be used for.

As a user, how to protect yourself and others

  • Do not upload someone else's private recording to infer their identity, health, or personality.
  • Do not use entertainment labels for major decisions in hiring, education, insurance, healthcare, or relationships.
  • Treat results from different times as snapshots of state; check recording conditions before interpreting changes.
  • Before sharing publicly, confirm audio rights, privacy, and the consent of the people involved.
CVoice Family principlesVoice can become a game, report, star, music, or narrative, but a technical appearance must not be used to package limited evidence as a final definition of a person.

Method references