Treat the report as a snapshot, not an ID card

A report answers, "What kind of perceptual impression does this recording present?" It does not permanently label a person or judge identity, age, or value.

Primary timbre and three-timbre combination

The primary timbre is the voice impression with the highest match for the current recording. It captures the strongest direction but not every detail. The second and third timbres add information about warmth, stability, activity, or vocal placement, so the three-timbre combination is usually more complete than a single name.

These names are entertainment-oriented terms designed to make acoustic features easier to read. It is normal for the same person's primary timbre to change when delivery, device, or recording environment changes.

Sample proportions and rarity titles

Sample proportions come from anonymous aggregate records: the primary-timbre proportion counts how often a term ranked first, while the combination proportion counts how often an unordered three-timbre combination appeared. They reflect frequency within system records, not real-world population proportions.

Rarity titles are calculated from the exact occurrence count divided by the total count. As the sample grows, a combination's percentage and title may change over time. Rare does not mean better, and common does not mean mediocre.

Why does it sometimes show 0%?For readability, the page limits decimal places. A combination that has occurred once may still round to 0% when the total sample is large. Rarity titles use the actual unrounded ratio.

Feminine, masculine, and neutral voice tendencies

These three values are a relative distribution of perceptual voice tendencies inferred from pitch, resonance structure, and multiple acoustic signals. The page uses the highest percentage as the displayed tendency, but when several values are very close, the result itself indicates a blurred boundary.

A "feminine voice tendency" does not mean the speaker is a woman, a "masculine voice tendency" does not mean the speaker is a man, and "neutral" is not an identity category. The algorithm cannot reliably determine a speaker's biological attributes from an ordinary recording.

Why age impression is only a perceptual impression

Age impression describes whether a voice sounds youthful, mature, or settled. Speaking rate, delivery, vocal placement, and recording quality can all change that impression. Adults may sound youthful and younger people may sound composed, so the report does not present it as actual age.

How to compare the eight voice qualities

Pitch tendency, brightness, softness, weight, activity, stability, warmth, and clarity are different projections of the same recording. There is no universal good or bad score for any one dimension; the stability and clarity useful for narration are not the same as the variation and tension useful in character voice acting.

The most meaningful comparison is to have the same person record different expressions under similar device and environmental conditions, then observe which dimensions change with delivery and which remain relatively stable.

Check recording quality before labels

Too little valid speech, loud background music, clipping, or excessive noise reduction can make a label look confident without a reliable basis. If the quality indicator is abnormal, re-record first rather than repeating the test until you get a result you like.