Where the Report Comes From
Insight uses a set of standardized speech-acoustic features and explicit local rules to organize statistics on pitch, energy, spectral shape, phonation quality, and rhythm. Scores are produced by rules; the language model only turns structured results into more readable explanations and cannot change the scores on its own.
The report covers cues related to emotion and activity, acoustic signals associated with tension, fatigue and vocal load, perceived voice-age ranges, masculine/feminine/neutral expression impressions, as well as clarity, stability, rhythm, and communication impressions.
Why You Must Consider the Reliable Range Too
More complex features do not mean medical truth has been obtained. Recording quality, valid speech duration, ambient noise, and device processing all limit the conclusions. The current rules have not been formally calibrated on large-scale real-world data from the target population, so they are suitable for observation and comparison, not clinical or employment decisions.
How to Make a More Meaningful Comparison
- Record in the same room, with the same device and distance.
- Use the same text and similar volume; do not read aloud in one recording and sing in the other.
- Treat the report as “a snapshot of this recording” and compare trends rather than chasing a single score.
- If recording quality is flagged as poor, re-record first instead of interpreting a low-confidence result.
Recording & Report Retention
Audio is deleted after analysis completes or fails; the maximum cleanup window for unexpected residual files is about 24 hours. Tasks and final structured reports are retained temporarily so results can be restored after a page refresh.
