Phrasing a Voice Score So It Stays a Wellness Claim
One word choice separates a useful wellness app from an FDA enforcement action.

A voice app that tells a user their "state of mind" score dropped six points since yesterday is making a wellness claim, and the difference from a medical claim lives entirely in word choice. Get the copy wrong and the FDA will eventually ask why an unregulated app is diagnosing anxiety. Get it right, and the product still ships something useful: a number that reflects vocal patterns tied to stress or fatigue, framed so it never crosses into telling someone what's wrong with them.
I've sat in the review meetings where a single verb got argued over for forty minutes. "Detect" versus "reflect." "Indicates" versus "may relate to." These aren't pedantic fights. They are the entire compliance strategy, because in FDA's framework, the claim is the product. Change the sentence, and you've changed what the software legally is.
Why the wording is the whole ballgame
FDA's General Wellness Policy, first issued in 2016 and still the reference point for this category, draws the line clearly: a device stays outside regulation if it's intended only for general wellness and presents low risk to a user's safety. General wellness means either promoting a healthy lifestyle or relating to a condition without claiming to diagnose, cure, treat, or prevent it.
A voice-derived score falls apart the moment the copy implies clinical function. "Your voice shows signs consistent with depression" is a diagnostic claim, full stop, regardless of what's happening in the model underneath. "Your voice patterns today suggest lower energy than your typical range" stays on the wellness side, because it describes a pattern relative to the user's own baseline rather than naming a disease state. Same underlying pitch, jitter, and shimmer measurements. Completely different regulatory exposure.
This is why the actual acoustic science barely matters in these reviews. The literature on vocal biomarkers, going back to the work on speech and mood tracking that groups like MIT's Media Lab and various digital health startups have published on, tells us that features like pitch variability, pause duration, and speaking rate correlate with mood states in aggregate. Correlation in a research paper, though, doesn't license a consumer product to say "we detect your mood." It licenses the product to say something far more modest, and the legal team's job is making sure the marketing team doesn't reach for the stronger version because it tests better in user interviews.
What the score copy actually says
Every wellness-framed voice score I've worked on ends up built from three ingredients: a relative comparison, a soft verb, and an explicit refusal to name a condition.
Take "state of mind" as the label itself. It's already doing work. It says state of mind rather than "mental health" or "mood disorder risk," a phrase closer to "energy level" or "focus" than to anything in the DSM-5. Under that label, the score copy might read: "Your state of mind score reflects patterns in your voice compared to your own recent history. It is not a measure of any diagnosed condition and should not replace guidance from a doctor or therapist."
Notice what's missing. No mention of anxiety, depression, or stress as named conditions. No claim that the score predicts anything. No comparison to a clinical population norm, because the moment you benchmark someone against a validated clinical scale, like the PHQ-9 for depression screening, you've imported that scale's diagnostic intent into your product. Comparing a user only to themselves, over time, keeps the tool in the same category as a sleep tracker showing you slept worse than your weekly average.
The disclaimer language does more lifting than most users ever read. Standard boilerplate across this product category looks something like: "This app is intended for general wellness purposes only. It is not intended to diagnose, treat, cure, or prevent any disease, and has not been evaluated by the FDA." That last clause matters legally; it's lifted nearly word for word from dietary supplement labeling, another category that lives under a similar general wellness carve-out, and it signals to a regulator on first read that the company knows exactly which side of the line it's standing on.
Where the review process actually catches problems
Every score-related feature I've seen ship goes through a review loop with legal, and increasingly a "clinical advisor," someone with a licensure background in psychology or psychiatry, brought on specifically to catch language creep before an update ships. The process runs in roughly four passes.
The first is a language audit on anything user-facing: push notifications, in-app copy, marketing pages, App Store screenshots. Screenshots matter more than people expect, because Apple and Google both review app store listings for health claims, and a screenshot showing "detect depression symptoms" can get an app rejected or pulled even if the in-app copy is clean.
The second is a check on comparison points. Does the score ever get compared to a named clinical instrument or population average? If yes, that's flagged immediately, because it reframes a wellness metric as a screening tool.
The third has someone running through the notification logic specifically, since push notifications are where marketing instinct sneaks diagnostic language back in. "You seem really low today, are you okay?" reads like concern, but it also reads like a company claiming to detect a mental state with confidence it hasn't earned. The safer version: "Your score is lower than usual today," full stop, no interpretation attached.
The fourth, and this one gets skipped more than it should, is a review of what happens when a score is persistently low over time. A single day's low score is a wellness data point. Fourteen straight days of decline starts looking, to a regulator, like the product is tracking a health trend and, by implication, suggesting the user has a condition worth monitoring. Some companies handle this by capping how the app talks about trends, showing the data but withholding interpretive language like "this may indicate" once a pattern crosses a certain length.
The tension nobody fully resolves
Here's what makes this genuinely hard rather than just a compliance exercise: the more useful the score is, the more it starts to sound diagnostic, because usefulness and diagnostic specificity are the same thing described from two different angles. A number that's just noise won't help anyone, but a number precise enough to matter to a user is precise enough to worry a regulator.
Companies in this space, and there are a growing number building on speech and audio signal processing for mental wellness, mostly resolve the tension by keeping the underlying model's actual claims quiet even from users. The algorithm might be trained on labeled clinical data, correlated against real diagnoses in a research setting, but the product surface never says so. What ships is a wellness score with a clean, boring name and copy that undersells what the model is actually doing under the hood.
I'd call that disciplined rather than dishonest. The alternative, a company overselling a voice score as diagnostic-grade without going through FDA clearance as a Software as a Medical Device, has real consequences: enforcement letters, forced product changes, and a credibility hit that follows the company into every future release. The wellness framing isn't a loophole exploited for convenience; it's the correct legal category for a product that generates a signal worth paying attention to but hasn't gone through the clinical validation that a diagnostic claim requires.
Get an FDA clearance eventually, with a validated clinical study behind the score, and the copy can change. Until then, the discipline holds: reflect, don't detect; compare to self, not to population; disclose the limits loudly, and let the score speak quietly.


