The Prosody Register

A Test-Retest Reliability Study for a Voice-Derived Score

A voice-based cognitive-load score needs test-retest reliability or it measures nothing but noise.

Editor at Large · · 8 min read · Updated
Cover illustration for “A Test-Retest Reliability Study for a Voice-Derived Score”
Features · August 19, 2026 · 8 min read · 1,751 words

A voice-derived cognitive-load score only earns its keep if it produces the same reading twice under the same conditions. That's the entire premise of test-retest reliability, and it's the step teams building these models rush past most often, or fake outright. Someone runs one retest, the number looks fine, and the slide gets built before anyone asks what "fine" was supposed to mean.

I've sat in those rooms. If a score swings by half a standard deviation when the same person records the same prompt twenty minutes apart, that instability undercuts the whole claim to measuring cognitive load in the first place, and no amount of downstream polish fixes it.

Voice is a noisy signal, and not in a way you can engineer around entirely. Pitch, jitter, shimmer, pause duration, speech rate: every one of these features moves in response to things that have nothing to do with cognitive load. Room acoustics, mic gain, a dry throat, too much coffee, the hour of the day, it all bleeds into the signal. A reliability study exists to separate the variance that reflects the construct you're measuring from everything else riding along with it. Skip that step and the noise stays fully intact. What you end up shipping is a product that can't tell the difference between someone under genuine strain and someone who just walked in from the cold.

What ICC actually tells you

Intraclass correlation coefficient is the standard statistic here, for good reason. It splits variance into between-subject and within-subject components, which is exactly the split that matters: a high ICC means most of the variance in your scores traces back to real differences between people, not noise within the same person across sessions. Cicchetti's guidelines from 1994, still widely cited, put values below 0.40 in "poor" territory, 0.40 to 0.59 as "fair," 0.60 to 0.74 as "good," and 0.75 and above as "excellent." Koo and Li, writing in 2016 for rehabilitation and clinical measurement research, push stricter, recommending a floor of 0.75 and treating 0.90 as the bar for anything feeding an individual clinical decision.

Consumer and clinical-adjacent teams blur that distinction constantly. An ICC of 0.65 might be perfectly fine for a research tool comparing group averages. It does not hold up for a product telling one specific user their cognitive load is elevated today versus last Tuesday. Population claims and individual claims need different bars. If your score drives a decision about one person, sit closer to 0.90 than to 0.75, and say so in the validation write-up rather than letting a "good" rating on Cicchetti's scale do work it was never built for. Burying that gap in a footnote is how products end up promising something the underlying number never supported.

There's also a choice between ICC forms, and this is where a lot of technically defensible papers quietly report the wrong number. ICC(2,1), the two-way random-effects model with absolute agreement, is what you want when sessions are supposed to be interchangeable, which is exactly the test-retest case. ICC(3,1), the consistency-only version, will almost always come out higher, because it forgives systematic shifts between sessions. Swap microphones between visit one and visit two and ICC(3,1) shrugs it off; ICC(2,1) does not. Report the consistency number if a reviewer wants it. The headline number belongs to absolute agreement, because that's the version that matches how the product actually gets used, on whatever device someone happens to be holding that day.

Designing the recording protocol

Two recordings per participant is a floor, not a target. A single test-retest pair gives you one estimate of within-subject variance with no way to check whether that pair was an outlier, and plenty of teams treat a two-session pilot as though it settled the question. It didn't. Three sessions, spread across different times of day and different days of the week, get you an estimate stable enough to trust. Four sessions across a two-week window is stronger still, budget and recruitment allowing, since that design catches both short-interval noise and longer-interval drift. That distinction matters if the product claims to track load over time rather than in one sitting.

The gap between sessions has to be long enough that participants aren't just repeating a memorized performance, but short enough that the trait you're measuring hasn't plausibly shifted underneath you. A 20-minute to 48-hour window between the first two sessions is standard in voice biomarker work, with a follow-up at one to two weeks to check longer-term stability. Too short, and you're measuring memory of the task. Too long, and you're measuring whether the person's baseline actually changed, a different question, and one your reliability study was never built to answer.

Sample size deserves a straight answer, not a hedge. ICC estimates get noisy fast at small n, so reliability studies need more participants than most people budget for. This is usually where the pushback starts, since recruiting 40 to 60 people for three sessions each costs real money and real weeks. Walter, Eliasziw, and Donner's 1998 sample size formulas, still the standard planning reference, suggest that distinguishing an ICC of 0.90 from a null of 0.70 with reasonable power lands somewhere in that range. Fall below 30 participants and the confidence interval around your ICC will likely swallow the point estimate whole. You might land on 0.85 with an interval running from 0.60 to 0.95, and that tells a reviewer nothing except go collect more data.

Prompt design is where studies quietly fail

The prompt has to hold the cognitive load task constant while still letting natural variation in performance come through. A fixed script read aloud mostly captures reading fluency and vocal tract mechanics, not cognitive load. Reading a familiar passage doesn't tax working memory the way spontaneous speech under constraint does. No amount of clever feature engineering downstream fixes a prompt that never stressed the right system to begin with.

Better designs draw on Sweller's cognitive load theory, first laid out in 1988, which splits load into intrinsic (the task's inherent difficulty), extraneous (difficulty added by clumsy task design), and germane (effort spent actually building understanding). A voice-based task needs to manipulate load in a way that's controlled and repeatable across sessions. Dual-task paradigms handle this well: have the participant run a working-memory task, backward digit span or serial subtraction, counting down from 100 by 7s, borrowed straight from the Mini-Mental State Exam, while speaking continuously about something neutral. The speech is the channel you sample; the arithmetic is the load manipulation riding underneath it.

Whatever task gets chosen, run it at matched difficulty across every session. Subtraction by 7s in session one and subtraction by 3s in session three confounds task difficulty with test-retest interval, and now any change in score can't be pinned on reliability one way or the other. Keep the load manipulation identical across sessions for a given participant, and vary only the surface stimuli, different starting numbers, different neutral topics, so practice effects don't quietly inflate your reliability estimate.

Recording length matters more than teams tend to assume. Most acoustic features used in cognitive load scoring, pitch variability and pause statistics especially, need 60 to 90 seconds of continuous speech before they stabilize. Cut that shorter and measurement noise climbs, dragging your ICC down for reasons that have nothing to do with the model and everything to do with not having enough data per sample. Thirty-second clips save time for the product team but cost the study its precision, and that trade rarely gets made on purpose.

Environmental controls you cannot skip

Microphone consistency is non-negotiable. Let participants use their own phone in session one and a study-provided headset in session two, and you've introduced a confound no statistical correction fully removes, because frequency response, gain, and noise floor differ systematically across hardware. Standardize on one recording setup. If the product is meant to run on consumer phones, standardize on device category instead, and test explicitly across a handful of common categories, rather than letting hardware vary freely inside your reliability sample and hoping it averages out. It won't.

Background noise, room size, and distance from the microphone all move the acoustic features that cognitive load models lean on. Run the study in a room where the acoustics shift session to session, and you're testing the reliability of a score under conditions that are themselves unreliable. That muddies the exact variance decomposition ICC was supposed to hand you cleanly. No amount of post-hoc filtering recovers what a controlled room would have given you for free.

The threshold to accept before shipping

Set the bar before the data comes in, not after. It's tempting to look at an ICC of 0.68, decide it's close enough, and write the validation report around the number you got instead of the number you needed. That rationalization happens in real rooms, with real deadlines attached, more often than anyone likes to admit; I've watched a perfectly good study get reframed after the fact because 0.75 was inconvenient. Following Koo and Li, the pre-registered threshold for any voice-derived score feeding an individual-level decision, clinical or otherwise, should be ICC(2,1) at or above 0.75, with 0.90 as the target for anything higher stakes, like tracking cognitive decline or fatigue in a safety-critical role.

Below 0.75, the honest move is to keep the score in a research-use-only bucket and say so plainly, in language no softer than the number warrants. A score below that line means a meaningful share of what users see session to session is measurement noise, not a real change in their state. No framing makes that acceptable in a product claiming to track something as consequential as cognitive load.

None of this proves the score measures what you say it measures, and it isn't supposed to. A high ICC only proves the score is stable; stability by itself says nothing about whether stable means correct. The score still needs convergent validity against an established measure, pupillometry, EEG-based indices, or a validated self-report instrument like the NASA-TLX. A perfectly reliable score that doesn't correlate with any independent measure of cognitive load is just a very consistent way of measuring something else, and consistency alone has fooled smarter teams than the one shipping next quarter. Run the reliability study first, because validating against an external criterion is pointless if the score can't reproduce on the same person twice. That first step is a foundation, nothing more. The harder work, the part that actually earns the word "validity," starts right after it.

More in Features