A speech BCI does not have one performance number. Word accuracy, sentence correctness, output speed, delay and the amount of help needed answer different questions. Comparing headline percentages without their units can make an ordinary difference in measurement look like a major scientific disagreement.

The June 2026 UC Davis account illustrates the problem directly: its controlled word-accuracy statement and its participant-rated sentence-correctness statement are not the same outcome. The 2025 streaming paper adds another dimension—how frequently the system generates audio. 1 2

A usable measurement dictionary

Measure The question it answers Common misreading
Word error rate How many word-level substitutions, deletions and insertions occur relative to the reference? The percentage of complete sentences a user finds useful
Sentence correctness How well whole outputs match the intended message under the stated scoring rule A directly interchangeable word-accuracy percentage
Words per minute How quickly the defined output is produced End-to-end productivity including every pause and correction
Latency How long a particular step or response takes Accuracy, or the time for the entire interaction
Vocabulary size How many words or outputs the system can select under the test Proof that all possible conversations were tested
Independent use What the participant can operate without the specified assistance Proof that no setup, caregiver or maintenance support is ever required

The definitions in this table are analytical distinctions. The studies must still specify how they operationalized and measured each item.

A fictional word-error calculation

Suppose a reference contains 100 words. A system substitutes five words, deletes three and inserts two. Under the conventional counting definition:

Word error rate = (5 + 3 + 2) ÷ 100 = 10%.

This is an illustrative calculation, not an observed result from a medical study. It explains why a word-error metric needs a reference and a defined scoring procedure. It also explains why inserted words can matter even when every intended word appears somewhere in the output.

A sentence with one wrong word might still be useful—or might communicate the wrong medication, name or negation. A single aggregate score cannot decide which kind of error occurred. That is why usability and error analysis belong beside raw decoding performance.

Controlled output and everyday usefulness

A controlled task gives a known target, making automated scoring possible. An open conversation may depend on the participant's own assessment of what they intended to say. The UC Davis report distinguishes those contexts rather than offering one universal accuracy number. 1

Our inference is that the two approaches are complementary. Controlled testing helps characterize a system reproducibly. Real-world use can reveal whether people continue to find it useful. Neither should be substituted for the other merely because its result looks more favorable.

A small vocabulary can change the meaning of speed

The NIH account of the 2025 streaming work describes different operating conditions for broader and restricted vocabularies. 3

A comparison that selects a restricted-task speed from one study and a broad conversational speed from another is not controlled for task difficulty. That does not invalidate either result; it invalidates the simple ranking.

A useful comparative table would retain at least output type, vocabulary constraint, prompts, participant count, error definition and the observation setting. A missing field should remain unknown rather than be filled with the most flattering assumption.

Concurrent tasks deserve separate testing

The September 2026 speech-and-gesture paper tests a problem not answered by a speech-only score: whether two channels still work when attempted together. Its results show why combining successful isolated decoders is not itself proof of successful concurrent communication. 4

For a future user, an interface may need to distinguish intended speech, intended cursor movement and inactivity. A result for one of those states does not automatically establish performance in all of them.

What a meaningful “best BCI” claim would need

It would first specify the user's task and the comparison setting. Is the goal rapid short messages, expressive audio, computer work, reliable all-day use or another outcome? It would then compare like with like and include the burden of operation.

The publications covered here do not supply a randomized head-to-head evaluation of every BCI against every alternative. We therefore report their distinct contributions rather than manufacture a winner.

Use the home-use report, streaming-voice report and multimodal report for the actual study-specific findings. This page is the measurement key that keeps those findings from being distorted when placed next to each other.