Table of Contents

Measuring Progress by Asking: What Can Speech Brain-Computer Interfaces Communicate?

Accuracy alone cannot tell us what a speech BCI can communicate. Open-Vocabulary Mutual Information measures both lexical coverage and decoding performance.

August 18, 2026•9 min read
By Dulhan Jayalath, Oiwi Parker Jones
Brain-Computer InterfacesSpeech DecodingInformation TheoryEvaluationOVMI

OVMI comparison across reference communication distributions.

How can we unify scores on a common scale to measure progress in speech BCIs?

The field of speech brain-computer interfaces (speech BCIs) is now entering an exciting era. Invasive speech BCIs – those that require surgically implanting a recording device in the brain – are reportably able to decode hundreds of thousands of words from patients’ attempted speech at an astounding level of accuracy (Card et al. 2024) which is better than the best WER across a set of standard benchmarks for SOTA ASR at the time of writing (see e.g. the Hugging Face Open ASR Leaderboard). What’s more, non-invasive devices, which do not necessitate the risk of surgical procedures, are beginning to be able to decode hundreds of words from brain data collected while subjects listen to speech. How can we continue to accelerate improvements in non-invasive BCIs? One strategy is to advance methods for speech decoding from brain data, in part, by comparing what works across as broad a set of studies as possible. But comparing results across studies can be surprisingly difficult.

Studies are not always easy to compare

Let’s start with a hypothetical: suppose two different speech BCI studies.

  • Study A reports an accuracy of 90% over a fixed set of 50 words.
  • Study B reports an accuracy of 70% over a fixed set of 1000 words.

Which system is better?

Try it yourself: Is a larger vocabulary worth lower accuracy? Compare Study A and Study B below, or try a real system, then adjust accuracy and lexical coverage—the share of words a user wants to say that the system supports—to see how both affect communication.

Interactive model

How much can the system communicate?

Reference · SUBTLEX-UK
Example studies

For these hypothetical examples, we assume each vocabulary contains the most frequent SUBTLEX-UK words. Different word choices would change coverage and OVMI.

Real systems
90.0%
45.5%
Vocabulary profile50 most frequent SUBTLEX-UK words50 decoder words

Open-Vocabulary Mutual Information

1.930bits
Coverage0.455
In-vocabulary MI4.238 bits
OVMI1.930 bits

Study A reports a higher accuracy (90% > 70%). But Study B covers a larger vocabulary (1000 > 50 words). If you only want to communicate a few words (e.g. yes, no), then you might prefer the higher accuracy of Study A. But if you want to speak in full sentences, then you may find a 50-word vocabulary to be restrictive: even basic conversational English involves thousands of words. How much accuracy would you be willing to give up for how many more words in your vocabulary?

To add another dimension, the choice of words matters too. 1000 obscure words (e.g. “agelast”, “borborygmus”, “fudgel”, “petrichor”, etc.) are probably less useful than more ordinary, everyday words (e.g. “help”, “thanks”).

  • Agelast: A person who never laughs
  • Borborygmus: A rumbling or gurgling sound from your stomach or intestines
  • Fudgel: Pretending to work when you are actually doing nothing at all
  • Petrichor: The pleasant, earthy smell of rain hitting dry soil

The 50 most common words in English are primarily composed of function words like “the”, which on their own don’t provide much scope for communication (and, depending on the set of function words, may not even be composable into grammatical sentences). To talk about most things requires content words which come from an open set (e.g. bird, cat, truck, airplane, emoji, deepfake, enshittification, situationship, rizz, mog). With “limited” vocabularies (e.g. fewer than tens or even hundreds of thousands of words), difficult choices need to be made about what you most want to talk about (if you haven’t seen it, check out xkcd’s Up Goer Five comic and the following amusing attempts to explain science with a restricted 1000-word vocabulary, as here and here).

Given these challenges, how are we to measure progress in non-invasive speech decoding when different studies report results on different vocabularies (e.g. d’Ascoli et al. 2025 and Jayalath et al. 2025)?

Benchmarks are not the whole answer

A natural solution for measuring progress in science is to standardise evaluation with a benchmark. This is what fields like Computer Vision managed with ImageNet, the famous image classification dataset that is now almost mandatory to report scores on for any paper that introduces a new method. The intention of benchmarks like this is that if all methods are evaluated on the same data under the same settings (e.g. fixed train/test splits), then we can identify which method performs best. Such benchmarks have begun to emerge in speech BCIs, for example in the structure of the LibriBrain (Ă–zdogan et al. 2025) and LibriBrain100 (Mantegna et al. 2026a) datasets and API for loading them in the pnpl library for our yearly PNPL competition (Landau et al. 2025, Mantegna et al. 2026b). But benchmarks are not the only way to address these challenges, and, despite our efforts to provide benchmarks and standards, it is natural for researchers in a field to explore other paths too.

There are many sources of heterogeneity in experimental settings within the field. Invasive and non-invasive systems use different recording technologies and necessitate different experimental protocols. Some studies decode attempted speech, where participants attempt to say sentences they are prompted with; others decode perceived speech, where participants perceive a visual or auditory sentence stimulus; and others even decode imagined speech, where participants silently think about a prompt (e.g. Ballyk et al. 2026). Study participants also come from different populations: sometimes paralysed patients in invasive studies, and healthy subjects in non-invasive ones. The prompts in those studies also differ, with some focusing on sentences or words related to clinical caregiving, while others focus on naturalistic narrative or conversational speech.

This raises another question:

What does performance on a speech BCI benchmark mean for communication?

As mentioned earlier, while a benchmark might tell us that model A outperforms model B on a 100-word classification problem, it does not tell us whether those words allow us to say 5%, 30%, or even 80% of the sentences we might actually wish to say. To give an example, if I wanted to say “In the beginning, the Universe was created,” but of the words in that sentence, only “the” were in the 100-word set, then the model would not be particularly useful.

Existing metrics are insufficient

Metrics like accuracy and word error rate (WER) are commonly reported on studies and benchmarks in the field. While these are excellent measures of decoding fidelity, they do not account for the size of the vocabulary or choice of supported words in the system. It is therefore possible to achieve very low WER, or equivalently high accuracy, while supporting very few of the words that are useful for communication. This motivates the development of a new metric, one that evaluates speech BCIs relative to a common and parameterisable distribution of words that are necessary for a user’s communication.

Introducing OVMI: Open-Vocabulary Mutual Information

OVMI maps study-specific scores onto a shared communication scale.

OVMI maps study-specific scores onto a common scale by measuring performance relative to a shared distribution over the words a user would want to say.

We recently introduced Open-Vocabulary Mutual Information (OVMI) designed to answer this question. If you have taken part in or followed the PNPL competition, you may have already noticed that OVMI appears as one of the official auxiliary metrics on the leaderboard. We have included it because we hope it can help make results from the competition more useful beyond this individual benchmark. With OVMI, future methods may be compared not only on how accurately they decode the competition vocabulary (via our official balanced accuracy metric), but also on how much information that performance would convey in another communication setting (e.g. clinical caregiving).

The intuition behind OVMI is that a useful speech BCI needs to do two things well:

  • Be expressive: does the vocabulary contain the words the user will need to communicate?
  • Be accurate: if a word is supported, then can the system decode it reliably?

OVMI is a single useful metric that measures both of these properties simultaneously. It can be thought of as

\[ \text{OVMI} = \text{Lexical coverage} \times \text{in-vocabulary mutual information} \]

Here, lexical coverage is the probability that a word drawn from a sentence that a user would want to say is actually supported by the speech BCI. If such a word is supported, we refer to the vocabulary as “covering” the intended communication distribution. In-vocabulary mutual information is derived from the decoder’s accuracy on its supported vocabulary, where mutual information is higher if the decoder is more accurate. Thus, increasing either lexical coverage or decoding performance can increase OVMI, but a useful system ultimately needs to be strong on both.

How do we define what sentences a user may want to communicate? This communication distribution specifies what we want a vocabulary to ideally cover. OVMI therefore evaluates a system relative to such a reference communication distribution over words. This could be words and their frequencies in everyday conversational English, or alternatively, the language used in a clinical care setting, or ideally, the things that a user would want to communicate.

We appreciate that users may not always have a specific vocabulary or communication distribution in mind. Therefore, we suggest using a standard reference distribution that provides broad coverage. In our work, we use SUBTLEX-UK (van Heuven et al. 2014), a list of word frequencies derived from a large corpus of film and television subtitles. As a result, it captures a lot of what a user may want to say in conversational English speech. While this particular choice is not important, we propose that the community converges on some standard, or set of standard reference distributions, that can be used for comparability.

OVMI complements existing measures by asking the broader question that we started with: given the words a user may wish to say, how much can this system communicate?

To see how we are using OVMI to benchmark progress in the field, check out our live leaderboard: https://neural-processing-lab.github.io/OVMI/

You can also select OVMI as a measurement on the 2026 PNPL competition’s leaderboard: https://neural-processing-lab.github.io/2025-libribrain-competition/editions/2026/leaderboard/

We have also released a simple Python package to help you compute OVMI: https://github.com/neural-processing-lab/OVMI

And lastly, we have a preprint on arXiv with all the mathematical, scientific, and technical details of OVMI: https://arxiv.org/abs/2609.02887

đź“„Cite the paper

If you enjoyed this blog post, please cite the paper.

@article{jayalath2026ovmi,
  title={A Common Measure of Communication for Speech Brain-Computer Interfaces},
  author={Jayalath, Dulhan and Ballyk, Benjamin and Parker Jones, Oiwi},
  journal={arXiv preprint arXiv:2609.02887},
  year={2026}
}

References

[1]

An Accurate and Rapidly Calibrating Speech Neuroprosthesis

Nicholas S. Card, Maitreyee Wairagkar, Carrina Iacobacci, Xianda Hou, Tyler Singer-Clark, Francis R. Willett, Erin M. Kunz, Chaofei Fan, Maryam Vahdati Nia, Darrel R. Deo, Aparna Srinivasan, Eun Young Choi, Matthew F. Glasser, Leigh R. Hochberg, Jaimie M. Henderson, Kiarash Shahlaie, Sergey D. Stavisky, David M. Brandman

New England Journal of MedicineVol. 391pp. 609--618(2024)
[2]

Towards decoding individual words from non-invasive brain recordings

Stéphane d’Ascoli, Corentin Bel, Jérémy Rapin, Hubert Banville, Yohann Benchetrit, Christophe Pallier, Jean-Rémi King

Nature CommunicationsVol. 16pp. 10521(2025)
[3]

Unlocking non-invasive brain-to-text

Dulhan Jayalath, Gilad Landau, Oiwi Parker Jones

arXiv preprint(2025)
[4]

LibriBrain: Over 50 Hours of Within-Subject MEG to Improve Speech Decoding Methods at Scale

Miran Ă–zdogan, Gilad Landau, Gereon Elvers, Dulhan Jayalath, Pratik Somaiya, Francesco Mantegna, Mark Woolrich, Oiwi Parker Jones

arXiv preprint(2025)
[5]

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Ă–zdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones

arXiv preprint(2026)
[6]

The 2025 PNPL competition: Speech detection and phoneme classification in the LibriBrain dataset

Gilad Landau, Miran Ă–zdogan, Gereon Elvers, Francesco Mantegna, Pratik Somaiya, Dulhan Jayalath, Luisa Kurth, Teyun Kwon, Brendan Shillingford, Greg Farquhar, Minqi Jiang, Karim Jerbi, Hamza Abdelhedi, Yorguin Mantilla Ramos, Caglar Gulcehre, Mark Woolrich, Natalie Voets, Oiwi Parker Jones

arXiv preprint(2025)
[7]

The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100

Francesco Mantegna, Gereon Elvers, Dulhan Jayalath, Gilad Landau, Tasha Kim, Miran Ă–zdogan, Luisa Kurth, Teyun Kwon, SungJun Cho, Benjamin Ballyk, Alex Fung, Anna Greer, Pratik Somaiya, Christian Herff, Yorguin Mantilla Ramos, Hamza Abdelhedi, Karim Jerbi, Greg Farquhar, Brendan Shillingford, Mark Woolrich, Oiwi Parker Jones

arXiv preprint(2026)
[8]

Physiological Noise Augmentation Improves Non-Invasive Brain-to-Speech

Benjamin Ballyk, Teyun Kwon, Miran Ă–zdogan, Oiwi Parker Jones

arXiv preprint(2026)
[9]

SUBTLEX-UK: A new and improved word frequency database for British English

Walter J. B. van Heuven, Pawel Mandera, Emmanuel Keuleers, Marc Brysbaert

Quarterly Journal of Experimental Psychology(2014)
[10]

A Common Measure of Communication for Speech Brain-Computer Interfaces

Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones

arXiv preprint(2026)