Imagine the frustration of not being able to say what’s on your mind, restrained by paralysis. For millions of individuals with diseases like ALS or disabling strokes, this is their reality every day. But new research at UC Berkeley and UC San Francisco is writing a new chapter, offering hope of naturalistic speech to the voiceless.

The breakthrough is in a brain-computer interface (BCI) that translates activity in the motor cortex, the part of the brain where speech is made, into spoken words in near-real-time. The breakthrough, outlined in Nature Neuroscience, solves one of the biggest hurdles for speech neuroprostheses: latency. Earlier systems took up to eight seconds to respond, halting conversation. And now researchers have cut that lag time to a fraction of a second, allowing for continuous, unbroken communication.
“Our streaming approach brings the same rapid speech decoding capacity of devices like Alexa and Siri to neuroprostheses,” said Gopala Anumanchipalli, Robert E. and Beverly A. Brooks Assistant Professor of Electrical Engineering and Computer Sciences at UC Berkeley and co-principal investigator of the study. This near-synchronous streaming of speech is an improvement in generating fluent, naturalistic speech synthesis.
The tech records the neural signals once the brain has determined what to communicate and how. Co-lead author Cheol Jun Cho told, “We are essentially intercepting signals where the thought is translated into articulation and in the middle of that motor control.” It takes a sample from motor cortex neural data and translates the signals using top-of-the-line AI algorithms into speech.
In a bid to get the machine trained, researchers collaborated with Ann, a participant whose capacity to produce sounds vocally was eliminated as a result of paralysis. Ann tried but without vocally producing silent sounds on a word screen like “Hey, how are you?” This helped the team gauge her neural response towards intended speech patterns. Since Ann was unable to make sounds using speech, the researchers completed through AI fill-ins with a pretrained text-to-speech system augmented by pre-injury recordings of Ann’s voice. The outcome: A voice that actually sounds remarkably like Ann’s own, not just making speech but self.
“This new technology has tremendous potential for improving quality of life for people living with severe paralysis affecting speech,” said UCSF neurosurgeon Edward Chang and senior co-principal investigator of the study. Chang directs a clinical trial at UCSF that employs high-density electrode arrays to capture brain surface neural activity.
What’s so surprising about this discovery is just how versatile it is. The approach is also applicable with other brain-sensing interfaces, such as invasive microelectrode arrays (MEAs) that break through the surface of the brain and non-invasive sensors such as surface electromyography (sEMG) that monitor activity of facial muscles. Co-lead author and graduate student Kaylo Littlejohn of UC Berkeley said, “By applying exact brain-to-voice translation to other silent-speech sets, we were able to demonstrate that this approach isn’t device-dependent.” Generalization of the system to rare words like “Alpha” and “Bravo” from the NATO phonetic alphabet was also investigated. Researchers discovered that the AI model could recover these words despite not being included in the training set. “We wanted to see if we could generalize to the unseen words and really decode Ann’s patterns of speaking,” Anumanchipalli said. “We found that our model does this well, which shows that it is indeed learning the building blocks of sound or voice.”
Ann herself delineated the revolutionary nature of the streaming synthesis method compared to previous methods for text-to-speech. “Hearing her own voice in near-real time increased her sense of embodiment,” Anumanchipalli reported. In this, psychological and emotional value of restoring speech in a natural, intimate way is quoted.
In the future, scientists continue to improve the technology. Producing synthesized speech that is more expressive—engaging tonal nuance, pitch, and emotional range is the objective. “That’s ongoing work, to try to see how well we can actually decode these paralinguistic features from brain activity,” Littlejohn said.
The application of this discovery goes far beyond the case. In merging the latest in AI and neural decoding, the scientists are making it available in a variety of applications in medicine and communication. As Anumanchipalli states, “It is exciting that the latest AI advances are greatly accelerating BCIs for practical real-world use in the near future.”
Supported by organizations such as the National Institute on Deafness and Other Communication Disorders (NIDCD) and the Science and Technology Agency of Japan, this book is a multi-disciplinary attack on one of the most disabling neuroprosthetic problems.
With ongoing innovation, hope for regaining normal speech in paralytic individuals is nearer than ever. To find out more, read the complete study on brain-to-voice neuroprosthesis, or see how AI is transforming speech synthesis. And conclusions drawn from real-time generated speech are a preview of what communication technology might become like in the future.

