AI-Powered Brain Implants Restore Expressive, Real-Time Speech for ALS Patients and Beyond

“Translating neural activity into text, which is how our previous speech brain-computer interface works, is akin to text messaging. It’s a big improvement compared to standard assistive technologies, but it still leads to delayed conversation. By comparison, this new real-time voice synthesis is more like a voice call.” said Sergey Stavisky, senior author at UC Davis, in a statement that highlights the scope of the latest advance in brain-computer interface.

Image Credit to depositphotos.com

A patient with amyotrophic lateral sclerosis (ALS), whose speech was almost undecipherable, was the first to speak not only words but also emotive speech tone, pitch, and even song through a surgically implanted BCI recently. This device, which is reported in Nature, utilizes artificial intelligence to interpret neural activity in 10-millisecond increments, converting attempted speech into audible, expressive voice output almost immediately.

At the center of this technology is deep-learning algorithms that have been trained to read the electrical signals created in the speech motor cortex of the brain. The subject’s implant consisting of 256 silicon electrodes, each 1.5 mm long reads these signals with very fine temporal resolution. As the man tries to speak, the system reads the neural patterns not just for the words he means to utter but also for the slight modulations that express emphasis, interrogation, and emotional intonation.

Previous BCI systems, though promising, had considerable latency occasionally as much as three seconds between user intent and system response. Delays of this sort interfere with conversation flow and restrict social engagement. The new method cuts this delay to under 10 milliseconds, a delay almost imperceptible from the response of natural speech. “With instantaneous voice synthesis, neuroprosthesis users will be able to be more included in a conversation. For example, they can interrupt, and people are less likely to interrupt them accidentally,” Stavisky elaborated.

One key technical breakthrough is that the system can decode neural activity in real time, feeding synthesized speech as the user tries to speak. It does this by correlating high-gamma power features 70 to 170 Hz oscillations across a network of electrodes spanning motor, premotor, and somatosensory cortices. Feature extraction is done in 10-millisecond frames so the deep-learning models are able to monitor the rapid, dynamic brain activity that accompanies planning and production of speech.

The decoding pipeline uses a mixture of unidirectional and bidirectional long short-term memory (LSTM) neural networks. The unidirectional model identifies the beginning of speech attempts, and the bidirectional model converts the neural signals to sequences of linear predictive coding (LPC) coefficients a type of acoustic feature that can be synthesized into audible speech with vocoder technology. The closed-loop design is optimized for low latency, with each unit running asynchronously to provide persistent, real-time output.

The performance of the system is not just technical. Synthesized words were identified correctly by human raters in listening tests at a rate of almost 80%, which is a huge improvement compared to earlier BCI-based communication systems. When listening in one study, almost 60% of the synthesized words were heard correctly, whereas only 4% when the subject was trying to speak unaided. The voice synthesized was also individualized, based on recordings of the participant’s pre-ALS speech, retaining his own distinctive vocal identity a change that researchers and patients alike characterize as groundbreaking.

Personalization is made possible by “voice banking,” which involves training AI models with archive recordings of the user’s original voice. The outcome is a computer-generated output that not just carries words, but also the user’s typical timbre, rhythm, and expressiveness. As BrainGate trial director Dr. Leigh Hochberg phrased it:”We had used an AI approach to taking his pre-ALS voice,” Dr. Hochberg said, “and allowing that voice to be the one that was heard through the computer.”

The wider context of this success is the quick development of BCI technologies for communication assistance. These early systems were based on text displays or sluggish, letter-by-letter spelling via eye-tracking or remaining muscle movement. Current breakthroughs in high-density electrocorticography (ECoG) and microelectrode arrays have allowed for the direct reading out of speech-related neural activity without bypassing the damaged motor pathways and restoring the potential for natural conversation.

Deep-learning neural decoding algorithms, and in particular recurrent neural networks, have been able to represent the intricate, time-varying patterns underlying fluent speech. These models are able to generalize from large datasets of imagined or attempted speech to novel words and phrases, even some not in the training set. There are some systems now at the fluency and rate of natural conversation, and they decode as high as 90 words per minute with accuracy.

One of the most dramatic abilities of the new system is that it can support expressive, paralinguistic functions. The subject could vary intonation to form questions, stress words, and even hum melodies over a range of pitches. “We don’t always use words to communicate what we want. We have interjections. We have other expressive vocalizations that are not in the vocabulary,” said Maitreyee Wairagkar, project scientist at UC Davis. This adaptability represents a breakthrough in the return of not only functional communication, but the entire range of human expression.

The potential for clinical practice and assistive technology is deep. For individuals with ALS, brainstem stroke, or locked-in syndrome, being able to communicate in real time through a voice recognized and expressive could be a return to significant social interaction and autonomy. As UCSF’s Dr. Edward Chang noted, “This new technology has tremendous potential for improving quality of life for people living with severe paralysis affecting speech.”

And although challenges persist in scaling such systems for wider clinical application, maintaining neural recordings over the long term, and applying the technology to patients with more severe speech loss or other neurological conditions, the accelerating intersection of neuroscience, engineering, and artificial intelligence is already transforming the field of assistive communication, giving us a glimpse of a future in which speech loss may not require voice loss.

spot_img

More from this stream

Recomended

Discover more from Modern Engineering Marvels

Subscribe now to keep reading and get access to the full archive.

Continue reading