One brain implant combines speech and gestures to control an avatar
An experiment with people with paralysis decoded two forms of communication at once. Training on combined attempts was essential, and the system still uses limited repertoires.

Leitura autorizada · 3 crédito(s) restante(s)
A greeting can bring together a word and a wave, but teaching a brain interface to recognize both at once is not simply a matter of joining two programs that work separately. A University of California, San Francisco team showed that a single implant can provide signals for decoding speech and gestures in parallel and controlling a digital representation of its user. Published on September 14 in Nature Neuroscience, the study by Samantha Brosler, Jessie Liu, Alexander Silva and colleagues demonstrates this possibility in tasks with predefined choices, still far from unrestricted conversation.
The researchers recorded brain activity from three people with different degrees of paralysis. In each person, a grid of 253 electrodes rested on the brain's surface, covering regions associated with speech and movement. This technique, called electrocorticography, captures electrical signals from populations of neurons. The avatar demonstrations and combined speech-and-gesture tasks involved two participants: one with paralysis following a brainstem stroke and another with amyotrophic lateral sclerosis, a disease that damages neurons responsible for movement.
In trials guided by on-screen cues, participants tried to produce phrases, perform or imagine gestures, or combine both actions, according to their abilities. Two machine-learning models received signals from the same implant: one classified phrases and the other classified gestures. The recognized phrase appeared as text, while the avatar waved, moved its head or performed another selected gesture. The commands went to the digital body; the experiment did not restore movement to the paralyzed body.
The difficulty emerged when models trained only on isolated speech or isolated gestures had to interpret simultaneous attempts. Some electrodes responded to both behaviors, and patterns recorded during combined attempts were not recognized as effectively. Including simultaneous examples in training improved performance. Another adjustment taught each model not to respond when only the other modality was present: for example, the speech classifier learned that an isolated attempt to wave should not become a phrase.
In the real-time simultaneous test with the participant with amyotrophic lateral sclerosis, the system correctly identified 70% of phrases and 66% of gestures, from ten alternatives of each type plus a rest class. The 95% confidence intervals, which express uncertainty in these estimates, were 63.5% to 76% for speech and 59.5% to 72.5% for gestures. Random selection among the eleven classes provided a reference of approximately 9.1%. These are separate measurements: they do not mean that both phrase and gesture were correct in 70% of trials.
The central contribution is showing how to train an interface for behaviors that coexist in communication. A gesture can complement a phrase or serve as a response on its own, expanding what an avatar allows someone to express. Turning this demonstration into an everyday tool will require larger repertoires, shorter response delays and performance testing in more people. The studied system also depends on a wired connection between the implant and external equipment.
Key points
- A single implant supplied signals for parallel speech and gesture decoding; two participants used the avatar.
- In the simultaneous test with ten phrases and ten gestures, one participant achieved 70% and 66% accuracy, respectively.
- Training with combined actions reduced interference, but repertoire, sample size and equipment still limit everyday use.

Comments
No comments have been published yet.
Sign in with a subscription to comment.