Forum Discussion
Improve Voice Access Accuracy for Atypical Speech
Improve Voice Access Accuracy For Users With Atypical Speech
Greetings,
First, I would like to say that I appreciate Microsoft's continued investment in accessibility and speech-based computer control. Voice Access has the potential to be an extremely important accessibility tool for people who cannot reliably use a mouse, keyboard, or touchscreen. However, when i have test it on Windows 11 Pro 25h2 last public version update, Voice Access misunderstands a large portion of what I say even when using a good microphone in a silent room.
So I recently joined the Windows Insider Program specifically to test the latest improvements to Windows Voice Access, especially the new Voice Isolation feature. I tested the new Voice Isolation feature and successfully completed the voice enrollment process. Voice Access was able to record my voice and create the voice profile without problems. However, after testing Voice Isolation, speech recognition accuracy remained very low, making me be sure to believe there is still an important accessibility gap for users with atypical speech.
I have atypical speech due to a health condition and I use a breathing machine. The machine itself is very quiet, but it produces a small amount of airflow while I speak. My voice is also relatively low in volume and does not always follow the speech characteristics that modern speech recognition systems appear to expect.
Even when I am in a quiet room with very little background noise, Voice Access misunderstands a large portion of what I say. This suggests that the primary problem in my situation is not simply environmental noise or interference from other speakers. The difficulty appears to be that the speech recognition model itself does not interpret my speech characteristics accurately.
I have also noticed an interesting difference between languages. The English version of Voice Access generally understands me better than the Brazilian Portuguese version, despite Brazilian Portuguese being my native language. This suggests that recognition quality may also vary significantly between language models for users whose speech differs from the typical speech patterns represented in the training data.
Voice Isolation is useful, but it is not speech training. I believe this distinction is extremely important. Voice Isolation appears to identify the user's voice and separate it from background noise, environmental sounds, or other people speaking nearby. That is valuable.
However, identifying which voice belongs to the user is different from teaching the speech recognition system how that user speaks. Generally for someone with typical speech, improving the signal quality may significantly improve recognition.
In the other hand for someone with atypical speech, however, the audio can be perfectly clean and still be interpreted incorrectly by the speech recognition model. This means that Voice Isolation can help the system hear my voice instead of the environment, but it does not necessarily help the speech recognizer understand my individual pronunciation and speech patterns. This is why I believe Voice Access would benefit enormously from an optional personalized speech recognition training system.
Older Speech Recognition Systems Already Solved Part Of This Problem
The application that has worked best for me so far is an accessibility program called Motrix, developed by UFRJ (Federal University of Rio de Janeiro):
https://intervox.nce.ufrj.br/motrix/
Motrix is a very old application that uses Microsoft's SAPI 4 speech recognition technology. Despite being based on technology from the Windows XP/Vista era, it provides significantly better results for my speech than many modern recognition systems. One major reason is that it includes a built-in voice training process. The user reads predefined text, and the speech recognition system uses this material to adapt itself to the way that individual user speaks. After training, the system becomes better at interpreting that person's speech.
This type of personalization was also available in older speech recognition products such as:
* IBM ViaVoice
* Dragon NaturallySpeaking
* older speech recognition systems based on Microsoft SAPI
These systems recognized that one universal speech model would not work equally well for every person.
Modern speech recognition models are much more powerful overall, but many modern systems seem to have moved away from explicit per-user acoustic training. For most users, that may be acceptable. For accessibility users with atypical speech, however, personalized training can be essential.
Why This Matters For Accessibility
For someone using speech recognition primarily for convenience, a recognition error may only be an annoyance. For someone who depends on Voice Access to operate the computer, recognition accuracy can determine whether the computer is independently usable at all.
Commands such as:
"click"
"scroll"
"move"
"open"
"close"
or application names must be recognized reliably.
If a user has to repeat a command five or ten times, or if Voice Access consistently interprets the same pronunciation incorrectly, speech control becomes exhausting and unreliable. This problem becomes even more significant for users who use Voice Access precisely because conventional input devices are difficult or impossible for them to use.
Accessibility software therefore needs to consider not only users with typical speech, but also users with:
* atypical speech patterns
* dysarthric speech
* reduced speech volume
* altered articulation
* respiratory support
* speech affected by physical disabilities
*other speech characteristics that differ from the speech patterns commonly represented in large speech recognition datasets
Not every user in these groups will have the same needs, and that is exactly why optional personalization would be so valuable.
Suggested Feature: Personalized Speech Training
I would strongly suggest adding an optional feature such as a personalized speech training and voice adaptation that could work similarly to older speech recognition training systems.
For example, a Voice Access personalized speech training would ask the user to read a medium size text (5.000~7.500 character lenght) from a example text list so the system could then build a personalized speech profile that adapts recognition to the user's pronunciation and acoustic characteristics. Training could happen gradually. A basic training session might take only a few minutes, while users who need additional adaptation could optionally provide more speech samples over time.
Personalized speech training seems to be the first priority to be implement in a future update. Ideally, this personalization would be stored locally on the device so it can work without internet connection and it could be reset, retrained, exported, or backed up by the user.
Voice Isolation And Personalized Training Could Complement Each Other
I do not believe Voice Isolation should be replaced. I believe it could become the first stage of a much more powerful accessibility system.
For example:
Microphone → Voice Isolation → Personalized Speech Model → Voice Access Recognition
* Voice Isolation could remove unwanted audio and identify the correct speakers;
* Personalized speech training could then adapt the recognition system to the way that specific person speaks.
These are two different problems, and solving both would potentially make Voice Access dramatically more accessible.
Command-specific Training Could Also Help
Even if full personalized speech recognition is technically difficult to implement, a smaller first step could be personalized training for Voice Access commands. A user could record commands several times, such as:
"click"
"scroll down"
"move left"
"open Start"
"show numbers"
and so on.
Voice Access could learn how that particular user pronounces those commands. For users who mainly depend on Voice Access for computer control rather than long-form dictation, even this limited form of personalization could make a major difference.
Microsoft and Nuance
Microsoft now owns Nuance, a company with decades of experience in speech recognition and voice accessibility technologies. Dragon NaturallySpeaking and related technologies demonstrated many years ago that speaker adaptation and personalized voice profiles can significantly improve recognition for individual users.
Modern machine learning technologies should make it possible to revisit this concept in a much more advanced way. I hope Microsoft can combine the strengths of modern speech recognition with the personalization capabilities that older systems provided.
Why I Am Submitting This Through The Insider Program
I joined the Windows Insider Program specifically because I wanted to test these new Voice Access improvements and provide meaningful accessibility feedback before features become widely deployed.
Voice Isolation is an interesting and potentially useful improvement, but my testing suggests that voice isolation alone does not address the main recognition problem experienced by some users with atypical speech.
The speech recognizer needs a way to adapt to the user, rather than requiring every user to adapt their speech to the recognizer. That distinction is extremely important.
Conclusion
Please consider adding optional personalized speech recognition training to Voice Access.
It could make Voice Access usable for people who are currently poorly understood by generic speech recognition systems. For users with atypical speech, this is not merely a request for slightly better recognition accuracy. It can determine whether voice control is a practical accessibility solution at all.
Voice Access already provides an excellent foundation for hands-free Windows control. Adding genuine per-user speech adaptation could make it accessible to a much broader group of people.
If other users have similar experiences, please consider supporting or upvoting this feedback so Microsoft can better understand the demand for personalized speech recognition.
Please also keep the discussion respectful and constructive. Accessibility needs vary greatly between individuals. A feature that may appear unnecessary to one user can be essential for another.
Thank you for considering this feedback and for continuing to improve accessibility in Windows.