The first time many sighted people hear an experienced screen-reader user’s synthetic speech, the reaction is disbelief. At very high speed, words can seem to collapse into a compressed stream of sound. Yet for the person using it, speech rate is not a novelty setting. It determines how quickly email, documents, menus, websites and work systems can be processed.
The problem is that there is no universal “best” speed. Faster speech can increase information throughput, but only until comprehension begins to fail. Slower speech may be easier, yet painfully inefficient. And until now, a low comprehension score could be difficult to interpret: was the speech too fast, was the sentence itself difficult, or had the listener simply not developed enough skill at high-speed synthetic speech?
A team led by Takahiro Miura of the National Institute of Advanced Industrial Science and Technology (AIST), with Masatsugu Sakajiri and colleagues at Tsukuba University of Technology and Ken-ichiro Yabu and colleagues at the University of Tokyo’s Research Center for Advanced Science and Technology, has developed a way to separate those factors statistically. The work was announced September 25 and is scheduled for presentation at Interspeech 2026 in Sydney, held September 27–October 1.[1][2]
Separate the difficulty of the sentence from the ability of the listener
The researchers borrowed a tool from educational measurement: Item Response Theory, or IRT. In testing, a low score can reflect either a difficult question or a less capable test-taker. IRT places item difficulty and person ability on a common statistical scale so they can be estimated separately.
The same logic can be applied to synthetic speech. If someone fails to understand a sentence read at high speed, the model does not automatically interpret that as a lack of listening ability. It estimates how difficult that particular spoken item was and how much comprehension ability the listener demonstrated across items.
The team used a Bayesian version of IRT. That matters because accessibility studies often cannot recruit hundreds or thousands of participants. Bayesian estimation can remain useful with relatively small samples and can explicitly represent uncertainty around the estimates.[2]
From 150 to 500 words per minute—up to roughly 3.5 times conversational speech
The experiment used Japanese sentences from the phonemically balanced ITA Corpus. Synthetic speech was created at five settings between 150 and 500 words per minute. In measured Japanese speech rate, that corresponded to roughly 6–27 mora per second. Typical conversation is around 7–8 mora per second, putting the experiment at approximately 0.8 to 3.5 times ordinary conversational speed.[2]
Eleven visually impaired screen-reader users listened to the material. Researchers recorded not only whether they understood the content but also how easy the speech was to hear and how understandable it felt.
One result was unsurprising: as speech became faster, item difficulty increased systematically. At 150 words per minute, comprehension was close to perfect across the sample. Performance fell at 300 and 500 words per minute. Overall, the researchers observed a marked transition around 250–300 words per minute.[2]
But 250 words per minute is not “the answer”
The more important result was the spread between people. Individual differences in modeled listening ability were larger than the difficulty changes produced by speech rate itself.
That means the study is not recommending that every blind or low-vision user set a screen reader to 250 words per minute. The opposite lesson is more important: a single default may underuse the skills of an experienced fast listener while overwhelming someone who has not developed the same ability.
It is a measurement framework that can estimate a different appropriate speed for each user by separating sentence difficulty from listening ability.
Blind users in this small sample maintained comprehension at higher speeds
The study also found a notable group difference. Participants who were totally blind tended to maintain comprehension at faster speech rates than participants with low vision.
Their everyday settings reflected that pattern. Totally blind users reported using screen-reader speeds averaging roughly 2.3 times the standard rate, compared with approximately 1.4 times among low-vision users.[2]
The researchers’ analysis suggested that the difference could not be explained solely by greater exposure to fast speech. Visual status and usage experience each appeared to have independent relationships with listening ability.
That finding needs restraint. The study included only 11 people. It cannot support claims that every totally blind user listens faster than every low-vision user, nor does it establish a biological cause. The team explicitly plans larger studies to test these patterns.[2]
Why speed matters: screen readers are not just audiobook players
A screen reader converts interface text and controls into synthetic speech. Narrator, VoiceOver, NVDA and Japan’s PC-Talker are among widely used examples. For a person operating a computer without relying on vision, the screen reader is not simply reading prose aloud; it is the primary interface to the operating system and applications.[2]
Sighted readers do not process a webpage character by character. They scan headings, whitespace, position and visual hierarchy, jumping quickly to relevant information. Screen-reader users recreate some of that efficiency through keyboard commands, heading navigation, landmarks—and high speech rates.
That is why a speed difference can have real economic and educational consequences. Reading dozens of workplace messages, reviewing documentation, searching the web or studying course material at 1.5 times versus 3 times speed can dramatically change how much information can be processed in a day.
Japan has about 273,000 registered people with visual disabilities
Japan’s Ministry of Health, Labour and Welfare estimated that about 273,000 people with visual impairments held a physical disability certificate in its 2022 Survey on Difficulties in Daily Life. The broader population experiencing significant vision loss is larger because the figure does not include everyone with low vision, age-related decline or no disability certificate.[6]
Accessibility therefore cannot end at “the text can technically be read aloud.” The speed, structure and usability of that audio interface affect education, employment and participation.
A long Japanese history of assistive speech interfaces
One of the study’s authors, University of Tokyo research adviser Tohru Ifukube, has spent roughly half a century working in welfare engineering. His research history includes assistive speech and hearing technologies as well as a Japanese screen reader known as the 95 Reader, ultrasonic glasses and tactile interfaces for visually impaired users.[7]
Seen in that context, the new research marks a shift in the maturity of the field. Early assistive-computing problems asked whether information could be converted into speech at all. Today’s problem is more refined: how quickly can a particular person consume that speech while preserving understanding?
Comprehension and comfort are not identical
Another subtle point is that a rate can feel comfortable without maximizing comprehension—or maximize comprehension while feeling frustratingly slow.
Experienced users may find ordinary synthetic speech painfully slow, yet pushing speed higher can eventually reduce accurate understanding even if the user feels familiar with the sound. By collecting ratings of listening ease and understandability alongside objective content questions, the research opens the possibility of optimizing more than one outcome.
The next step: a short test that recommends your speed
The team plans larger studies to improve estimation accuracy and to examine how vision status and screen-reader experience relate to listening ability. One priority is following relatively new users over time to measure how high-speed listening skill develops with practice.
The researchers also want to build a tool that can assess a person’s synthetic-speech comprehension in a short session and then recommend an appropriate rate immediately.[2]
If that becomes practical, accessibility setup could eventually move beyond a generic speed slider. An operating system or screen reader might ask a user to complete a brief listening assessment and produce a personalized starting point—with room to adapt as skill improves.
Accessibility is moving from “can use” to “can use efficiently”
Digital accessibility is often measured in binary terms: can a screen reader reach the text, is an image labeled, can the interface be controlled without a mouse? Those questions are essential, but they do not capture the whole experience.
If a person can technically access the same information but needs several times longer to process it, a meaningful productivity gap remains. The new work approaches that gap through one deceptively simple parameter: speed.
Synthetic-speech technology has spent decades trying to sound increasingly natural. Screen-reader use poses a different engineering goal. The best voice may not be the voice that sounds most like casual human conversation. It may be the voice that a particular listener can understand as quickly and accurately as possible.
Sources
- Japan Science and Technology Agency: Quantifying comprehension of synthetic speech, Sept. 25, 2026
- Joint release by AIST, Tsukuba University of Technology, University of Tokyo and JST
- AIST: Research release
- University of Tokyo RCAST: Research summary
- Tsukuba University of Technology: Research summary
- Ministry of Health, Labour and Welfare: 2022 Survey on Difficulties in Daily Life
- University of Tokyo RCAST: Tohru Ifukube profile and assistive-technology history

