Tired? We Can Hear It in Your Voice #ASA190

Physical exertion affects the pitch, intensity, and temporal characteristics of speech, making speech recognition difficult for systems used by emergency response personnel and wearable devices.

PHILADELPHIA, May 14, 2026 — The “talk test” is often used as a low-tech way to measure exercise intensity: If you can easily talk or even sing, your workout is fairly light, but if conversation is difficult, you are exercising vigorously.

Physical task stress affects the coordination between breathing and speaking. Zahra Omidi from the University of Texas at Dallas studies this relationship and will present her work Thursday, May 14, at 11:15 a.m. ET as part of the 190th Meeting of the Acoustical Society of America, running May 11-15.

Group of runners participating in an outdoor race on a cloudy day, with two women in the foreground wearing athletic gear with an inset  audio waveform and spectrogram showing sound intensity with a highlighted segment in red.

Vocal pitch, intensity, and pause structure are the vocal characteristics most impacted by changes in breathing and exercise. Credit: Zahra Omidi and Presidio of Monterey (CC0)

“Physical exertion directly alters respiration and phonation, and because speech shares the same respiratory system, these changes propagate into pitch, timing, and voice quality,” Omidi said.

Vocal pitch, intensity, and pause structure are the vocal characteristics most sensitive to changes in breathing and effort. Pitch and intensity both increase, while intensity also becomes less stable. Because speakers need to allocate more time to breathing, their speech rate slows down and becomes more segmented with longer and more frequent pauses.

Some of these changes might not be so noticeable to a listener, but the measurements clearly indicate a physiological difference.
“Features like pitch, intensity, and timing show clear and consistent changes, even when those differences are not immediately obvious by listening,” Omidi said. “This suggests that physical stress may operate below the threshold of perceptual salience in some cases but still induces measurable changes in the production mechanism.”

Understanding exactly how physical stress causes changes to vocal patterns can help train speech recognition systems, which often struggle with speech that differs from the average.

“Examples include emergency response, military operations, aviation under workload, and wearable voice interfaces, where people are speaking while physically active,” Omidi said. “In all these cases, speech deviates from neutral conditions due to respiratory and vocal effort constraints, leading to reduced intelligibility and system performance.”

In order to better represent real-world speech behavior, Omidi hopes researchers will adapt a more holistic view of speech variation as a reflection of a speaker’s characteristics rather than focusing solely on linguistics. Task stress is just one of the many physiological variables that can affect these variations.

“Human speech is inherently shaped by the body, and physical task stress provides a clear example of how physiological factors influence speech production,” Omidi said.

###

For more information:
AIP Media
1 301.209.3090
media@aip.org


Main Meeting Website: https://acousticalsociety.org/philadelphia/
Technical Program: https://eppro01.ativ.me/web/planner.php?id=ASASPRING2026

ASA PRESS ROOM
In the coming weeks, ASA’s Press Room will be updated with newsworthy stories and the press conference schedule at https://acoustics.org/asa-press-room/.

LAY LANGUAGE PAPERS
ASA will also share dozens of lay language papers about topics covered at the conference. Lay language papers are summaries (300-500 words) of presentations written by scientists for a general audience. They will be accompanied by photos, audio, and video. Learn more at https://acoustics.org/lay-language-papers/.

PRESS REGISTRATION
ASA will grant free registration to the in-person conference at the Philadelphia Marriott Downtown for credentialed and professional freelance journalists. If you are a reporter and would like to attend the meeting and/or press conferences, contact AIP Media Services at media@aip.org. For urgent requests, AIP staff can also help with setting up interviews and obtaining images, sound clips, or background information.

ABOUT THE ACOUSTICAL SOCIETY OF AMERICA
The Acoustical Society of America is the premier international scientific society in acoustics devoted to the science and technology of sound. Its 7,000 members worldwide represent a broad spectrum of the study of acoustics. ASA publications include The Journal of the Acoustical Society of America (the world’s leading journal on acoustics), JASA Express Letters, Proceedings of Meetings on Acoustics, Acoustics Today magazine, books, and standards on acoustics. The society also holds two major scientific meetings each year. See https://acousticalsociety.org/.

AI Voices Are Easier to Understand than Human Voices

Only a few seconds of sampling is enough to create an AI copy of a person’s voice, and researchers are unsure why they are so intelligible.

Illustration of a female singer projecting sound waves with glowing stars symbolizing vocal excellence.

Voice clones, which can recreate a human’s speech using only a few seconds of recorded speech, are more intelligible in noisy environments, research finds. AIP

WASHINGTON, April 21, 2026 — Synthetic voices are increasingly a part of our lives, from digital assistants like Siri and Alexa to automated telemarketers and answering machines. With the expansion of generative AI, a new type of synthetic voice has been developed: voice clones, which can recreate a facsimile of a person’s voice from only a few seconds of recorded speech.

In JASA, published on behalf of the Acoustical Society of America by AIP Publishing, a pair of researchers from University College London and the University of Roehampton evaluated the…click to read more

From: The Journal of the Acoustical Society of America
Article: Voice clones are easier to understand in noise than their human originals: the voice cloning intelligibility benefit
DOI: 10.1121/10.0043094

The Science of Screaming

Karen Perta – karen.perta@elmhurst.edu

Instagram: @karenperta
Elmhurst University, Elmhurst, IL, 60126, United States

Zhaoyan Zhang, UCLA School of Medicine, Los Angeles, CA, United States.
Donna Erickson, Haskins Laboratories, New Haven, CT, United States.
Ryoko Hayashi, Kobe University, Kobe, Japan.
Toshiyuki Sadanobu, Kyoto University, Kyoto, Japan.

Popular version of 1pSC9 – Physiologic and acoustic characteristics of the angry scream
Presented at the 189th ASA Meeting
Read the abstract at https://doi.org/10.1121/10.0040197

–The research described in this Acoustics Lay Language Paper may not have yet been peer reviewed–

Most people can recall a day so bad that ended with screaming into a pillow. Emotional vocalization is a critical part of human communication. People scream when having fun at sporting events and theme parks, for safety, or to be heard in noisy environments. However, not all screaming and yelling is the same. Some may lose their voice after one night at a concert; others can protest on the picket lines for days without a problem. Why is this?

The purpose of this study is to analyze and compare angry, emotional screaming with trained, “healthy” yelling using magnetic resonance imaging (MRI) and acoustic measures. The MRI shows movements inside of the vocal tract so we can understand exactly how these sounds are created. In this study, a single vocally trained female participant produced angry screaming versus “healthy” belting. Here is a look inside the vocal tract during these sounds:

Figure 1. MRI images of Scream versus Belt (courtesy of authors).

Acoustic measures help characterize the differences between the sounds and provide further insight into how they are produced. Both MRI and acoustic analyses help determine the features that are harmful to the vocal folds versus the features that allow the voice to be heard safely. Here is a power spectrum view that shows frequency (x-axis) and intensity (y-axis) of the sounds as one snapshot in time:

Figure 2. Power spectrum of Scream versus Belt (courtesy of authors).

Based on the MRI measures, we determined that Scream was produced with 1) the highest position of the larynx 2) the largest mouth opening 3) the smallest throat space. Belt was produced with 1) a high larynx position though to a less extreme degree 2) a smaller mouth opening 3) more open space in the throat. Compared to Belt, Scream was also produced with an extremely high pitch – twice that of Belt.

During Scream, the tight throat space led to prolonged contact and strong compression of the vocal folds. This allowed Scream to produce higher intensity (stronger harmonic peaks in the spectrum) at high frequencies (above 6kHs) in Scream as compared to Belt. However, this high intensity production came at the cost of vocal fold injury. The Scream caused the participant to develop small vocal fold lesions that took about two weeks to resolve:

Figure 3. Participant vocal fold lesions following scream (courtesy of authors).

In conclusion, Scream is a primitive vocalization that is produced with a very constrictive action that is similar to swallowing. During swallowing, the vocal tract and vocal folds squeeze and compress in order to keep food and liquid from going into the airway. In contrast, Belt is a learned, trained behavior that is less constrictive and “overrides” innate tendencies for squeezing the vocal tract and pressing the vocal folds. During screaming, the highly constrictive actions of the vocal tract put extra strain and force on the vocal folds that contribute to vocal fold injury. Though it may take some practice, safe yelling should not be tight, feel painful, or cause voice loss. Use caution. Happy yelling!

A Twangy Timbre Cuts Through the Noise

Among loud noise, a brassy and bright voice can help speakers be understood.

A study by Tsai et al. showed that twangy, female voices are best understood amongst plane and train sounds. Credit: AIP

A study by Tsai et al. showed that twangy, female voices are best understood amongst plane and train sounds. Credit: AIP

WASHINGTON, July 29, 2025 — Twangy voices are a hallmark of country music and many regional accents. However, this speech type, often described as “brassy” and “bright,” can also be used to get a message across in a noisy environment.

In JASA Express Letters, published on behalf of the Acoustical Society of America by AIP Publishing, researchers from Indiana University found that it was easier to understand twangy female voices compared to neutral voices when…click to read more

From: JASA Express Letters
Article: How vocal timbre impacts word identification and listening effort in traffic-shaped noises
DOI: 10.1121/10.0037043

Why is it easier to understand people we know?

Emma Holmes – emma.holmes@ucl.ac.uk
X (Twitter): @Emma_Holmes_90

University College London (UCL), Department of Speech Hearing and Phonetic Sciences, London, Greater London, WC1N 1PF, United Kingdom

Popular version of 4aPP4 – How does voice familiarity affect speech intelligibility?
Presented at the 186th ASA Meeting
Read the abstract at https://doi.org/10.1121/10.0027437

–The research described in this Acoustics Lay Language Paper may not have yet been peer reviewed–

It’s much easier to understand what others are saying if you’re listening to a close friend or family member, compared to a stranger. If you practice listening to the voices of people you’ve never met before, you might also become better at understanding them too.

Many people struggle to understand what others are saying in noisy restaurants or cafés. This can become much more challenging as people get older. It’s often one of the first changes that people notice in their hearing. Yet, research shows that these situations are much easier if people are listening to someone they know very well.

In our research, we ask people to visit the lab with a friend or partner. We record their voices while they read sentences aloud. We then invite the volunteers back for a listening test. During the test, they hear sentences and click words on a screen to show what they heard. This is made more difficult by playing a second sentence at the same time, which the volunteers are told to ignore. This is like having a conversation when there are other people talking around you. Our volunteers listen to many sentences over the course of the experiment. Sometimes, the sentence is one recorded from their friend or partner. Other times, it’s one recorded from someone they’ve never met. Our studies have shown that people are best at understanding the sentences spoken by their friend or partner.

In one study, we manipulated the sentence recordings, to change the sound of the voices. The voices still sounded natural. Yet, volunteers could no longer recognize them as their friend or partner. We found that participants were still better at understanding the sentences, even though they didn’t recognize the voice.

In other studies, we’ve investigated how people learn to become familiar with new voices. Each volunteer learns the names of three new people. They’ve never met these people, but we play them lots of recordings of their voices. This is like when you listen to a new podcast or radio show. We’ve found that people become very good at understanding these people. In other words, we can train people to become familiar with new voices.

In new work that hasn’t yet been published, we found that voice familiarization training benefits both older and younger people. So, it may help older people who find it very difficult to listen in noisy places. Many environments contain background noise—from office parties to hospitals and train stations. Ultimately, we hope that we can familiarize people with voices they hear in their daily lives, to make it easier to listen in noisy places.