AI Content Moderation Takes a Lesson from Economics #ASA190

Economic theory of attention can help understand and increase reliability of AI models searching for online hate speech.

PHILADELPHIA, May 12, 2026 — Spend enough time on the internet, and you’ll likely encounter some pretty appalling content. Hate speech tends to flourish on social media and in online communities, particularly those with little to no moderation. Even on sites with strict community standards, the volume of content makes effective moderation nearly impossible.

Large language models (LLMs) may be able to solve this problem. These AI algorithms can rapidly analyze both the content and the context of large volumes of text, filter hate speech automatically, and provide feedback to human reviewers. However, LLMs are expensive to run at scale, especially when asked to provide explanations for each piece of content they flag.

Yuan Zhao from the New Jersey Institute of Technology will present his research on creating an interpretable and low-cost method for evaluating LLMs’ hate speech classification Tuesday, May 12, at 1:30 p.m. ET as part of the 190th Meeting of the Acoustical Society of America, running May 11-15.

Man wearing glasses typing code on a Lenovo laptop in a focused programming session.

Researchers used an economic model to understand how large language models classify hate speech. Credit: Kowalski7cc on Wikimedia (CC0)

Zhao’s framework relies on the Rational Inattention (RI) model, an economic idea developed to explain human behavior. The model describes how humans act when their attention is limited and assigns a cost to that attention. According to the model, people tend to reserve their attention for high-reward decisions, spending it where it would have the greatest effect.

And while LLMs are not humans, these ideas of attention and decision-making can still be applied.

“LLMs are different from people, but we envision them as decision-makers facing some trade-off between performance and computational cost,” said Zhao. “Our approach uses the RI model as a simple yet interpretable tool to understand how LLMs make decisions.”

Zhao tested LLMs in a range of conditions to determine whether they behave like rational decision-makers. Then, he used the RI model to mimic the behavior of those LLMs, finding that it accurately predicts how LLM performance changes in different conditions.

This analysis can be utilized to guide digital communities using LLMs as part of their content moderation efforts.

“LLMs are already widely used, but there are still concerns about their reliability. Models like Rational Inattention can help make them more trustworthy by showing how their performance changes when text becomes ambiguous or intentionally disguised,” said Zhao. “This helps online platforms identify when human review is needed and where the system needs improvement.”

###

For more information:
AIP Media
1 301.209.3090
media@aip.org


Main Meeting Website: https://acousticalsociety.org/philadelphia/
Technical Program: https://eppro01.ativ.me/web/planner.php?id=ASASPRING2026

ASA PRESS ROOM
In the coming weeks, ASA’s Press Room will be updated with newsworthy stories and the press conference schedule at https://acoustics.org/asa-press-room/.

LAY LANGUAGE PAPERS
ASA will also share dozens of lay language papers about topics covered at the conference. Lay language papers are summaries (300-500 words) of presentations written by scientists for a general audience. They will be accompanied by photos, audio, and video. Learn more at https://acoustics.org/lay-language-papers/.

PRESS REGISTRATION
ASA will grant free registration to the in-person conference at the Philadelphia Marriott Downtown for credentialed and professional freelance journalists. If you are a reporter and would like to attend the meeting and/or press conferences, contact AIP Media Services at media@aip.org. For urgent requests, AIP staff can also help with setting up interviews and obtaining images, sound clips, or background information.

ABOUT THE ACOUSTICAL SOCIETY OF AMERICA
The Acoustical Society of America is the premier international scientific society in acoustics devoted to the science and technology of sound. Its 7,000 members worldwide represent a broad spectrum of the study of acoustics. ASA publications include The Journal of the Acoustical Society of America (the world’s leading journal on acoustics), JASA Express Letters, Proceedings of Meetings on Acoustics, Acoustics Today magazine, books, and standards on acoustics. The society also holds two major scientific meetings each year. See https://acousticalsociety.org/.

AI Voices Are Easier to Understand than Human Voices

Only a few seconds of sampling is enough to create an AI copy of a person’s voice, and researchers are unsure why they are so intelligible.

Illustration of a female singer projecting sound waves with glowing stars symbolizing vocal excellence.

Voice clones, which can recreate a human’s speech using only a few seconds of recorded speech, are more intelligible in noisy environments, research finds. AIP

WASHINGTON, April 21, 2026 — Synthetic voices are increasingly a part of our lives, from digital assistants like Siri and Alexa to automated telemarketers and answering machines. With the expansion of generative AI, a new type of synthetic voice has been developed: voice clones, which can recreate a facsimile of a person’s voice from only a few seconds of recorded speech.

In JASA, published on behalf of the Acoustical Society of America by AIP Publishing, a pair of researchers from University College London and the University of Roehampton evaluated the…click to read more

From: The Journal of the Acoustical Society of America
Article: Voice clones are easier to understand in noise than their human originals: the voice cloning intelligibility benefit
DOI: 10.1121/10.0043094

Understanding Why Engine Noise Feels Loud in Hybrid Vehicles with AI

Shinichi Suganuma – shinichi_suganuma@camal.mech.chuo-u.ac.jp

Graduate School of Science and Engineering
Chuo University
1-13-27 Kasuga
Bunkyo-ku, Tokyo, 112-8551
Japan

Shimpei Nagae
Nissan Motor Co., Ltd.
Kanagawa, Japan

Takeshi Toi
Chuo University
Tokyo, Japan

Popular version of 4aNSa2 – Development of a Machine Learning Model to Predict Engine Noise Perception Considering Regional and Driving Environment Differences
Presented at the 189th ASA Meeting
Read the abstract at https://doi.org/10.1121/10.0041106

–The research described in this Acoustics Lay Language Paper may not have yet been peer reviewed–

When driving a hybrid vehicle, many people notice the moment when the quiet electric drive suddenly switches to the engine — and the engine can feel “loud,” even when the actual sound level is modest. Why does this happen? And does the way drivers perceive this noise differ across countries? In this study, we used machine learning to predict how people judge engine noise annoyance and to uncover insights that may help make future hybrid vehicles more comfortable.

Figure 1. AI Model for Predicting Engine Noise Perception
Video 1. On-Road Driving Example for Data Collection

We conducted on-road evaluations in Japan, the United States, and the United Kingdom. During each test, we simultaneously recorded in-cabin sound, vehicle parameters, and drivers’ ratings of engine noise on a three-level scale (“Not noisy,” “Noisy,” “Very noisy”), creating a dataset for AI training. In Japan, we used the series-hybrid Nissan Note e-POWER. In the U.S., where this model is not sold, we reproduced its engine sound on the Nissan Ariya EV, and in the U.K. we used the Qashqai e-POWER engine sound played on the Ariya. Because vehicles, drivers, and road environments differed across regions, the study provided a stringent test of model generality.

Figure 2. On-Road Evaluation Conditions in Japan, the U.S., and the U.K.

First, we tested how accurately AI could predict annoyance using only in-cabin sound data such as loudness and sharpness etc. In Japan, prediction accuracy reached about 57%. When we added three vehicle parameters — engine speed, acceleration torque, and vehicle speed — accuracy increased to 67%, demonstrating that driving conditions, not just sound, play an important role in annoyance perception. The same trend was observed in the U.S. and the U.K.

Figure 3. Prediction Accuracy Improvements Using Vehicle Data and Time History

However, the relative importance of the three vehicle parameters differed by region. In Japan and the U.S., engine speed contributed most strongly to predictions. In contrast, in the U.K., acceleration torque was the most influential factor. This likely reflects the presence of many roundabouts in the U.K. test route, where frequent acceleration and deceleration lead drivers to value the coherence between engine sound and vehicle motion. This aligns with the author’s own experience living in the U.K. for three years.

Next, we incorporated several seconds of engine-speed history into the vehicle parameters. In all regions, adding this short-term history improved prediction accuracy. Although the optimal history length differed slightly — around 5.5 seconds in Japan and 6.5 seconds in the U.S., — the common finding was clear: people judge engine noise not from a single moment but from the pattern of change over several seconds.

Figure 4 Prediction Improvement When Engine-Speed History Is Added

Despite differences in vehicles, traffic environments, and evaluation routes, considering “vehicle operating conditions” together with “recent temporal changes” consistently improved the AI’s ability to predict annoyance across all regions. These findings provide valuable clues for designing hybrid vehicles that feel smoother and more comfortable for drivers around the world.

Can Artificial Intelligence Accurately Clone Dysphonic Voices?

Pasquale Bottalico – pb81@illinois.edu

University of Illinois at Urbana-Champaign
Champaign, Illinois, 61801
United States

Additional Authors
Charles J. Nudelman
Daniel Fogerty
Virginia Tardini
Keiko Ishikawa

Popular version of 2aSCa8 – Can Artificial Inteligence Accurately Clone Dysphonic Voices? A Perceptual and Intelligibility Assessment
Presented at the 189th ASA Meeting
Read the abstract at https://doi.org/10.1121/10.0040365

–The research described in this Acoustics Lay Language Paper may not have yet been peer reviewed–

Artificial intelligence is now remarkably good at cloning human voices, but can it convincingly imitate a disordered voice? Our findings suggest that while AI excels at copying healthy speech, it still struggles to capture the acoustic complexity of dysphonia, a condition that makes the voice sound rough, strained, or breathy.

Dysphonia affects millions of people and often reduces speech intelligibility, especially in noisy environments. Because collecting large amounts of patient data can be difficult, researchers wondered whether AI voice-cloning technologies might one day help them simulate disordered speech for training, education, or early-stage clinical research.

To test this idea, the team recorded 12 speakers (six with healthy voices and six with dysphonia)  and used a commercial AI system to create a digital “voice clone” of each person. These AI voices were trained using about one minute of recorded speech for each speaker. More than 60 listeners participated in three online experiments designed to evaluate whether the AI-generated voice clones truly preserved the qualities of disordered speech.

Watch the short video below to see exactly how the experiment worked.

In the listening tasks, participants heard pairs of sentences. Sometimes both sentences were from the real speaker, sometimes both were AI-generated, and sometimes one was real and one was AI. In some trials, listeners tried to decide whether the two voices came from the same person. In others, they had to identify which sentence (if any) was produced by AI. A third task tested how well listeners understood real and AI-generated dysphonic speech in background noise.

In the first experiment, as shown in Figure 1, listeners were very accurate when both samples were real. Here, accuracy refers to the proportion of trials in which listeners correctly judged whether the two voice samples were from the same or different speakers. Accuracy dropped slightly when both samples were AI-generated. But when one sample was real and the other AI-generated, performance fell sharply, especially for healthy voices, where the AI clones often sounded strikingly similar to the real person.

Figure 1. Bar plot showing the percentage of correct AI identification responses across conditions for normal and dysphonic voices. Bars represent mean percentages with 95% confidence intervals. Note: RL = real speech; AI = AI-generated speech.

Figure 2. Bar plot showing the percentage of correct AI identification responses across conditions for normal and dysphonic voices. Bars represent mean percentages with 95% confidence intervals. Note: RL = real speech; AI = AI-generated speech.

A second experiment asked listeners to identify which sentences were AI-generated. For healthy voices, AI was difficult to detect. For dysphonic voices, however, listeners were more successful — suggesting the AI system smoothed out or failed to reproduce key features of dysphonia. The results are shown in Figure 2.
The final experiment delivered the strongest finding: AI-generated dysphonic voices were significantly more intelligible than real dysphonic voices when played in background noise. In other words, the AI unintentionally “cleaned up” the voice disorder, creating speech that sounded clearer and easier to understand than the real dysphonic voices. The results are shown in Figure 3.

These results demonstrate that while AI voice cloning is impressively realistic for healthy speech, it does not yet capture the natural irregularities of disordered voices. For now, real patient recordings remain essential. However, this research highlights the exciting potential of improved AI tools in the future.

Figure 3. Mean intelligibility scores (IS) of normal and dysphonic groups in real and AI-generated voice conditions. The IS values vary from 0 to 1. Error bars indicate standard errors. Note: RL = real speech; AI = AI-generated speech.

Laying the Groundwork to Diagnose Speech Impairments in Children with Clinical AI #ASA188

Laying the Groundwork to Diagnose Speech Impairments in Children with Clinical AI #ASA188

Building large AI datasets can help experts provide faster, earlier diagnoses.

Media Contact:
AIP Media
301-209-3090
media@aip.org

PedzSTAR
pedzstarpr@mekkymedia.com

NEW ORLEANS, May 19, 2025 – Speech and language impairments affect over a million children every year, and identifying and treating these conditions early is key to helping these children overcome them. Clinicians struggling with time, resources, and access are in desperate need of tools to make diagnosing speech impairments faster and more accurate.

Marisha Speights, assistant professor at Northwestern University, built a data pipeline to train clinical artificial intelligence tools for childhood speech screening. She will present her work Monday, May 19, at 8:20 a.m. CT as part of the joint 188th Meeting of the Acoustical Society of America and 25th International Congress on Acoustics, running May 18-23.

Children at a childcare center. Credit: GETTY CC BY-SA

AI-based speech recognition and clinical diagnostic tools have been in use for years, but these tools are typically trained and used exclusively on adult speech. That makes them unsuitable for clinical work involving children. New AI tools must be developed, but there are no large datasets of recorded child speech for these tools to be trained on, in part because building these datasets is uniquely challenging.

“There’s a common misconception that collecting speech from children is as straightforward as it is with adults — but in reality, it requires a much more controlled and developmentally sensitive process,” said Speights. “Unlike adult speech, child speech is highly variable, acoustically distinct, and underrepresented in most training corpora.”

To remedy this, Speights and her colleagues began collecting and analyzing large volumes of child speech recordings to build such a dataset. However, they quickly realized a problem: The collection, processing, and annotation of thousands of speech samples is difficult without exactly the kind of automated tools they were trying to build.

“It’s a bit of a catch-22,” said Speights. “We need automated tools to scale data collection, but we need large datasets to train those tools.”

In response, the researchers built a computational pipeline to turn raw speech data into a useful dataset for training AI tools. They collected a representative sample of speech from children across the country, verified transcripts and enhanced audio quality using their custom software, and provided a platform that will enable detailed annotation by experts.

The result is a high-quality dataset that can be used to train clinical AI, giving experts access to a powerful set of tools to make diagnosing speech impairments much easier.

“Speech-language pathologists, health care clinicians and educators will be able to use AI-powered systems to flag speech-language concerns earlier, especially in places where access to specialists is limited,” said Speights.

——————— MORE MEETING INFORMATION ———————
Main Meeting Website: https://acousticalsociety.org/new-orleans-2025/
Technical Program: https://eppro01.ativ.me/src/EventPilot/php/express/web/planner.php?id=ASAICA25

ASA PRESS ROOM
In the coming weeks, ASA’s Press Room will be updated with newsworthy stories and the press conference schedule at https://acoustics.org/asa-press-room/.

LAY LANGUAGE PAPERS
ASA will also share dozens of lay language papers about topics covered at the conference. Lay language papers are summaries (300-500 words) of presentations written by scientists for a general audience. They will be accompanied by photos, audio, and video. Learn more at https://acoustics.org/lay-language-papers/.

PRESS REGISTRATION
ASA will grant free registration to credentialed and professional freelance journalists. If you are a reporter and would like to attend the meeting and/or press conferences, contact AIP Media Services at media@aip.org. For urgent requests, AIP staff can also help with setting up interviews and obtaining images, sound clips, or background information.

ABOUT THE ACOUSTICAL SOCIETY OF AMERICA
The Acoustical Society of America is the premier international scientific society in acoustics devoted to the science and technology of sound. Its 7,000 members worldwide represent a broad spectrum of the study of acoustics. ASA publications include The Journal of the Acoustical Society of America (the world’s leading journal on acoustics), JASA Express Letters, Proceedings of Meetings on Acoustics, Acoustics Today magazine, books, and standards on acoustics. The society also holds two major scientific meetings each year. See https://acousticalsociety.org/.

ABOUT THE INTERNATIONAL COMMISSION FOR ACOUSTICS
The purpose of the International Commission for Acoustics (ICA) is to promote international development and collaboration in all fields of acoustics including research, development, education, and standardization. ICA’s mission is to be the reference point for the acoustic community, becoming more inclusive and proactive in our global outreach, increasing coordination and support for the growing international interest and activity in acoustics. Learn more at https://www.icacommission.org/.