Listening for ice: Teaching AI to detect ice using sound

Leora Robinson – leorarobinson13@gmail.com

Brigham Young University, Provo, UT, 84602, United States

Tracianne Neilsen

Popular version of 3pUW6 – Acoustic binary classification of ice cover conditions using deep learning
Presented at the 190th ASA Meeting
Read the abstract at https://eppro01.ativ.me/web/index.php?page=session&project=ASASPRING2026&id=4082831

–The research described in this Acoustics Lay Language Paper may not have yet been peer reviewed–

Some Arctic animals don’t need to see ice to find it—they can hear it. Species like the beluga whale use sound to navigate through icy waters where visibility is limited, finding breathing holes in the ice without ever seeing them. This project asks a simple question: Can a computer learn to do the same? By analyzing acoustic signals, we show that a neural network can detect ice without relying on visual information.

Initial experiments were conducted in a laboratory tank (Figure 1) at the Brigham Young University Department of Physics and Astronomy. We took sound recordings when ice was and was not present on the surface of the water. Then, we trained a machine learning classifier to label the recordings as ‘ice’ or ‘no ice.’

Robotic arms mounted above a large transparent tank filled with water in a laboratory setting with control screens attached.Figure 1. Laboratory tank (side view).

For these experiments, we placed an underwater loudspeaker (transmitter) and an underwater microphone (hydrophone) in the tank. The transmitter produced ultrasonic chirps of increasing frequency when ice was and wasn’t present. We added about 600 pounds of block ice to the tank and took one-second recordings before ice was added, while it was present, and after it melted. We took two additional sets of recordings for testing the neural network: one using block ice and one using pebble ice.

After we acquired the recordings, we needed to label them. We did this using camera footage of the tank (Figure 2). Recordings with about 5% or more ice coverage between the transmitter and the hydrophone were labeled ‘ice,’ and recordings with less than 5% coverage were labeled ‘no ice.’ We chose this 5% threshold to differentiate between negligible and non-negligible ice cover. We converted each labeled recording into a time-frequency spectrogram and used the spectrograms to train a machine learning classifier.

Two robotic arms manipulating multiple white ice cubes floating in a clear water tank from an overhead view.Figure 2. Camera footage of the laboratory tank for labeling.

For the machine learning classifier, we selected a convolutional neural network (CNN) because it can detect important features indicating the presence of ice. We passed the spectrograms and their associated labels through the classifier for training, where the CNN learned to associate certain features of spectrograms with their labels. Ten classifiers were trained to provide a statistical representation of performance.

Diagram showing audio signals converted to spectrograms, processed by a CNN classifier to label presence or absence of ice.Figure 3. Roadmap of how each audio recording was processed and classified.

Once the ten classifiers were trained, we tested their performance on two other datasets that they were not trained on. We did this to see how well the CNN could generalize to other conditions. This generalizability is important because, in practical applications, the ocean environment is always changing: no two recordings will ever have identical conditions. The mean labeling accuracy across the ten classifiers on the testing block ice dataset was 93.5% ± 0.9%. On the pebble ice dataset, the classifiers achieved 94.3% ± 1.4% accuracy. These tests show that the CNNs can generalize well to new conditions.

The high accuracy of these initial experiments indicates that a CNN can use sound to detect the presence of ice. Just as the beluga whale listens for audio cues to find breathing holes in the ice, the neural network extracts important information from the sound to determine whether ice is present.

Using Deep Learning to Enhance Photoacoustic Brain Images

Matthew Olmstead – mjo5585@psu.edu

Instagram: @mattomatty707
Graduate Program in Acoustics, The Pennsylvania State University, University Park, PA, 16802, United States

Hyungjoo Park – hpp5133@psu.edu
MD Rizwanul Kabir – rizwanulkabir@vt.edu
Aiguo Han – aiguohan@vt.edu
Yun Jing – yqj5201@psu.edu

Popular version of 1aBAa7 – Improving Photoacoustic Imaging through the Skull using Deep Learning: Considering 3D Effects
Presented at the 190th ASA Meeting
Read the abstract at https://eppro01.ativ.me/web/page.php?page=Session&project=ASASPRING2026&id=4082496&nohistory&nohistory=true

–The research described in this Acoustics Lay Language Paper may not have yet been peer reviewed–

Did you know that there’s potentially a better way than ultrasound to image your brain? Say hello to photoacoustics, which combines the benefits of optics and sound, giving us the “best of both worlds.” Instead of an acoustic signal, we’re sending a laser signal into the brain, then receiving an acoustic signal back thanks to the phenomenon of the photoacoustic effect! When we’re dealing with a mixed medium like human tissue, the ultrasound signal often gets distorted by the time it’s picked up by the receiving probe. Thankfully in photoacoustics, the acoustic signal only travels one way, so compared to what happens in traditional ultrasound imaging, it goes through less distortion by the time it reaches the ultrasound probe. One example of a mixed medium is the skull, which has a very porous layer in the middle (see Figure 1).

Diagram showing pulse laser and ultrasound array detecting blood vessels inside a human skull.

Figure 1. Schematic of the photoacoustic imaging process. Image courtesy of Hyeonu Heo.

Over the past several years, it has been a challenge trying to get a good acoustic signal when imaging through the skull. However, as you’re probably aware, AI has lately become popular in enhancing different applications, and the biomedical field is no exception to that. This project proposes using a deep learning model to improve photoacoustic images distorted by the skull, by training it on these images and comparing them to their “ground truth” counterparts. By the time the model is finished training, it will be able to improve the quality of images that it has never seen before! The model we’re working with is called U-Net, named after its literal shape of a U. The two main parts of U-Net are the encoder on the left side and the decoder on the right side (see Video 1). The encoder takes an image, lowers its resolution, and extracts important features out of it to learn from. Later, the decoder restores the image’s resolution and is able to pinpoint down to each pixel which parts of the image are what (for example, a blood vessel or background).

So far in our research, we’ve noticed how the types of images that we feed the model are very important. For instance, if we train it on images that only consider 2D wave effects, it isn’t going to perform as well when tested on images with more realistic 3D wave effects. It is crucial that the training data for the model is as realistic as possible, before it gets deployed out into the real world to be used in clinical settings. Fortunately, our U-Net model has proven to be very robust, and the results that we’ve obtained thus far have pointed us toward ways to further improve it. The future in this field is exciting, since several categories of photoacoustic imaging tasks can benefit from deep learning enhancements, such as monitoring stroke diseases.

5aEA2 – What Does Your Signature Sound Like?

Daichi Asakura – asakura@pa.info.mie-u.ac.jp
Mie University
Tsu, Mie, Japan

Popular version of poster, 5aEA2. “Writer recognition with a sound in hand-writing”
172nd ASA Meeting, Honolulu

We can notice a car approaching by noise it makes on the road or can recognize a person by the sound of their footsteps. There are many studies analyzing and recognizing these noises. In the computer security industry, studies have even been proposed to estimate what is being typed from the sound of typing on the keyboard [1] and extracting RSA keys through noises made by a PC [2].

Of course, there is a relationship between a noise and its cause and that noise, therefore, contains information. The sound of a person writing, or “hand writing sound,” is one of the noises in our everyday environment. Previous studies have addressed the recognition of handwritten numeric characters by using the resulting sound, finding an average recognition of 88.4%. Based on this study, we seek the possibility of recognizing and identifying a writer by using the sound of their handwriting. If accurate identification is possible, it could become a method of signature verification without having to ever look at the signature.

We used the handwriting sounds of nine participants, conducting recognition experiments. We asked them to write the same text, which were names in Kanji, the Chinese characters, under several different conditions, such as writing slowly or writing on a different day. Figure 1 shows an example of a spectrogram of the hand-writing sound we analyzed. The bottom axis represents time and the vertical axis shows frequency. Colors represent the magnitude – or intensity – of the frequencies, where red indicates high intensity and blue is low.
handwriting

The spectrogram showed features corresponding to the number of strokes in the Kanji. We used a recognition system based on a hidden Markov model (HMM) – typically used for speech recognition –, which represents transitions of spectral patterns as they evolve in time. The results showed an average identification rate of 66.3%, indicating that writer identification is possible in this manner. However, the identification rate decreased under certain conditions, especially a slow writing speed.

To improve performances, we need to increase the number of hand writing samples and include various written texts as well as participants. We also intend to include writing of English characters and numbers. We expect that Deep Learning, which is attracting increasing attention around the world, will also help us achieve a higher recognition rate in future experiments.

 

  1. Zhuang, L., Zhou, F., and Tygar, J. D., Keyboard Acoustic Emanations Revisited, ACM Transactions on Information and Systems Security, 2009, vol.13, no.1, article 3, pp.1-26.
  2. Genkin, D., Shamir, A., and Tromer, E., RSA Key Extraction via Low-Bandwidth Acoustic Cryptanalysis, Proceedings of CRYPTO 2014, 2014, pp.444-461.
  3. Kitano, S., Nishino, T. and Naruse, H., Handwritten digit recognition from writing sound using HMM, 2013, Technical Report of the Institute of Electronics, Information and Communication Engineers, vol.113, no.346, pp.121-125.