Across Acoustics
Across Acoustics
POMA Student Paper Competition: Honolulu
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
In this episode, find out what the next generation of acousticians is researching! In this episode, we talk to the latest round of POMA Student Paper Competition winners, from the joint 189th meeting of the ASA and the Acoustical Society of Japan, held in Honolulu in December 2025. Their topics include:
- Perception of the PIN/PEN merger in speakers from the US South (Irene Smith, McGill University)
- Using distance to better inform models for multi-channel sources separation (Rajesh Rameshbabu, University of Illinois, Chicago)
- Designing resin reeds for oboes (Fumihiko Kurosawa, University of Tsukuba)
- Modelling cochlear synaptopathy, or "hidden hearing loss" (Kai Burian, Technical University Munich)
Associated papers:
Irene Smith and Meghan Clayards. "Merged perception of PIN and PEN: Interspeaker variation or partial merger in perception?" Proc. Mtgs. Acoust. 60, 060013 (2025) https://doi.org/10.1121/2.0002242
Rajesh Rameshbabu; Rashen Fernando; Ryan Corey. "Learning distance-dependent spatial structure for multichannel source separation." Proc. Mtgs. Acoust. 60, 055004 (2025) https://doi.org/10.1121/2.0002298
Fumihiko Kurosawa; Naoto Wakatsuki; Tadashi Ebihara. "Control of characteristics through topological optimization of 3D-printed oboe reeds." Proc. Mtgs. Acoust. 60, 035003 (2025) https://doi.org/10.1121/2.0002307
Kai Burian, Ahsan J. Cheema, and Sunil Puria. "Modeling loudness contours to determine human auditory nerve fiber distribution and cochlear nerve degeneration." Proc. Mtgs. Acoust. 60, 050006 (2025) https://doi.org/10.1121/2.0002265
Read more from Proceedings of Meetings on Acoustics (POMA).
Learn more about Acoustical Society of America Publications.
Kat Setzer (00:27)
Today we're featuring a new round of POMA Student Paper Competition winners, this time from the 189th joint meeting of the Acoustical Society of America and the Acoustical Society of Japan. The meeting was held in December 2025 in Honolulu. First up, I'm talking to Irene Smith, who authored the paper “Merge perception of PIN and PEN: Interspeaker variation or partial merger and perception?” Congrats on the award and thanks for taking the time to speak with me today. How are you?
Irene Smith (00:54)
I'm great. Thanks.
Kat Setzer (00:56)
So, first, tell us a bit about your research background.
Irene Smith (00:58)
So I recently finished my PhD in linguistics at McGill University. I'm a phonetician, so that means that I study the production and perception of speech sounds. And in my research, I'm particularly interested in how acoustic cues associated with speech vary both across and within speech communities, and how this kind of variation affects speech perception as well. I also have a bachelor's degree in both linguistics and electrical engineering, and I worked previously as an engineer doing signal processing for ocean acoustics applications, but I always knew I wanted to do a PhD. So I ended up deciding to study linguistics because that's where the burning questions were for me. But my signal processing and acoustic modeling skills have definitely helped with my work in phonetics.
Kat Setzer (01:43)
Yeah, yeah, I was gonna say I feel like there's a lot of signal processing that can be used in terms of perception of speech and everything. So…
Irene Smith (01:50)
For sure, yeah.
Kat Setzer (01:51)
Speaking of speech perception, why are linguists interested in the pronunciation of the words “pin” and “pen” and what is the PIN-PEN merger?
Irene Smith (01:50)
So linguists don't actually care about the words “pin” and “pen,” per se. They're really just a stand-in or placeholder for a whole set of words involving two specific vowels in a specific context. So the vowels we're interested in are /ɪ/, like in the word “kit” or “pin,” and /ɛ/, like in the word “dress” or “pen.” And linguists are interested in how these two vowels relate to each other before nasal consonants, so m or n, like in the words “pin” and “pen.”
So the PIN-PEN merger is a pattern of speech specific to the US South and to African American English, where pin and pen are pronounced the same. So that's where we get the notion of merger between two otherwise distinct sounds. But the PIN-PEN merger doesn't just apply to these two words. It also applies to words like him and hem, or even image or gentle. So I actually would pronounce the word “gentle” like “gintle” because I have the PIN-PEN merger.
Kat Setzer (02:57)
Oh, interesting. Where are you from?
Irene Smith (03:00)
I'm from Texas.
Kat Setzer (03:02)
Okay, yeah, there you go. I'm from Atlanta, so…
Irene Smith (03:05)
Nice.
Kat Setzer (03:06)
So how does it show up both in speech production and speech perception?
Irene Smith (03:11)
So the way I just described it was very production-centric, right? I said these two vowels are pronounced the same in prenasal context. But mergers are also said to exist in perception. So somebody might not be able to hear the difference between pin and pen, even when those two words are produced distinctly. But the empirical picture for perception of pin and pen is much less well established than it is for merged production of pin and pen.
Kat Setzer (03:38)
Okay, okay, got it. This study was actually an extension of some previous work you did. Can you give us an overview of that study and its findings?
Irene Smith (03:47)
So the goal of the first study was to establish the empirical facts about merged perception of pin and pin in individuals from the US South. So to do this, we created stimuli on a continuum between /ɪ/ and /ɛ/. So vowels are acoustically characterized by their first two spectral peaks, which we call “formants.” These peaks are formed by two subresonances in the oral cavity, so where we put our tongue changes where those resonances fall. And /ɪ/ and /ɛ/ are very close together in formant space, which is why they are susceptible to merger. We can synthesize a continuum between two vowels by applying a series of filters that change the formant frequencies of the vowels. So the middle of the continuum should be an ambiguous vowel that falls kind of halfway between /ɪ/ and /ɛ/.
So we took our vowel continuum stimuli and we placed them in two different environments. We put them in prenasal environments to create a continuum between the words “bin” and “ben.” And also as a control, we put them in an environment before an oral consonant, so between “bid” and “bed.” Then we recruited American English speakers and divided them into two groups. So the first group are individuals from the US South. And the other group is Americans from outside of the South. And the subjects listened to stimuli and said what word they heard. So if they heard a stimulus that was in a pre-oral context, they would say whether they heard “bid” or “bed.” And if they heard a stimulus that was from a pre-nasal context, they would say whether they heard “bin” or “ben.”
We also collected a speech sample from each participant and asked them some questions about their language and dialectic background, but we didn't use this in the first experiment.
We found that Southern listeners are more perceptually merged prenasally than pre-orally, and also that they are more perceptually merged than non-Southerners pre-nasally. So to some extent, this confirms that perceptual merger is real, but we found that on average Southern listeners still performed above chance at the continuum endpoints, meaning that as a group, they were not fully merged in perception.
Kat Setzer (05:57)
Okay, okay, so like, basically the idea is because Southerners don't speak these two vowels differently in this situation, they also learn to not hear it because it's not something they need to hear. Right?
Irene Smith (06:12)
Exactly. Yeah.
Kat Setzer (06:14)
Okay, okay, got what were the goals of this new study?
Irene Smith (06:16)
So the first study left us with two main lingering questions. The first question was whether this group-level finding that Southerners are partially merged in perception extends to individuals, or does it reflect the fact that maybe there are two subgroups of Southerners, one of which is fully merged in perception, and the other of which behaves like the non-Southern group, so is not merged in perception. And then the second question was whether there are additional factors that can help predict merged perception. So we have good reasons to believe that producing the PIN-PEN merger or familiarity with southern accents, or familiarity with African American vernacular English or AAVE, all might be expected to correlate with merged perception. So we gave each participant a southern accent score and AAVE score and also rated how merged they sound in their production and used these as predictors of merged perception as well.
Kat Setzer (07:20)
Okay, okay. So how southern they sound well to predict how southern they hear.
Irene Smith (07:26)
Exactly. Yeah. How, kind of holistically, how southern are they?
Kat Setzer (07:30)
Yeah. So what did you actually do to assess people's sensitivity to the vowel continua?
Irene Smith (07:37)
We wanted to quantify sensitivity to the continuum as basically how good are individuals at hearing the difference between two adjacent continuum steps. So if you were to plot the probability of responding /ɛ/ as a function of the continuum step, you would expect in general responses to look like a sigmoid curve. So individuals with high sensitivity will have a sharp transition from giving /ɪ/ responses to giving /ɛ/ responses somewhere in the middle of that continuum. And individuals with lower sensitivity will have a more gradual transition. And in in extreme cases, they could even have a completely flat response curve.
Kat Setzer (08:16)
So then you did some modeling to help you analyze the data from the participants. What did you do?
Irene Smith (08:22)
We were interested in the slope of the transition between /ɪ/ and /ɛ/ responses. And to do this, we fit a Bayesian logistic regression model to predict the probability of an /ɛ/ response as a function of continuum step, as well as a couple of other predictors. So the other predictors were following context, so nasal or oral, like we talked about, listener region, so southern or non-southern, merged production score, southern accent score, AAVE score, and all of the interactions.
And we also fit maximal by-speaker random effects in order to model individual level responses. We then used the fitted model in order to quantify the contribution of each of these predictors to the slope with respect to continuum step for each individual. So it's the slope with respect to continuum step that gives us a way to quantify sensitivity to the continuum.
Kat Setzer (09:12)
Okay, okay. So let's get back to your research questions. Are individual Southern listeners partially merged in perception?
Irene Smith (09:20)
Yes. So the vast majority of Southern listeners were partially merged in perception as opposed to fully or not merged. So they were less sensitive to the prenasal continuum than to the preoral continuum, but they still performed above chance at the continuum endpoints. If we contrast that to the non-southern group, the vast majority of non-southerners had no difference in their sensitivity to the prenasal and pre-oral continua.
Kat Setzer (09:47)
What besides listener region predicts merged perception?
Irene Smith (09:51)
So none of the other predictors were significant. There could be a lot of reasons for this, though. So there were about seven or eight AAVE speakers in each group. So for this predictor, we might have just lacked power. However, the direction of this effect was the opposite of what we would have expected. So maybe there's something more going on there.
For the southern accent score, there's a lot of collinearity with listener region, so there was just a lot of model uncertainty. And for merged production, we used an entirely perceptually coded measure of this, but we know that there are much more objective gradient ways to quantify merged production, so this might have just been kind of a noisy predictor.
Kat Setzer (10:36)
What was the most exciting, surprising, or interesting aspect of this research for you?
Irene Smith (10:42)
Simply being from the South is still the best predictor of merged perception, despite everything we might expect. So even though canonical southern accent features are disappearing, even though young people are producing the merger less and less, or young Southerners specifically, and even though we looked for other likely predictors of merged perception, we still found that being from the South is the best predictor of merged perception. And that's not necessarily expected.
Kat Setzer (11:12)
Yeah, that is very, very interesting. What are the next steps in the research?
Irene Smith (11:16)
So, like I alluded to before, there are better ways to quantify merger in production. So implementing that on the speech that we have already collected and using that in the model would be a very sensible next step. If we were to rerun an experiment and gather more data, I would do a few things differently. So for the experimental setup, I would also, want to ask individuals to rate how confident they are in their binary responses. And this would just open up new research questions in terms of what sensitivity to the continuum, as we have defined it, actually means. We'd also want to ask more specific questions about language background in order to elicit higher quality and more fine-grained responses, and this could also be incorporated into the modeling and hopefully yield a little bit better modeled results. So these additional steps would all help answer some broader theoretical questions like what is the nature of the relationship between speech perception and production? And also what does it mean to be perceptually merged?
Kat Setzer (12:23)
Interesting, interesting. Yeah. well, thank you again for sharing your research into the phenomena of the PIN-PEN merger. Ss I mentioned as someone from who grew up in the south, I'm now kind of curious if I have a merged perception of this vowel-nasal consonant combination. This is a very fun concept to hear about. Congratulations again on winning your award, and I wish you the best of luck in your future research.
Irene Smith (12:47)
Thank you. Thank you for having me.
Kat Setzer (12:51)
Our next POMA Student Paper Competition Winner from the Honolulu meeting is Rajesh Rameshbabu. We'll be discussing his article, “Learning distance-dependent spatial structure for multi-channel source separation.” Thank you so much taking the time to speak with me today, and congratulations on the award. How are you?
Rajesh Rameshbabu (13:07)
Thank you so much. Yeah, I'm good.
Kat Setzer (13:09)
So first tell us a bit about your research background.
Rajesh Rameshbabu (13:11)
So I am basically from India and I have completed my MS by research program in Indian Institute of Technology Mandi. That's how my research actually started. So I first started working on music source separation, and that experience really sparked my interest in how machine can understand complex acoustic scenes. So in 2024, I joined the listening technology lab at the University of Illinois, Chicago, where I work with Professor Ryan Corey. Since then, my research has expanded into multi-channel source separation, interference reduction in live recordings, distance-based source separation, and distributed acoustic sensor networks. So, if we look across all these projects, there is a common thread, right? So I am fascinated by how the physics of sound, like distance, room acoustics, and microphone geometry, can be combined with machine learning to build audio systems that are not only more robust but to better reflect how sound behaves in real world.
Kat Setzer (14:11)
Very cool. So this article is related to source separation. To give our listeners a bit of background, what is source separation? Why is it important? And what are the challenges that arise when trying to separate sound sources?
Rajesh Rameshbabu (14:25)
Source separation is all about extracting individual sound sources from a mixture of sound. Traditionally, that means separating different speaker in a conversation or different instrument sources in a piece of music. So what's fascinating is that the humans do this effortlessly every day. Imagine you are sitting in a busy restaurant. Even with lots of conversation happening around you, you could usually focus on the person sitting across from you. So that is known as the “cocktail party effect.” And one of the long-term goals of the source separation research is to give machines that same ability. So this technology has many applications, including music production, speech enhancement, robotics, virtual reality, video conferencing, and hearing aids. So personally, I find hearing aids to be one of the most exciting applications.
So traditional hearing aids mainly amplify everything, which can actually make noisy environments even harder to understand. But if we separate the speech or some sounds from the background sounds, the conversation becomes much clearer. Of course, the much easier than said, but microphone doesn't hear individual sources. They only record the combined mixture. Different sounds overlap in time and frequency. So the problems even becomes more challenging because of room reverberation, background noise, and room acoustics. So our work explores a slightly different perspective from the traditional approach. So instead of separating sounds only based on who is producing them, we ask whether we can also separate them based on how far away they are from the microphones. So that's a main goal of our work.
Kat Setzer (16:15)
Okay, okay. So can you explain how attenuation, spectral color, direct-to-reverberation ratio, and coherence can be used to estimate distance? How sensitive are each of these measures?
Rajesh Rameshbabu (16:17)
So actually distance change sounds in several ways. And each of those changes give us some clues about how far the source is from the microphones that we record. So the simplest clue is the loudness or what we call attenuation. Just like someone's voice becomes quieter as they walk far away from you, so the sound source reaching from the microphone also becomes weaker.
The distance also changes the spectral balance. Sometimes we call it spectral coloration. Higher frequencies are absorbed more easily by the air and by the interaction with the room other than the low frequencies. So as the source moves far away, the sound gradually loses some of its high-frequency content and becomes slightly more muffled.
And another important cue is the direct-to-reverberation, what we call is DRR. So, when someone is close to you, you mostly hear their direct voice. As they move far away, the reflection from the walls, floors, ceilings become relatively stronger. So, a high DRR usually suggests a nearby source, while a low DRR suggests the source is far away.
Finally, we have spatial coherence cue, which simply measures how similar the signals recorded by different microphones are. Nearby sources produce stronger direct sound while distant sources contains relatively more reflected sound arriving from many directions—a situation often called as “diffuse sound field.” So as the balance between the direct and the reflected sound changes, the similarity between the microphone signal changes as well. So these exact patterns also depend on how the microphones are arranged, since the different microphones array absorbs the sound field differently.
Of course, like all these spatial cues gives some kind of information about the distance, but none of these cues are perfect on its own. For example, like loudness depends on how loudly someone speaks, spectral balance depends on the sound itself, and DRR depends on the room, and coherence depends on the microphone array. But if we combine all these cues together, they complement one another and give us much more reliable estimate of the source distance.
That's exactly the idea behind the DistanceNet. Instead of relying on a single cue, it learns how all of these acoustic cues work together to understand the distance.
Kat Setzer (18:55)
Okay, okay. How is the distance between a sound source and a microphone array typically handled when one is trying to separate sound sources?
Rajesh Rameshbabu (19:03)
Most multi-channel source separation methods rely primarily on the directional information. For example, if they use tiny differences in when a sound reaches each microphone, known as “time differences of arrival,” or small differences in the phase of the sound waves between the microphones, called “interchannel phase differences.” So these differences help estimate where a sound is coming from actually.
Another common approach is the beamforming, which combines the microphone signal to emphasize sound arriving from one direction while suppressing sounds from others. You can think it as of electronically pointing the microphone toward one particular speaker. But here is the interesting part: those same recordings also contain information about how far away the sound is. As we discussed earlier, distance affects loudness, reverberation, spectral balance, and the relationship between the microphone signals. Yet in most existing systems, the information is either learned only implicitly or not explicitly used it at all. So instead the model mainly focuses on where the sound is coming from, rather than how far away it is.
So that's where our work takes a different perspective. Traditional source separation aims to recover different speakers or instruments using directional cues to distinguish them. We ask a different question. Can distance itself become a more useful way of organizing and separating sounds. So to explore this idea, we explicitly learn a representation of distance and use it to guide the separation process.
Kat Setzer (20:36)
Okay, okay, that makes sense. So like if you can differentiate like how far something is from the microphone, that can help you understand which things the person actually wants to be listening to, versus what they don't necessarily want to be listening to.
Rajesh Rameshbabu (20:49)
Yes, exactly.
Kat Setzer (20:51)
Researchers typically use travel times to estimate source distances. How does this differ from that?
Rajesh Rameshbabu (20:56)
Traditional approaches estimate distance geometrically by measuring tiny differences in the time it takes for sound to reach different microphones. As I said, all the spatial cues, like it can be derived using signal processing methods directly. So these methods work very well in many situations, especially for sound localization. But they can become less reliable when the microphones are close together, or the sources are more far away, or the room is highly reverberant.
So our approach takes a different direction. Instead of trying to calculate distance explicitly from the geometry, we let the neural network learn how distance changes the sound directly from the multi-channel recordings. So I like to think of this way. Traditional methods try to measure distance, whereas our networks try to understand the distance. It learns from multiple acoustic cues like loudness, reverberation, spectral balance, and coherence, all at the same time and combines them into a compact representation. And most importantly, our goal isn't to estimate the distance as accurately as possible. The goal is to learn a representation of a distance that helps the network to distinguish and separate the sources based on how far away they are from the microphones.
Kat Setzer (22:12)
Okay, okay. So what was the goal of this work?
Rajesh Rameshbabu (22:15)
Okay, when we started project, we asked ourselves a simple question: Can distance itself become a useful way of separating sounds? Traditionally, as I said, like source separation focuses on recovering different speakers or different musical instruments. We wondered if we could look at the problem from a different perspective. So instead of asking who is making the sound, we could also separate sounds based on how far away they are from the microphones.
To explore that idea, we developed a network that learns a compact representation of a distance directly from the multi-channel recordings and use that information to guide the separation process. At a broader level, our hypothesis was that the distance isn't just a byproduct of sound propagation. It's a meaningful information that a neural network can learn and use it for the distance-based source separation. And fortunately, our experiment shows that the idea really does work.
Kat Setzer (23:08)
Very exciting. So tell us about your distance embedding network. How did it work?
Rajesh Rameshbabu (23:14)
We developed a network called a DistanceNet. Its job is to learn a distance-aware representation directly from multi-channel recordings. So during training the network learns to estimate the distance between source and the microphones. But along the way, it also learns what we call an “embedding.” You can think of an embedding as a compact summary of what network has learned. So instead of just a single number, say like the source is three meter away or four meters away, it learns a vector, a set of numbers that captures how a distance has shaped the sound through things like loudness, reverberation, and other acoustic cues.
So here is the interesting part: like we don't actually use the final distance estimate during the source separation at all. Instead, we use the embedding because it contains more richer information than a single distance value. So rather than simply telling the separator how far away the source is, the embedding tells how the distance has affected the sound. So that's the information that ultimately guiding the separation process.
Kat Setzer (24:16)
Okay, okay. So it's almost like, rather than using the final result of the equation, you're kind of using the equation itself to train the model.
Rajesh Rameshbabu (24:25)
Yes, exactly.
Kat Setzer (24:26)
Okay. How did you use the distance embeddings produced by DistanceNet to modulate the multi-channel separator?
Rajesh Rameshbabu (24:34)
So like that once the DistanceNet learns the embedding, we use it to guide the source suppression network. So as I said, like instead of relying only on the microphone recordings, the separator also receives this embedding, like a compact summary of how distance has affected the sound. So that gives the network additional context about the acoustic scene itself. To do this, we used a technique called feature-wise linear modulation or commonly called “FiLM layers.” So, you don't need to think of FiLM as a changing the input to the network. So instead it acts more like a guide, helping the network to decide which acoustic pattern are most relevant for the current sound. For example, if a embedding suggests that a source is far away, the network pays more attention to the characteristics that are typical of distant sounds, like stronger reverberation and weaker direct sound. If the source is very closer, it focuses more on patterns associated with nearby sounds. In that way, the separator adapts its processing to the source distance rather than treating every sound in exactly the same way.
Kat Setzer (25:38)
So how did you test your distance aware source separation framework?
Rajesh Rameshbabu (25:42)
So, to evaluate this framework, we generated a simulated multi-channel recording in reverberant rooms where we could precisely control the source locations and microphone array and the room acoustics. So everything is like being controlled. That allowed us to create many different near-far source scenarios while always knowing the correct answer or what we call the ground truth information.
We then compared to three systems. One was the conventional multi-channel source separator. The second one was the distance-aware separated, but the distance embedding was like obtained through various spatial cues as I explained. Then the third was like our neural network learned distance-aware separator. To make this comparison fair, like all the systems use the same network architecture and were trained on the same data and received the same microphone recordings. The only difference was that the proposed system also received the learned distance embedding. So that was important because it meant we could isolate the contribution of distance representation itself. If we see an improvement, we could confidently attribute to the learned distance information rather than the differences in the model or the training procedure. So finally, we evaluated all the systems using the standard source separation metrics called scale-invariant source-to-distortion ratio and consistently found that the learned distance representation improved separation performance.
Kat Setzer (27:05)
Okay, okay. So how did your framework end up performing?
Rajesh Rameshbabu (27:08)
So the results were very encouraging actually. So across our experiments, the distance-aware approach consistently outperformed both the conventional separator and a baseline that used the handcrafted spatial features. So in our near-far separation experiments, we observed the improvements of around two to three dB in SISDR, showing that the learned distance representation provides useful information beyond the traditional acoustic feature itself.
But for me, the most important result wasn't the numerical improvement. It was the experiments confirmed our original idea that the distance isn't the byproduct of sound propagation. It's meaningful information that a neural network can learn and use the distance to organize and separate a scene. So I think that's the most exciting takeaway from this work. It suggests that, like if we could teach the neural network to understand how physics shaped the world rather than just learning it from the raw data, we can build like more capable and more interpretable audio systems.
Kat Setzer (28:07)
Right, right, right. That totally makes sense. So what are the next steps for this research?
Rajesh Rameshbabu (28:12)
So like this work answered an important first question we asked like does learning a representation of distance actually help with the source operation? So to answer that cleanly as possible, we trained the model using isolated source recordings. That allowed us to focus purely on the value of distance representation without introducing any other sources of uncertainty.
But like practically having isolated sound sources is forbidden, right? Like so typically for practical system we should operate within the mixture recordings. So as the next step, to make it more practical, after completing the work, we have developed an end-to-end framework that learns distance representation directly from the mixed recording itself without recording of isolated source signals. That brings the method much more closer to how it would operate in the real-world application.
So looking ahead, we would like to evaluate this framework in more realistic acoustic environments, extend it beyond two-source, near-far scenarios, and explore whether other physical properties of sound, such as not just distance, like it can also be learned and used to improve the source separation.
Kat Setzer (29:20)
Okay, okay. So you kind of touched on this already, but what was the most exciting, surprising, or interesting aspect of this research for you?
Rajesh Rameshbabu (29:28)
Yeah, so for me the most exciting realization was the distance isn't something we measure. It can actually become a useful way of understanding and organizing an acoustic scene. So traditionally, like as I said, the source separation focus on separating who is making the sound, whether the different speaker or different musical instrument, we asked a different question, like can we use how far away a sound is to help to separate it?
So, what surprised us most was how much information distance actually carries. So, we often think if it has a simply affecting loudness, but it also changes the reverberation, spectral balance, and the relationship between the microphones. So, when a neural network learns all of these effects together, distance becomes much more than a number. So it becomes a richer description of an acoustic scene. So for me, the most exciting part is that like distance isn't just one of the physical properties of the sound. I hope this work encourages people to think about other physical principles we can teach our models to understand. So ultimately, I think the future of audio AI isn't about building a bigger neural network models. It's also about like building models that can understand like why sound behaves the way it does.
Kat Setzer (30:41)
Right, right. So it's kind of like having models that can essentially think about the world in a different way, whereas previously we had to keep it fairly simple. Now since they can understand more complex information about the audio scene, they can yeah, right.
Rajesh Rameshbabu (30:55)
Yeah, exactly like instead of like treating it as a black box or like some traditional approaches, like now neural network and learn like how the spatial properties of sound actually behaves itself, yeah.
Kat Setzer (31:08)
Yeah, yeah. Well, not to use the word exciting again, but it is exciting to hear about this new option for source separation and hopefully it will become useful for some of the applications you mentioned, like hearing aids. Congratulations again on winning the award and I wish you the best of luck in your future research.
Rajesh Rameshbabu (31:25)
Yeah, thank you so much for this opportunity and like yeah, thanks for the award.
Kat Setzer (31:32)
Yeah, you're welcome.
Our next POMA student paper winner takes us to the world of musical acoustics, specifically in relation to the oboe. With me is Fumihiko Kurosawa, who wrote the article “Control of characteristics through topological pptimization of 3D-printed oboe reeds. Thank you for taking the time to speak with me today, and congratulations on the award. How are you?
Fumihiko Kurosawa (31:54)
Thank you so much. I feel so happy today.
Kat Setzer (31:55)
We're happy to have you! So first just tell us a bit about your research background.
Fumihiko Kurosawa (32:00)
I'm a first-year PhD student at the University of Tsukuba in Japan. My research focuses on making resin oboe reeds, especially using new technology like 3D printing. When I was a middle/high-school student, I picked up the oboe for the first time as part of my school’s co-op activities. The oboe is sometimes said to be the most difficult instrument in the world, and I certainly struggled quite a bit at first because I couldn't play it well. That memory is one of the driving forces behind my current research.
Kat Setzer (32:46)
That's really interesting. I don't think I ever knew that the oboe was considered one of the most difficult instruments to play in the world, but I guess that makes sense. So tell us about oboe reeds. How do they work, and how are they normally produced?
Fumihiko Kurosawa (32:53)
The oboe is a woodwind instrument, and it is classified as a double-reed instrument. This means the reed has two vibrating blades made from natural cane facing each other. When the player blows breath into it, then the reed vibrates and it excites the air column inside the oboe to make sound. Normally making these reeds requires very advanced technique and experience. The players have to adjust them to match their performance, and this takes much time and effort.
Kat Setzer (33:35)
Right, okay. I can imagine that might be very tedious and require a lot of skill. So what challenges arise with cane and resin reeds?
Fumihiko Kurosawa (33:42)
For cane there is a problem that supply and price become unstable because production areas are limited. Also, there are quality variations because it is a natural material. And recently it is reported that cane can be an allergen for some players.
Resin reed have advantage of stable quality and longer life, but their sound quality and playing feel are generally considered inferior to cane. The biggest problem is physical properties. Cane is a hard and lightweight material, but polymer is a soft and heavy material. Therefore, we simply copy the shape of cane reed with resin, we cannot get good acoustic characteristics.
Kat Setzer (34:39)
Right, right, that makes a lot of sense. So what were the goals of this study?
Fumihiko Kurosawa (34:44)
Our objective was to maximize bending stiffness while reducing the weight of the resin reed. We aim to shift the reed design from a traditional empirical approach to our systematic computer-aided optimization framework. Specifically, as I mentioned earlier, polymer is a heavy and flexible material. For this reason, this is necessary to remove unnecessary solid material while maintaining rigidity in order to reduce weight. This is also related to natural frequency of the reed.
Kat Setzer (35:28)
Right, right, that makes sense. So like the resin reed can't look or be built exactly the same way as a cane reed to create the same sound.
Fumihiko Kurosawa (35:36)
Yes, exactly.
Kat Setzer (04:02)
So what is the relationship between the natural frequency of the reed and the sound produced? How did you model and manufacture reeds for the study?
Fumihiko Kurosawa (35:45)
That is an important point. In the self-sustained oscillation of woodwind instruments like the oboe, the natural frequency of the reed must be higher than the playing frequency because of the phase condition. For manufacturing, we created a parameterized model using 3D cache software based on measurements of an actual cane reed. After that we use an SLA or stereolithographic 3D printer, and this printer irradiates UV light to cure the liquid resin layer by layer.
Kat Setzer (36:26)
Okay, okay. So you propose using topology optimization in reed design. How does that work?
Fumihiko Kurosawa (36:33)
Topology optimization is a computer-based design method of optimize the distribution of material. In this case, we applied this only to the internal support structure of the reed. The computer calculates where to remove material to reduce weight. And at the same time it will use material in the place needed to maximize stiffness. In our model, we reduced the internal mass to 36% of the original value.
Kat Setzer (37:10)
Oh, okay, interesting. So how did you evaluate the effectiveness of topology optimization in reed design?
Fumihiko Kurosawa (37:17)
We evaluated in two ways. First we used numerical simulation, or finite element analysis, to check the natural frequencies. Next we did experiment. We compared a 3D-printed reed and natural cane reeds through a human blowing test in our soundproof room. We compared acoustic spectra for the G4and G5 notes.
Kat Setzer (37:48)
Okay, okay. So how did your methodology end up performing?
Fumihiko Kurosawa (37:50)
In the simulation, the support structure successfully shifted the natural frequency upward; the first mode frequency increased by about 17% and the second mode increased by about 10%. In the human blowing test in the experiment, the high-frequency harmonic components of the 3D-printed reed were slightly emphasized, but the overall spectral shape was not largely deviated from the natural cane reed.
Kat Setzer (38:25)
That's really neat. So you ended up with a sound that was pretty similar to the cane reed’s sound. So what was the most exciting, surprising, or interesting aspect of this research for you?
Fumihiko Kurosawa (38:34)
Actually, I was most excited when I played the musical instrument sound with a reed I designed myself. Manufacturing a team super like an oboe reed with a 3D printer is very difficult but we fine-tuned the material and the manufacturing parameters many times to make it possible to manufacture the model design on the 3D printer exactly. And when looking at the envelope of the playing sound, it was surprising that only the reed we designed this time showed distinctive features compared to our previous design reed. I think this is a cue suggesting that our method this time affects that characteristics of the oboe’s playing sound.
Kat Setzer (39:32)
Yeah, yeah. So what are the next steps in the research?
Fumihiko Kurosawa (39:35)
Yes, there are several next step. We need to create a model that considers the double-reed contact and the reed constraints of the players to further define the design grid lines. Also we need to systematically evaluate how different printing conditions and material properties affects the result. Finally, we want to evaluate the reproducibility with multiple players and we want to quantify subjective evaluations like timbre and playing field.
Kat Setzer (40:16)
Okay, okay. Yeah, all those steps make sense. Well, it's so cool that you found a way to adapt the design of resin reeds to better replicate the sound and performance of cane reeds. This technique sounds like it'd be quite helpful for many musicians, and save folks some time and money. Thank you again for taking the time to speak with me today, and I wish you the best of luck in your future research.
Fumihiko Kurosawa (40:35)
Thank you so much.
Kat Setzer (40:37)
Our last POMA student paper competition interview is with Kai Burian, who wrote the paper, “Modeling Loudness Contours to Determine Human Auditory Nerve Fiber Distribution and Cochlear Nerve Degeneration.” Thanks for taking the time to speak with me today, Kai, and congrats on the award. How are you?
Kai Burian (40:53)
Hi, thank you for having me. I'm great. I actually just had my graduation ceremony a couple of days ago, so now I'm done with this chapter in my life and looking forward to the next.
Kat Setzer (41:03)
That's super exciting. Congratulations. So first tell us a bit about your research background.
Kai Burian (41:09)
Yeah, so I did my bachelor's in electrical engineering, but I was focusing on automation technology. And at some point I realized that I wanted to have a more positive and direct impact on people, and I found a specialization within electrical engineering called Neuro and Bioengineering at the Technical University of Munich, where I thought I could apply my engineering mind to the human body And that's where I went and I ended up focusing on the human auditory system in Werner Hemmert's lab at the Technical University of Munich. And then I had the amazing opportunity to further explore this direction in Sunil Puria's otobiomechanics lab at the Eaton Peabody Laboratories at Massachusetts Eye and Ear Infirmary as a visiting graduate student at Harvard Med School.
Kat Setzer (42:00)
Very cool. That's a nice little, you know, path to where you ended up. What is cochlear synoptopathy and how is it different from other causes of hearing loss? Is it just a difference in perceiving loudness levels?
Kai Burian (42:12)
So cochlear synaptopathy is, in Haggerty's words, the loss of synapses between the inner hair cell and the auditory nerve, despite survival of the sensory hair cells. It is thought to be one of the potential mechanisms behind “hidden hearing loss”.
If I put it more simply, the inner hair cells convert sound-driven movement within the cochlea into an electrical response, and then they release neurotransmitters onto the auditory nerve fibers. This then leads to the elicitation of action potentials within those fibers that then carry the information further towards the brain. In cochlear synaptopathy, the hair cells can still be functional, but some of the synaptic connections that allow them to communicate with with the auditory nerve have been lost.
This is different from other forms of sensory neural hearing loss. Traditionally, damage to the sensory hair cells has received the most attention. In particular, the outer hair cells contribute to the active amplification of sound within the cochlea. And when they are damaged, the cochlea becomes less sensitive, meaning that sound has to be louder before can be detected. This results in elevated hearing thresholds, which are visible on a standard audiogram, which you can do at a doctor's office. Yeah, and experimental studies have shown a clear relationship between outer hair cell loss and permanent threshold elevation.
Now, cochlear synaptopathy, however, can remain hidden within this audiogram because the audiogram primarily measures the quietest pure tone that a person can detect at different frequencies, and even when many synapses have been lost, the remaining auditory nerve fibers may still be sufficient for detecting these very quiet tones. That's why the hearing threshold can remain normal, or close to normal, even though fewer nerve fibers are available to represent sounds at levels above the threshold.
So, no, cochlear synoptopathy is not simply a difference in how loudly a person perceives sound. It is fundamentally a loss of neural connections and therefore a reduction in the information transmitted from cochlea towards the brain. It has been proposed that this reduced neural input could contribute to difficulties with more demanding listening tasks, such as understanding speech in environments with a lot of background noise.
Kat Setzer (44:55)
Okay, okay, like the cocktail party problem, right? And it's basically you don't have as many clues being sent to the brain to help you deduce what's actually being said.
Kai Burian (45:03)
Yes, that's exactly one of the hypotheses.
Kat Setzer (45:05)
Okay. So what are the current methods for assessing cochlear synaptopathy, and what are their limitations in terms of diagnosing the condition?
Kai Burian (45:13)
So cochlear synaptopathy is very difficult to assess in humans. In animal studies, researchers can examine the cochlea directly and count the synapses between the inner hair cells and auditory nerve fibers. And, well, I say that so easily, but this is a very difficult process. But still, it is possible in animals, and this allows them to confirm synaptic loss, but it, as I said, requires examination of the cochlear tissue and cannot be done in vivo. In humans, that's why we have to rely on indirect measurements.
One commonly studied method is the auditory brainstem response or ABR. Sounds are played through headphones or earphones, whilst electrodes record the resulting electrical activity. Yeah, researchers are usually interested in in wave one, which mainly reflects activity from the auditory nerve. And a reduced wave one amplitude may indicate reduced auditory nerve output and could therefore be consistent with cochlear synaptopathy. other methods include the envelope following response, electrocochleography, the middle ear muscle reflex, and/or speech-in-noise tests.
The main limitation is that none of these measures is specific to cochlear synaptopathy. For example, ABR wave one can also be affected by age, sex, anatomy, electrode placement, and recording conditions, while similarly poor speech-in-noise performance can be influenced by attention, memory, language ability, for example, if you're speaking in a different language that is not your mother tongue, and yeah, central auditory processing. So although we have several promising research tools, there's currently no widely accepted clinical test that can reliably diagnose cochlear synaptopathy in an individual person.
Kat Setzer (47:17)
Which is where you come in! So let's start with the big picture. What were you trying to find out in this study?
Kai Burian (47:22)
The main goal of our study was to investigate how cochlear synaptopathy might affect loudness perception. We used a computational approach because, as I've mentioned before, synaptic loss cannot currently be measured directly or selectively manipulated in living humans. So a model would allow us to systematically remove different auditory nerve fiber populations and predict how this would change the neural activity and ultimately how it would change perceived loudness. For simulating that fiber loss, however, we first needed to establish a relationship between the model's auditory nerve activity and human loudness perception.
Kat Setzer (48:09)
So to establish this relationship between loudness and auditory nerve activity, you used an auditory nerve model to create something called “equal spike contours.” What are those and how do they relate to equal loudness contours?
Kai Burian (48:21)
Yes, you mentioned two very important parts. So let me start with the equal loudness contours. The equal loudness contours show the sound pressure levels required at different frequencies for tones to be perceived as equally loud. This is necessary because our hearing is not equally sensitive at every frequency, meaning that two tones presented at the same sound pressure level may therefore not sound equally as loud. For our study, we wanted to create a sort of neural equivalent of these contours. We worked with the neural count hypothesis, which assumes that perceived loudness is directly related to the total number of spikes across the auditory nerve.
At this point, I think it is important to say that this is a simplified hypothesis and it is disputed as a complete explanation of loudness perception. However, it gave us a clear starting point for linking the output of the auditory nerve model to this perceptual loudness data.
To create an equal spike contour, we used the model to find the sound level needed at each frequency to produce the same overall amount of auditory nerve activity. Connecting those levels gave us the neural equivalent of the equal loudness contours. So an equal loudness contour represents equal perceived loudness, while an equal spike contour represents equal model predicted auditory nerve activity. Under the simplified neural count hypothesis, we would expect the two contours to be reasonably similar.
Kat Setzer (50:02)
Okay. So when you first compared the model with normal human loudness perception, how well did it work and how did you improve it?
Kai Burian (50:10)
So initially the model did not reproduce normal human loudness perception particularly well. When we compared the equal spike contours with the equal loudness contours, the average difference was around 10 dBSPL, with especially large differences at the very low and very high frequencies. One possible reason was that the assumed distribution of auditory nerve fiber subtypes. Auditory nerve fibers are commonly divided into high, medium, and low spontaneous rate fibers, which respond differently to sound. However, their exact proportions across the human cochlea are not known. And the original model used the same proportions at every location within the cochlea, which this distribution was kind of derived from animal models, but it may be an oversimplification. We therefore allowed the proportions of these three fiber types to vary across the cochlea and optimized them using a numerical optimization algorithm called particle swarm optimization to make the equal spike contours best match the equal loudness contours. This reduced the average difference from around 10 dBSPL to 2.5 dBSPL.
Kat Setzer (51:34)
Okay, okay. So the optimization produced a strongly varying distribution of auditory nerve fiber subtypes. How do you interpret that result?
Kai Burian (51:43)
Yeah, that's a good question. So the optimized distribution should be interpreted as a mathematical best fit, not as a true biological distribution of auditory nerve fiber subtypes. some of the changes between neighboring cochlear regions were quite abrupt, whereas in biology, distributions would usually be expected to vary more gradually. This kind of suggests that the optimization may partly be compensating for simplifications or missing mechanisms within the model.
Kat Setzer (52:17)
Okay, okay. So once you optimized this normal hearing model, how did you use it to simulate cochlear synaptopathy and hair-cell-related hearing loss?
Kai Burian (52:26)
Once we had optimized the normal hearing model, we used it as our baseline. And to simulate cochlear synaptopathy, we then progressively reduced different auditory nerve fiber populations, mainly in the high-frequency region, where synaptic loss is often expected to be the strongest. We then tested three patterns of increasing severity, removing different combinations of low, medium, and high spontaneous rate fibers to see how both the extent and the type of fiber loss affected the predicted loudness.
In a second experiment, we then added age-related pattern of hair cell damage based on a typical age-related audiogram. And as is commonly done in this type of modeling, two-thirds of the resulting hearing threshold shift was attributed to outer hair cell dysfunction and one-third to inner hair cell impairment. We then simulated the hair cell damage both by itself and in combination with the three synaptopathy profiles, allowing us to compare the separate effects of hair cell damage and nerve fiber loss as well as how they interacted.
Kat Setzer (53:40)
Okay. So how did the simulation of synaptopathy perform? And how did the loss of hair cells affect the result?
Kai Burian (53:47)
The synaptopathy simulations produced a clear severity-dependent effect. As more auditory nerve fibers were removed, the same sound pressure level produced less predicted loudness, particularly at the higher frequencies where we had applied the fiber loss.
In the most severe condition, which was 50% loss of all three fiber types, the difference reached up to around 16 dB SPL. In our simulations the hair cell damage alone mainly shifted the loudness growth curves. This means that roughly the same additional sound pressure level was required across all tested loudness levels, while the shape of the curve remained largely unchanged. When synaptopathy was added on top of that, it also changed the slope of the curves, resulting in an abnormal loudness growth like curves, meaning that the perceived loudness changed at an unusual rate as sound level increased.
Kat Setzer (54:58)
Okay. So, like, if a person were listening to something, even though it might gradually increase in loudness, they might hear it be like super, super quiet, and then it's all the way really loud. Like for instance, that kind of classic comedy routine of somebody keeps on being like, “What are you saying? I can't hear you,” and the person says something a little louder, and then they say it a little louder, and then they say it a little louder. The listener keeps on being like, “I still can't hear you! Speak up, speak up!” and then the speaker screams the answer or screams whatever they were saying and the listener is like, “Why are you talking so loud? You don't have to yell!” Is that the kind of an idea?
Kai Burian (55:35)
Yes that is a very intuitive example, and synaptopathy is thought to be a potential factor in this phenomenon.
Kat Setzer (55:41)
Yeah, okay. So what was the most exciting, surprising, or interesting aspect of this research for you?
Kai Burian (55:46)
That's a good question. For me, I think the most exciting part was experiencing the American research culture. Everyone was incredibly open and collaborative, and there was this real willingness to exchange ideas and help each other, which I've never experienced in at this level, and it was just amazing to see. Also, it was quite surreal to meet and kind of work alongside researchers whose papers I've been reading for years now. And seeing the people behind the research and being able to discuss ideas with them directly was one of the most memorable parts of my experience. And well, of course attending my first conference in Hawaii was not exactly a hardship either.
Kat Setzer (56:37)
Yeah, you can't really complain about that part. It is kind of funny. It's like kind of like meeting celebrities when you see these researchers in real life, and you're like, “Oh! Like I think about your work all the time, but now I get to meet you in person.”
Kai Burian (56:51)
Exactly.
Kat Setzer (56:53)
What are the next steps in the research?
Kai Burian (56:56)
I think one of the most natural next steps would be to test whether our results can be reproduced using a different auditory nerve model. That would help us determine whether these findings are robust or depend strongly on the assumptions of the particular model that we used.
Another thing we would also need to address are some of the limitations introduced by our approach. For example, the simplified neural count hypothesis and the biologically unlikely fiber distribution produced by the optimization. And then much, much further down the line, once the modeling framework is more robust, one of the next steps would be the experimental validation in human listeners.
Kat Setzer (57:38)
Okay, okay. Well, it's really interesting learning about this other mechanism for hearing loss, and it's so cool that your work is helping us understand you know, so-called “hidden hearing loss.” I wish you the best of luck in your future research and congratulations again on the award and graduating.
Kai Burian (57:54)
Thank you very much. I would also like to use this opportunity to once again thank Sunil Puria and Ahsan J. Cheema for all their guidance and support without which this research would never have been possible. Thank you.
Kat Setzer (58:08)
So before we wrap up, I also want to give a quick shout out to our fifth POMA Student Paper Winner, Mizuki Takahashi, who also won an award for the article, “Urgency Impression and Driving Behavior Evaluation for Warning Sound Inside Car Cabin,” but was unable to do an interview. And student listeners, if you're presenting at the December meeting in Baltimore, please make sure to enter the competition. Not only can you boost your CV with an editor reviewed paper, but you can also win three hundred dollars and appear on this podcast.