Back to all posts

Blog post

Listening Is an Interaction

SoundWellbeingInteraction Design

From 1–3 September, Asian Sound Cultures 3 brought researchers, artists, and practitioners to Tohoku University under the theme Art Meets Science, Performance Meets Technology. I joined a roundtable on sound and wellbeing in Japan.

The conversation ranged from nature sounds and spatial audio to public announcements, AI-generated soundscapes, and the politics of deciding what a “good” environment should sound like. One idea connected almost every topic:

Listening is a closed-loop interaction between attention, body, and environment.

That may sound abstract, but it changes some very practical design decisions.

A sound is not a soundscape

We can measure a sound pressure level, frequency spectrum, or reverberation time. Those measurements matter, especially for safety. But they do not tell us the whole experience.

Imagine the same loud bass in two places: a festival where you chose to stand near the stage, and a hospital room where you cannot leave. The signal may be similar. Its meaning, and its effect on the listener, is not.

This distinction is built into the ISO framework for soundscape research, which treats a soundscape as an acoustic environment as it is perceived or experienced in context. In other words, a soundscape is not simply what reaches the microphone. It also involves who is listening, what they are doing, where they are, and what control they have.

Listening is something we do

At a crowded reception, many voices reach our ears at once, yet we can often follow one conversation. Researchers call this the cocktail party problem. The auditory system is continually grouping, selecting, and predicting—not passively receiving a finished scene.

The body participates too. Small head movements change the timing and level of sound at each ear, giving us additional clues about where a source is located. Experiments on active listening in 3D sound localisation show that allowing listeners to move can improve localisation compared with a fixed posture.

For interaction design, this suggests three useful listening modes:

  1. Background listening — sound is present but not the focus.
  2. Active listening — attention is deliberately directed toward it.
  3. Interactive listening — a person acts, hears the result, and adjusts again.

The third mode is especially important to my work. In an interactive music system, listening is part of the control loop. The sound changes what the performer does next, just as their action changes the sound.

Measure the relationship, not only the signal

If listening is an interaction, evaluating a sonic experience requires more than one number. I find it useful to think across four layers:

The last layer is easy to miss. A system might measure heart rate or movement and adapt its output, but those signals are ambiguous. A rising heart rate could mean stress, excitement, exercise, or simply standing up. Physiology is an input—not emotional truth.

An adaptive sound system should therefore show uncertainty, use as little personal data as possible, and keep the listener in the decision-making loop. The goal is not to build a machine that claims to know how someone feels. It is to give that person another way to notice and regulate their own experience.

Wellbeing may begin with subtraction

It is tempting to approach sonic wellbeing by adding something pleasant: birdsong in a station, ambient music in an office, or a “healing” frequency in an app. Sometimes that helps. A synthesis of research on natural sounds and health outcomes found benefits across measures including stress, annoyance, and positive affect, although the effects depend on the sound and the study context.

But a pleasant layer does not repair an unhealthy environment underneath it. The World Health Organization’s environmental noise guidance treats unwanted noise as a public-health issue, not merely an aesthetic inconvenience.

So the first question should often be: what can we remove?

Repeated announcements, mechanical noise, competing media, and unnecessary alerts all make claims on limited attention. Removing them can create room for conversation, concentration, recovery, and the sounds that already give a place its character. This is not a call for sterile silence. Silence itself has context: it can feel restorative, uncanny, isolating, or ecologically alarming.

Do not add beautiful sound to hide an unhealthy environment.

The same caution applies to spatial audio. Immersion can intensify comfort, but it can also intensify confusion, threat, or sensory overload. Realism is a technical parameter; wellbeing is an outcome. Spatial audio creates possibilities—it does not guarantee them.

Personalisation should increase agency

The line between a therapeutic intervention and manipulation is not determined by the technology. It depends on purpose, evidence, transparency, and control.

Can the listener understand why the sound is changing? Can they adjust it, refuse it, or leave? Who owns the data used to adapt it? These questions matter precisely because sound can work in the background, shaping attention and behaviour without demanding conscious focus.

In headphones, personalisation can be individual. In a school, hospital, workplace, or public space, it becomes a question of governance. Acoustic and health experts can establish safety limits, but designers, staff, residents, and people with different sensory needs should all be able to shape the result. A shared soundscape should be negotiated with the people affected by it.

Japan as a sonic laboratory of contradictions

Japan is often described from outside as unusually attentive to sound. That image contains real practices, but it can easily become a flattering stereotype.

Everyday life in Japanese cities includes an unusually dense layer of designed audio: train melodies, pedestrian signals, appliance feedback, shop music, safety announcements, and spoken instructions. Each sound may be useful on its own. Together, they can become cognitively demanding. Local optimisation does not necessarily produce a healthy overall soundscape.

At the same time, seasonal festivals, temple bells, jazz kissaten, forests, coastlines, and devices such as the suikinkutsu offer very different relationships with listening. These are not evidence of one national sensitivity. They are a plurality of practices—sometimes complementary, sometimes contested.

That is what makes Japan interesting to study. It makes tensions visible: coordination and personal preference, efficiency and contemplation, technology and tradition. Japan is not the answer; it is a sonic laboratory of contradictions.

What I took home

Three working principles stayed with me after the discussion:

  1. Listen before adding. Study the existing environment and its cumulative demands.
  2. Keep people in the loop. Personalisation should make control more visible, not hide it.
  3. Evaluate relationships. Measure context, behaviour, and differences between people—not only averages or acoustic signals.

This also suggests a role for a Sonic Lab beyond a conventional laboratory. Controlled experiments remain essential, but listening can also connect research to streets, forests, hospitals, cultural spaces, and communities. A lab can become a living network where sound is not simply delivered to people, but investigated and shaped with them.

That feels like the most useful shift: from sound as instruction to sound as interaction.

Further reading