Category: Audio Perception

  • How Do You Make ADR Sound Like It Was Never Replaced? Paul Carden and Chris Navarro on Performance, Technology, and the Art of Dialogue Replacement

    Paul Carden and Chris Navarro

    How do you make ADR sound like it was never replaced?

    A line of dialogue may last only a few seconds, yet replacing it convincingly can require an extraordinary combination of preparation, performance, technology and judgement. The original production recording must first be identified as unusable, the replacement carefully documented and prepared, the actor returned to the emotional and physical circumstances of a performance recorded months earlier, and the new dialogue captured so that it matches the timing, vocal quality, microphone perspective and acoustic character of the original scene. If the process succeeds, the audience should never know that any of this work happened. During their joint online guest lecture for Edinburgh Napier University, ADR supervisor Paul Carden and ADR mixer Chris Navarro demonstrated this complete process from beginning to end. Rather than discussing Automated Dialogue Replacement only in theory, they created a deliberately compromised line of production dialogue, prepared it for replacement, recorded it on an ADR stage and evaluated the result. Their demonstration revealed a process in which meticulous preparation and technical fluency serve a deceptively simple objective: allowing everyone involved to concentrate upon the performance.

    Carden began not in a recording studio, but outside beside a car. His objective was to demonstrate how an apparently simple piece of dialogue could become unusable during production. Wearing a lavalier microphone, he performed the five-word line “I’m late for work” while getting into the vehicle. The exercise immediately revealed how vulnerable production dialogue can be. Clothing could obscure the microphone, a zipped jacket could change the sound, movement could complicate recording and the closing car door could mask the line itself. The example was deliberately simple. Carden was standing almost still, knew exactly what he was going to say and had constructed the situation specifically for the demonstration. On a real production, actors may be walking, running, interacting with objects and performing emotionally demanding scenes while the production sound team manages multiple radio microphones and one or more booms. From this perspective, the surprise is often not that some lines require replacement, but that so much production dialogue survives at all.

    The demonstration also established one of the central tensions within ADR. Production dialogue is not simply speech recorded on location. It contains a performance created at a particular moment, within a particular physical environment, as part of an interaction with other performers. Months later, an actor may arrive on an ADR stage while already immersed in an entirely different project and be asked to recreate a few seconds from that earlier performance. The recording environment offers little of the original context. The actor stands in a comparatively neutral room, watches an image on a screen, listens for cues and attempts to reproduce not merely the words, but the emotional and physical conditions under which those words were originally spoken. Carden emphasised that this is one reason actors can find ADR difficult. Matching a line involves returning to a performance that may no longer feel immediate or familiar.

    Before any actor reaches the stage, however, the line must become an ADR cue. Carden demonstrated this preparation process by locating the damaged dialogue within the picture, defining the cue and documenting the reason for replacement. The cue was assigned a unique identifier, linked to its scene information and accompanied by notes explaining that the car door had closed over the line. He also demonstrated the value of professional flexibility. Where a production line was imperfect but potentially usable, he might identify the replacement as optional rather than forcing an unnecessary argument over whether it must be replaced. The cue sheet then carried the information required by the recording stage, including project details, version information, character and actor names, cue number, timecode and recording information. A five-word performance therefore arrived at the stage supported by an extensive information system designed to ensure that everybody was working on the correct material.

    Version control was particularly important. Carden described production teams as living and dying by version dates, reflecting the practical reality that picture changes can quickly make carefully prepared cues inaccurate. Cue numbering performs a similarly essential role. Each take is voice-slated so that the recording itself retains its identity even if paperwork becomes separated from the audio. These details may appear administrative when viewed from outside professional production, yet they protect the continuity of the entire process. An ADR session may contain hundreds of replacement lines recorded across multiple days, actors and facilities. The actor should not have to think about whether the cue is correctly identified or whether the stage is working from the right picture version. Preparation creates the stability within which performance can happen.

    Carden also offered one deceptively simple instruction for aspiring ADR mixers: keep recording until somebody clearly indicates that the take has ended. An actor may continue the performance, a director may decide to record another version immediately or the session may move into wild recording without a formal interruption. Stopping too early can lose useful material for no meaningful benefit. Storage is cheap; an unrecoverable performance is not. The advice reflected a wider principle that would continue through the lecture. Good ADR practice depends upon remaining attentive to what is happening in the room rather than allowing the machinery of recording to dictate the session.

    The second stage of the demonstration moved into Navarro’s ADR facility. Even though the project consisted of a single line created for the lecture, he established the session as though it were a conventional production. Folder structures, documents, Pro Tools sessions and media were organised according to the same principles he would use for a project involving one session or hundreds. His ADR template was already configured for the ordinary demands of recording, while remaining adaptable when unusual situations arose, such as multiple actors performing together with several microphones each. A good template does not prescribe every session in advance. It removes predictable technical work so that attention remains available for situations that cannot be predicted.

    Even the imported picture introduced a practical lesson. Navarro noticed that the system was responding sluggishly and identified the compressed H.264 picture as a likely cause, since the computer had to decode it continuously during playback. The moment was minor, but revealing. Professional technical fluency often appears not through dramatic troubleshooting, but through the ability to recognise quickly why a system is behaving unexpectedly and continue without allowing the problem to dominate the room. Earlier, Carden had offered similarly pragmatic advice about computer failure: save the work, restart the system and continue rather than allowing panic to consume valuable time. Both speakers treated technical expertise as calm familiarity rather than technological display.

    Microphone selection then revealed how ADR matching differs from conventional voice recording. The objective is not to capture the most beautiful possible version of the actor’s voice. It is to create a recording that can inhabit the existing production soundtrack without attracting attention. Navarro showed that the stage had previously been configured for voice-over work, with microphones positioned closely and directly to produce the clear, present sound required for narration. ADR demanded a different approach. Carden had recorded the original line using a lavalier, so the replacement needed to reproduce the qualities of that production perspective rather than simply offering a technically superior recording.

    Carden explained that ADR sessions commonly record both boom and lavalier microphones, ideally using the same models employed during production. Even when the original line appears to come predominantly from one microphone, alternatives can prove unexpectedly useful. Navarro demonstrated this by positioning a shotgun microphone close to Carden but substantially off-axis. Pointing it directly at the performer would have produced a cleaner and more extended sound, but that was not necessarily desirable. The off-axis position rejected particular frequencies and created a tonal character that could potentially sit closer to the lavalier recording. The result was counterintuitive: the boom microphone could, under those conditions, sound more like the production lavalier than the replacement lavalier itself. Matching therefore began not with assumptions about microphone categories, but with listening.

    This distinction between technical quality and contextual suitability runs throughout professional ADR. A beautifully recorded line can fail if it sounds too clean, too close, too rich or too controlled for the image and surrounding production dialogue. Conversely, a microphone position that would appear unconventional in another recording context may provide exactly the spectral character required for a convincing match. The mixer must understand microphones well enough to use their imperfections deliberately. The question is never simply which microphone sounds best. It is which recording can become part of the scene without revealing the process that created it.

    Once Carden stepped in front of the microphone, the demonstration moved from equipment towards performance. The familiar system of three audible beeps established the timing, with the actor beginning the line where an imaginary fourth beep would occur. In principle, the task appeared straightforward. In practice, Carden repeatedly entered early, became self-conscious about the timing and discovered how quickly an apparently trivial five-word line could become difficult once performance, synchronisation and technical awareness competed for attention. His reaction gave the students an unusually useful demonstration of the psychological demands placed upon actors during ADR. Knowing the line is not enough. Understanding the timing is not enough. Even achieving synchronisation is not enough. The replacement still has to sound as though it belongs to the original performance.

    At one point, Carden produced a take that was synchronised correctly but immediately recognised that the performance itself was wrong. Navarro’s response was subtle. He did not ask for shouting or a dramatically louder delivery. He suggested only slightly more projection. The difference was not simply level. A small change in the physical production of the voice altered its tone and gave the line the additional presence heard in the original recording. When the replacement was compared directly with production dialogue, the improvement became obvious. The exercise demonstrated that ADR matching cannot be reduced to waveform alignment, pitch correction or equalisation. Performance changes the spectrum of the voice before any microphone or processor becomes involved.

    Navarro later developed this point in greater depth. Two performances can have similar apparent loudness and pitch while differing substantially in vocal quality. An actor performing naturally on set may have a relaxed throat and a particular physical relationship with the surrounding scene. On the ADR stage, tension, self-consciousness or the effort to satisfy technical instructions can change the voice. A performer may become tighter, brighter or less natural even while reproducing the words and timing accurately. Matching therefore requires attention to qualities that are difficult to describe numerically. Projection, pitch, volume and rhythm matter, but so do muscular relaxation, breath and the physical origin of the voice.

    This presents the ADR mixer with a delicate problem. The mixer may hear precisely what is preventing a line from matching, yet communicating every technical observation to the actor may make the performance worse. Navarro warned that performers can absorb only so many notes before they begin thinking about the mechanics of speech rather than the character. A useful intervention must therefore translate technical listening into language that supports performance. Sometimes the correct decision is to offer a suggestion. Sometimes it is to communicate through the director. Sometimes it is to recognise that an imperfection is not important enough to justify disturbing the creative balance of the room. Hearing a problem and knowing whether to mention it are separate professional skills.

    The lecture also demonstrated an older approach to ADR recording that remains remarkably useful. Navarro sampled the original production line and played it repeatedly into Carden’s headphones. Instead of concentrating simultaneously upon picture, beeps and performance, Carden could hear the original line and immediately reproduce it, repeating the process several times. Navarro then edited the resulting takes into position and compared them with the production recording. The method offered a direct reference for timing, intonation, volume and vocal character, allowing the performer to respond to the sound of the original performance rather than attempting to reconstruct every detail intellectually.

    Navarro’s explanation of this technique revealed an important insight into perception. Picture can sometimes become a distraction. An actor watching for a mouth movement may wait until it becomes visually apparent, by which time the correct moment to begin has already passed. If the original audio is correctly synchronised to picture, matching its rhythm and timing can naturally reproduce picture sync. The performer can therefore concentrate upon hearing and responding rather than continually monitoring several streams of information at once. Despite the age of the sampling technique, Navarro regarded it as one of the most useful tools available on the stage. Its continued value comes not from technological sophistication, but from the way it simplifies the performer’s task.

    The same objective shaped Navarro’s control room. His system contained extensive routing, multiple microphone inputs, separate monitoring paths and a heavily customised control surface, yet the purpose of this complexity was to make the session itself feel simple. He might need to manage separate mixes for the control room, recording stage, actor headphones, supervisor headphones and remote participants, with each requiring different material during rehearsal, recording and playback. Attempting to reconfigure every route manually between passes would slow the session and create repeated opportunities for error. Navarro therefore designed systems that allowed complex changes to happen immediately.

    His customised keypad provided a particularly revealing example. Individual buttons could trigger sequences of macros that armed tracks, initiated recording and performed repetitive editing operations once a take had finished. Navarro had spent months developing the core system and continued refining it whenever he noticed himself repeating unnecessary keyboard operations. He had begun his career as an ADR recordist and understood the division of labour on a two-person stage, where one person could manage recordings and files while the mixer concentrated upon the performers. Working alone, he used automation to reproduce some of that support, delegating repetitive technical operations to macros so that his attention could remain directed towards the stage.

    This led to one of Navarro’s most important ideas: the ADR mixer must work at the speed of creativity. An actor or director may suddenly discover a new approach to a line and want to record it immediately. If the mixer responds by asking everyone to wait while tracks are configured, routes changed or files prepared, the idea may lose its immediacy. A delay of only a few seconds can alter the atmosphere of a performance. Technical speed therefore has a human purpose. The mixer learns the system thoroughly enough that machinery does not interrupt thought.

    Navarro connected this principle to advice he had received about respected ADR mixer Tommy O’Connell. Asked what distinguished O’Connell’s work, a sound editor gave a simple answer: he anticipates. Useful actions are completed before anyone needs to request them. Navarro interpreted this not as a mysterious talent, but as the result of attention. If the mixer is watching the stage, listening to conversations and understanding the direction in which a session is moving, preparation for the next action can begin before a formal instruction arrives. Anticipation therefore depends upon both technical readiness and social awareness. A mixer whose attention is buried in the workstation may complete every requested task correctly while still remaining one step behind the session.

    Carden’s preparation of the cue and Navarro’s management of the stage reveal two sides of the same professional process. Carden reduces uncertainty before the session begins through accurate cueing, documentation, version control and communication. Navarro reduces friction during the session through templates, routing, automation and anticipation. Neither form of preparation is intended to make the process more rigid. Both preserve the possibility that an actor, director or mixer can respond immediately when something unexpected and valuable occurs.

    Microphone signal flow offered another example of this balance. Navarro described a deliberately clean recording path from microphone through preamplifier into Pro Tools, avoiding unnecessary outboard processing. Within the workstation, compression and equalisation could be used sparingly, but he warned against making irreversible decisions without good reason. Recording completely flat preserves maximum flexibility for the re-recording mixer, while careful decisions made during the session can still produce useful, committed tracks. Heavy processing may be difficult to undo later. A technically impressive recording decision is not valuable if it reduces the options available to the production.

    Navarro’s session was configured to make microphone comparison immediate. Several microphone inputs could remain available, feeding dedicated record tracks at appropriate levels. A boom and lavalier could be recorded simultaneously as discrete channels, preserving both perspectives for later evaluation. The workflow reflected the uncertainty inherent in matching. The microphone expected to work best may not produce the most convincing result once the line is placed against production. Recording alternatives gives later editors and mixers material from which to construct the most believable transition.

    His discussion of recording format was equally pragmatic. For conventional ADR, he worked at the established production standard of 48 kHz, 24-bit audio. Higher sample rates could be valuable for sound effects intended for extensive manipulation, where recordings might later be slowed or processed heavily. ADR serves a different purpose. Recording at unnecessarily high rates would increase storage requirements before the material was eventually converted to the format used by the production. The appropriate technical choice depends upon what will happen to the sound. More data is not automatically more useful.

    As the lecture progressed, it became increasingly clear that the most difficult parts of ADR were not contained within any equipment specification. Navarro estimated that learning the technical fundamentals and becoming comfortable with Pro Tools took years, yet he distinguished that competence from the broader ability to run a session. Recording dialogue in sync with picture is ultimately a technical process that can be learned. Managing actors, directors, supervisors, performance, uncertainty and the emotional energy of the room represents a different level of expertise.

    Actors may arrive nervous. Directors may be highly collaborative, completely self-sufficient or resistant to suggestions. A performer may struggle with a line while becoming increasingly aware of the difficulty. The mixer must understand how much intervention the situation can support. Navarro described the importance of creating a calm and comfortable environment, particularly when clients are unfamiliar with the process. Confidence can be communicated without dominance. If someone is uncertain, the mixer can explain what will happen, answer questions and demonstrate that the session is under control. Technical authority becomes useful when it reduces anxiety rather than displaying superiority.

    Carden and Navarro also acknowledged the professional judgement involved in deciding when not to intervene. A mixer may recognise an aspect of a performance that could be improved, yet nobody has asked for technical input and the issue may not be important enough to disrupt the session. Navarro framed this as a balancing act. Would the intervention materially improve the line? Is the director open to suggestions? Will another note distract the actor from a performance that is already working? Expertise includes recognising problems, while professional maturity requires distinguishing consequential problems from imperfections that do not matter.

    The ADR stage therefore brings technical precision and emotional sensitivity into the same process. The mixer must hear minute changes in vocal quality, understand microphone behaviour, maintain synchronisation, manage several monitoring environments and operate the recording system almost instinctively. At the same time, attention must remain on body language, conversation, confidence, frustration and creative momentum. Mastery of the technology is essential precisely so that the technology no longer consumes the attention needed elsewhere.

    Carden’s demonstration also placed the scale of professional ADR into perspective. Their five-word line required production recording, cue preparation, documentation, stage setup, microphone selection, rehearsal, multiple takes, alternative recording methods, editing and comparison. A skilled actor might complete around ten or twelve conventional cues in an hour, while a feature film may contain 150 or 200 ADR lines. Complicated performances naturally take longer. What appears to an audience as a few moments of seamless dialogue can therefore represent days of concentrated work.

    Yet the success of that work is measured largely through its invisibility. When Navarro compared Carden’s replacement line with the production recording, the evaluation concerned relationships: volume, tone, vocal quality, perspective and the way the replacement interacted with the surrounding scene. The door slam that had originally damaged the line was restored around the new performance, returning the replacement to the event from which it had temporarily been separated. Heard alone, the ADR recording was merely a voice on a stage. Placed back into the scene, it became part of an action.

    This transformation captures the philosophy shared by both speakers. ADR is often described as the replacement of unusable dialogue, but their demonstration showed that replacement is only the beginning of the problem. The objective is reconstruction. The actor reconstructs a performance. The mixer reconstructs a microphone perspective. Editors reconstruct timing and continuity. The final soundtrack reconstructs the relationship between voice, action and environment so successfully that the audience experiences a single uninterrupted moment.

    The joint nature of the lecture made this especially clear. Carden approached ADR through the complete production process, showing how a line travels from location problem to documented cue and finally to the recording stage. Navarro approached the same process from inside the control room, revealing the technical systems, listening skills and interpersonal judgement required to capture a convincing replacement. Their perspectives met at the point where professional preparation serves human performance.

    By the end of the session, the deliberately damaged line had become something much larger than a technical demonstration. It revealed why ADR demands more than synchronisation, why the cleanest microphone is not always the correct microphone, why a performer can match timing while missing the voice, why an old sampler can remain useful in a modern digital workflow and why a highly automated control room can make a session feel more human rather than less. Above all, Carden and Navarro showed that successful ADR depends upon where attention is directed. The actor should be thinking about performance rather than machinery. The director should be thinking about the scene rather than routing. The mixer should be watching and listening to the room rather than fighting the workstation. The audience, finally, should be thinking about none of these things. When every part of the process works together, they simply hear a character speak.

  • How Do You Mix Television Sound Under Pressure? Frank Morrone on Dialogue, Workflow, and the Art of Re-Recording

    Frank Morrone

    How do you mix television sound under pressure?

    A television soundtrack may contain hundreds of dialogue recordings, sound effects, Foley performances, backgrounds, ADR takes and music stems, all competing for space within a mix that must remain clear, emotionally convincing and technically suitable for broadcast. The audience should never become aware of that complexity. They should simply understand every line, believe every environment and remain absorbed in the story. During his online guest lecture for Edinburgh Napier University, re-recording mixer Frank Morrone explored the craft behind achieving that apparent simplicity. Drawing upon a career in film and television that began in 1979, and projects including Lost, The Strain, Sleepy Hollow and Criminal Minds: Beyond Borders, he revealed a discipline shaped equally by technology, organisation, collaboration and judgement. Throughout the session, one principle emerged repeatedly. The most effective mixing workflows allow enormous technical complexity to disappear behind the story.

    Morrone began by tracing a career that developed across several different areas of professional audio. His earliest work took place in music studios, recording jazz and orchestral film scores before following those recordings into the dubbing theatre and becoming increasingly interested in the way complete soundtracks were assembled. A move into post-production allowed him to work across dialogue editing, music editing and Foley recording before concentrating upon re-recording mixing. That breadth of experience shaped the collaborative philosophy running throughout the lecture. Unlike a music recording session, where one engineer may remain closely involved from recording through to the final mix, film and television sound brings together work created by many different specialists. The dub stage is where those contributions finally meet. Successful mixing therefore depends upon understanding not only the material itself, but the people, processes and decisions that produced it.

    Television makes this collaboration particularly demanding. Morrone described an industry in which track counts have continued to increase while schedules and budgets have become progressively tighter. Sophisticated surround mixes must be created rapidly, and emerging formats add further complexity without removing the need to support conventional playback systems. His response is not simply to work faster. It is to design workflows that remove unnecessary decisions from the mixing stage. As soon as he joins a project, he communicates with the supervising sound editor about track requirements and provides a starting template so that incoming material already fits an established structure. Organisation begins before the mixer enters the room. Under severe time pressure, the ability to find, control and compare material immediately becomes part of the creative process itself.

    This reveals something deeper about Morrone’s understanding of expertise. His templates, early conversations with sound editors, knowledge of production microphones, preparation of alternative takes and habit of printing completed passes all anticipate problems before they are allowed to interrupt the mix. The same thinking extends beyond the dubbing theatre. He considers how broadcast processing will react to dynamics, how a mix will translate to domestic systems and how material created for one format will behave when heard through another. Professional experience, in this sense, is not simply the ability to solve problems quickly. It is the ability to recognise where problems are likely to emerge and construct a workflow in which many of them have already been addressed before they become urgent.

    The scale of Lost provided a striking illustration. Morrone showed the students sessions containing extraordinary numbers of elements, including dozens of tracks dedicated to the Smoke Monster alone. Its identity emerged from a deliberately ambiguous combination of animal voices, pneumatic machinery, roller-coaster wheels and other contrasting sources, creating something that resisted being understood as either entirely organic or entirely mechanical. Those effects existed alongside hard effects, backgrounds, Foley, production dialogue, ADR, group recordings and substantial music deliveries. The technology available at the time imposed strict limits upon voices and processing, requiring careful decisions about resource allocation as well as creative balance. Complexity could not simply be solved by adding more processing. The session itself had to be organised so that the mixers could navigate it instinctively.

    Custom fader layouts and VCA groups became essential to that process. Morrone described arranging controls so that principal dialogue could immediately be balanced against ADR, group recordings and music, while different categories of material remained independently accessible. Music sources could be separated from score, while dialogue in different languages could be isolated for international deliverables. Every layer of organisation reduced the time between hearing a problem and solving it. This became especially important during pilot season, when mixers might receive unusually elaborate material without first having time to develop a workflow around the programme. A sufficiently flexible template must already be capable of accommodating whatever arrives. Preparation, in this context, creates the conditions in which creative decisions can still be made under pressure.

    The production of Lost also demonstrated the tension between creative ambition and delivery requirements. Morrone recalled the exceptional resources devoted to the programme, including a large pilot budget and Michael Giacchino’s insistence upon recording a live orchestra for each episode. Yet the soundtrack still had to survive the restrictions of television broadcast. The team therefore created a more dynamic version for DVD before producing a contained broadcast mix designed to survive transmission processing. A similar approach was later adopted on The Strain. The distinction mattered. A mix can remain technically within specification and still behave poorly when subsequent broadcast processing responds to excessive dynamics. Re-recording therefore requires mixers to think beyond the dubbing theatre. They are mixing for every system through which the programme will eventually reach its audience.

    Dialogue occupied the centre of Morrone’s approach. His first objective is always to preserve the production performance wherever possible. When ADR has been recorded, he wants to know why. A line replaced for performance reasons presents a different problem from one replaced to solve a technical fault. If the director wanted a different performance, Morrone respects that decision while keeping the original available as an alternative. If the problem was technical, he first explores whether the production recording can be repaired. Modern restoration tools have dramatically expanded what can be rescued, reducing the need to replace performances that may possess subtleties difficult to recreate months later in an ADR studio.

    When ADR is necessary, matching involves far more than applying equalisation and reverb. Morrone keeps production dialogue available alongside the replacement so that he can compare transitions directly, matching tone, acoustic environment, pacing and performance. He prefers to receive several strong takes and recordings from both boom and lavalier microphones, recognising that the microphone apparently closest to the production perspective is not always the easiest to integrate. Understanding which microphones and wireless systems were used during production can also provide valuable clues, particularly where transmission systems have imparted their own sonic characteristics. ADR matching consequently becomes a process of reconstructing relationships rather than searching for a single corrective setting.

    Technology has transformed that work, though Morrone repeatedly warned against allowing powerful restoration tools to encourage excessive processing. Noise reduction, spectral repair, ambience matching, EQ matching and dereverberation can rescue material that would previously have required replacement. Yet he deliberately removes less noise than might appear necessary when dialogue is heard in isolation. Once backgrounds, effects and music return, much of the remaining noise may be perceptually masked. Processing that sounds impressively clean in solo can leave dialogue lifeless and constricted in the finished mix. Morrone therefore keeps copies of original material and sometimes returns to less processed versions during the final mix. The objective is not the cleanest possible dialogue track. It is dialogue that remains natural and convincing within the complete soundtrack.

    His attitude towards restoration reveals a broader philosophy of technology. Morrone began working with magnetic tape, a Cat 43 noise reduction unit and a notch filter, and he clearly values the extraordinary capabilities available to contemporary mixers. Yet greater technical power has not removed the need for judgement. In many respects, it has increased it. The ability to remove more noise does not mean that more noise should be removed, just as the ability to place sound almost anywhere within an immersive field does not mean that every available position should be used. New tools expand the range of possible decisions. They do not determine which decisions are appropriate. Throughout Morrone’s lecture, technical capability remained subordinate to perception.

    This distinction between isolation and context became one of the lecture’s most important ideas. Morrone described television mixing as beginning with dialogue and music, establishing the foundation against which the effects mixer can develop backgrounds and action. Once those elements come together, the mixers decide what should drive the scene. Some moments belong to music, others to effects, while still others require a more subjective perspective or a deliberate reduction of material. Experienced mixing partners develop an almost instinctive understanding of one another’s decisions. If an effect obscures a line during an early pass, Morrone knows that a trusted colleague will create space for it during refinement. Collaboration is therefore more than the division of tracks between two people. It is a shared process of deciding where the audience should listen.

    The idea that sounds must be judged in context reaches far beyond balancing dialogue against effects. Dialogue that appears too noisy when soloed may become entirely convincing once the world of the scene surrounds it. ADR that attracts attention after twenty repeated comparisons may pass unnoticed when the audience encounters it once within a continuous performance. Group recordings that sound absurd under isolated scrutiny may perform their role perfectly when placed at the correct distance behind principal dialogue. Morrone’s examples repeatedly challenged the assumption that individual elements should be perfected independently. The meaningful unit of evaluation is ultimately the audience’s experience of the scene. A soundtrack succeeds through relationships between sounds rather than the isolated perfection of its components.

    Normally, the priority within those relationships remains dialogue. Morrone described it as both the foundation of the mix and the principal carrier of storytelling. His concern with intelligibility, however, extends beyond maintaining a technical hierarchy between dialogue, music and effects. It is fundamentally about preserving attention. His most revealing test is not a meter reading but a listener asking what somebody has just said. At that moment, the audience member has been pulled out of the story and required to think about the failure of the soundtrack. Clear dialogue therefore supports immersion precisely by avoiding attention to itself. The better the mix communicates, the less the audience needs to think about the process of communication.

    Yet the rule is not absolute. For one sequence in The Family, a child was being used as bait to attract a kidnapper in a crowded shopping centre. The environment needed to feel genuinely busy, yet the available collection of separate crowd recordings and reverberant elements never created a convincing whole. Morrone took a six-channel recorder into a real shopping centre, captured the food court from different perspectives and brought the recordings back into the mix. The result allowed the environment to crowd the dialogue slightly, which was precisely what the scene required. Clarity remains fundamental, but realism sometimes depends upon controlled difficulty.

    The shopping-centre sequence illustrates an important tension within Morrone’s approach. His practice is built upon strong principles, but those principles do not become inflexible rules. Dialogue normally takes priority, except when allowing the environment to interfere with it makes the dramatic situation more believable. Restoration should preserve intelligibility, except when excessive cleaning destroys naturalness. Acoustic simulation is valuable, except when recording the real physical relationship between people and space produces a more convincing result. Expertise therefore involves knowing the rules well enough to understand what they protect, then recognising the moments when the needs of the scene justify bending them.

    This willingness to leave the dubbing theatre and record real spaces appeared repeatedly throughout the session. Morrone discussed impulse responses, convolution reverbs and carefully developed presets for rooms, vehicles and other environments, all of which provide useful starting points for worldizing sound. Yet he remained pragmatic about their limitations. A real location does not necessarily sound like the space suggested by the finished image, and even a carefully captured impulse response may require additional reflection, delay or reverberation before it feels convincing. Cars present an especially difficult problem, combining strong early reflections from glass with highly absorptive surfaces elsewhere. Experience gradually produces a library of useful starting points, though listening still determines the final result.

    Sometimes no simulation is as convincing as returning to the physical situation itself. Morrone described a scene in which a group of foster children were supposed to be creating chaos upstairs while a conversation took place below. Studio-recorded group voices did not reproduce the peculiar combination of footfalls, structural transmission and reflections travelling down a staircase into another room. His solution was direct. He gathered children on the upper floor of a house, recorded from downstairs and captured the complete acoustic event as it occurred. No increasingly elaborate chain of processing was required. The physical relationship between performers, building and microphone provided what the scene needed.

    The pressures of television production make such judgement particularly important. Morrone compared television mixing to boot camp. Schedules leave little room for hesitation, and mixers must develop workflows capable of producing strong results quickly. Once a scene has been successfully mixed, he prefers to print it rather than trusting that automation will remain untouched throughout later work. Accidental writes, changed sends and technical errors can occur even in experienced hands. Printing completed work provides security and allows later changes to be punched into established stems. Efficiency does not mean rushing blindly. It means reducing the opportunities for avoidable problems to consume the limited time available.

    Long sessions also introduce a more human limitation: hearing fatigue. Morrone described working on the highly dynamic soundtrack of The Strain, where sustained exposure to loud material forced him to think deliberately about auditory recovery. His solution was simple but important. He left the room for short periods, walked and allowed his ears to recover while his mixing partner continued working. The two mixers could alternate demanding passes, giving each other opportunities for rest without stopping the session. After decades of mixing, one of the useful discoveries was not another plug-in or processor, but the value of leaving the chair for ten minutes. Professional listening depends upon recognising the limits of the listener.

    Morrone’s discussion of client relationships revealed another dimension of the re-recording mix that students may rarely encounter in technical demonstrations. Mixers are working with directors, producers and other clients who may have strong preferences that differ from their own. Morrone described situations in which clients wanted music loud enough to compete with dialogue. His responsibility was to explain the likely consequences, demonstrate how the mix translated at lower levels and on smaller monitors, and search for a compromise that preserved the client’s intention while protecting intelligibility as far as possible. The mixer offers expertise, but does not own the programme. Knowing which decisions are worth challenging and which require accommodation is part of the craft.

    His description of deciding whether an issue represents a “hill to die on” reveals a sophisticated understanding of professional authority. Expertise does not give the mixer unlimited control over the work, nor does collaboration require the abandonment of professional judgement. The mixer must advocate for the audience, explain likely consequences and make alternatives audible, while recognising that the final creative intention belongs to the client. Professional judgement therefore includes negotiation. Sometimes expertise means defending a decision. Sometimes it means finding a compromise that neither side initially imagined. Occasionally, it means implementing a choice that remains contrary to personal taste while ensuring that it works as successfully as possible.

    Technical fluency plays an important role in maintaining those relationships. When a client requests a change, Morrone wants to make it immediately, play it once and continue. Searching through tracks or repeatedly troubleshooting a familiar process changes the atmosphere of the room and interrupts attention to the programme itself. A well-designed session keeps the conversation focused upon storytelling and communication. The deeper the mixer’s command of the tools, routing and session layout, the less those systems intrude upon the creative discussion.

    This is another form of transparency. Morrone repeatedly returned to the idea that technology should disappear from the client’s experience, yet transparency does not mean that technology has become unimportant. The opposite is closer to the truth. Considerable technical knowledge is required to make complex systems feel immediate. Templates, custom layouts, routing, monitoring, printed stems and intimate familiarity with the workstation create an environment in which a creative request can become an audible result without breaking the flow of the session. Mastery becomes visible through the absence of friction.

    The growth of immersive sound has expanded this challenge further. Morrone described mixers as working between extremes, from Dolby Atmos and sophisticated home theatres to stereo playback and mobile devices with earbuds. His philosophy was to begin with the best mix possible in the most capable format, then ensure that it translates successfully into simpler ones. Immersive technology may offer extraordinary spatial possibilities, but Morrone remained cautious about using novelty without considering perception. In particular, he defended dialogue remaining anchored to the centre. Moving voices between speakers can alter timbre and create distracting changes that audiences may notice without understanding their source. New formats create possibilities, though they do not invalidate principles developed through decades of listening.

    His argument about dialogue placement is particularly revealing. Audiences do not need to identify the technical source of a problem in order to experience discomfort. Morrone described viewers sensing that something was wrong when dialogue moved around an immersive field, even when they could not explain precisely what disturbed them. This places an unusual responsibility upon the mixer. Professional listening must sometimes diagnose experiences that ordinary listeners can feel but cannot name. The purpose of expertise is not to dismiss those responses as technically uninformed, but to understand the perceptual conditions that produced them.

    The same concern for translation shaped his approach to low frequencies. Subwoofers vary enormously between listening environments, and domestic listeners frequently adjust them far beyond calibrated levels. Morrone therefore warned against depending entirely upon the LFE channel for the weight of a soundtrack. Low-frequency energy can also be carried through the main channels, creating a result that remains powerful across a wider range of playback systems. His experience of hearing a domestic subwoofer struggle with low-frequency material from Sleepy Hollow reinforced the point. A soundtrack must survive real listening environments, not merely sound impressive on a perfectly calibrated dubbing stage.

    As the discussion widened, Morrone considered the future possibility of mixes that adapt more intelligently to different devices and contexts. Streaming, immersive audio, virtual reality and personalised playback were already creating pressure for soundtracks to function across radically different systems. Yet his underlying philosophy remained remarkably consistent. Whatever the format, begin with the strongest possible mix, preserve the storytelling hierarchy and understand how human perception responds to the result. Technology changes quickly. The responsibility to guide attention and communicate narrative does not.

    Questions from the students returned the discussion to practical preparation. Morrone strongly supported editors delivering material that is already sensibly balanced before it reaches the stage. Dialogue editors who use clip gain to create consistent levels save valuable mixing time, while backgrounds and effects arriving close to useful operating levels allow the mixer to begin creatively rather than first correcting avoidable problems. He described requesting particular editors for demanding productions precisely for this reason. Good preparation is noticed. In a professional environment where a pilot may need to be mixed in only a few days, the person who consistently delivers well-organised, intelligently balanced material becomes someone mixers actively want on the next project.

    His discussion of ADR offered another deceptively simple lesson about perception. Morrone sometimes works on difficult replacement lines privately through headphones while the effects mixer is making a pass. The client then hears the finished line only once in context rather than listening to it repeated dozens of times during adjustment. Repetition directs attention towards the repair and teaches the listener exactly where to expect it. The same awareness shaped his humorous rule about never soloing loop group in front of a client. Background conversations that work perfectly as part of a scene may sound absurd when isolated and scrutinised. What matters is not whether every element survives examination on its own, but whether it performs its intended role in the scene.

    These examples reveal that Morrone is not simply mixing sound. He is managing attention, expectation and knowledge. Once listeners have been taught where an edit exists, they may hear it differently. Once an element has been isolated, they may judge it according to criteria that have little relevance to its actual purpose. Perception is shaped not only by acoustic information, but by what listeners have been encouraged to notice. Part of the mixer’s craft therefore lies in protecting the audience’s experience from unnecessary awareness of the mechanisms used to construct it.

    Yet group recording could also become a powerful storytelling tool. On Criminal Minds: Beyond Borders, episodes moved between international locations while much of the production remained based in Los Angeles. Carefully performed local-language group recordings, combined with music and other environmental elements, became essential to establishing each location convincingly. Here, group material could be brought forward rather than hidden. There was no universal rule governing how loudly an element should be mixed. Its appropriate level depended upon what the scene needed to communicate.

    By the end of the session, Morrone’s account of re-recording mixing had moved far beyond faders, plug-ins and delivery specifications. The technology matters enormously, as do templates, routing, restoration tools, monitoring and control surfaces, but those things serve a larger process. A mixer must understand performance, storytelling, perception, collaboration, translation and the subtle politics of working with clients under pressure. The session may contain hundreds of tracks, yet the audience should hear a coherent world rather than the complexity required to construct it. Morrone’s lecture revealed a craft built upon anticipation, contextual judgement and the careful management of attention. Preparation preserves the possibility of creativity under pressure. Technical knowledge allows technology to disappear from the conversation. Rules provide essential foundations, while experience reveals when the needs of a scene require them to bend. Great television sound is not created by making every element impressive or every recording perfect in isolation. It emerges from understanding what the audience needs to hear, recognising what they should never need to notice, and making hundreds of individual decisions feel like one continuous experience.

  • How Does Sound Affect Us? Julian Treasure on Listening, Wellbeing, and Designing with Our Ears

    Julian Treasure

    How does sound affect us?

    Most people think about sound only when it becomes a problem. We notice the neighbour’s loud music, the traffic outside a bedroom window, the distracting conversation in an open-plan office or the shrill alarm that interrupts an otherwise quiet day. Far less attention is paid to the countless sounds that quietly shape our emotions, influence our behaviour and affect our health from one moment to the next. During his online guest lecture for Edinburgh Napier University, Julian Treasure argued that this oversight represents one of the greatest shortcomings of modern design. Buildings, products and public spaces are often designed primarily for the eye, while the ear receives remarkably little attention. Yet sound continually influences the way people think, work, communicate and feel. Throughout the presentation, one message emerged repeatedly. If we wish to design better experiences, we must learn to design with our ears.

    Treasure began by asking what sound actually affects. The answer, he suggested, is surprisingly simple. Sound influences our happiness, our effectiveness and our wellbeing, along with those of everybody who shares the environments we create. This observation immediately shifts the discussion away from traditional concerns about noise control or acoustic specifications. Sound becomes a human issue rather than merely a technical one. The quality of an acoustic environment influences far more than whether a room sounds pleasant. It affects how effectively people communicate, how comfortably they work, how safely they respond to hazards and how they experience the spaces in which they spend their lives. Sound therefore deserves to be regarded as one of the fundamental materials of design rather than an afterthought considered once construction has already been completed.

    To explain why sound exerts such profound influence, Treasure described four principal ways in which it affects human beings. The first is physiological. Unlike vision, which depends upon the direction in which we happen to be looking, hearing continuously monitors the environment around us. Human beings cannot close their ears in the way they close their eyes. Throughout evolution, this has made hearing our primary warning system, continually searching for signs of danger beyond the limits of our vision. As a consequence, sound reaches deeply into the nervous system with remarkable immediacy. Sudden or unpleasant sounds trigger hormonal responses associated with stress and vigilance, while calmer acoustic environments encourage relaxation. Treasure illustrated this contrast using familiar examples. An unexpected loud noise immediately increases physiological arousal, whereas gentle natural sounds such as breaking waves often slow breathing and encourage a sense of calm. These responses are not matters of personal preference alone. Sound influences heart rate, hormone secretion, breathing patterns and even patterns of brain activity, quietly shaping the body’s internal rhythms throughout the day.

    The discussion then moved beyond physiology towards psychology. Music provides perhaps the most familiar illustration of this relationship. People instinctively choose particular music to celebrate, to concentrate, to relax or to reflect, recognising that different sounds evoke different emotional states. Treasure argued that natural sounds often produce similarly powerful responses. Birdsong, for example, tends to create feelings of safety and reassurance. Rather than being arbitrary preferences, these reactions may reflect deep evolutionary associations developed over thousands of generations. Birds sing when environmental conditions are relatively safe, allowing those sounds to become unconsciously associated with security. Although listeners rarely analyse these processes consciously, they nevertheless influence emotional experience in subtle yet persistent ways. For sound designers, this observation carries important implications. Designing an acoustic environment involves much more than controlling sound levels. It also requires understanding the emotional associations that different sounds naturally evoke.

    Treasure’s argument became even more relevant to contemporary workplaces when he turned to the cognitive effects of sound. Human attention is a limited resource. People often imagine that they can listen to several conversations simultaneously, though the reality proves rather different. Treasure observed that the human brain possesses only a limited capacity for processing speech, making it extremely difficult to concentrate when nearby conversations compete for attention. Open-plan offices provide a familiar example. Designers frequently value openness, flexibility and visual communication, yet the resulting soundscape often undermines the very productivity these environments seek to encourage. Relevant speech continually draws attention away from the task at hand, interrupting concentration and increasing mental effort. Research cited during the presentation suggests that productivity can fall dramatically under these conditions, illustrating that acoustic design contributes directly to cognitive performance rather than simply influencing comfort. Decisions about the sonic character of workplaces therefore become decisions about how effectively people can think.

    By this stage, a broader pattern had already become clear. Sound is not simply something that accompanies our activities. It shapes them. Physiological responses, emotional reactions and cognitive performance all depend, to varying degrees, upon the acoustic environments within which people live and work. This perspective challenges a long-standing tendency to regard sound as secondary to visual design. Treasure instead presented listening as a central consideration for architects, designers, engineers and sound professionals alike. Before deciding how a space should look, he suggested, we should also ask how it will sound, and how those sounds will influence the people who experience them every day. That question would remain at the heart of the remainder of the discussion.

    Having established that sound influences our physiology, psychology and cognition, Treasure turned to its effect upon behaviour. This influence often operates below the level of conscious awareness, making it particularly easy to overlook. Most people assume they make decisions independently of their acoustic surroundings, yet evidence suggests otherwise. Unpleasant environments encourage people to leave sooner, while attractive soundscapes invite them to remain longer. Treasure illustrated this with a striking study of consumer behaviour. In a supermarket displaying French and German wines with identical visual presentation, researchers changed nothing except the background music. On days when French music was played, French wine substantially outsold German wine. When German music replaced it, purchasing patterns reversed. Customers generally remained unaware that the music had influenced their choices, demonstrating that sound can shape behaviour without requiring conscious attention. For designers, retailers and architects alike, this example reinforced an important point. The acoustic environment is never simply a backdrop. It actively participates in shaping human decisions.

    The implications extend far beyond retail spaces. Treasure argued that every designed environment communicates through sound, whether intentionally or otherwise. A restaurant may create an atmosphere that encourages relaxed conversation, while another overwhelms diners with reverberation and competing voices. A hospital waiting room may reduce anxiety through carefully considered acoustics, or increase it through intrusive alarms and mechanical noise. An office may support concentration, or continually undermine it through poorly managed speech privacy. In each case, the acoustic environment becomes part of the overall design, influencing how people behave within the space. Designers therefore make decisions about human experience whenever they make decisions about sound, even if those decisions consist of ignoring it altogether.

    Treasure observed that this neglect reflects a broader imbalance within contemporary design practice. Buildings are routinely judged according to their appearance, products are evaluated through their visual form, and digital technologies devote enormous attention to graphical interfaces. Comparatively little thought is often given to how these same environments sound. This imbalance is surprising when one considers that hearing operates continuously. We can choose where to look, though we cannot simply decide to stop hearing the world around us. Sound therefore accompanies every activity, continually influencing perception in ways that visual design alone cannot achieve. Rather than treating acoustics as a specialist concern addressed late in a project, Treasure encouraged students to recognise listening as a fundamental design consideration from the very beginning.

    This perspective resonates strongly with professional sound design. Whether creating a film soundtrack, designing interface sounds, producing a virtual instrument or developing an interactive game, practitioners rarely add sound simply to occupy silence. Every sound communicates information, guides attention or influences emotional response. Treasure’s presentation broadened this principle beyond media production into everyday life. The same questions that sound designers ask while constructing a soundtrack also apply to architecture, product design and urban planning. What should the listener notice? Which sounds deserve emphasis? Which should remain unobtrusive? How can sound support rather than distract from the intended experience? The boundaries between sound design and environmental design begin to blur once listening itself becomes the central concern.

    Perhaps the most compelling aspect of Treasure’s argument lay in its optimism. If sound can undermine wellbeing, productivity and behaviour, it can equally improve them. Pleasant acoustic environments encourage relaxation, reduce physiological stress and support clearer thinking. Appropriate sound can strengthen communication, promote social interaction and make public spaces more welcoming. Rather than presenting acoustics as a matter of reducing unwanted noise, Treasure reframed the discussion in positive terms. The objective is not simply to remove bad sound, but to create environments in which good sound actively contributes to human wellbeing. This shift in perspective encourages designers to think creatively about the role sound can play rather than treating it solely as a problem to be controlled.

    These ideas naturally led towards a broader discussion of listening itself. If sound exerts such profound influence over human experience, then the ability to listen carefully becomes an essential professional skill rather than an incidental personal habit. Treasure suggested that hearing and listening are not the same activity. Hearing occurs automatically, while listening demands conscious attention, intention and practice. In an increasingly noisy world filled with competing sources of information, the ability to listen thoughtfully may be becoming more valuable rather than less. This distinction between passive hearing and active listening would ultimately form the foundation of his concluding message, not only for sound designers but for anyone responsible for creating environments in which other people live, work and communicate.

    Having demonstrated that sound influences physiology, emotion, cognition and behaviour, Treasure turned towards a more practical question. If sound has such profound effects upon human experience, what should designers actually do differently? His answer was strikingly optimistic. Rather than treating acoustics as a problem to be solved, he encouraged students to think of sound as a resource that can be shaped deliberately to improve people’s lives. Well-designed soundscapes do more than reduce unwanted noise. They encourage particular patterns of behaviour, support communication and create environments in which people feel healthier, calmer and more engaged. Designing with sound therefore becomes an act of positive intervention rather than damage limitation.

    Treasure illustrated this philosophy through a series of real-world projects. Airports, shopping centres and public spaces all benefited from carefully designed soundscapes that considered not only what people heard, but how those sounds influenced the way they behaved. Introducing natural sounds and thoughtfully composed musical environments increased customer satisfaction, encouraged visitors to remain longer and, in several cases, improved commercial performance. In one public space, the introduction of a biophilic soundscape was even associated with a measurable reduction in crime. These examples reinforced a central point running throughout the presentation. Sound does not merely accompany human activity. It shapes it. Decisions about the acoustic environment therefore become decisions about wellbeing, behaviour and social experience rather than simply matters of technical acoustics.

    Although these examples came from architecture and environmental design, their relevance extends directly to professional sound design. Every soundtrack contains foreground and background elements competing for the listener’s attention. Treasure encouraged students to think carefully about the role each sound should play within that wider acoustic picture. Not every sound deserves prominence, and not every moment benefits from additional music or greater complexity. Like a visual composition, an effective soundscape depends upon balance, hierarchy and clarity. He also encouraged designers to draw inspiration from natural environments, particularly through the thoughtful use of biophilic sound and adaptive or generative soundscapes that evolve over time rather than repeating mechanically. Different spaces support different activities, and their sonic character should reflect those differing purposes. Designing for concentration requires different acoustic decisions from designing for relaxation, learning or social interaction.

    The discussion naturally led back to listening itself. Treasure argued that hearing should never be confused with listening. Hearing is automatic. Listening is intentional. It requires attention, effort and continual practice. In an age characterised by constant distraction and increasingly complex acoustic environments, the ability to listen carefully becomes one of the most valuable professional skills a sound designer can develop. Technical expertise undoubtedly remains important, though it cannot substitute for careful listening. The most sophisticated recording equipment or software offers little value if the designer fails to recognise what listeners actually experience. Listening therefore becomes both a creative skill and an ethical responsibility. Before changing the sound of the world, designers must first learn to hear it properly.

    Treasure concluded by describing what he called the four foundations of effective listening: being conscious, committed, compassionate and curious. Conscious listening requires recognising that listening is an active process rather than a passive consequence of hearing. Commitment acknowledges that good listening demands time, attention and intention. Compassion encourages genuine understanding of other people through careful listening, particularly when viewpoints differ from our own. Curiosity reminds us that every sound and every conversation offers an opportunity to learn something new. Although these principles were presented in the context of listening, they also describe many of the qualities that distinguish thoughtful sound designers. Successful practitioners remain attentive, purposeful, empathetic and continually curious about how people experience the acoustic world around them.

    Treasure’s final appeal brought together everything that had preceded it. He encouraged students to become champions of listening and, above all, to “design with your ears.” This simple phrase encapsulated the wider philosophy running throughout the presentation. Sound should never be regarded as an afterthought added once visual design has been completed. It is one of the primary ways in which people experience the world. Every building, product, public space and interactive system possesses an acoustic identity that influences those who encounter it. Whether designing a film soundtrack, a hospital, a mobile application or a railway station, the same principle applies. The sounds we create shape the lives of the people who hear them.

    Taken together, Treasure’s presentation offered a compelling vision of contemporary sound design. It challenged the traditional tendency to regard sound as secondary to vision and instead positioned listening at the centre of human experience. Physiological responses, emotional wellbeing, cognitive performance, behaviour and communication all depend, to varying degrees, upon the acoustic environments we inhabit. For sound designers, this represents both an opportunity and a responsibility. Every decision about sound has consequences extending beyond aesthetics alone. Designing well therefore means more than creating compelling audio. It means understanding how people listen, recognising how profoundly sound affects everyday life and applying that knowledge to create environments in which individuals and communities can genuinely flourish.

  • What Can Sound Communicate That Words Cannot? Jim Metzner on Memory, Listening, and Going Places That Words Cannot Go

    Jim Metzner

    Jim Metzner began the lecture with a mystery.

    A sound was played. Students suggested possible explanations. Some heard machinery. Others heard something else entirely. For a few minutes the recording remained unresolved. Much of Metzner’s work inhabits that moment before a sound settles into a clear explanation. Before it becomes a bird, a vehicle, a voice, or a machine, it exists as an experience. During his online guest lecture for Edinburgh Napier University, discussions of field recording, travel, documentary production, family history, and memory repeatedly returned to this idea. How can sound communicate aspects of experience that are difficult to convey in any other way?

    Listening, in Metzner’s view, is not simply a way of gathering information. It is a way of encountering people, places, and experiences. Much of his work begins from a deceptively difficult question. How can a sound recording help somebody experience something they have never encountered for themselves?

    That challenge appeared repeatedly as students discussed their own recordings. Several described recording parks, public events, city streets, and everyday environments. Similar observations emerged from each example. Carrying a recorder changes the way people move through the world. Sounds that normally fade into the background suddenly become noticeable. Distant traffic acquires texture. Birds occupy distinct locations within a soundscape. Conversations, machinery, weather, and footsteps separate themselves into layers. The microphone becomes a reason to pay attention. One student described attempting to record ambience in a local park while aircraft repeatedly passed overhead. The interruptions were frustrating. Each time the environment seemed to settle, another aircraft arrived. Metzner responded with a story from his own work. While recording in the Great Swamp near a major airport, he encountered a similar situation. Waiting for silence would have meant waiting forever. Rather than treating the aircraft as a problem, he began treating it as part of the environment itself.

    Metzner’s answer reflected a recurring theme throughout the session. Recording is not always about removing the world. Sometimes it involves allowing the world to remain present. Sounds that initially appear intrusive may become important parts of the story. The aircraft was not simply interfering with the student’s recording. It was also shaping the student’s experience of being in that place. Standing in a park, looking upwards, waiting for the noise to pass, became part of the memory. In that sense, the aeroplane belonged to the story as much as the birds or the wind.

    The conversation then moved towards a problem that confronts many documentarians. The person who makes a recording remembers far more than the recording itself contains. They remember the weather, the location, the circumstances, and their own reactions. Future listeners possess none of this knowledge. How, then, can an experience be shared with somebody who was never there? A recording alone rarely provides a complete answer. Context becomes necessary. Yet explanation creates its own difficulties. Too little information leaves listeners uncertain about what they are hearing. Too much information can overwhelm the recording itself. Over the course of his career, Metzner has carried microphones through deserts, cities, forests, festivals, religious ceremonies, and countless other environments. Yet the purpose of these recordings has never been simply to build an archive of unusual sounds. Instead, they function as forms of communication. During the discussion, he compared recordings to postcards. A postcard never contains everything about a place. It presents only a fragment. Yet that fragment can still communicate something meaningful. Sound recordings operate in a similar way. They do not reproduce entire experiences. They provide partial access to them. Listeners complete the picture through imagination, memory, and interpretation.

    Listeners themselves become part of the process. No recording contains everything. Microphones record sound pressure variations. They do not record temperature, light, smell, movement, or the countless other details that contribute to an experience. Yet listeners rarely encounter recordings as collections of isolated sounds. They actively construct meaning from what they hear. A few seconds of ambience may be enough to suggest an entire environment. A familiar voice may evoke a person more vividly than a photograph. A distant church bell, footsteps in a corridor, or voices heard from another room can suggest a much larger world than the recording itself contains. Documentary production frequently relies upon this relationship between recording and imagination. Rather than attempting to communicate everything, the producer provides enough material for listeners to begin constructing their own understanding of a place, event, or experience. Recordings do not simply transmit information from one person to another. They create opportunities for participation. Listening becomes an active process through which people assemble impressions, associations, and memories from fragments of sound.

    Rain on a conservatory roof. Crickets during summer evenings. A vacuum cleaner moving through a family home. Songs sung by parents. Early computer games. Calls to prayer heard while travelling. When Metzner asked students to think about sounds they remembered from childhood, the answers arrived quickly. Few of the sounds were unusual. Their importance had little to do with acoustics. What mattered was everything attached to them. The examples revealed how deeply sound can become woven into personal history. Many of the memories were linked to recurring experiences rather than singular events. The sound of rain returning night after night. A family member singing repeatedly over many years. Household sounds that seemed insignificant at the time. Their importance emerged gradually through repetition. Long after specific conversations or individual days had been forgotten, the sounds remained. Several contributions also highlighted how difficult it can be to predict which sounds will become meaningful. People rarely decide in advance that a particular sound will become a lifelong memory. More often, significance emerges retrospectively. A sound that once seemed entirely ordinary acquires importance through later experience. Hearing a familiar sound years later can reactivate memories, emotions, and associations that extend far beyond the recording itself. What returns is rarely just the sound. People remember places, relationships, circumstances, and feelings connected to it. A recording therefore preserves more than an acoustic event. It can preserve pathways back towards experiences that might otherwise feel increasingly distant. A sound that appears entirely ordinary to one listener may carry decades of meaning for another. Hearing is rarely confined to the present moment. Certain sounds seem capable of collapsing time. A familiar voice, a piece of music, or an environmental sound can reconnect listeners with people, places, and relationships that might otherwise feel distant.

    While still in high school, Metzner began recording conversations with his grandfather. There was no documentary project in mind. He was not gathering material for publication. He simply wanted to preserve conversations with somebody he loved. Years later, those recordings became something entirely different. After his grandfather had died, the tapes acquired a significance that would have been impossible to recognise when they were first made. What had once seemed routine became irreplaceable. The story resonated with many listeners precisely because it involved no grand plan. Had Metzner waited until the recordings appeared important, it would already have been too late. Their value emerged from the simple decision to record ordinary conversations while the opportunity existed. From that experience came one of the clearest pieces of advice offered during the lecture. Record parents. Record grandparents. Record the people whose voices matter. Many recordings appear ordinary when they are made. Their value often becomes visible only later. The suggestion was not motivated by nostalgia alone. Voices contain forms of information that are difficult to preserve in any other way. Speech patterns, accents, pacing, humour, hesitation, and personality all become embedded within a recording. Written transcripts can preserve words. Recordings preserve presence. As Metzner reflected on these recordings, the discussion broadened into a larger point about time. Much of everyday life feels too ordinary to document. Conversations happen. People tell stories. Family members describe events that seem familiar and unremarkable. Yet these moments often become increasingly valuable as years pass. Recording provides a way of preserving details that might otherwise disappear unnoticed. Metzner’s reflections on these recordings returned repeatedly to the differences between memory and recording. Human memory is selective. Certain details remain while others disappear. Recordings preserve details indiscriminately. Accents. Hesitations. Laughter. Breathing. The rhythm of a voice. Background sounds that seemed unimportant at the time. Small details that might otherwise have been forgotten can later become deeply meaningful. A recording preserves more than information. It preserves traces of presence.

    Metzner has never been entirely comfortable with the phrase “capturing sounds”. The word suggests possession. It implies that a sound has somehow been seized and stored away. Throughout the discussion he returned to a different idea. Sounds are given rather than captured. Once a recording has been made, the challenge becomes helping somebody else experience what made that sound meaningful in the first place. Context matters. Stories matter. Yet explanation has limits. Documentary production often involves helping listeners approach an experience and then stepping aside so that the sounds can speak for themselves. A successful recording does not simply tell listeners what to think. It creates conditions in which they can form their own relationship with what they hear. The idea sits comfortably alongside much of his work. Recordings are not trophies collected from the world. They are invitations to listen more closely to it.

    Expensive microphones appeared surprisingly rarely in the lecture. Recording technology was never dismissed, though it was rarely placed at the centre of the discussion. Microphones matter. Recording techniques matter. Editing tools matter. Yet none of them can substitute for curiosity. A person who pays close attention to the world will often discover interesting sounds regardless of equipment. Conversely, expensive equipment cannot compensate for a lack of attention. Many of the examples discussed during the session pointed towards the same conclusion. Meaningful recordings often emerge from moments that other people would simply pass by. A sound heard while travelling. A conversation with a grandparent. Rain on a roof. An aircraft passing overhead. None of these experiences appear remarkable at first glance. Their significance emerges through listening.

    Near the end of the session, the lecture’s title, Going Places That Words Cannot Go, felt increasingly apt. Certain experiences resist straightforward description. The sound of rain on a roof. A grandparent’s voice. A crowded street in a distant city. A celebration, a conversation, or a moment of quiet. Words can describe such things. Sound can sometimes bring listeners closer to experiencing them. For Metzner, that possibility lies at the heart of listening. Sound does not simply tell us about the world. Under the right circumstances, it can preserve traces of people, places, and experiences long after the original moment has passed. More importantly, it can allow those experiences to be shared with somebody else. A recording offers only a fragment. A voice. A place. A conversation. A few seconds of sound preserved from a particular moment in time. Yet those fragments can remain meaningful for decades. They can reconnect people with memories, places, and relationships that might otherwise fade. Listening, as Metzner reminded students throughout the session, is not simply a way of gathering information about the world. It is one way of remaining connected to it.

  • How Do We Know What We Are Hearing? Professor Albert S. Bregman on Auditory Scene Analysis and Perceptual Organisation

    Albert Bregman

    How do we know what we are hearing?

    The question sounds simpler than it is. A voice is heard as a voice. A violin is heard as a violin. A passing vehicle is recognised almost immediately. Everyday listening creates the impression that sound sources reveal themselves directly. Most people rarely stop to consider how much processing has already taken place before recognition becomes possible. Professor Albert S. Bregman’s research begins from the observation that sound sources do not arrive at the ears. Acoustic mixtures do. By the time vibrations reach a listener, contributions from many different events have already combined. Voices, musical instruments, footsteps, ventilation systems, birdsong, machinery, and countless other sources may all contribute to the same signal. The auditory system must somehow determine which parts of that mixture belong together and which do not. Before Bregman’s work, hearing research had developed detailed accounts of pitch, loudness, masking, localisation, and frequency analysis. Considerably less attention had been paid to a more fundamental question. How does a listener determine what produced a sound?

    Bregman did not begin with a theory. He began with a puzzle. During memory experiments involving sequences of short sounds, he noticed that listeners often perceived groupings that were not physically present within the stimulus itself. Sounds sharing similar characteristics appeared to organise themselves into separate perceptual streams. The observation recalled ideas from Gestalt psychology, where visual elements combine into structures that cannot be understood simply by examining their individual parts. What began as an unexpected observation gradually became a larger problem. If listeners organise sounds into streams, how does that organisation occur? More importantly, what role does it play in perception itself? Bregman often approached the issue through analogy. Imagine standing beside a lake while observing only two floating markers moving up and down on the water’s surface. The movement provides evidence that something has happened, though many explanations remain possible. A boat may have passed nearby. Several boats may be moving in different directions. Wind may be disturbing the surface. Something may have fallen into the water. The available evidence does not identify the cause. Any conclusion depends upon inference. According to Bregman, hearing presents a similar challenge. The ears receive information about acoustic activity, though they do not receive direct information about the events that produced it. From patterns of pressure variation reaching two eardrums, listeners somehow infer the existence of voices, instruments, machines, animals, and other sound-producing events. Nothing in the signal arrives labelled. The auditory system must determine which acoustic components belong to the same source.

    Questions of speech perception, localisation, attention, and communication all depend upon this process. Before speech can be understood, before a melody can be followed, and before a sound source can be identified, the auditory system must first determine which acoustic components belong together. Organisation is therefore not one stage among many. It provides the conditions under which many other aspects of perception become possible. This perspective led Bregman towards what became known as auditory scene analysis. The term reflects an analogy with vision. Just as visual perception involves identifying objects within a visual scene, auditory perception involves identifying sound-producing events within an acoustic scene. The challenge lies in the fact that sound sources combine before reaching the listener. The auditory system therefore faces a decomposition problem. It must separate a complex mixture into components that plausibly belong to distinct events. A central claim running throughout the lecture was that perception involves more than detecting acoustic information. It also involves organising that information. Bregman’s demonstrations repeatedly returned to this point. Listeners often assume that qualities such as rhythm, melody, pitch, loudness, timbre, and location belong directly to sounds themselves. His examples suggested a more complicated picture.

    Auditory stream segregation provides one illustration. Under certain conditions, listeners stop hearing a single sequence of sounds and begin hearing multiple independent streams. Once this occurs, rhythms that were previously obvious may disappear. New rhythmic structures emerge. Melodic patterns change. The acoustic signal remains unchanged, though the perceptual outcome does not. Bregman’s demonstrations suggested that the consequences extend much further than rhythm or melody alone. Again and again, he returned to the idea that many perceptual properties depend upon how sounds are grouped. Listeners often assume that pitch, loudness, timbre, and spatial location belong directly to sounds themselves. Yet these properties can also be influenced by the way acoustic components are assembled into perceptual objects. When those groupings change, perception may change even when the underlying stimulus remains constant. This claim sits near the centre of auditory scene analysis. The framework is not simply concerned with separating one sound source from another. It is concerned with how perceptual objects are formed in the first place. Before listeners can judge the loudness of a sound, identify its pitch, recognise its timbre, or determine its location, the auditory system must first decide which components belong together. The resulting structure shapes many of the properties that listeners subsequently experience. From this perspective, perception becomes a problem of interpretation. Faced with an acoustic mixture, the auditory system must determine which explanation is most plausible. What listeners hear is not a direct copy of the physical world. It is the outcome of a process through which the auditory system attempts to reconstruct the events most likely to have produced the available evidence.

    Bregman argued that listeners exploit regularities commonly found in the physical world. Certain acoustic relationships provide evidence that components are likely to originate from the same source. Harmonicity offers one example. Many naturally occurring sounds contain frequency components related by simple numerical ratios. When such relationships are detected, the auditory system often groups those components together. Similar reasoning underlies what Bregman described as common fate. Components that begin together, change together, or move together over time frequently appear to belong to the same event. These principles do not guarantee correct interpretation. Rather, they provide strategies that usually correspond with the structure of the physical world. Auditory scene analysis is therefore concerned with probability rather than certainty. The auditory system rarely knows exactly what caused a sound. It generates interpretations that are likely to account for the available evidence. Most of the time those interpretations correspond closely enough to events in the environment that listeners remain unaware that any interpretation has occurred at all. Throughout the lecture, Bregman emphasised that these organisational processes usually pass unnoticed. Listeners rarely experience themselves as constructing interpretations. The world appears already divided into voices, instruments, footsteps, vehicles, and other familiar sources. Auditory scene analysis directs attention to the work required to produce that impression. The apparent simplicity of hearing may be one reason the problem remained difficult to recognise. Successful perception conceals many of the processes that make it possible.

    Music occupied an interesting position within the lecture. Bregman suggested that composers had discovered practical consequences of auditory organisation long before psychologists attempted to explain them. Counterpoint, orchestration, and performance practice frequently involve maintaining distinctions between perceptual streams or encouraging sounds to fuse into larger structures. Musical traditions therefore provide a long record of experimentation with the same organisational tendencies that auditory scene analysis later sought to describe. Music also offers situations in which these processes become unusually apparent. Changes in perceptual organisation can alter the melodies and rhythms listeners hear, making it possible to observe principles that often remain hidden during everyday listening. Bregman was not suggesting that composers were unconsciously applying psychological theory. Rather, centuries of musical practice had encountered many of the same perceptual constraints that later became objects of scientific investigation.

    Yet music represented only one instance of a broader phenomenon. Following a conversation in a crowded room, recognising a familiar voice over the telephone, locating a sound source in a busy environment, distinguishing one instrument from another, and understanding speech in noise may appear to involve different problems. Bregman’s framework suggested that each depends upon a prior act of organisation. Auditory scene analysis altered the relationship between many areas of hearing research by drawing attention to this common foundation. Rather than treating speech, music, localisation, and auditory attention as entirely separate domains, the framework highlighted organisational processes upon which they all depend. Seen in this way, auditory scene analysis is not merely a theory about particular auditory illusions or laboratory demonstrations. It addresses a question that sits beneath much of auditory perception research. How does a listener move from an undifferentiated acoustic mixture to a world populated by distinct events, objects, and sources?

    The framework also shifted attention away from sound as a purely physical phenomenon and towards perception as a process of inference. Earlier approaches often focused on the contents of the acoustic signal. Bregman drew attention to a prior question. Before a listener can recognise a voice, identify an instrument, understand speech, or respond to a warning signal, the auditory system must first decide what probably caused the sound.

    The answer is usually reached so quickly that the problem remains unnoticed. Voices appear as voices. Instruments appear as instruments. Bregman’s work suggests otherwise. Listening depends upon a continual process through which the auditory system constructs explanations from incomplete evidence. Most of the time those explanations correspond closely enough to the surrounding environment that hearing feels direct and effortless.

  • Why Is Data So Quiet? Hugh McGrory on Sonification, Accessibility, and the Future of Information

    Hugh McGrory

    Why is data so quiet?

    Modern life is shaped by data. Governments collect it. Businesses depend upon it. Scientists analyse it. Social media platforms generate vast quantities of it every second. Increasingly, decisions about healthcare, transportation, education, finance, climate, and public policy are informed by information that exists primarily as data. Yet despite its growing importance, most people encounter data in remarkably similar ways. We see charts, graphs, dashboards, spreadsheets, maps, and visualisations. We are expected to look at information rather than listen to it.

    During his online guest lecture for Edinburgh Napier University, Hugh McGrory challenged this assumption. Drawing upon a career that has spanned animation, virtual reality, software development, data storytelling, and sonification, McGrory described a field that sits between sound, design, accessibility, and communication. Across a wide-ranging discussion that moved from GPS navigation to astronomy, podcasting, climate data, artificial intelligence, and urban infrastructure, he repeatedly returned to a deceptively simple question. If data is everywhere, why do we still experience almost all of it through our eyes?

    McGrory’s route into sonification was anything but conventional. Beginning in computer animation and experimental digital media, he later worked with medical imaging researchers at Yale University, where he encountered scientific data in a completely new way. Rather than using cameras to create images, researchers were transforming data into visual representations that scientists could study and interpret. Later work in virtual reality continued this fascination with information and how people interact with it. Yet throughout these experiences, one issue kept resurfacing. Data communication was overwhelmingly visual. Even the language reflects this bias. The field is known as data visualisation. Information is generally assumed to become meaningful once it has been converted into something that can be seen.

    For McGrory, this raises an obvious question. Why should vision carry so much of the burden? The field of sonification attempts to address this imbalance by exploring how information can be communicated through sound. Yet McGrory encouraged students to think beyond narrow academic definitions. The challenge is not simply turning data into audio. The challenge is deciding when sound might be a more useful way of communicating information. GPS navigation provides a useful example. Rather than forcing drivers to consult maps continuously, navigation systems deliver information precisely when it becomes relevant. A driver does not need every detail about the surrounding road network. They need to know when to turn left. Effective sonification follows the same principle. Its purpose is not to communicate everything. Its purpose is to communicate what matters. Throughout the lecture, McGrory repeatedly argued that the modern problem is rarely a lack of information. More often, it is an excess of information. The real design challenge lies in deciding what should reach people, when it should reach them, and how it can be communicated without demanding unnecessary attention.

    This perspective places sonification within a much broader discussion about interface design. Many of the tools through which people still interact with information were developed for a world in which data was far less abundant than it is today. Screens, keyboards, menus, and dashboards remain remarkably successful technologies, though they are not always suited to situations in which people are moving through complex environments while simultaneously performing other tasks. McGrory pointed towards cities as a particularly interesting example. Vast quantities of public information now exist concerning transportation systems, environmental monitoring, infrastructure, weather, traffic, and public services. Much of this information could potentially support everyday decision-making, yet most people never encounter it. One reason is that visual interfaces demand attention. Looking at a screen competes with countless other activities. Sound offers different possibilities. It can accompany movement, coexist with visual tasks, and communicate information without constantly demanding that people stop and look elsewhere.

    Questions of accessibility revealed why these issues matter so much. Much of McGrory’s work has involved collaboration with blind and visually impaired communities, experiences that challenged many of his assumptions about information design. Designers often assume that providing access to information is sufficient. The reality is considerably more complicated. Screen readers can successfully read large quantities of information aloud, though understanding the overall structure of that information remains difficult. McGrory compared this experience to attempting to complete a jigsaw puzzle without ever seeing the image on the box. Individual pieces are available, yet the broader picture remains difficult to grasp. A spreadsheet containing thousands of values can be read sequentially, though identifying patterns, relationships, and trends becomes far more challenging. Sonification offers one possible response to this problem. Rather than replacing detailed exploration, it can provide rapid overviews that help listeners understand the shape of information before investigating individual details.

    These experiences also led McGrory to question some established assumptions within sonification itself. Many projects focus heavily on transforming data into sound while paying comparatively little attention to context. Listeners are often presented with unfamiliar sounds and expected to interpret them independently. For McGrory, this represents a significant limitation. Communication rarely functions through raw information alone. Context, explanation, and narrative play equally important roles. Podcasting, journalism, and storytelling all demonstrate how audiences use framing to understand unfamiliar material. Projects such as the BBC’s Audiograph series combine sonification with narration, allowing listeners to understand not only what they are hearing but why it matters. This approach shifts attention away from sonification as a technical exercise and towards sonification as communication. Sound becomes one element within a broader process of explanation rather than an isolated solution expected to function independently.

    A particularly memorable example involved astronomy. McGrory discussed the work of a blind astronomer who uses sonification to explore stellar data, challenging assumptions about both astronomy and accessibility. At first glance, astronomy appears inseparable from visual observation. Popular images of galaxies, stars, and nebulae reinforce the assumption that astronomical knowledge depends upon sight. Yet the example revealed something much more interesting than accessibility alone. Rather than simply compensating for an inability to see, sonification provided an alternative way of engaging with information.

    This distinction became important throughout the lecture. Discussions of accessibility sometimes assume that non-visual approaches exist primarily to reproduce experiences that sighted people already have. McGrory encouraged a different perspective. Listening is not merely a substitute for seeing. It is a different mode of perception with its own strengths and limitations. Patterns that may be difficult to recognise visually can sometimes become more apparent through sound. Temporal relationships, repetition, variation, rhythm, and change often lend themselves naturally to auditory interpretation. For this reason, sonification can function as more than an accessibility tool. It can become an analytical tool. Different sensory approaches reveal different aspects of information. A graph may highlight one set of relationships. A sonification may reveal another. Rather than competing with visualisation, the two approaches can complement one another. The question is not whether data should be seen or heard. The question is what becomes possible when both options are available.

    Climate data provided another revealing case. Contemporary societies generate vast quantities of information about environmental change, though much of that information remains difficult for non-specialists to engage with meaningfully. Charts and graphs may communicate trends accurately, yet they do not necessarily encourage engagement. McGrory discussed projects that transform environmental datasets into musical or sonic experiences, creating opportunities for audiences to encounter information differently. Such approaches are not intended to replace scientific analysis. Rather, they create alternative routes into understanding. A sonification may not communicate every statistical detail, though it may encourage curiosity, emotional engagement, or reflection in ways that conventional visualisations struggle to achieve. Throughout the lecture, McGrory repeatedly returned to the idea that communication is not simply a matter of transmitting information. It is a matter of helping people connect with information in the first place.

    Underlying many of these examples was a broader argument about innovation and interdisciplinary thinking. McGrory repeatedly suggested that some of the most interesting developments emerge when disciplines that rarely interact begin to overlap. He cited a definition of innovation that particularly resonated with him: innovation happens when things that are separate become mixed. Sonification itself reflects this principle. The field draws simultaneously upon sound design, music, journalism, accessibility research, user experience design, data science, software development, artificial intelligence, and communication studies. No single discipline possesses all the answers. Progress emerges through collaboration between people approaching similar problems from different directions. This perspective also explains why McGrory consistently framed sound as part of larger design conversations rather than as an isolated specialism. Questions about how people receive information, understand systems, and engage with the world are not purely visual problems. They are design problems.

    These ideas become increasingly significant as emerging technologies continue to reshape everyday life. Artificial intelligence, conversational interfaces, spatial audio, augmented reality, and wearable computing all point towards futures in which information may become less dependent upon screens. While immersive visual technologies continue to evolve, audio already occupies a privileged position within contemporary life. Millions of people carry headphones throughout much of the day. Voice assistants have become commonplace. Podcasts reach global audiences. The infrastructure required for sophisticated auditory experiences already exists. For McGrory, the challenge is no longer technological feasibility. The challenge is learning how to design auditory experiences that are genuinely useful.

    Looking back across the lecture, perhaps the most striking aspect was the extent to which sonification emerged as a human problem rather than a technical one. Questions about mapping data to sound, selecting parameters, or designing auditory displays certainly matter, though they consistently led back to larger concerns about accessibility, communication, attention, understanding, and engagement. Again and again, McGrory encouraged students to think less about sound as an isolated discipline and more about what sound can contribute when combined with other ways of understanding the world.

    Data has become one of the defining materials of contemporary society. Increasingly, our institutions, technologies, economies, and daily lives depend upon it. Yet much of that information remains hidden behind interfaces designed primarily for looking rather than listening. For Hugh McGrory, this represents an enormous opportunity. Whenever information exists without sound, there is the possibility of creating new ways for people to experience, understand, and engage with it. More importantly, there is the possibility of discovering things that might otherwise remain hidden. Listening does not simply provide another route to the same destination. Sometimes it changes what can be found along the way.

    Perhaps that is why his central question continues to resonate long after the lecture ends. In a world increasingly shaped by data, why should our ears be left out of the conversation?

  • How Do Mobile Games Sound Bigger Than They Are? George Vlad on Game Audio, Field Recording, and Creative Constraints

    George Vlad

    How do mobile games sound bigger than they are?

    Many people associate game audio with large development studios, lengthy production schedules, and vast teams of specialists. The image is often one of blockbuster productions involving hundreds of developers working over several years. During his online guest lecture for Edinburgh Napier University, sound designer, field recordist, and Edinburgh Napier alumnus George Vlad offered a rather different perspective. Drawing on a career that has included audio for hundreds of mobile games, Vlad described a world in which sound designers are frequently asked to achieve ambitious creative goals under severe practical constraints. Development schedules may last only weeks. Budgets are often limited. Storage space can be measured in megabytes rather than gigabytes. Yet players still expect games to feel rich, engaging, and alive.

    Across the lecture, Vlad repeatedly demonstrated that successful audio design is rarely about having unlimited resources. More often, it is about learning how to achieve more with less.

    Vlad’s own route into the industry reflects this philosophy. Long before he entered formal education, he was fascinated by sound itself. Childhood memories centred on listening to objects resonate, experimenting with makeshift instruments, and becoming absorbed by the sonic characteristics of everyday materials. At the same time, video games became an equally important influence. These parallel interests eventually converged after several years spent working across Europe, saving money to build a small studio and gradually developing the skills needed to pursue audio professionally.

    The path was far from conventional. Without immediate access to formal training, Vlad relied heavily upon experimentation, books, online communities, and practical experience. Early work editing podcasts and audiobooks gradually led to opportunities in games, particularly during the rapid growth of smartphone applications in the early 2010s. Later, after moving to Edinburgh in 2013, he enrolled on Edinburgh Napier University’s Sound Design programme, where formal study helped fill many of the gaps he had identified in his own knowledge. Rather than describing graduation as the end of a learning process, however, Vlad suggested that education had mainly revealed how much more there remained to learn.

    Looking back, many of these experiences involved similar challenges. Whether teaching himself new skills, building a freelance business, or learning how to work within the realities of mobile development, progress depended less upon ideal circumstances than upon adaptability. This theme would recur throughout the lecture.

    The realities of mobile game development provide a particularly clear illustration of this challenge. Unlike major console or PC titles that may take years to complete, many mobile games operate on remarkably compressed schedules. A developer might contact a sound designer only days before release, requiring dozens of sound effects and music assets within a very short period. Under these circumstances, efficiency becomes essential.

    What emerged from Vlad’s description was a picture of sound design that differs considerably from popular perceptions of creative work. Inspiration certainly plays a role, though much of the process involves practical decision-making. Developers provide lists of required sounds, visual references, gameplay footage, or playable builds. From these materials, the sound designer develops an understanding of how the game should feel. This emphasis on feeling proved particularly important. Before focusing on individual sounds, Vlad explained that he first tries to understand the intended player experience. Should the game feel exciting, relaxing, humorous, energetic, or mysterious? These broader emotional goals help shape countless later decisions.

    This approach reflects an important aspect of game audio more generally. Sounds do not exist independently. Their purpose is to support gameplay, reinforce feedback, communicate information, and contribute to the overall experience. A technically impressive sound that conflicts with the desired mood may ultimately be less effective than a simpler alternative.

    Over the course of his career, Vlad has contributed audio to hundreds of games. Working at this scale demands a different way of thinking about sound design. Rather than approaching every project as a completely unique undertaking, practitioners develop workflows, libraries, recording practices, and decision-making strategies that allow them to work efficiently without sacrificing quality. Consistency, organisation, and adaptability become just as important as creativity.

    The lecture provided numerous examples of how these principles operate in practice. Casual mobile games aimed at younger audiences often require sounds that are immediately understandable and emotionally positive. Designers frequently request what they describe as “cartoony” sounds, a term that may initially appear vague but which often carries fairly specific expectations. Sounds should be simple, clear, playful, and easily interpreted. Complex or highly realistic effects may actually prove less effective if they distract from the intended experience.

    Such decisions become particularly important when working on long-term projects. Vlad described his involvement with Adventure Smash, a mobile title developed by PeopleFun, the studio founded by several of the developers behind Age of Empires. What began as a relatively modest project gradually expanded into a much larger undertaking involving thousands of individual sound assets.

    One of the most interesting aspects of this discussion concerned iteration. Many sounds were revised repeatedly as the game evolved. New characters appeared. Design priorities changed. Playtesting revealed unexpected problems. Audio that seemed appropriate at one stage later required substantial modification. Rather than treating this as a failure, Vlad presented iteration as a normal and essential part of development.

    Playtesting proved especially valuable. Watching players encounter a game for the first time often revealed issues that were invisible to the development team. After listening to the same sounds hundreds or even thousands of times, designers naturally become accustomed to them. New players bring fresh perspectives. Their reactions can highlight confusing feedback, excessive repetition, or sounds that no longer fit the overall direction of the game.

    Listening to these examples, it became clear that game audio involves much more than creating sounds. It requires understanding how those sounds function within a larger interactive system. The effectiveness of an audio asset depends not only upon its quality but also upon when it appears, how frequently it occurs, and how players interpret it.

    Technical constraints provide one of the clearest examples of this mindset. Mobile games often operate within strict memory limitations. Vlad described projects containing thousands of audio assets while occupying only a few dozen megabytes of storage. Achieving this requires more than compression. Designers must think carefully about how sounds are structured, reused, combined, and implemented. Rather than viewing constraints as obstacles, the lecture suggested that they often become catalysts for creativity. Limited resources encourage solutions that are more elegant, efficient, and flexible than those developed under less restrictive conditions.

    Alongside game audio, Vlad discussed another major aspect of his professional practice: field recording. Over the years he has become increasingly involved in recording natural environments, wildlife, ambiences, and unusual sound sources. Although these activities initially developed alongside his game work, they have gradually become an important creative outlet in their own right.

    Field recording might appear separate from game development, though the lecture revealed numerous connections between the two. Recording environments, wildlife, machinery, weather, and unusual sound sources continually expands the palette available for future projects. A recording captured for no particular purpose may later become the foundation of a game sound effect, a commercial sound library, or an entirely different creative project. Field recording therefore functions not only as a creative activity in its own right but also as a long-term investment in future possibilities. This relationship between recording and design reflects the broader philosophy running throughout Vlad’s work. Resources are rarely available precisely when they are needed. Building libraries, developing skills, and collecting recordings creates opportunities that may not become useful until years later. Much of professional audio involves preparing for problems that have not yet appeared.

    What was particularly striking was the way field recording complements game audio. Time spent outdoors often provides opportunities for reflection that are difficult to find within a studio environment. Vlad described discovering solutions to creative problems while sitting quietly in a car monitoring microphones placed hundreds of metres away. Distance from the immediate pressures of production sometimes creates the mental space necessary for new ideas to emerge. The discussion of recording techniques revealed another dimension of creativity. Recording is often presented as a technical process involving microphones, recorders, and acoustic environments. Vlad acknowledged the importance of these factors while emphasising that microphone placement, recording strategies, and listening perspectives can fundamentally alter the resulting material. Small changes in approach frequently produce dramatically different outcomes.

    Perhaps the most interesting aspect of the lecture was the way it challenged simplistic distinctions between technical and creative work. Audio professionals are sometimes portrayed as belonging to one category or the other. Vlad’s experiences suggest that the reality is considerably more complicated. Technical decisions influence creative outcomes. Creative ambitions depend upon technical understanding. Success often emerges through the interaction between both.

    Questions about freelancing reinforced this point. Building a sustainable career requires skills extending far beyond audio production. Client communication, project management, marketing, financial planning, networking, and professional development all become part of daily practice. Creative expertise alone is rarely sufficient.

    Freelancing introduced another form of constraint. Unlike permanent employment, freelance work rarely provides complete stability or predictability. Projects arrive unexpectedly. Workloads fluctuate. Technologies change. Client requirements evolve. Vlad spoke candidly about periods of uncertainty throughout his career, though these experiences reinforced the same lesson visible elsewhere in the lecture. Long-term success depends less upon avoiding change than upon learning how to respond to it effectively.

    Looking back across the lecture, what emerges most clearly is a picture of audio work defined by adaptability. Technologies change. Projects evolve. Clients revise their requirements. Storage limits impose restrictions. Budgets create compromises. Development schedules compress. Yet creative ambitions remain.

    Throughout his career, George Vlad has repeatedly encountered situations in which the available resources were smaller than the desired outcome. Mobile games needed to feel larger than their budgets suggested. Limited memory had to support rich sonic worlds. Tight schedules still demanded professional results. Field recordings gathered in remote locations eventually found new purposes years later. Again and again, progress emerged through resourcefulness rather than abundance.

    For students considering careers in game audio, this may be the lecture’s most valuable lesson. Technical knowledge matters. Creative ability matters. Yet neither guarantees success on its own. Professional practice involves solving problems, working within constraints, adapting to change, and finding opportunities where others might see limitations.

    George Vlad’s career demonstrates that there is no single route into professional audio. His journey has included self-directed learning, formal education, freelance practice, field recording, game development, experimentation, and continual adaptation. Across all these experiences, one principle remained remarkably consistent. Creative work is rarely about having unlimited resources. More often, it is about recognising possibilities that remain invisible until constraints force new ways of thinking.

  • Can Sound Quality Be Measured? David Bowen on Psychoacoustics, Product Design, and Human Perception

    David Bowen

    Can sound quality be measured?

    For engineers, the question seems perfectly reasonable. Modern acoustic analysis can measure sound pressure levels, frequency content, vibration, loudness, roughness, sharpness, tonal components, and countless other characteristics. Faced with such an abundance of data, it is tempting to assume that product sound quality can ultimately be reduced to a collection of numbers. If we can measure a sound accurately enough, surely we can determine whether it is good or bad.

    During his online guest lecture for Edinburgh Napier University, David Bowen spent much of his time explaining why the answer is not nearly so simple. Across more than three decades working in acoustics, vibration, psychoacoustics, and product sound quality, Bowen has helped organisations understand how people respond to the sounds products make. Throughout a career spanning industrial research, consultancy, and product development, he has worked at the intersection of acoustics, psychoacoustics, engineering, and product design. Again and again, his examples pointed towards the same conclusion. Sound can be measured. Sound quality cannot.

    This distinction formed the foundation of the lecture. Sound quality, Bowen argued, is not a property of a product. It is a response of people. A microphone does not experience annoyance. A sound level meter does not perceive quality. Only listeners do. Understanding product sound quality therefore requires understanding both the physical sound and the human beings who hear it. Difficulties emerge as soon as engineers attempt to connect measurements to human responses. Bowen illustrated this challenge through examples in which sounds with similar measured levels produced dramatically different subjective reactions. A pure tone, broadband noise, an organ note, or a piece of industrial machinery may all produce similar sound levels, yet listeners often describe them in very different ways. Some sounds are judged pleasant. Others are irritating. Some feel powerful. Others feel weak. Part of the difficulty lies in the way human hearing operates. Psychoacoustics has demonstrated repeatedly that listeners do not experience sound in a simple or linear fashion. Sensitivity varies across frequencies. Loudness does not increase proportionally with sound pressure. Perception depends not only upon what reaches the ears but also upon how the brain interprets it. Measuring the sound itself is only part of the problem.

    Bowen illustrated this point through several examples that challenge common assumptions about listening. Human memory for loudness is surprisingly limited. When listeners hear two sounds separated by even a relatively short interval, their ability to compare levels accurately begins to deteriorate. Judgements become influenced by expectation, context, and interpretation rather than purely acoustic characteristics. Even when measurements are reliable, the perceptual processes through which listeners experience those sounds remain considerably more complex.

    For decades, researchers attempted to bridge this gap through increasingly sophisticated metrics. If sound pressure level proved insufficient, perhaps loudness would provide a better predictor. If loudness proved inadequate, perhaps perceived noisiness, roughness, sharpness, or other psychoacoustic measures would help. Each new metric offered valuable insights, yet each also revealed new limitations. Bowen discussed how the arrival of jet aircraft exposed weaknesses in existing approaches to noise evaluation, prompting the development of measures intended to capture perceived noisiness more effectively. Those measures improved predictions in some contexts while proving less successful in others. Similar challenges emerged across industrial machinery, transportation systems, and consumer products. As soon as one perceptual factor appeared understood, another emerged. Listening proved stubbornly resistant to simple description.

    Bowen’s career spans a period during which acoustics increasingly recognised that physical measurements alone could not explain human responses. Successive generations of psychoacoustic metrics attempted to narrow the gap between measurable sound and perceived quality. Each represented an improvement upon what came before, though none provided a complete solution. Human perception remained influenced by context, expectation, memory, meaning, and experience. The history of product sound quality therefore became, in part, a history of increasingly sophisticated attempts to understand how people listen. Similar problems emerge elsewhere. A piano recording played backwards retains many of its measurable characteristics, yet listeners immediately perceive something fundamentally different. What sounds like a piano becomes something closer to an organ. Human listeners detect meaningful changes that conventional measurements often struggle to explain. Again and again, perception proves more complicated than measurement.

    If measurements alone cannot fully predict how people will respond, a difficult question follows. How should products be designed?

    For Bowen, the answer lies in listening. Much of the lecture focused on sound quality jury testing, a methodology that places human listeners at the centre of the evaluation process. Rather than asking which sound measures best, researchers ask which sound people prefer, which sound communicates particular qualities, and which sound supports the intended experience of a product.

    This creates an interesting tension. Engineers naturally seek measurements. Manufacturers want targets that can be specified, monitored, and improved. Product development processes favour quantities that can be compared and optimised. Yet listeners remain the ultimate judges of quality. No matter how sophisticated a measurement becomes, a product succeeds or fails according to how people experience it. Jury testing therefore emerged not as a rejection of engineering but as a recognition that engineering alone could not answer every question.

    Carefully designed listening tests provide information that measurements alone cannot. This approach complements rather than replaces traditional acoustic analysis. Measurements help researchers understand what a product is doing acoustically. Listening tests help them understand how people respond. Product sound quality emerges through the relationship between these two perspectives. Designing listening tests of this kind is far from straightforward. Participants must be selected carefully. Stimuli need to be prepared consistently. Presentation order can influence responses. Questions must be designed in ways that avoid leading participants towards particular conclusions. Statistical analysis becomes essential if meaningful patterns are to emerge from the resulting data. Throughout the lecture, Bowen emphasised that listening tests require as much methodological care as any engineering measurement.

    One particularly interesting aspect of this work involves the creation of what Bowen described as virtual products. Rather than constructing numerous physical prototypes, researchers can isolate individual sound components and manipulate them independently. Motor noise, airflow noise, pump sounds, valve sounds, and other elements can be adjusted before being recombined into new versions of the product. Listeners can then evaluate these alternatives, allowing researchers to explore how specific design decisions influence perceived quality without repeatedly redesigning the product itself.

    One of the lecture’s most illuminating examples involved front-loading washing machines. Modern washing machines generate a wide range of sounds, including motor noise, water movement, pumping systems, valves, and the movement of clothes within the drum. Traditional noise control might focus simply on reducing these sounds wherever possible. Bowen’s research adopted a different approach. Rather than treating the machine as a single noise source, the different sounds produced during filling, washing, draining, and spinning were analysed separately. Each stage introduced its own acoustic characteristics and potential design challenges. Water movement, pump operation, motor behaviour, valve activity, and the interaction between clothes and the drum all contributed differently to listener perceptions. Individual sound components were isolated and manipulated. Participants evaluated these variations through listening tests, allowing researchers to identify which sounds influenced acceptability most strongly.

    The resulting data could then be analysed using statistical models that linked changes in specific sound components to listener ratings. One of the most interesting aspects of this work involved the creation of response-surface models that allowed engineers to visualise how perceived quality changed as different sound characteristics were adjusted. Rather than producing a simple pass-or-fail result, the models created maps of possible design outcomes. Engineers could explore how increasing one characteristic while reducing another might influence listener responses. Product sound quality rarely involves finding a single perfect solution. Designers must balance acoustic quality against manufacturing constraints, performance requirements, reliability considerations, and cost limitations. Statistical modelling provides a way of navigating these trade-offs while retaining a clear understanding of how design decisions influence perception.

    Similar principles appeared in Bowen’s work on vacuum cleaners. Consumers often claim that they want quieter products, yet a vacuum cleaner that becomes almost silent introduces a different problem. Users may begin to question whether it is working properly. Certain sounds communicate power, airflow, and cleaning effectiveness. Eliminating every sound is not necessarily desirable. In this case, the challenge is not simply reducing noise but preserving those aspects of the sound that contribute positively to the user’s perception of performance. What emerges from Bowen’s examples is a view of product sound that differs significantly from traditional approaches to noise control. Sounds are not merely by-products of mechanical systems. They communicate information about performance, condition, reliability, quality, and identity. A washing machine, a vacuum cleaner, a refrigerator, and an aircraft each occupy different places in people’s lives. Listeners bring different expectations to each. A refrigerator should not sound like a lawnmower. Equally, a lawnmower should not sound like a refrigerator. The challenge is therefore not simply reducing sound, but designing sounds that make sense within a particular context.

    Seen in this light, product sound quality becomes a remarkably human problem. Engineers can measure sound with extraordinary precision. Researchers can develop increasingly sophisticated psychoacoustic models. Statistical techniques can reveal relationships between acoustic characteristics and listener preferences. Yet none of these tools removes the need to understand people. Sound quality emerges not from products alone but from the relationship between products and listeners. What emerged from the lecture was a challenge to a familiar engineering instinct. Faced with a difficult problem, engineers naturally seek better measurements. Bowen’s work suggests that measurements remain essential, though they are not enough on their own. Product sound quality exists at the point where physical acoustics meets human perception. This creates an unusual situation. Few areas of engineering depend so heavily upon subjective judgement while simultaneously demanding rigorous measurement. Product sound quality requires microphones, analysers, statistical models, listening tests, psychoacoustic theory, and human listeners. Remove any one of these elements and the picture becomes incomplete.

    Perhaps this is why the question that opened the lecture remains so difficult to answer. Can sound quality be measured? David Bowen’s career suggests that the answer is both yes and no. Sounds can be measured with extraordinary precision. Human responses can be studied, modelled, and predicted. Yet quality itself ultimately emerges through experience. The most successful products are not necessarily the quietest products, nor the products with the best acoustic measurements. They are the products whose sounds make sense to the people who use them. In the end, product sound quality is not really about sound at all. It is about understanding listeners.

  • Who Are We Designing For? Professor Bruce Walker on Sound, Accessibility, and Human-Centred Design

    Bruce Walker

    Who are we designing for?

    At first glance, the answer appears obvious. Designers create products for users. Engineers build systems to help people accomplish tasks. Technology exists to solve problems. Yet during his online guest lecture for Edinburgh Napier University, Professor Bruce Walker repeatedly returned to examples suggesting that the answer is often more complicated than it first appears. Again and again, he described situations in which technically impressive systems failed to account for the realities of the people expected to use them. A solution might work perfectly according to engineering specifications while proving frustrating, distracting, or simply undesirable in practice. Across projects involving navigation, education, museums, accessibility, sonification, and auditory interfaces, Walker argued that successful design begins not with technology but with understanding human needs.

    Sound provided the central thread running through the lecture. Many people associate auditory interfaces with simple alerts and alarms. Computers beep. Phones ring. Vehicles issue warning tones. Yet Walker demonstrated that sound can serve far more sophisticated purposes. It can guide navigation, communicate data, support education, enable accessibility, assist decision-making, and provide entirely new ways of interacting with technology. These possibilities emerge not from adding sounds indiscriminately but from carefully considering what information people need, when they need it, and how they can use it effectively.

    This emphasis on understanding users before designing solutions appeared throughout the lecture. These broader questions became particularly visible in Walker’s discussion of the SWAN project, the System for Wearable Audio Navigation. The goal was deceptively simple. Could sound help people navigate when they could not rely on vision? Blind users are an obvious example, though Walker deliberately framed the problem more broadly. Firefighters moving through smoke-filled buildings, soldiers operating at night, divers underwater, and workers engaged in visually demanding tasks may all find themselves unable to look or unable to see. Rather than designing exclusively for a particular group, the project focused on a shared perceptual challenge.

    Approaching the problem in this way immediately changed the nature of the design process. Navigation might seem like a familiar activity, something most people perform every day without conscious thought. Yet once the team began analysing the task carefully, they discovered layers of complexity hidden beneath the surface. People need to know where to go, though they also need information about obstacles, changes in terrain, landmarks, hazards, and points of interest. Navigation is not simply a matter of moving from one location to another. It involves understanding an environment while continuously making decisions within it.

    The resulting system employed spatialised audio beacons that users could follow through space. Walker compared them to a virtual carrot suspended ahead of the listener. Rather than receiving a sequence of verbal instructions, users simply moved towards a sound source. When a turn became necessary, the sound shifted position accordingly. The concept appears elegant partly because it exploits abilities listeners already possess. Humans are remarkably good at locating sounds. Rather than teaching an entirely new interaction technique, the system builds upon existing perceptual skills.

    What makes the project particularly revealing, however, is the extent to which users shaped its development. Early assumptions frequently proved incorrect. Designers initially believed they should describe surface transitions in detail, informing users when they were moving from pavement to grass or from one surface to another. Users quickly pointed out that such information was often redundant. They already knew where they were standing. What mattered was knowing what was coming next. By listening carefully to users rather than insisting upon their original assumptions, the team produced a more efficient and less intrusive system.

    This pattern appeared repeatedly throughout the lecture. Successful auditory interfaces emerge through collaboration, observation, and evaluation rather than through technological enthusiasm alone. Walker repeatedly emphasised the importance of testing systems with real users performing real tasks. Maps of navigation paths, performance data, error rates, and subjective feedback all played important roles in understanding whether a design genuinely worked. The question was never whether a system could be built. The question was whether it improved people’s ability to accomplish what they were trying to do.

    Perhaps nowhere was this philosophy more apparent than in the team’s work with bone-conduction audio. Navigation systems often rely upon stereo or spatial audio, which typically requires headphones. Yet many users rejected conventional headphones immediately. Blind people rely heavily upon environmental sounds. Firefighters need to hear what is happening around them. Covering the ears solved one problem while creating another. Rather than treating this as an unavoidable limitation, the researchers reconsidered the entire system. Bone-conduction devices allowed spatial audio to be delivered without blocking environmental awareness. Once again, the solution emerged not from pursuing technology for its own sake but from understanding the realities of users’ lives.

    Walker extended these ideas beyond navigation into the design of auditory menus and interfaces. Modern computer systems contain large amounts of information organised visually through menus, icons, scroll bars, and navigation structures. Translating these elements into sound presents significant challenges. Simply reading everything aloud through text-to-speech quickly becomes inefficient and frustrating. Long contact lists, extensive music libraries, and complex software systems require more sophisticated approaches.

    Many of the solutions developed by Walker and his collaborators demonstrate an intriguing combination of technical ingenuity and psychological insight. Spearcons, for example, compress spoken words into highly abbreviated auditory cues that users rapidly learn to recognise. Spindex systems provide indexing sounds that help listeners move efficiently through large collections without hearing every item in sequence. Whispered speech can indicate unavailable menu items while preserving the overall structure of an interface. These innovations are clever, though their importance lies less in their novelty than in their effectiveness. Each emerged through extensive experimentation designed to determine what users actually understood and preferred. What makes this work particularly interesting is that it was never simply about making visual interfaces audible. Instead, it explored how auditory interfaces might exploit the strengths of listening itself. Users do not experience sound in the same way they experience vision, and successful interfaces acknowledge that difference rather than treating audio as a substitute for a screen.

    The lecture repeatedly highlighted a distinction between engineering and design. Engineering can produce systems that function correctly. Design concerns whether people will actually want to use them. Walker discussed several examples of technically impressive systems that failed precisely because they neglected this distinction. Some provided too much information. Others demanded excessive attention. Many ignored the broader contexts within which people operate. A system may be capable of delivering enormous amounts of information, though that does not necessarily mean people want to receive it.

    Questions of accessibility broadened these concerns further. Much of Walker’s work focuses on enabling participation rather than merely providing access. He described projects supporting education for blind students, making scientific information more accessible, and improving experiences within museums and aquariums. These examples revealed another recurring theme. Accessibility is not simply about removing barriers. It is about ensuring that people can engage meaningfully with experiences, opportunities, and ideas.

    His work within museums and aquariums illustrates this particularly well. Modern cultural institutions often provide physical access while leaving informational and emotional access largely unresolved. A blind visitor may be able to enter an aquarium, though the experience remains fundamentally visual. Standing before an enormous tank filled with whale sharks and rays offers little if the most important aspects of the exhibit remain inaccessible. Walker’s team explored ways of using tracking systems, sonification, narration, and auditory displays to communicate not only what was present but what was happening. The objective was not merely to describe the environment. It was to create opportunities for engagement, curiosity, and wonder.

    Educational projects revealed similar concerns. Throughout the lecture, Walker repeatedly returned to the distinction between access and participation. Providing information is only one part of inclusion. Students also need opportunities to explore, question, discover, and develop understanding independently. Whether working with scientific data, classroom materials, museum exhibits, or public spaces, the challenge remained remarkably consistent. How can information be presented in ways that support meaningful engagement rather than passive reception? Accessibility, in this sense, becomes a question of design quality rather than a specialised feature introduced at the end of a project.

    Underlying many of these projects is the field of sonification, the practice of representing information through sound. Walker described sonification as both a design challenge and a research problem. Any attempt to translate data into sound requires decisions about mapping, scaling, timing, context, and interpretation. Should temperature be represented through pitch, loudness, or tempo? How should complex information be organised so that listeners can understand it? These questions have no universal answers. Effective solutions depend upon understanding both the data and the people expected to interpret it.

    One reason sonification remains challenging is that many listeners have relatively little experience interpreting information through sound. Graphs, charts, and maps are familiar cultural forms. Auditory representations are far less common. Designers therefore need to balance learnability with expressiveness. A system may communicate information accurately while remaining difficult to interpret. Conversely, a system may sound appealing while conveying very little. Walker’s research repeatedly demonstrated that effective sonification emerges through iterative testing with users rather than through theoretical assumptions alone.

    Such challenges reveal why Walker remains sceptical of simplistic approaches to auditory design. Throughout the lecture, he criticised the tendency to reduce audio interfaces to collections of arbitrary beeps and alerts. Sound possesses enormous communicative potential, though realising that potential requires careful thought. Designers must consider attention, context, usability, aesthetics, cultural expectations, and human behaviour. A sound that performs well in a laboratory may fail completely in everyday life. A technically accurate representation may prove ineffective if users cannot interpret it.

    The phrase “beeps and bops” became a useful shorthand for this problem. Many technologies employ sound only at the most superficial level, relying upon alerts, warnings, and notifications while overlooking the richer possibilities of auditory interaction. Walker’s work points towards a broader conception of sound, one capable of supporting navigation, exploration, learning, communication, and discovery. The challenge is not simply adding sound to technology. It is designing meaningful auditory experiences.

    Towards the end of the discussion, Walker reflected on what he described as a “failure of imagination” in technology design. Sometimes designers struggle to imagine how people actually live with technologies. At other times, users struggle to imagine possibilities that do not yet exist. Successful innovation requires navigating both challenges simultaneously. Revolutionary technologies rarely emerge through user requests alone. Yet genuinely useful technologies also cannot emerge through engineering in isolation. Design becomes a process of bridging these perspectives.

    Looking back across the lecture, what emerges most clearly is not a story about auditory interfaces but a broader philosophy of design. Sound happens to be the medium through which Professor Walker explores these questions, though the underlying principles extend much further. Navigation systems, auditory menus, museum exhibits, educational technologies, sonification projects, and accessibility tools all reveal the same challenge. Technologies succeed not when they demonstrate technical sophistication but when they become meaningful parts of human activity.

    Perhaps this is why Walker repeatedly resisted framing accessibility as a specialised concern. The challenges faced by blind users, firefighters, drivers, students, museum visitors, and countless others often reveal broader truths about human interaction with technology. Designing for specific needs frequently produces insights that benefit everyone. When designers stop asking what a system can do and start asking what people need, entirely new possibilities begin to emerge.

    Throughout the lecture, examples ranging from spatial navigation to aquarium exhibits pointed towards the same conclusion. Successful technologies rarely begin with devices, software, algorithms, or interfaces. They begin with people. Understanding how people listen, learn, move, explore, communicate, and make decisions provides the foundation upon which everything else is built.

    For Professor Bruce Walker, the future of auditory interfaces does not lie in adding more sounds to the world. It lies in understanding how sound can help people navigate, learn, communicate, discover, and participate more fully in the experiences around them. The technology matters. The research matters. The engineering matters. Yet each ultimately serves a more fundamental question, one that quietly shaped the entire lecture from beginning to end: who are we designing for?

  • What Did They Say? Gary Bourgeois on Dialogue, Attention, and the Art of Film Mixing

    Gary Bourgeois

    What happens when an audience misses a line of dialogue?

    At first glance, the consequences seem relatively minor. A viewer leans towards a friend. Someone quietly asks for clarification. A sentence is repeated. Yet during his online guest lecture for Edinburgh Napier University, veteran re-recording mixer Gary Bourgeois suggested that this moment reveals something important about the relationship between sound and storytelling. The audience has stopped following the narrative and started thinking about the soundtrack. For Bourgeois, whose career spans more than five decades across film, television, music, and streaming media, preventing that moment has remained one of the central responsibilities of a mixer.

    This might appear surprising. Popular discussions of film sound often focus on spectacle. We talk about explosive action sequences, immersive surround sound systems, powerful musical scores, and increasingly sophisticated technologies. Yet Bourgeois repeatedly returned to a much simpler idea. Sound exists to support communication. Every creative and technical decision ultimately serves the story. If audiences cannot understand what matters at the moment it matters, even the most technically impressive soundtrack has failed in its primary task.

    Throughout the lecture, Bourgeois described film mixing as a process of guiding attention. A finished soundtrack may contain dialogue, Foley, ambience, music, effects, backgrounds, transitions, and countless other elements. These sounds do not all demand equal attention simultaneously. Their relationships are constantly shifting. During a conversation, dialogue may occupy the foreground while music retreats slightly into the background. During a dramatic reveal, music may briefly become the dominant element. An action sequence may allow effects to take centre stage before returning attention to character and narrative. Mixing therefore involves much more than balancing levels. It involves shaping the audience’s experience of a story.

    This perspective helps explain why Bourgeois places such importance on dialogue. Writers spend months or years developing scripts. Actors devote enormous effort to performance. Directors construct scenes around the communication of information, emotion, and character. If a crucial line becomes unintelligible, the audience loses access to part of that work. More importantly, they momentarily leave the fictional world. Instead of thinking about the characters, they begin thinking about the soundtrack. The illusion is interrupted.

    One of the most interesting aspects of the lecture concerned the relationship between film mixing and human perception. During the discussion, we explored the idea that many mixing decisions effectively replicate forms of selective attention that listeners perform naturally. In everyday life, people can focus on a particular voice within a crowded room, follow a conversation in a noisy taxi, or attend to one sound source while ignoring dozens of others. The auditory system constantly prioritises information. Bourgeois agreed that much of professional mixing involves recreating these perceptual priorities for audiences. The mixer helps listeners focus on what matters without drawing attention to the process itself.

    Seen in this light, many familiar audio tools acquire a different significance. Equalisation is not simply a way of adjusting frequencies. Compression is not merely a method of controlling dynamics. Reverb is not only about creating a sense of space. These processes become valuable insofar as they help establish relationships between sounds. A dialogue track may require subtle equalisation to distinguish it from surrounding ambience. A sound effect may need certain frequencies reduced so that speech remains intelligible. A reverberant environment may need careful shaping to preserve clarity. The technical operations matter, though their ultimate purpose remains perceptual. Ultimately, they help prevent the audience from asking the question that opened the lecture. What did they say?

    Several examples from Bourgeois’ career illustrated this philosophy particularly well. Large-scale productions such as Transformers are often associated with spectacle, scale, and sonic intensity. Audiences remember giant robots, enormous impacts, and dense layers of sound. Yet Bourgeois described how even the most elaborate action sequences depend upon careful control of attention. One memorable example involved introducing a single frame of silence immediately before an explosion. The audience never consciously notices this interruption. Nevertheless, the brief absence of sound creates a perceptual contrast that makes the subsequent impact feel considerably larger. The effect depends not on additional volume but on the way listeners perceive change.

    Examples such as this reveal a recurring principle running throughout the lecture. Effective sound design often depends less upon adding material than upon managing relationships between existing elements. A soundtrack filled continuously with dramatic gestures eventually loses its ability to surprise. Contrast becomes difficult. Emphasis becomes impossible. Restraint therefore plays an important role within the mixer’s craft. Sometimes the most effective decision is deciding what not to hear.

    This concern with attention also shapes Bourgeois’ attitude towards immersive audio formats such as Dolby Atmos. The technology provides extraordinary creative possibilities. Sounds can move through three-dimensional space with remarkable precision. Environments can become more detailed and immersive than ever before. Yet Bourgeois consistently framed these capabilities in relation to storytelling rather than technology. An Atmos mix succeeds when it helps audiences engage more deeply with a scene. It fails when the technology becomes the focus of attention itself. More speakers do not automatically produce better storytelling. The same principles still apply. Audiences need to understand what matters and why it matters.

    A particularly revealing section of the lecture explored Bourgeois’ lifelong curiosity about listening. Long before spatial audio became a major industry topic, he was conducting informal experiments with binaural recording, environmental acoustics, and perceptual phenomena. Rather than treating recording purely as a professional necessity, he approached it as an opportunity to investigate how sound behaves.

    One story involved recording a stream in rural Canada. Expecting to capture clear differences between close, medium, and distant perspectives, he recorded the same source from multiple locations. When he returned to the studio, however, the recordings sounded remarkably similar. What initially appeared disappointing became an important lesson. Distance is often communicated less by direct sound than by reflections, environmental interactions, and contextual cues. The stream itself had changed very little. The surrounding environment had provided most of the information listeners normally use to judge distance. Stories such as this reveal another dimension of Bourgeois’ approach. Technical expertise emerges not only from formal training but also from observation. Throughout the lecture, he repeatedly emphasised the importance of listening carefully to the world. Many of the insights that shaped his professional practice originated in moments of curiosity rather than commercial necessity. A recording experiment, an unusual acoustic environment, or an unexpected perceptual effect could become the foundation for future creative decisions.

    His reflections on Canada extended this theme further. Bourgeois noted that a surprisingly large number of Hollywood film mixers originate from Canada. While partly humorous, the observation led into a broader discussion about listening environments. Growing up in quieter surroundings encouraged attention to subtle acoustic details, spatial relationships, and environmental sounds. Whether or not this fully explains the phenomenon, the anecdote reinforced a larger point. Listening is not a passive activity. It is a skill developed through experience, practice, and sustained attention.

    The conversation eventually turned towards emerging technologies, particularly artificial intelligence. Here again, Bourgeois adopted a perspective shaped by decades of professional experience. Throughout his career he has witnessed repeated technological transformations. Analogue workflows gave way to digital systems. New recording formats emerged. Distribution platforms changed. Entire production processes evolved. Each transition created uncertainty alongside opportunity.

    Rather than treating AI as fundamentally different from earlier technological developments, Bourgeois viewed it as another stage within a continuing process of change. New tools will inevitably alter professional practice. Some tasks may become easier. Others may disappear entirely. Yet the underlying challenge remains remarkably consistent. Practitioners must learn how new technologies work, understand their limitations, and identify meaningful ways of applying them. Avoiding change rarely proves productive. Understanding it usually does.

    Looking back across the lecture, what emerges most clearly is a conception of mixing rooted in attention. Compressors, equalisers, reverbs, Atmos systems, loudness standards, recording technologies, and AI tools all matter. Yet they matter only insofar as they help audiences remain connected to a story. Bourgeois repeatedly returned to the same fundamental question. Can the audience understand what matters at the moment it matters?

    Many discussions of sound focus primarily on technology. Gary Bourgeois offered a useful reminder that technology is ultimately a means rather than an end. The purpose of a soundtrack is not to demonstrate technical sophistication. Its purpose is to support communication, emotion, and narrative understanding. The most successful mixes often pass unnoticed precisely because they allow audiences to remain fully absorbed in the world unfolding before them.

    Perhaps that is why the simple question that opened the lecture remains so revealing. What happens when an audience misses a line of dialogue? For Bourgeois, the answer extends far beyond a few misunderstood words. It represents a brief fracture in the relationship between story and listener. Much of the mixer’s craft is devoted to preventing that fracture from occurring. Every adjustment, every balance decision, every technical process ultimately serves the same goal: helping audiences hear not merely the sounds of a film, but the story those sounds are trying to tell.