Category: Audio Perception

  • How Do You Give a Monster a Voice? Matthew Collings on Performance, DSP, and Creature Sound Design

    Matthew Collings

    Every audience knows what a dinosaur sounds like. Dragons roar, aliens snarl and giant monsters shake cinemas with impossibly deep voices. Yet none of these creatures has ever existed. Every sound associated with them has been created from scratch, yet audiences instinctively accept them as real. Creating that illusion is one of the most fascinating challenges in sound design.

    This was the starting point for Matthew Collings’ guest lecture on creature vocalisation and real-time sound design. As an audio programmer at Crotos, the company behind Dehumanizer, Collings explored how advances in digital signal processing are changing the way creature voices are created. The lecture, however, was about much more than a single piece of software. It invited the audience to consider a broader question: how can technology extend creative expression without diminishing the artistic judgement that remains central to sound design?

    Creating convincing creature voices has traditionally been one of the most demanding areas of audio production. Designers layer recordings of animals, manipulate pitch, combine multiple processing techniques and painstakingly synchronise every vocalisation with an animated character. The results can be extraordinary, but they depend upon considerable time, technical expertise and countless creative decisions made after the original recording has taken place.

    To illustrate this established approach, Collings showed a video in which legendary sound designer Ben Burtt described the creation of Chewbacca’s voice for Star Wars. Rather than relying on a single animal recording, Burtt assembled the character’s voice from bears and numerous other animals, selecting each recording because it conveyed a particular emotional quality. Some sounds suggested affection, others frustration, aggression or excitement. Through careful editing and synchronisation, these individual elements became the distinctive voice of a character that had never existed. The example serves as a reminder that audiences respond to emotion and personality as much as acoustic realism.

    Collings argued that this creative principle remains unchanged, even as production workflows evolve. Traditionally, creature voices were assembled through careful editing once recording had finished. Increasingly, however, designers can manipulate voices in real time, hearing the transformed result immediately as they work. Instead of constructing every roar, growl or vocal gesture afterwards, they can shape those sounds during the act of performance itself. Digital signal processing therefore becomes more than a post-production tool. It becomes an expressive instrument, allowing sound design to move beyond editing and towards live creative interaction.

    That shift has implications far beyond creature effects. It changes the relationship between performer, sound designer and technology, bringing them together within a single creative process. The remainder of the lecture explored how real-time processing makes this possible, why it offers advantages over traditional workflows, and what this evolution might mean for the future of sound design in film, games and immersive media.

    If real-time processing changes the way creature voices are produced, what actually makes this approach different from traditional sound design? At first glance, the technology itself seems familiar. Pitch shifting, convolution, granular processing and modulation have all been part of the sound designer’s toolkit for many years. As Collings demonstrated, the real innovation is not the individual processes but the way they can be combined and controlled during a live performance. Instead of waiting until recording has finished, designers can hear the transformed voice immediately and respond to it in the moment.

    This represents a significant departure from established workflows. Traditionally, creature vocalisations have been built through accumulation. Designers layer recordings of animals, manipulate pitch, combine multiple processing techniques and synchronise every sound with an animated performance. The results can be extraordinarily rich and expressive, but they are achieved through careful editing and refinement after the original recording has taken place. Every decision is made retrospectively.

    Collings demonstrated an alternative way of working. Rather than treating the recorded voice as material to be reconstructed later, the performer hears the transformed character immediately while responding to the animation. Every breath, hesitation and change in vocal expression influences the processed sound as it happens. Recording, listening and refinement become part of a continuous creative cycle rather than a sequence of separate production stages.

    Watching this process unfold was particularly revealing. It felt less like observing a conventional editing session and more like watching a musician perform with an unfamiliar instrument. The software undoubtedly transformed the incoming voice, but the character emerged through timing, expression and continual adjustment. Technology extended the performer rather than replacing them.

    The individual processing techniques shown during the lecture each contributed something different to the final result. Pitch shifting altered the apparent scale and physical presence of a creature. Granular processing introduced texture and unpredictability, while convolution blended characteristics borrowed from recordings of the natural world. None of these techniques was presented as a complete solution. Instead, convincing creature voices emerged through the interaction of multiple subtle transformations, all shaped by a responsive vocal performance and careful listening.

    This also explains why recordings from the natural world remain so important. Animal vocalisations contain acoustic qualities that listeners instinctively recognise, even when heavily transformed. Rather than imitating individual species directly, designers borrow elements from familiar sounds and reshape them into voices that feel plausible despite belonging to imaginary creatures. The result is neither entirely natural nor entirely artificial. It occupies a convincing space somewhere between the two.

    Perhaps the most striking aspect of the demonstration was how quickly the technology faded into the background. Audiences do not hear convolution, granular synthesis or pitch shifting. They hear a creature reacting to its environment. For the practitioner, the individual processing modules matter far less than the expressive possibilities they create. The objective is not to showcase sophisticated digital signal processing, but to shape a voice that feels authentic within the fictional world it inhabits.

    By reframing digital signal processing as a creative instrument rather than simply a collection of effects, Collings presented real-time sound design as more than a technical improvement. It represents a different way of thinking about the craft itself, where listening, experimentation and creative judgement become increasingly intertwined with the act of performance.

    The implications of this approach extend well beyond creature vocalisation. Although Collings’ demonstrations centred on monsters and dinosaurs, the underlying principles apply wherever sound must respond dynamically to performance. Film, games and immersive media increasingly require audio that can evolve alongside the action rather than being fixed during post-production. As production workflows become more interactive, sound design is beginning to shift from constructing sounds after the event towards shaping them as creative decisions unfold.

    This evolution also changes the relationship between performer and sound designer. Traditionally, these roles have been separated. An actor delivers a vocal performance, while the sound designer interprets and transforms it later through editing and processing. In Collings’ demonstrations, however, those boundaries became less distinct. The transformed voice could be heard immediately, allowing every vocal gesture to influence the next. Performer, practitioner and software formed part of a continuous dialogue in which listening, adjustment and expression happened simultaneously.

    Watching this process unfold felt surprisingly different from observing a conventional recording session. Rather than painstakingly constructing every vocalisation afterwards, Collings responded to the animation as it played, continually refining the sound in real time. The software functioned as an expressive instrument, yet the musicality came from the person using it. Timing, pacing and emotional intent remained far more influential than any individual processing technique.

    Perhaps the most revealing aspect of the lecture was its emphasis on listening. Numerous processing techniques were demonstrated, but none was presented as a formula for creating convincing creature voices. Instead, expressive results emerged through continual refinement, balancing multiple transformations until the voice felt appropriate for the character and the scene. The demonstrations reinforced a familiar truth within sound design: technical knowledge provides possibilities, but careful listening determines which possibilities are worth pursuing.

    This way of working also encourages a more exploratory creative process. Because ideas can be evaluated immediately, designers are free to experiment with subtle variations in delivery, processing and timing without committing to lengthy post-production workflows. Some ideas will inevitably be discarded, but others may reveal unexpected qualities that would have been difficult to discover through a more linear editing process. Real-time processing therefore supports experimentation not by making creative decisions easier, but by making them easier to explore.

    Looking beyond creature vocalisation, Collings’ demonstrations suggest a broader evolution in the practice of sound design. As digital signal processing becomes increasingly responsive, practitioners are no longer confined to refining performances after recording has ended. They can participate in the creative act itself, shaping sound as it develops while drawing upon the same critical listening, imagination and editorial judgement that have always defined exceptional work. The tools may be changing, but the craft remains firmly rooted in human perception and creative decision-making.

    Although Dehumanizer provided the focus for the demonstrations, the lecture ultimately explored a much broader evolution in sound design practice. Throughout the session, Collings showed that the most effective technology is rarely the most conspicuous. Whether refining a creature’s vocalisation, combining different processing techniques or responding to an animated sequence in real time, the objective remained the same: to create voices that audiences accepted instinctively as part of a believable world.

    One consequence of this approach is that experimentation becomes far more immediate. Presets provide useful starting points, while more advanced controls allow practitioners to shape every aspect of the resulting sound. Instead of investing time constructing complex processing chains before hearing the outcome, designers can move quickly between ideas, evaluating and refining them as they work. The software therefore accelerates exploration without diminishing the value of experience or technical understanding.

    The lecture also reinforced an enduring lesson about creature vocalisation itself. Convincing voices rarely emerge from a single recording or a single processing technique. They develop through the careful interaction of vocal performance, recordings from the natural world, digital signal processing and critical listening. Each contributes something different, but none is sufficient on its own. The illusion succeeds because these elements are brought together in support of character, emotion and storytelling.

    Taken together, Collings’ demonstrations point towards a broader change in the discipline. As real-time processing becomes increasingly sophisticated, the traditional boundaries between recording, editing and performance begin to blur. Sound designers are no longer limited to refining material after a recording session has ended. Increasingly, they are able to shape expressive performances as they unfold, responding directly to performers, animation and narrative in the moment.

    This does not redefine the purpose of sound design, but it does reshape the role of the practitioner. The craft continues to depend upon careful listening, imagination and aesthetic judgement. What is changing is the point at which those skills are applied. Rather than waiting until production has finished, they become part of the performance itself.

    That is perhaps the lecture’s most enduring contribution. Rather than presenting another collection of digital effects, Collings demonstrated how real-time processing is reshaping the practice of sound design while leaving its creative foundations unchanged. If the opening question was how imaginary creatures can sound so believable, the answer lay not in software alone, but in giving sound designers increasingly expressive ways to listen, respond and shape performances as they unfold.

  • How Does Sound Make Virtual Reality Believable? Varun Nair on Presence, Perception, and Designing for Immersion

    Varun Nair

    Why does virtual reality sometimes feel astonishingly real, while other experiences never quite convince us? Advances in display technology have made virtual environments increasingly convincing visually, yet realism depends upon far more than what we see. The sounds that surround us help define where we are, how large a space feels and whether the events unfolding around us seem believable. Even a visually impressive virtual world can lose its sense of presence if the accompanying audio fails to match the experience.

    Varun Nair explored this relationship between sound and perception in his guest lecture on immersive audio and virtual reality. Rather than concentrating solely on the technologies behind virtual reality, he focused on a more fundamental question: what makes people believe that a virtual world is real? His answer was not a single piece of software or hardware, but the careful combination of perceptual cues that allow listeners to interpret an artificial environment as though they were genuinely present within it.

    A simple example illustrates the point. Imagine standing beside a helicopter as its rotors begin to turn. Even without looking directly at the aircraft, the changing pitch, movement and intensity of the sound immediately communicate what is happening. If those acoustic cues fail to match the visual scene, the illusion quickly begins to collapse. The helicopter may still look convincing, but the experience no longer feels authentic. Sound is not simply an accompaniment to the image; it helps establish whether the world itself is believable.

    This relationship becomes even more significant in virtual reality because the listener is no longer observing events from a fixed position. Every movement of the head changes the perspective, and the soundtrack must respond immediately if the illusion is to be maintained. Instead of accompanying a predetermined sequence of images, audio becomes an active part of the experience, continuously adapting to the listener’s actions.

    Sound helps listeners understand where they are, what is happening around them and where their attention should move next. These functions have always been part of effective sound design, but virtual reality places them at the centre of the experience. Spatial relationships that might previously have enhanced realism now become essential for maintaining the illusion that the listener occupies a coherent physical environment.

    This is one reason why binaural audio has become such an important topic in immersive media. By recreating the subtle differences between the sounds arriving at each ear, binaural rendering allows listeners wearing headphones to perceive sound sources as existing around them rather than inside their heads. Yet, as Nair emphasised throughout the lecture, convincing immersion depends upon much more than simply positioning sounds in three-dimensional space.

    Believable virtual worlds are not built from individual technologies. They emerge when every perceptual cue supports the same experience. The challenge for immersive audio is therefore not simply to reproduce sound accurately, but to persuade listeners that the environment surrounding them behaves as a coherent and convincing world.

    The remainder of the lecture explored how this illusion is constructed, why the human brain depends upon multiple acoustic cues rather than any single technology, and what these ideas mean for the future of sound design in virtual reality.

    The idea that immersive audio depends upon many different perceptual cues naturally raises another question. What are those cues, and why does the brain rely on so many of them at once? Popular discussions of virtual reality often reduce three-dimensional audio to binaural rendering, as though placing sounds around the listener is enough to create convincing immersion. Nair argued that this captures only part of the picture. Positioning a sound correctly is important, but convincing listeners that it genuinely belongs within a virtual world requires a much richer combination of acoustic information.

    Binaural audio provides the foundation for this process because it recreates the way sound naturally reaches the ears. Every sound arriving at the listener is subtly shaped by the head, shoulders and outer ears before reaching the eardrum. Over a lifetime, the brain learns to interpret these tiny differences, allowing people to judge whether a sound is above them, behind them or approaching from one side. Virtual reality exploits these learned perceptual cues so that sounds appear to occupy positions within the surrounding environment rather than simply playing through a pair of headphones.

    One of the more surprising points in the lecture was that none of this is particularly new. Researchers have investigated binaural hearing for decades, and commercial implementations existed long before today’s virtual reality headsets. What has changed is the wider technological landscape. Modern processors are now capable of performing the necessary signal processing in real time, while affordable head-mounted displays continuously track the listener’s movements. As a result, ideas that were once largely confined to specialist research have become practical tools for everyday production.

    Yet Nair was equally clear that accurate positioning alone does not create a believable acoustic environment. Standing inside a cathedral, people hear much more than a nearby voice. Reflections from distant walls, the gradual decay of reverberation and subtle changes in tone all contribute to an immediate understanding of the surrounding space. Remove those cues and the voice may still come from the correct direction, yet it no longer feels connected to the environment. A sound can be positioned accurately without ever feeling present.

    This distinction led to one of the lecture’s central ideas: externalisation. Successful spatial audio should persuade listeners that sounds exist outside their heads, occupying the same world as the visual scene. That impression emerges from many different cues working together. Early reflections reveal the boundaries of a room. Reverberation communicates its size and acoustic character. Air absorption subtly alters distant sounds, while Doppler shifts reinforce movement through space. No single cue creates the illusion on its own. Together, however, they allow the listener to accept the environment as a coherent whole.

    Nair also highlighted another element that is easy to overlook because people perform it almost unconsciously: head movement. Humans rarely remain perfectly still while listening. A slight turn of the head provides additional information that helps resolve ambiguities about where a sound is located. This is particularly important when distinguishing between sounds in front of and behind the listener, a challenge that has long accompanied binaural reproduction. By continuously updating the soundtrack as the listener moves, head tracking becomes far more than a technical feature of virtual reality headsets. It mirrors the way people naturally explore the acoustic world, allowing perception itself to become part of the localisation process.

    Many of these mechanisms operate almost entirely beneath conscious awareness. Few listeners actively notice early reflections, frequency-dependent air absorption or subtle changes in reverberation while exploring a virtual environment. Instead, these cues quietly reinforce one another, allowing the brain to construct a coherent interpretation of the surrounding space. Remove one or two and the experience may still appear convincing. Remove several, however, and the illusion begins to weaken, even if listeners cannot explain precisely why.

    Taken together, these ideas challenge one of the most common assumptions about immersive audio. It is tempting to search for a single breakthrough technology capable of making virtual reality sound convincing. Nair suggested that no such shortcut exists. Believable experiences emerge because numerous perceptual cues remain consistent with one another, each reinforcing the listener’s interpretation of the environment. The role of the sound designer is therefore not simply to position sounds accurately, but to ensure that every acoustic detail supports the same perceptual story.

    Seen in this way, immersive audio becomes much more than an exercise in spatialisation. It is an exercise in environmental design, where every decision contributes to the listener’s understanding of the world around them. The objective is not simply to convince people that a sound exists. It is to convince them that the world itself exists.

    Understanding how people perceive space is only part of the challenge. Once those principles are understood, the next question becomes considerably more practical. How should sound designers actually use them? Knowing that reflections, localisation and head movement contribute to immersion does not automatically produce a compelling soundtrack. Technology provides the possibilities; design determines how successfully those possibilities are translated into an experience.

    One of the recurring themes throughout Nair’s lecture was the importance of scale. Although the term is often associated with visual design, he argued that sound is equally responsible for establishing the perceived size of a virtual world. A room may appear life-sized, but if sounds travel too far, decay unnaturally or fail to attenuate in believable ways, the illusion quickly begins to weaken. Listeners may not consciously identify the problem, yet the environment no longer feels internally consistent. Sound and vision must reinforce the same interpretation of the world if presence is to be maintained.

    Nair also distinguished carefully between scale and size. The overall scale of a virtual environment establishes the listener’s relationship with the world, while the apparent size of individual objects determines how those objects should behave acoustically. A mosquito and a passing truck may both move away from the listener, but they do not disappear into the distance in the same way. The truck remains audible over a much greater distance, while the mosquito rapidly fades from perception. These differences are shaped not only by level but also by pitch, frequency content and attenuation. Together, they allow listeners to make surprisingly sophisticated judgements about the physical properties of objects without consciously thinking about them.

    This observation extends well beyond realistic simulations. Sound designers frequently create worlds that bear little resemblance to everyday life, whether they are fantastical landscapes, stylised games or abstract interactive experiences. The objective is therefore not to reproduce reality exactly, but to establish a world whose internal logic remains believable. A listener who believes they are inhabiting the body of a giant should hear the environment differently from someone exploring the same space at the scale of a child. Perception is always relative to the world the experience creates.

    One of the most thought-provoking moments in the lecture came when Nair challenged a common assumption surrounding immersive audio. It is easy to believe that adding a binaural renderer or spatial audio plug-in will automatically transform an ordinary mix into an immersive one. His argument was almost the opposite. These technologies do not replace sound design; they expose it. Spatial audio recreates many of the same perceptual constraints that exist in everyday listening, meaning that weak decisions about balance, distance or attention become more apparent rather than less. Better technology cannot compensate for poor creative judgement.

    That perspective places even greater emphasis on the fundamentals of sound design. Decisions about pitch, dynamics, attenuation and spectral balance continue to shape the listener’s experience, not because virtual reality introduces entirely new principles, but because it allows established principles to influence perception in more direct and immediate ways. Spatial audio therefore broadens the range of creative decisions available to the sound designer rather than replacing established practice. It provides additional ways of communicating meaning, directing attention and shaping emotional response, while relying upon the same careful listening and critical judgement that have always underpinned excellent sound design.

    It also explains why Nair resisted presenting immersive audio as an all-or-nothing proposition. Some elements of a soundtrack benefit greatly from full spatialisation, while others may work more effectively in conventional stereo. Music, interface sounds and other non-diegetic elements all raise creative questions that remain largely unresolved. Rather than prescribing fixed rules, he encouraged sound designers to experiment, allowing the needs of each project to determine how technology should be applied rather than assuming that every sound must be treated in the same way.

    This willingness to experiment reflects the broader stage of development that virtual reality currently occupies. Stereo, surround sound and film mixing each developed their own conventions over decades of experimentation. Virtual reality is still at an earlier stage in that journey. Many of its techniques, workflows and creative conventions remain open questions. For today’s sound designers, that uncertainty represents an opportunity rather than a limitation. Instead of inheriting an established language, they have the opportunity to help create one.

    As the lecture drew to a close, Nair returned to a theme that had quietly connected every part of his discussion. Virtual reality is often presented as a technological revolution, yet its greatest challenges are no longer defined by processing power or display resolution. They concern perception, creativity and the ways in which people make sense of unfamiliar experiences. The technology continues to evolve rapidly, but the language of immersive sound is still being written.

    This places today’s sound designers in an unusual position. Many areas of audio production have inherited decades of established practice. Recording engineers, film mixers and game audio professionals all work within conventions that have gradually matured through years of experimentation and shared experience. Virtual reality offers far fewer certainties. Questions surrounding spatial music, interface sounds, narration, transitions between perspectives and the balance between realism and artistic expression remain active areas of exploration rather than settled practice.

    Nair suggested that this uncertainty should not be viewed as a weakness. It provides an opportunity to rethink long-standing assumptions about how audiences experience sound. Conventional stereo and surround formats are largely designed around listeners who remain outside the action, observing events from a fixed perspective. Virtual reality places the listener inside the experience. They choose where to look, what to investigate and how quickly to move through the environment. Audio therefore becomes less about presenting a carefully controlled sequence of events and more about supporting exploration without overwhelming the listener.

    These ideas also extend beyond entertainment. Training, education and professional simulation all depend upon helping people interpret unfamiliar environments naturally and confidently. The same understanding of attention, localisation and environmental perception that strengthens a virtual world can also improve learning, communication and decision-making wherever immersive experiences are used.

    Another message emerged just as clearly. Progress in immersive audio is unlikely to come from a single technological breakthrough. Better rendering algorithms, more accurate head tracking and increasingly sophisticated hardware will undoubtedly improve future systems, but believable experiences will continue to depend upon thoughtful design. Every new technical capability simply provides another opportunity for sound designers to apply their judgement. Creativity remains the defining ingredient.

    That perspective also reinforces an idea that extends well beyond virtual reality itself. Throughout the history of audio production, advances in technology have repeatedly expanded the possibilities available to practitioners without replacing the need for careful listening. Multitrack recording, digital editing, surround sound and object-based audio all transformed production workflows, yet each ultimately depended upon people deciding how those tools should be used. Immersive media appears to be following the same trajectory. The challenge is not learning a completely new discipline, but applying enduring principles of sound design within a medium that offers new creative possibilities.

    Perhaps this explains why Nair’s lecture felt less like a demonstration of emerging technology and more like a discussion about perception. Virtual reality undoubtedly introduces new tools and workflows, but its success continues to depend upon familiar questions. How do people understand the spaces around them? Which sounds deserve attention? What makes an environment feel coherent enough that listeners stop analysing it and simply accept it as real? These are questions that sound designers have explored for decades. Virtual reality provides a new context in which to ask them.

    The lecture therefore leaves an encouraging conclusion for anyone entering the field. The future of immersive audio will certainly be shaped by advances in computing, displays and signal processing, but it will be shaped just as much by the people who decide how those technologies should be used. As virtual reality continues to mature, its greatest innovations may prove to be neither new algorithms nor new devices, but new ways of making virtual worlds feel believable enough that listeners simply accept them as real.

  • Why Does Game Audio Matter? Professor Lennart Nacke on Playability, Player Experience, and the Psychology of Sound

    Professor Lennart Nacke

    How much would a game change if you turned the sound off?

    Many players might be tempted to answer very little. Graphics dominate marketing campaigns, technical demonstrations and online discussion. Higher resolutions, realistic lighting and increasingly detailed worlds are often presented as the defining characteristics of modern games. Sound, by comparison, can seem almost secondary. It accompanies the experience, but rarely receives the same attention. If forced to prioritise one feature over another, many players would probably choose visual quality long before they chose audio.

    Everyday experience suggests something rather different. Silence a racing game and judging speed suddenly becomes more difficult. Remove the weapon sounds from a first-person shooter and every encounter feels strangely detached from the player’s actions. Eliminate footsteps, environmental ambience or warning cues and even familiar games begin to feel oddly incomplete. The mechanics have not changed. The graphics remain identical. Objectives, controls and level design are exactly as they were before. And yet something fundamental has disappeared. If sound is supposedly one of the least important elements of a game, why does its absence change the experience so profoundly?

    That paradox formed the starting point of Professor Lennart Nacke’s online guest lecture for Edinburgh Napier University. Drawing on research spanning human-computer interaction, game studies and psychophysiology, Professor Nacke explored a deceptively simple question: what role does sound actually play in the experience of playing a game? At the time of the lecture, he was Research Director of the HCI Games Group and preparing to join the University of Waterloo, where his work continued to investigate player experience, physiological interaction and the design of engaging interactive systems.

    One of the first findings he discussed appeared, at first glance, to reinforce the common assumption that audio matters relatively little. Survey research suggested that players ranked sound below features such as graphics and, perhaps most significantly, playability when asked to identify the most important qualities of a game. Many observers interpreted this as evidence that audio simply was not a priority. Professor Nacke saw something quite different. His question was not whether sound mattered. It was whether players had misunderstood what sound was actually contributing.

    Viewed from that perspective, playability becomes much more than responsive controls or balanced mechanics. It is the continuous conversation between player and game. Every action produces a consequence. Every consequence communicates something back to the player, guiding the next decision. Much of that conversation takes place through sound. A weapon firing, footsteps approaching from behind, an engine changing pitch during acceleration or the subtle cue confirming a successful interaction all provide information that helps players understand both the game world and their own actions within it. Sound does considerably more than create atmosphere. It helps make games playable.

    This also explains why players often underestimate its importance. Rhythm games and audio-only games make sound impossible to ignore, placing it at the centre of the experience. Conventional games achieve almost the opposite effect. Audio works so seamlessly alongside graphics, animation and interaction that its contribution becomes largely invisible. Only when it disappears do players recognise how much guidance it had been providing all along. Feedback becomes less immediate. Decisions become less confident. Successes feel less satisfying, while failures become harder to interpret. Sound is not simply something players hear. It is one of the ways games teach players how to play.

    By the time Professor Nacke moved from the history of game audio towards his own research, the lecture had already shifted onto much firmer ground. The question was no longer whether sound made games more entertaining. It was whether sound quietly underpinned many of the qualities players described as good game design. That possibility led naturally towards psychology, player experience and a series of experiments designed to investigate whether sound changes not only what players hear, but also how they understand and navigate interactive worlds.

    If sound helps make games playable, the next question becomes unavoidable. How does it do that? A racing game does not become easier simply because its engine sounds more realistic, nor does a first-person shooter become more engaging merely through louder explosions. Something more fundamental is taking place. Sound continually provides information that helps players interpret the world around them, anticipate events and judge the consequences of their own actions. Without that information, interaction becomes less certain, even when every visual element remains exactly the same.

    Feedback lies at the heart of this process. Every time players press a button, move a character or perform an action, the game responds. Some responses are visual, others are physical through a controller, and many arrive through sound. Together they form a continuous exchange between player and system. Rather than simply confirming that something has happened, effective feedback helps players understand what happened, why it happened and what they should do next. Viewed in this way, playability becomes inseparable from communication.

    Professor Nacke argued that this relationship between sound and feedback has often been underestimated. Discussions of game audio frequently focus on music, realism or cinematic atmosphere, all of which undoubtedly influence the player’s experience. Yet those aspects represent only part of audio’s contribution. Equally important are the sounds that players rarely think about at all. Footsteps revealing an approaching opponent, the subtle cue confirming that an object has been collected, the changing rhythm of a weapon, the shift in an engine’s pitch or the warning that danger is just beyond the edge of the screen all guide behaviour long before players consciously reflect upon them. The most valuable sounds are often the ones that disappear into the act of playing itself.

    Some games make this communicative role impossible to miss. Rhythm games such as PaRappa the Rapper, Dance Dance Revolution, Guitar Hero and Rock Band place sound at the centre of interaction. Players succeed only by listening carefully and responding with precise timing. Performance is inseparable from audio. Professor Nacke observed that these games introduced a strongly performative dimension in which rhythm, musical timing and vocal control become the mechanics of play rather than decorative additions to it.

    Most conventional games appear very different, yet the distinction is not quite as large as it first seems. Action games, strategy games and multiplayer titles rarely ask players to perform music, but they still depend upon continual interpretation of auditory information. Experienced players often react to sounds almost automatically, recognising threats, opportunities and changing situations without consciously analysing each cue. Learning to play therefore involves learning a game’s sonic language alongside its visual and mechanical systems.

    Audio-only games expose another assumption that many players rarely question. Remove the graphics from most games and many people expect the experience to collapse. Yet audio games communicate entire worlds through hearing alone. Developed partly to improve accessibility for players who cannot rely on conventional visual interfaces, they demonstrate that sound can convey location, movement, interaction and progress with remarkable effectiveness when designed carefully enough. Eliminating graphics does not eliminate gameplay. It simply requires designers to think much more carefully about how information is communicated.

    Taken together, these examples gradually reshape the original question. Earlier, the lecture asked whether sound contributes to playability. By this point, the discussion suggested something stronger. Playability may itself depend upon the quality of communication between player and game, and sound is one of its most important languages. If that proposition is correct, it should also be possible to investigate it scientifically. Professor Nacke’s own research set out to discover whether those ideas could be measured as rigorously as they could be experienced.

    A persuasive theory is only the beginning. Can the contribution of sound actually be measured? It is one thing to argue that players rely upon audio when navigating a virtual world. Demonstrating that relationship through rigorous research is considerably more demanding. Professor Nacke devoted the latter part of his lecture to exactly this challenge, describing an experiment that attempted to move beyond intuition and examine whether sound produces measurable changes in the player experience.

    Rather than relying solely upon interviews or personal opinion, the study combined subjective evaluations with physiological measurement. Participants played Half-Life 2 under different audio conditions while researchers recorded indicators such as heart rate and skin conductance alongside established player experience questionnaires. Bringing these different forms of evidence together reflected a broader ambition within human-computer interaction. If engagement, immersion or enjoyment genuinely change when sound changes, perhaps those differences should be visible not only in what players report afterwards, but also in the body’s physiological responses during play.

    The initial results appeared disappointing. The physiological measures revealed remarkably little difference between the experimental conditions. Heart rate and skin conductance remained far more stable than expected, suggesting that the presence or absence of audio did not produce the clear biological distinctions the researchers had anticipated. Judged solely on those measurements, it would have been tempting to conclude that sound contributed relatively little to the player’s experience.

    The investigation, however, did not end there. When participants described their experiences directly, a rather different picture emerged. Questionnaire responses consistently indicated that sound influenced immersion, enjoyment and the overall quality of play. Players recognised differences that the physiological measurements had failed to capture. Far from undermining the original hypothesis, the contrasting results raised a much more interesting question. The problem may have lain not with the importance of sound, but with the methods used to measure its effects.

    Professor Nacke reflected on several possible explanations. Physiological signals fluctuate for many reasons, many of them unrelated to the specific design features being investigated. More importantly, the experiment averaged measurements across relatively long periods of gameplay. Memorable moments rarely unfold in that way. A sudden enemy encounter, an unexpected sound cue or the brief confirmation that an action has succeeded may last only a fraction of a second. Averaging responses across several minutes risks smoothing away precisely the moments that matter most.

    That observation carries implications extending well beyond this individual study. Human experience is rarely distributed evenly across time. Games are composed of countless moments requiring players to notice, decide and respond. Audio often contributes most powerfully at those precise instants, directing attention, confirming actions or warning of imminent danger. A research method designed to capture overall levels of physiological arousal may therefore overlook the highly localised effects that make sound so valuable during interaction.

    In many ways, this became one of the most revealing aspects of the lecture. Scientific research is often presented as a straightforward progression from hypothesis to confirmation. Professor Nacke instead demonstrated something much closer to the reality of research. Unexpected findings do not necessarily invalidate an idea. Sometimes they expose limitations in the way questions have been framed or measurements have been collected. Negative or ambiguous results become opportunities to design better experiments rather than reasons to abandon promising theories.

    By this point, the discussion had moved well beyond the simple question of whether sound matters. The evidence already suggested that players experience games differently when audio changes. Research was no longer asking whether sound mattered. It had begun asking how games should communicate with their players. That shift naturally pointed towards the future of game audio research, where the challenge is no longer proving that audio influences experience, but understanding which sounds matter most, which forms of feedback guide behaviour most effectively and how interactive systems can communicate with players more clearly, naturally and intelligently.

    By the end of the lecture, the original question had changed almost completely. It no longer seemed particularly interesting to ask whether sound matters in games. Few experienced players would seriously argue that it does not. A more revealing question had emerged instead. How should games communicate with their players? Once viewed through that lens, sound becomes much more than an aesthetic choice. It becomes one of the primary ways interactive systems explain themselves.

    Such a shift carries important consequences for the future of game design. Advances in graphics have often dominated discussions of technological progress, while audio has frequently been treated as a complementary layer added once the visual experience has been established. Professor Nacke’s work suggests that this sequence deserves to be reconsidered. If sound continually guides attention, confirms actions and shapes decision-making, then it should be regarded as part of the interaction itself rather than something applied afterwards to increase realism or atmosphere.

    Sound designers also occupy a rather different position within this way of thinking. Creating convincing effects and emotionally engaging soundtracks remains an essential part of the craft, yet interactive media asks for something more. Every cue becomes part of an ongoing dialogue between player and game. A well-designed sound does not merely create excitement or reinforce mood. It helps players understand where they are, what has changed and how they should respond next. Good game audio therefore succeeds not simply when it sounds impressive, but when it communicates clearly while remaining almost invisible.

    Lessons emerging from game audio extend far beyond entertainment. Interactive systems increasingly shape everyday life, from educational software and medical training to vehicle interfaces, industrial control systems and virtual reality. Each depends upon users making accurate decisions within changing environments. Many of those decisions rely upon information that can be communicated more quickly and more intuitively through sound than through visual displays alone. Games therefore provide an unusually rich environment in which to explore how people receive, interpret and act upon information under constantly changing conditions.

    Equally revealing was Professor Nacke’s willingness to embrace uncertainty. The physiological experiment did not produce the clear confirmation the researchers had hoped for, yet that outcome ultimately opened more interesting questions than it closed. Which sounds matter most? At what moments do they influence behaviour? How should researchers measure effects that may last only fractions of a second? Scientific progress rarely follows a straight line. Careful experiments often reveal that the next question is more valuable than the original answer.

    Looking back to the opening paradox makes that progression easier to appreciate. Players may continue to rank graphics above audio when asked which aspects of a game matter most. Yet the lecture suggests that such judgements overlook the extent to which sound quietly supports many of the qualities they value most. Confident interaction, satisfying feedback, effective learning and a strong sense of presence all depend upon communication between player and system. Much of that communication happens through sound, even when players scarcely notice it.

    Future interactive technologies will make these questions even more important. As games and other digital systems become increasingly adaptive, personalised and intelligent, designers will need to think less about individual sensory channels and more about the complete experience of interaction. Sound will remain one of the fastest, richest and most flexible ways of guiding attention without demanding it. Used thoughtfully, it can inform, reassure, warn, encourage and teach, often within fractions of a second.

    Professor Nacke’s most enduring contribution may be a simple change in perspective. Game audio is not fundamentally about making virtual worlds louder, more cinematic or more realistic. It is about making interaction intelligible. Every carefully designed cue, every subtle confirmation and every moment of auditory feedback helps players understand a world that exists only through continual exchange between person and system. Sound is not simply something games produce. It is one of the ways games reveal themselves to their players, allowing them to understand, navigate and ultimately master the worlds they inhabit.

  • How Can Sound Change What We Taste? Charles Spence on Crossmodal Perception, Multisensory Design, and the Future of Experience

    Charles Spence

    How much of flavour actually comes from the food itself?

    Most people would probably answer almost all of it. Sweetness belongs to sugar, bitterness belongs to coffee, freshness belongs to mint and carbonation belongs to sparkling water. Sound certainly accompanies these experiences, but it seems difficult to imagine that it could fundamentally change them. Whether a room is silent or filled with music, whether somebody eats alone or in a crowded restaurant, surely the food itself remains exactly the same.

    Everyday experience, however, quietly points in another direction. Coffee often tastes less satisfying on an aircraft than on the ground. Crisps seem fresher when they produce a louder crunch. Champagne feels more celebratory when its bubbles sparkle audibly in the glass, while the atmosphere of a restaurant can transform the enjoyment of a meal without a single ingredient changing. None of these observations seems especially surprising on its own. Taken together, however, they raise a rather uncomfortable question. If the food has not changed, what exactly has?

    That question formed the starting point for Professor Charles Spence’s online guest lecture for Edinburgh Napier University. As Head of the Crossmodal Research Laboratory at the University of Oxford, Spence has spent more than two decades investigating how the senses interact to construct experience. His research spans psychology, neuroscience, design and consumer behaviour, leading to collaborations with chefs, airlines, manufacturers, advertisers, perfumers and technology companies. Food featured prominently throughout the lecture, yet it soon became clear that gastronomy was simply one expression of a much broader scientific question.

    For much of modern science, the senses were treated as though they operated independently. Vision belonged to the eyes, hearing to the ears, taste to the tongue and smell to the nose. Perception appeared to be assembled almost like a jigsaw, with each sense contributing its own separate piece before the brain combined them into a complete picture. Spence’s work challenges that assumption. Information arriving through one sensory pathway immediately begins influencing the interpretation of information arriving through another. What people hear alters what they believe they taste. What they see changes what they expect to smell. The texture of a surface influences impressions of quality before conscious reasoning has even begun. Rather than operating as isolated systems, the senses appear to cooperate continuously, constructing experience through their interaction rather than through their independence.

    Sound therefore occupies a rather different role from the one most people imagine. Rather than serving simply as accompaniment, decoration or atmosphere added after an experience has already been designed, it becomes one of the materials through which perception itself is shaped. A carefully chosen sound can influence whether food seems fresher, sweeter, more bitter or more luxurious. It can establish expectations before the first mouthful, alter emotional responses during consumption and even affect the memories people later form of an experience. Sound does not simply accompany flavour. Under the right conditions, it helps create it.

    As the lecture unfolded, experimental psychology sat comfortably alongside fine dining, product packaging, aircraft cabins, advertising, perfume, digital interfaces and some of the world’s most celebrated restaurants. At first sight these subjects seemed to have little in common beyond an occasional reference to sound. Gradually, however, a consistent pattern emerged. Experiences that appear to belong almost entirely to one sense are often shaped by several others at the same time.

    Food was only the beginning. If hearing can influence taste, perhaps many experiences usually treated as purely visual, tactile or auditory are also products of continual interaction between the senses. The challenge for designers is therefore no longer simply to create attractive sounds, images or objects in isolation. It is to understand how they work together to shape perception as a whole. That broader question ultimately became the central theme of Spence’s lecture, pointing towards a future in which sound is understood not as an accessory to experience, but as one of the materials from which experience itself is constructed.

    If sound can influence flavour at all, the obvious question is how. At first sight the idea seems almost impossible. Taste depends upon chemical receptors inside the mouth, while hearing begins with vibrations entering the ears. One system detects molecules, the other detects pressure waves. They appear to have almost nothing in common. Yet the experiments Charles Spence presented suggest that the brain pays surprisingly little attention to these traditional boundaries. Instead of treating each sense as an isolated source of information, it continually searches for relationships between them.

    Every meal illustrates this process. Before food even reaches the mouth, its appearance has already created expectations about freshness, sweetness, richness or quality. Aroma begins shaping anticipation, while the weight of cutlery, the texture of a plate and the surrounding environment all contribute further information. Sound enters that process from the beginning. The clink of a glass, the crack of a crisp crust, the fizz of a carbonated drink and the background atmosphere of a restaurant all become additional clues from which the brain constructs a single interpretation. People do not consciously separate these sensations before combining them again. They simply perceive flavour.

    For many years, perception was often described as though the senses operated independently. Vision belonged to the eyes, hearing to the ears, taste to the tongue and smell to the nose, while the brain merely assembled their separate outputs into a complete picture. Increasingly, evidence points towards something much more dynamic. Information arriving through one sense immediately begins shaping the interpretation of information arriving through another. What people hear alters what they believe they taste. What they see changes what they expect to smell. Texture influences impressions of quality before conscious reasoning has even begun. Human perception emerges through continual interaction rather than through the addition of independent sensory streams.

    Spence explored these relationships through what he describes as crossmodal correspondences. Although the terminology sounds specialised, the underlying phenomenon is surprisingly familiar. Many people instinctively associate higher-pitched sounds with sweetness and lower pitches with bitterness. Other combinations of pitch, timbre, rhythm and texture consistently evoke impressions such as freshness, creaminess or sharpness. None of these relationships is consciously taught, yet they appear repeatedly when large groups of people are asked to match sounds with tastes.

    Consistency is what makes these findings especially significant. Individual preferences naturally vary, but the broader patterns remain remarkably stable. Participants who have never met one another frequently make similar associations, suggesting that these relationships are not simply matters of personal preference or cultural coincidence. Human perception appears to organise sensory information in surprisingly consistent ways. Designers therefore gain something unusually valuable: perceptual tendencies that can be anticipated rather than guessed.

    Crossmodal correspondences are sometimes confused with synaesthesia, despite important differences. People with synaesthesia may genuinely experience colours when hearing music or perceive letters as possessing particular tastes. Those experiences are real, but they are also highly individual. Crossmodal correspondences operate differently. They describe tendencies shared across many people rather than unique experiences belonging to particular individuals. That distinction allows them to move from psychological curiosity to practical design principle.

    Laboratory findings soon begin to acquire practical significance. Spence described collaborations in which composers and sound designers created musical soundscapes intended to reinforce particular taste qualities. Pitch, rhythm, timbre, consonance and articulation were carefully manipulated to suggest sweetness, bitterness, sourness or saltiness. Participants consistently matched particular soundscapes with particular tastes. More remarkably, under appropriate conditions those same soundscapes subtly altered the way the food itself was perceived. Music associated with sweetness could increase perceived sweetness without any change to the ingredients. Chemistry remained exactly the same, yet perception shifted in a predictable direction.

    Taste, in other words, emerges through interpretation rather than chemistry alone. Molecules reaching the tongue remain essential, but they represent only one source of information among many. Every sound surrounding a meal, from the crunch of food itself to the atmosphere of the room, contributes another piece of evidence that the brain may incorporate into its final judgement. Designers have often exploited these relationships intuitively for decades. Spence’s research provides a scientific explanation for why they work.

    Viewed from this perspective, familiar experiences begin to appear rather different. The satisfying crack of a crisp packet, the hiss of a freshly opened bottle or the sound of coffee being ground no longer seem like incidental by-products of physical events. They become active components of perception itself. Some establish expectations before tasting even begins. Others reinforce qualities already present or quietly direct attention towards particular sensations. Sound no longer sits outside flavour. It has become one of the ingredients from which flavour itself is assembled.

    Once the relationship between sound and flavour becomes scientifically plausible, an even more interesting question begins to emerge. What happens when designers start working with that knowledge? Crossmodal perception is no longer simply an explanation for a series of intriguing laboratory experiments. It becomes a different way of thinking about the creation of experiences. If perception is constructed through continual interaction between the senses, then every sound surrounding a product becomes part of the product. Designers are no longer shaping isolated objects or individual sensory channels. They are shaping the conditions under which people construct experience.

    That shift in perspective explains why Charles Spence’s research extends so far beyond psychology. His collaborations have involved chefs, advertisers, manufacturers, retailers, technology companies and designers working across remarkably different industries. At first sight those partnerships appear unrelated. Yet each asks essentially the same question. If sound changes the way people perceive quality, freshness, luxury or value, how should products and experiences be designed differently?

    Many organisations have traditionally approached multisensory design by attempting to stimulate as many senses as possible. Spence suggested that this is only part of the challenge. More sensory information does not necessarily create a better experience. Success depends upon coherence rather than quantity. Vision, sound, touch, smell and taste need to reinforce one another, guiding perception towards a shared interpretation. A beautifully designed sound can strengthen what people already expect to encounter. An inconsistent one can quietly undermine the entire experience.

    Small design decisions therefore begin to acquire unexpected significance. Opening a packet of crisps, hearing the hiss of a carbonated drink or listening to coffee beans being ground may seem like incidental moments within a much larger experience. Yet those sounds establish expectations before tasting even begins. They encourage the brain to anticipate freshness, quality or intensity long before chemistry has an opportunity to contribute. The product remains physically unchanged. What changes is the perceptual framework through which the product is interpreted.

    Few environments demonstrate these relationships more clearly than restaurants. Every aspect of a meal contributes information beyond the food on the plate. Lighting, tableware, acoustics, conversation and music all participate in shaping expectations and emotional responses. Fine dining therefore becomes an exercise in multisensory design rather than culinary technique alone. Preparing exceptional food remains essential, yet the overall experience emerges from the interaction of many carefully orchestrated elements.

    One of the lecture’s most celebrated examples came from Heston Blumenthal’s Sound of the Sea. Diners listened to recordings of waves, seabirds and the sounds of a coastal landscape while eating a seafood course. Initially the idea appeared almost whimsical. Why should listening to the sea alter the taste of food served indoors? Yet diners consistently described the dish as fresher, more vividly maritime and more immersive when accompanied by the soundscape. Even members of Spence’s own research group admitted that the strength of the effect had surprised them. What initially appeared to be a theatrical flourish turned out to reveal something far more fundamental about the way perception operates.

    Experiments of this kind naturally led towards increasingly sophisticated forms of experience design. Spence described cocktail events in which each drink was paired with its own carefully composed soundscape. Participants were invited to compare different combinations, discovering how subtle changes in the auditory environment altered the character of the same drink. Attention gradually shifted away from asking whether one cocktail tasted better than another. Instead, the question became how the interaction between sound and flavour could produce the most satisfying overall experience. Perception itself had become the object of design.

    Once viewed from this perspective, the implications become difficult to contain within food and drink alone. Mobile devices, consumer products, retail spaces, packaging and digital interfaces all communicate through sound, whether intentionally or accidentally. Notification tones, confirmation signals, mechanical clicks and interface feedback continually influence users’ impressions of quality, reliability and identity. Those sounds are not decorative additions arriving after the creative process has finished. They help determine how the product itself is understood.

    Perhaps the most important lesson from this part of the lecture is that food was never really the subject. Food simply makes multisensory perception unusually easy to observe. Once the underlying principles become visible, they begin appearing almost everywhere. Products, services, environments and digital technologies all rely upon the same interactions between the senses. Sound therefore ceases to be something added to an experience. It becomes one of the materials from which perception itself is constructed.

    By the end of the lecture, it became increasingly difficult to think about sound in quite the same way as before. Traditional approaches to sound design often begin by asking what something should sound like. Charles Spence’s work quietly suggests a different question. What should people perceive? Those two objectives frequently overlap, but they are not identical. A sound can be technically accurate while contributing little to the intended experience, or it can depart from physical realism yet guide perception in ways that feel entirely convincing. The designer’s task therefore extends beyond producing sounds. It becomes one of shaping interpretation.

    That shift carries important implications for the future of design. Emerging technologies make these ideas increasingly relevant. Artificial intelligence can already personalise recommendations, anticipate preferences and generate new forms of interaction in real time. Yet many of these developments continue to treat the senses as largely separate channels through which information is delivered. Spence’s research points towards a different possibility. Systems might instead learn how combinations of sound, vision, touch and other sensory cues influence perception as a whole, adapting experiences rather than simply adapting individual outputs.

    Such ideas extend naturally into interaction design, healthcare, education, transport and entertainment. A navigation system might reduce stress by carefully coordinating spoken instructions, interface sounds and visual feedback. Assistive technologies could combine auditory and tactile information to improve confidence and accessibility. Retail environments may shape perceptions of quality without relying solely upon visual presentation, while museums, exhibitions and virtual environments can construct richer experiences through carefully orchestrated sensory relationships. None of these applications depends upon one extraordinary technological breakthrough. They emerge from a deeper understanding of how people already perceive the world.

    Sound designers occupy a particularly interesting position within this changing landscape. Traditionally, sound has often entered the creative process after many other design decisions have already been made. Music, ambience and effects enrich an experience that largely exists in visual or physical form. Multisensory research challenges that sequence. If sound actively participates in constructing perception, it deserves consideration from the earliest stages of design rather than being treated as a finishing touch. Decisions about audio become decisions about cognition, emotion and behaviour.

    This perspective also encourages greater humility. Human perception remains remarkably complex, and no designer can predict every response with complete certainty. Cultural experience, personal memory, expectation and context all continue to influence how people interpret the same sensory information. Spence repeatedly acknowledged that multisensory design is not a collection of universal formulas guaranteeing identical outcomes for every individual. Instead, it offers evidence-based principles that increase the likelihood of particular perceptual responses while recognising that people remain active participants in constructing their own experiences.

    Perhaps that is why the lecture remained so compelling. It never suggested that science could replace creativity or reduce design to a set of predictable rules. On the contrary, scientific understanding expanded creative possibility rather than limiting it. Once designers understand how the senses cooperate, they gain entirely new materials with which to work. Sound becomes capable of shaping flavour, texture, atmosphere, expectation and memory alongside its more familiar roles in communication and emotion. Creativity is not constrained by such knowledge. It is given a richer foundation upon which to build.

    Returning to the opening question reveals how much has changed. How much of flavour actually comes from the food itself? Chemistry remains indispensable, but chemistry alone is no longer a sufficient answer. Flavour emerges through continual interaction between taste, smell, vision, touch, sound and expectation, each contributing to a perceptual whole that cannot easily be reduced to its individual parts. Food becomes one component of a broader multisensory experience rather than its sole determinant.

    That conclusion reaches well beyond gastronomy. Every designed experience, whether a film, game, product, concert hall, mobile application or public space, is ultimately encountered through the cooperation of the senses. Designers therefore shape more than objects, interfaces or soundtracks. They shape the perceptual conditions through which people understand the world around them.

    Perhaps the most enduring contribution of Charles Spence’s work is not the demonstration that sound can make food taste sweeter or fresher, remarkable though those findings remain. It is the reminder that perception itself is creative. The brain does not passively record reality. It continually interprets, predicts and combines information arriving from many different sources, constructing an experience that feels immediate, coherent and effortless. Sound is not merely something that accompanies that process. It is one of the materials from which that process is built.

  • How Do You Know When a Sound Is Right? Nigel Christensen on Storytelling, Instinct, and Designing for Feeling

    Nigel Christensen

    How do you know when a sound is right?

    Two recordings may represent exactly the same physical event. One feels convincing, inevitable and emotionally satisfying. The other feels strangely artificial, despite being technically more accurate. A sound designer can spend hours recording, editing, layering and processing material, refining transients, balancing frequencies and positioning every element within an increasingly sophisticated spatial environment, yet the finished scene can still seem somehow incomplete. At other times, the smallest adjustment suddenly allows everything to settle into place. A barely perceptible atmosphere makes a location believable. A slight change in perspective gives an object convincing weight. Removing a sound proves more effective than adding another. Technical expertise explains how a soundtrack has been assembled. It does not entirely explain why one version feels right.

    That question lay quietly beneath Nigel Christensen’s online guest lecture for Edinburgh Napier University. Although he never reduced his work to a single principle, almost every example returned to the same underlying challenge. Across more than three decades in film and television, Christensen has worked on productions as different as Neighbours, The Promise, Mad Max: Fury Road, The LEGO Movie, The LEGO Ninjago Movie, Colin from Accounts, Boy Swallows Universe and Sentient. The productions differ enormously in genre, scale and audience, yet each required him to solve a remarkably similar problem. Faced with countless possible creative decisions, how does a sound designer recognise the one that truly belongs?

    Christensen’s answer was never framed as a technique or a method. Instead, it emerged through the way he spoke about listening. Recording, Foley, editing, restoration, sound effects and mixing remain indispensable parts of the profession, yet audiences never experience them separately. They experience conversations unfolding within spaces, footsteps revealing character, environments establishing place, music shaping emotional direction and silence becoming as expressive as any sound surrounding it. Individual recordings rarely possess dramatic meaning on their own. They acquire it through the way they interact with dialogue, image, performance and everything surrounding them. A beautifully crafted sound can weaken a scene once placed alongside other elements, while something almost imperceptible can quietly transform the emotional weight of an entire sequence.

    As the lecture unfolded, Christensen repeatedly returned to this idea from different directions. One project demanded months of experimentation to create the sound of an event that could never occur in reality. Another relied upon extensive recordings of heavily modified vehicles before the smallest mechanical detail felt convincing. Comedy rewarded restraint where action encouraged energy, while documentary required a different form of responsibility altogether. The practical solutions changed from production to production. The process of listening did not.

    Perhaps the most revealing aspect of Christensen’s lecture was his insistence that this process cannot be reduced to a formula. Professional experience gradually develops a way of hearing that extends beyond technical knowledge. Designers learn to recognise when attention is drifting towards the wrong element, when a soundtrack feels unexpectedly crowded or when apparently unrelated sounds suddenly begin supporting one another in ways that make the entire scene feel more convincing. Christensen described constructing an internal three-dimensional framework in which dialogue, music, ambience and effects occupy relationships of frequency, perspective, depth and dramatic importance rather than existing as isolated events. Dialogue often provides the centre of gravity, while every other element continually negotiates its place around it. Success rarely depends upon making every sound equally noticeable. More often, it depends upon recognising what deserves attention and what should remain quietly in support.

    His own fascination with sound began long before he imagined a professional career. One of his earliest memories involved his grandfather’s Acme whistle, manufactured during the Second World War. Years later, while watching a Warner Bros. cartoon, he recognised exactly the same sound emerging from the television. The experience left a lasting impression. An everyday object had somehow become part of a completely different world, attached to movement, timing and character until audiences accepted it without question. Around the same period he discovered that sounds could also be recreated. Having become adept at imitating the distinctive trill of his school’s classroom intercom, he once produced such a convincing imitation that his teacher crossed the room, answered the handset and began speaking to an office that had never called. Looking back, both experiences pointed towards the same observation. Sounds are never interpreted through acoustics alone. Context, expectation and previous experience continually influence what listeners believe they are hearing.

    Those early discoveries did not immediately suggest a future in sound design. Christensen’s interests ranged across music, photography and writing until one of his teachers recognised that filmmaking offered a place where those interests might converge. Entering the Australian post-production industry in 1989, he arrived during a period of considerable technological change. Analogue workflows still dominated everyday practice, while digital editing systems were only beginning to appear. Film workprints, synchronised magnetic tape and physical editing remained routine, and Christensen recalled demonstrations of early digital systems being met with understandable scepticism. Looking back, however, the technological transformation itself seemed less significant than another lesson emerging from those years. Every advance expanded the range of available creative possibilities. None removed the need to listen carefully, understand the dramatic needs of a scene and decide what genuinely belonged.

    That understanding developed rapidly during Christensen’s work on the long-running television series Neighbours. Part of his role involved preparing international music and effects tracks from which dialogue had been removed, allowing episodes to be dubbed into other languages. Working alongside Foley artists, editors and mixers, the team completed five episodes every day. Such a schedule left little opportunity for prolonged experimentation. Sounds had to be found or created, evaluated within the scene and completed before attention immediately shifted to the next problem.

    Working at that pace encouraged a discipline that remained visible throughout Christensen’s later career. Every additional sound had to earn its place. A convincing footstep, a carefully judged atmosphere or a subtle mechanical detail succeeded not through individual brilliance but through the contribution it made to the scene as a whole. Television therefore became an education in priorities. Decisions needed to be made quickly and confidently, yet always in service of the film rather than the sound itself.

    As Christensen’s career expanded into larger productions, the budgets, schedules and creative possibilities changed dramatically. The underlying challenge, however, remained remarkably familiar. More recording time, more sophisticated processing and more powerful technology did not simplify creative decision-making. They simply increased the number of possible solutions. Learning how to create a sound was only the beginning. The more demanding task was recognising which of those possibilities truly belonged within the world of a particular story.

    The increasing range of creative possibilities became particularly apparent as Christensen moved into feature filmmaking. Television had taught him to make confident decisions under relentless production schedules. Larger productions presented a different challenge. More time, larger budgets and increasingly sophisticated technology did not necessarily make the work easier. Instead, they multiplied the number of possible solutions. Every additional recording, every new processing technique and every creative option demanded another decision about whether it genuinely strengthened the film or merely demonstrated what could be done.

    One of the clearest examples came from Chen Kaige’s The Promise. Christensen was asked to create the sound of a character running through the wall of time, an event with no equivalent in the physical world. There was nothing to record, no existing reference and no audience memory against which the result could be measured. The task therefore became one of invention rather than recording reality. Months were spent exploring combinations of recordings, processing techniques and layered textures, searching for sounds capable of suggesting impossible movement without disrupting the internal logic of the film.

    Christensen’s account revealed something fundamental about cinematic realism. Audiences do not judge a soundtrack solely by comparing it with everyday experience. They judge whether it feels convincing within the world the film has established. Nobody has heard a person move through time, yet viewers recognise when such a moment appears believable. Credibility emerges from consistency. Every element supports the same dramatic proposition until the impossible begins to feel strangely inevitable.

    At first glance, Mad Max: Fury Road appears to present exactly the opposite problem. Rather than inventing sounds for imaginary events, George Miller’s production provided an abundance of authentic material. Vehicles were extensively recorded under different operating conditions, capturing engines, transmissions, exhaust systems and countless mechanical details from multiple perspectives. Christensen’s responsibilities included many of the motorcycles, the Vuvalini vehicles and several machines surrounding the War Rig, giving him access to recordings that reflected the unique personalities of the actual vehicles.

    Authenticity, however, proved to be only the beginning of the creative process. Christensen recalled George Miller focusing intently on the sound of a chain swinging during one sequence. Numerous chains were recorded, performed repeatedly and compared before one finally possessed the character the scene demanded. Nothing about the exercise suggested that a genuine recording automatically produced the most convincing result. A chain needed to do more than resemble a chain. Its movement had to reinforce rhythm, weight, performance and dramatic tension without distracting from everything else unfolding on screen.

    Listening in this way gradually changes the role of the sound designer. Rather than collecting impressive recordings, Christensen described himself as continually organising a shifting network of interactions. Throughout the lecture he referred to an internal three-dimensional framework within which dialogue, ambience, music and effects occupy different positions according to their dramatic importance. Dialogue often anchors the listener’s attention, while surrounding sounds establish scale, movement, physicality and emotional colour. Introducing a single element alters the function of every other element nearby. Nothing exists independently for very long. This also explains why Christensen rarely discussed recordings in isolation. A beautifully captured sound can become unexpectedly intrusive once dialogue, performance and music are added around it, while an apparently ordinary recording may become indispensable after finding its place within the larger dramatic structure.

    The same principle extends across the soundtrack as a whole. Dialogue, music, ambience, Foley and effects continually reshape one another. Raising one element alters the significance of another. Simplifying a sequence can sometimes reveal emotional details that had previously been hidden beneath unnecessary complexity. Christensen’s descriptions repeatedly suggested that the soundtrack behaves less like a collection of separate tracks than like a living system in which every adjustment influences the whole.

    Experience gradually transforms this way of listening into instinct. Christensen described recognising physically when a scene remained unresolved. A sequence could feel crowded, awkward or strangely unsettled long before he could identify the precise technical cause. Rather than reaching immediately for another process or another recording, he would continue listening, adjusting perspective, emphasis and texture until everything seemed to settle naturally into place. Only afterwards would it become possible to explain why the solution worked. Analytical understanding and intuition were not opposing ways of working but different stages of the same creative process.

    Professional development therefore means more than acquiring new technical skills. Software evolves, workflows change and tools become increasingly powerful, yet none of these developments removes the need to listen critically. Over time, designers become less concerned with demonstrating what technology can achieve and more interested in recognising when a scene has reached the point where nothing further needs to be added. Restraint becomes a sign of growing confidence rather than limited imagination.

    Animation appears to encourage limitless invention. Worlds such as The LEGO Movie and The LEGO Ninjago Movie overflow with movement, visual detail, rapid editing and comic incident, inviting sound designers to populate every available moment with activity. Christensen described arriving at almost the opposite conclusion. As visual complexity increases, audiences become less able to attend consciously to every sound. Organisation therefore becomes more important than accumulation. The objective is not to make everything audible, but to ensure that listeners instinctively recognise what matters at any particular moment.

    This approach shaped the soundtracks of both films. Dialogue, music, ambience and effects coexist in remarkable density, yet they rarely compete for attention. Some sounds carry narrative information, others reinforce physical movement or establish the scale of an environment, while many remain almost unnoticed until repeated viewings reveal their contribution. Rather than treating every recording as an opportunity to impress, Christensen spoke about allowing individual elements to emerge and recede according to the changing needs of the scene. Richness becomes a consequence of careful organisation rather than sheer quantity.

    Comedy makes this balancing act especially demanding. A joke succeeds as much through timing as through content, and sound participates in that timing just as fully as dialogue or editing. A visual action may need to arrive a fraction before its accompanying effect. Elsewhere, delaying a sound by only a moment can strengthen surprise or allow a performance to land more naturally. Christensen’s observations suggested that comic sound design depends less upon creating amusing noises than upon understanding precisely when audiences are ready to hear them. Rhythm, expectation and restraint become inseparable.

    Those same principles reappear in Colin from Accounts, although expressed in a completely different register. Here humour grows from recognisable people navigating awkward situations rather than exaggerated cinematic spectacle. Sound therefore works quietly. Everyday actions, body movements and small environmental details were adjusted repeatedly, not to attract attention but to preserve the fragile credibility upon which the comedy depends. An effect that feels too carefully emphasised risks reminding audiences that somebody has designed the moment. Christensen repeatedly returned to the importance of preserving the illusion that events are simply unfolding before the viewer.

    This distinction reveals something broader about realism itself. Literal accuracy and perceptual credibility are not necessarily the same thing. Everyday life rarely presents sounds in perfect isolation or with equal clarity. Attention shifts continually between conversations, environments and passing events, while the brain filters enormous amounts of information without conscious effort. A soundtrack attempting to reproduce every available detail with equal prominence would often feel less convincing than one carefully shaped around the way people actually experience the world. Christensen’s decisions consistently reflected this understanding. The aim was not to reproduce reality mechanically, but to recreate the experience of listening within it.

    Bottle Top Bill and His Best Friend Corky offered another unexpected perspective on this philosophy. Rather than removing every trace of the performers manipulating the miniature characters, Christensen allowed subtle handling noises and hints of human presence to remain within the finished soundtrack. Conventional post-production might regard these as imperfections. Here they became part of the programme’s identity, quietly reminding viewers that imagination begins with physical performance. The result feels neither unfinished nor technically compromised. Instead, those traces contribute warmth and authenticity, strengthening the audience’s relationship with the world rather than diminishing it.

    A similar sensitivity informed Christensen’s work on Boy Swallows Universe. The series moves naturally between everyday realism, memory and moments approaching magical realism without drawing sharp boundaries between them. Less experienced designers might have emphasised those transitions through conspicuous effects or dramatic sonic transformations. Christensen preferred a subtler approach. Small changes in perspective, texture and atmosphere gently alter the audience’s emotional orientation without announcing themselves as technical interventions. By the time viewers become aware that the world has shifted, they have already accepted the transition.

    Across these productions, Christensen continually returned to the same way of listening. Individual sounds mattered, but never as isolated achievements. Their significance emerged through timing, proportion and context. A footstep, an atmosphere or a carefully judged silence gained expressive power only through its contribution to everything surrounding it. Attention remained directed towards characters and narrative rather than towards the soundtrack itself.

    That philosophy also shaped Christensen’s understanding of collaboration. Every sound designer accumulates recordings that required considerable effort to obtain and solutions that feel particularly satisfying. Filmmaking has little interest in personal attachment. Directors, editors, composers and producers respond to complete scenes, not individual technical accomplishments. Christensen spoke openly about learning when to defend an idea and when to let it go. Removing an effect that no longer served the film was not a creative defeat but another way of improving the work. The objective was never to preserve favourite sounds. It was to strengthen the audience’s experience.

    Professional confidence often reveals itself through flexibility rather than certainty. It allows practitioners to contribute decisively while remaining willing to abandon their own solutions whenever a better one emerges. Cinematography, editing, music, production design and sound each illuminate different aspects of the same film. Christensen’s lecture suggested that the strongest collaborations develop when specialists maintain deep expertise within their own discipline while remaining attentive to the wider dramatic picture.

    This balance between technical possibility and creative restraint became particularly clear in Christensen’s discussion of the documentary Sentient. Fiction allows designers considerable freedom to invent or reshape sonic worlds, provided audiences continue believing them. Documentary introduces different responsibilities. Recordings often carry evidential weight alongside dramatic value, and modern restoration tools make it increasingly tempting to remove imperfections or reconstruct missing material. Christensen approached those possibilities with notable caution.

    Background noise, recording limitations and environmental interruptions were not automatically problems waiting to be solved. In many cases they formed part of the circumstances under which the events had actually occurred. Eliminating them indiscriminately risked altering the audience’s understanding of the material itself. Christensen therefore described restoration as an exercise in careful discernment rather than technical demonstration. Clarify what genuinely prevents communication, preserve what belongs to the recorded reality and resist interventions that merely produce a cleaner soundtrack. Sometimes the most accomplished decision is recognising that nothing more should be done.

    Placed alongside the productions discussed throughout the lecture, Sentient brings Christensen’s wider philosophy into particularly sharp focus. Fantasy, action, animation, comedy and documentary all present different technical challenges, yet none ultimately asks a different artistic question. Every project depends upon understanding how audiences listen, recognising what deserves their attention and shaping sound so naturally that the process itself quietly disappears behind the experience.

    Seen across more than three decades of work, Christensen’s career reveals a remarkably consistent philosophy. Technologies have changed, production methods have evolved and audiences now encounter films through an ever wider variety of formats and devices, yet the central responsibility of the sound designer has altered very little. Every project begins with the same challenge: to understand how sound can shape an audience’s experience without ever distracting from the experience itself.

    It is striking how little Christensen spoke about technology for its own sake. His professional life has encompassed one of the most transformative periods in the history of film sound, from analogue editing and magnetic tape through digital workstations, immersive formats and increasingly sophisticated restoration tools. Each innovation expanded the range of available creative possibilities. None simplified the underlying task. More powerful tools inevitably produce more possible solutions, yet no technology can determine which of those possibilities belongs within a particular film. That decision continues to depend upon careful listening, experience and an understanding of dramatic purpose.

    This perspective feels increasingly relevant as machine learning and artificial intelligence become more deeply integrated into post-production. Processes that once required hours of specialised work can now be completed within minutes. Dialogue can be separated from background noise, damaged recordings restored, voices synthesised and entirely new sounds generated with remarkable accuracy. Christensen’s lecture quietly suggested that such developments change the mechanics of production far more profoundly than they change its artistic purpose. As technical barriers continue to diminish, creative discernment becomes even more valuable. Producing convincing sounds remains an essential skill. Recognising which sounds genuinely strengthen a scene becomes the defining one.

    Every production discussed during the lecture reinforces that conclusion from a different direction. The Promise demonstrated that audiences willingly accept impossible events when every aspect of the soundtrack supports the same dramatic world. Mad Max: Fury Road showed that authentic recordings still require interpretation before they acquire dramatic force. The LEGO Movie illustrated how increasing complexity demands greater clarity rather than greater density, while Colin from Accounts revealed that realism often depends upon carefully controlled understatement. Bottle Top Bill and His Best Friend Corky questioned conventional assumptions about technical perfection, and Sentient demonstrated that documentary sometimes requires consciously resisting interventions that modern technology makes entirely possible. Different productions posed different creative problems, yet each ultimately returned to the same discipline of listening carefully before deciding how, or whether, to intervene.

    Listening, in Christensen’s work, extends well beyond identifying frequencies or evaluating recording quality. It becomes a way of understanding narrative itself. Designers gradually learn to recognise where attention naturally settles, when emotional emphasis begins to overwhelm a performance or when apparently minor adjustments transform the balance of an entire sequence. Such understanding develops through repeated encounters with practical problems rather than abstract rules. Experience teaches practitioners not simply how to manipulate sound, but how to recognise when a scene has found its natural equilibrium.

    From that perspective, sound design becomes less concerned with creating memorable individual sounds than with shaping the conditions in which those sounds acquire meaning. Dialogue, music, ambience, Foley, effects and silence continually influence one another, while every adjustment subtly reshapes the audience’s perception of the whole. A soundtrack resembles an ecological system more than a collection of recordings. Once a new element enters the scene, every existing element is heard a little differently. What finally reaches the audience is not a series of isolated sounds but a carefully organised pattern of interactions.

    Collaboration reinforces this understanding. Editors, cinematographers, composers, production designers and sound teams each bring different forms of expertise to the same film, and none can fully determine the finished experience alone. Christensen’s reflections suggested that specialist knowledge becomes most valuable when it remains responsive to the broader needs of the production. Strong ideas deserve thoughtful advocacy, yet they also require enough flexibility to be abandoned whenever another solution serves the film more effectively. Professional maturity lies as much in recognising what can be removed as in discovering what can be added.

    By the end of the lecture, the opening question has acquired a rather different meaning. Knowing when a sound is right is not primarily a technical problem. It is an interpretative one. A sound belongs when it reinforces performance, clarifies dramatic intention and helps audiences inhabit the world unfolding before them without becoming conscious of the craftsmanship involved. Technical expertise makes such moments possible. Careful listening allows them to emerge.

    Audiences rarely become aware of the countless decisions that shape this experience. They never hear the recordings that were rejected, the alternative mixes that were abandoned or the hours spent adjusting perspective, texture and timing until everything settled into place. Instead, they remember the weight of a machine crossing the desert, the awkward silence between two characters, the energy of an animated world or the quiet authenticity of a documentary moment. Those memories contain the work of the sound designer, even though the individual decisions have disappeared into the larger experience.

    Perhaps that is Christensen’s most enduring insight. Sound design has never been an exercise in demonstrating technical virtuosity. Recording, editing, restoration, Foley and mixing remain indispensable crafts, yet they all exist in service of something larger. Their purpose is to help audiences believe in a world that does not exist until sound, image and performance begin working together.

    The most successful soundtrack, then, is not necessarily the one containing the greatest number of remarkable sounds. It is the one that allows those sounds to disappear into the experience so completely that audiences stop noticing them altogether. Their attention remains with the characters, the unfolding drama and the emotional movement of the film rather than the techniques used to construct it. By that point, the soundtrack has achieved something more valuable than technical perfection. It has become inseparable from the world it was created to support.

  • How Do You Preserve the Sound of a Place? Damian Murphy on OpenAir, Acoustic Heritage, and Reconstructing Lost Spaces

    Damian Murphy

    How do you preserve the sound of a place?

    A building can survive through photographs, architectural drawings, maps and written descriptions. Its dimensions can be measured, materials catalogued and appearance reconstructed long after the original structure has disappeared. Yet places are not experienced through vision alone. A cathedral changes the sound of a choir, a tiled chamber changes the sound of a voice, snow alters the acoustic behaviour of a forest, and the hard surfaces of a mausoleum can allow sound to continue long after its source has stopped. Architecture surrounds every action with reflections, resonances and reverberation, but this part of a place can disappear without leaving anything visible behind.

    During his online guest lecture for Edinburgh Napier University, Professor Damian Murphy of the University of York explored almost two decades of work investigating how the acoustics of places can be measured, preserved, reconstructed and experienced. Much of this work centres upon OpenAir, the Open Acoustic Impulse Response Library, which contains acoustic measurements gathered from buildings, landscapes and other environments. Its contents range from cathedrals, churches and theatres to industrial buildings, caves, forests and vehicles, connecting acoustic science with music production, spatial audio, games, archaeology, heritage and historical research. Murphy moved between these different places and applications through a question that became larger as the examples accumulated. What can the sound of a place tell us that its image cannot?

    Murphy began with a ruin. The Temple of Decision stands on a hill overlooking the Falkland Estate in Fife, where artists David Chapman and Louise K. Wilson had been commissioned to explore the landscape through sound. Archive material could offer clues about the temple’s former appearance, while historical research could provide fragments of information about its use, but the surviving structure could no longer reveal how the intact building had sounded. What would it have been like to speak inside the room? How might voices have behaved around a table? How would conversation, movement or a fire have interacted with its surfaces? The artists posed a question that provided a starting point for Murphy’s lecture: in the absence of clear echoes, how can we know a place?

    Answering that question first requires an understanding of what an acoustic environment contributes to anything heard within it. Imagine a short, sharp sound produced inside a room. A listener initially receives the direct sound travelling along the shortest path from source to receiver. Early reflections arrive shortly afterwards from nearby surfaces, followed by increasingly complex patterns as energy travels through the room, interacting repeatedly with walls, floor, ceiling and objects before gradually decaying. Together, these components form a room impulse response for a particular relationship between a source and receiver. Direct sound carries information about the source and its distance, while early reflections contribute to perceptions of geometry and position. Later reverberation communicates qualities associated with volume, materials and enclosure. Change the architecture and the response changes. Move the source or listener and it changes again. An impulse response is therefore not a complete acoustic identity for a building, but a record of how sound travelled between particular positions under particular conditions.

    Once captured, that relationship can be used for more than numerical analysis. Through convolution, a recording made without significant room acoustics can be combined with an impulse response measured elsewhere. Murphy demonstrated the process using a four-part vocal ensemble recorded in an anechoic chamber. The singers had never performed in York Minster, yet convolution with a measured response allowed their dry voices to acquire characteristics of the cathedral. York Minster has a reverberation time of approximately eight seconds through part of the mid-frequency range, compared with around half a second for a typical living room. Voices behave very differently in each environment. Notes overlap, transitions blur and the building continues sounding after the performers have stopped producing sound. The acoustic is not simply decoration placed around a performance. It changes the temporal relationships through which that performance is heard.

    Auralisation, however, introduces a distinction between recreating acoustic conditions and recreating experience. Singers performing in an anechoic chamber do not behave as they would inside a highly reverberant cathedral. Performers hear themselves and adapt. Tempo, articulation, phrasing, dynamics and pauses can change in response to sound returning from the room, while musicians continually adjust to one another through the acoustic environment they share. Convolution can reproduce the effect of a measured response upon a recording, but it cannot retrospectively create the performance that might have developed inside that space. Murphy acknowledged this limitation directly when discussing the York Minster example. The anechoic performance was not the performance the singers would have given in the cathedral, and even the spacing of phrases in the demonstration had been altered to allow the reverberation to emerge more clearly.

    The distinction matters beyond the simulation of reverberation. A room impulse response can describe how energy travels between defined positions, but people are not passive sound sources or microphones. They move, listen selectively, change their behaviour and respond to what they hear. Preserving a response gives researchers evidence about the acoustic conditions of a place. What people did in response to those conditions remains a different question.

    OpenAir developed from an ambition to preserve such evidence and make it available for others to explore. Murphy traced one important influence to Angelo Farina’s work on recording concert halls for posterity. Improvements in measurement techniques and the emergence of practical convolution reverberation created an opportunity to document significant spaces not merely through reverberation times and other summary values, but through impulse responses that could be analysed, reproduced and used creatively. OpenAir extended this principle into a growing archive, with a measurement system designed to collect spatially rich data that could remain useful beyond the immediate research question.

    Early measurement sessions used a Genelec S30D loudspeaker to excite the space while microphones captured its response. A computer-controlled turntable allowed measurements at regular angular intervals, and an ambisonic Soundfield microphone was combined with a Neumann KM100 cardioid microphone to provide spatial information alongside material suitable for different forms of analysis and production. Measurements could be repeated across several source and receiver positions, preserving a set of relationships rather than presenting each building through one supposedly definitive response. Ambisonics was particularly valuable for an archive whose future applications could not be predicted. A first-order ambisonic recording represents a three-dimensional sound field through an omnidirectional component and three directional components, separating the captured information from one fixed loudspeaker arrangement. Material can later be decoded for different reproduction systems or manipulated in ways that may not have been anticipated when the recording was made.

    Flexibility matters when access to a significant site may last only a few hours. Researchers need to gather material rich enough to support questions that have not yet been formulated and technologies that may change long after a measurement session has ended. During the discussion after the lecture, Murphy described more recent work at St Paul’s Cathedral, undertaken with a composer who wanted impulse responses from the building. Three researchers had only three hours to move equipment through the enormous space and capture responses from locations including the nave, a stairwell and the Whispering Gallery. Practical decisions about where to measure become part of preservation itself. As Murphy observed, there is no single sound of a large building. Different positions offer different acoustic experiences, and any archive necessarily records selected relationships within a much larger field of possibilities.

    As OpenAir expanded, its growing range of places made the idea of acoustic preservation less straightforward. York Minster was an obvious candidate, since its long reverberation is closely connected with experiences of worship, tourism and musical performance. Other sites raised different questions about what deserves to be preserved and why. When the former Terry’s chocolate factory in York closed, Murphy and his colleagues gained access before redevelopment and measured spaces including a warehouse and the former typists’ room, a striking interior dominated by glass and wood. The activities for which these spaces had been designed had already disappeared. An empty typists’ room could still be photographed, but its appearance prompted another question: what might it have sounded like when filled with the overlapping mechanical activity of typewriters?

    A subterranean reactor hall beneath the Royal Institute of Technology in Stockholm preserved another relationship between architecture and former activity, while measurements in historic churches allowed acoustic theories to be tested rather than merely repeated. At St Andrew’s Church in Lyddington, the team examined jars embedded within the walls, architectural features sometimes interpreted through theories of resonant vessels extending back to the Roman architect Vitruvius. Measurement provided little evidence that the jars were making a substantial contribution to the acoustic character of the church. Their form did not correspond closely with the behaviour expected of effective Helmholtz resonators. Acoustic research could challenge explanations attached to historic architecture as well as document spaces admired for their sound.

    York Theatre Royal shifted attention from buildings as fixed objects towards places in changing states. Murphy’s team measured the auditorium before refurbishment and returned afterwards to document its altered acoustic. They also captured measurements with an audience present during a pantomime, recognising that an occupied theatre does not behave acoustically like the same room when empty. Seats, clothing and bodies absorb and scatter sound, making occupancy part of the acoustic system rather than simply a group of listeners placed within it. Materials age, spaces are repurposed and environmental conditions change. Preserving an acoustic environment may therefore involve documenting several states of the same place rather than searching for one definitive response.

    Outdoor and semi-outdoor locations created different practical problems, and the original measurement system could not simply be carried everywhere. On the Falkland Estate, the Bottle Dungeon could be reached through a trapdoor, but access was sufficiently awkward that the usual equipment was impractical. A balloon was attached to one stick, a pin to another and a microphone lowered into the space. Bursting the balloon remotely provided the excitation needed to capture a response. The improvised arrangement was far removed from the computer-controlled turntable used elsewhere, yet it addressed the same underlying need: introduce a suitable sound into an inaccessible environment and record how the space transforms it.

    Landscapes demanded further adaptation. Equipment had to become portable and independent of mains electricity as researchers travelled into the Yorkshire Dales to measure caves and gorges. Work in Finland examined the same forest under different seasonal conditions. Researchers used GPS alongside ribbons tied around trees to return to the same positions after deep snow had transformed the landscape. Geographically, it remained the same forest. Acoustically, it had changed. Snow altered the interaction between sound, ground and surrounding environment, demonstrating that acoustic character can change while location remains constant.

    A collaboration with Codemasters carried that thinking into interactive media. Looking at an archive rich in churches and historic interiors, the developer asked a practical question: what about the environments needed for games? The collaboration encouraged further work on landscapes and led to experiments for GRID Autosport involving vehicle interiors, where the acoustic problem was unusually complex. Codemasters wanted to represent a car gradually falling apart during a race, so the team needed measurements capable of describing changing states, including doors opening or disappearing and the boot being open. The experience of being inside a racing car also comes from more than airborne sound. Engine vibration travels mechanically through the structure and contributes to what an occupant hears and feels.

    Conventional room measurement could not fully represent that relationship, so Murphy and his colleagues experimented with using the engine itself as part of the measurement process. Revving the car caused the structure to vibrate as it would during use, after which signal-processing methods were applied to separate the excitation from the resulting cabin response and derive an approximation of the impulse response. The approach sought to preserve something more specific than the reverberation of a small enclosure. It attempted to capture the interaction between a vibrating machine, its structure, the enclosed air and the listener inside it.

    Games also demonstrated how preservation can become creative infrastructure. An impulse response gathered for research may later help construct a virtual environment, become part of a music-production tool or support a question that had not existed when the measurement was made. Murphy described OpenAir material finding its way into software used by musicians and audio practitioners, allowing measurements collected years earlier to acquire new purposes. Distribution under a Creative Commons licence reflected this wider ambition. Researchers can analyse the acoustic behaviour of a building, composers can use the same response creatively, sound designers can place fictional events inside measured environments and developers can incorporate selected material into new tools.

    OpenAir originally allowed members of the wider community to upload their own measurements. As contributions accumulated, variations in recording quality became difficult to ignore. The team eventually reviewed the existing material, retained the strongest contributions and moved towards a more curated model in which new contributors contact the team directly. Open access expands what an archive can become, but reuse also depends upon confidence in how the material was produced.

    These measurements deal with places that still exist, even when they are changing. The Temple of Decision presents a different problem. Its original interior has already gone. Once a room has disappeared, there is nothing left to measure. Its former acoustic behaviour has to be approached indirectly through surviving evidence and modelling.

    Using information about a lost structure, researchers can construct a three-dimensional geometric representation and simulate the propagation of sound within it. Virtual rays travel through the model, reflecting between surfaces to generate impulse responses for selected source and receiver positions. Those responses can then be analysed or used to process voices and other recordings, allowing listeners to hear an interpretation of how sound might have behaved inside architecture that no longer survives. Such a result differs fundamentally from measuring an existing building. Geometry may be uncertain, material properties need to be estimated and every modelling method introduces limitations. Historical auralisation produces an evidence-based acoustic proposition rather than a recording recovered from the past.

    Work on St Mary’s Abbey in York made both the possibilities and limitations of this approach audible. The former church survives as a ruin, while archaeological and architectural evidence allowed Murphy’s team to construct a three-dimensional representation suitable for acoustic modelling. Simulated impulse responses could then be compared with measurements from York Minster, a surviving building with some comparable characteristics. Estimated reverberation times occupied a similar range, but the reconstructed St Mary’s sounded noticeably brighter. The difference did not necessarily reveal a historical distinction between the buildings. Murphy explained that the ray-tracing model used for St Mary’s was less effective at reproducing low-frequency behaviour than the physical measurement system used in York Minster. Part of what listeners heard therefore belonged to the method of reconstruction itself.

    A model may sound convincing while still containing audible consequences of the technique used to create it. Plausibility has to emerge from evidence, comparison and methodological transparency rather than from the apparent realism of the result alone. Once those boundaries are understood, a model can make relationships perceptible in ways that drawings and numerical data cannot. During a public performance among the ruins of St Mary’s Abbey, a live choir was captured and processed through impulse responses derived from the reconstructed church, then projected to an audience gathered at the site. Present-day voices sounded among the physical remains while the model returned an interpretation of the missing acoustic architecture. Research data became part of an experience connecting a surviving place with a vanished interior.

    Reconstructing an abbey or temple can help audiences imagine how a lost building might have shaped music and speech. Murphy’s more recent work moved towards a question with wider historical consequences. If architecture changes what people can hear, could reconstruction help investigate who had access to speech in the past?

    Working with historians and art historians, Murphy’s team investigated spaces associated with the historic House of Commons within the former Palace of Westminster. Among the most revealing parts of that history was the experience of women who listened to parliamentary debate from a roof space above the chamber. Known as the Ventilator, this space allowed women excluded from formal political participation to gather above the House of Commons and listen through the architectural structure separating them from the debate below. Historical evidence could establish that they were there. Acoustic reconstruction allowed another question to be asked: what could they actually understand?

    Answering it required more than recreating the debating chamber as an isolated room. Researchers had to consider the chamber, roof void, Ventilator and routes through which speech travelled. The women listening above did not have a direct line of sight to the speaker, making their experience a problem of acoustic transmission through a complex architectural arrangement rather than ordinary listening within a single enclosure. Murphy connected this project with broader questions of directionality, speech transmission and listening position, while acknowledging that objective measures cannot reproduce every aspect of human attention or historical experience.

    No surviving recording can reveal exactly what those listeners heard. The team therefore combined historical reconstruction with comparative measurement, examining surviving spaces connected with parliamentary history or comparable in geometry, scale, period or use. Measurements from rooms at the University of Oxford, York Guildhall and the present House of Commons chamber provided contexts against which aspects of the model could be considered. None could prove how the lost chamber sounded, but together they allowed reverberation and speech intelligibility to be examined across different positions and occupancy conditions. A question about architectural acoustics had become inseparable from a question about political access.

    Hearing is a form of access. Architecture, distance, reverberation, occupancy and barriers influence whether speech remains intelligible, while attention, familiarity and expectations affect what can be understood from imperfect information. A person may be physically close to political debate while remaining acoustically separated from it. Reconstructing the conditions of listening can therefore contribute to historical questions about participation and exclusion that visual records alone cannot answer.

    The Temple of Decision now appears less like an isolated case and more like the beginning of a much larger enquiry. In the absence of clear echoes, how can we know a place? We can measure what survives, compare different conditions, preserve spaces before they change and build models from evidence when the original architecture has disappeared. We can listen to those models while remaining clear about what they can and cannot establish. Most of all, we can ask questions that become difficult to formulate when architecture is treated only as something seen.

    A photograph can preserve the appearance of a parliamentary chamber from one position. An architectural plan can show where walls, doors, galleries and roof spaces were located. Written testimony can tell us that people gathered somewhere to listen. Acoustic research adds another layer by asking how sound travelled between those positions, how reverberation affected speech and whether architecture enabled or obstructed understanding. The same thinking can ask how a ruined abbey shaped musical performance, how a theatre changed during refurbishment or how winter transformed the acoustic behaviour of a forest.

    Preserving the sound of a place is not simply an attempt to save an attractive reverberation before it disappears. Places participate in human activity. They change how people speak, perform and listen. Their surfaces and geometry influence whether voices remain intimate or become collective, whether musical phrases overlap or remain distinct, and whether somebody standing beyond a barrier can understand words spoken elsewhere. Acoustic conditions can shape behaviour, access and participation without leaving a visible trace.

    The Temple of Decision can still be visited. St Mary’s Abbey remains visible as a ruin. The old House of Commons chamber has disappeared, while the former typists’ room at Terry’s chocolate factory no longer performs its original function. A forest changes when snow arrives and changes again when it melts. Even a building that survives intact contains many acoustic relationships, only some of which can be measured during the limited hours when researchers have access.

    Sound is especially vulnerable to disappearance. Once a room changes, an audience leaves or a building is lost, its former acoustic behaviour cannot simply be photographed. Murphy’s lecture showed that it is not entirely beyond preservation. What remains may be a measurement, a model, a comparison or a carefully documented uncertainty, each offering a different way of understanding a place as more than a visual container for history.

    In the absence of clear echoes, we may never know a place completely. We can, however, preserve evidence of how sound moved through it, reconstruct plausible relationships when the original has disappeared and ask what those acoustics meant for the people who performed, spoke and listened there. An impulse response may last only a few seconds, yet within it can remain an acoustic trace of a cathedral, a factory, a theatre, a cave, a forest under snow or a room about to change forever. When the room itself has already disappeared, a model can return a possibility rather than a certainty: a voice reflecting from lost walls, a choir inhabiting a ruined abbey or political speech travelling towards listeners hidden above a chamber from which they were excluded.

    Preserving the sound of a place means preserving another way of understanding what happened there. Walls determine more than what people can see. They shape what can be heard, how clearly it can be understood and who is able to listen.

  • How Do You Design Sound for an Experience You Cannot Control? Wylie Stateman on Storytelling, Simplicity, and the Future of the Soundtrack

    Wylie Stateman

    How do you design sound for an experience you cannot control?

    A filmmaker can frame an image. The edges of the screen define what the audience sees, while composition, focus, lighting and editing direct attention within it. Sound is less obedient. It extends beyond the frame, surrounds the audience, enters rooms with different acoustics and reaches listeners through systems ranging from enormous cinema installations to headphones, televisions, laptops and tiny mobile-phone speakers. A soundtrack may be created with extraordinary precision, yet nobody involved in its production can completely control how, where or at what level it will finally be heard.

    During his online guest lecture for Edinburgh Napier University, supervising sound editor and sound designer Wylie Stateman explored the creative and professional consequences of working with such an elusive medium. Drawing upon a career spanning more than four decades and collaborations with filmmakers including Quentin Tarantino, Oliver Stone and John Hughes, he described sound as an art form positioned between science and subjectivity. Its physical behaviour can be measured, yet its meaning depends upon perception, attention, expectation and context. His work has extended from feature films and television to advertising, audiobooks and theme-park attractions, but a consistent question connects these different forms: how can sound professionals create an experience that remains clear, emotionally purposeful and dramatically coherent when both production and listening contain so much uncertainty?

    For Stateman, answering that question requires sound practitioners to think beyond individual sounds. A designer may create remarkable material, but the audience experiences relationships: dialogue against music, effects within an environment, silence before an impact and the entire soundtrack through a particular playback system at a particular level. Creating those relationships at scale requires collaboration, while maintaining them requires somebody to understand the complete experience. Throughout the lecture, Stateman moved repeatedly between these levels, from the organisation of large creative teams to the placement of a single sound and from the controlled environment of the mixing stage to the unpredictability of a listener pressing play somewhere else. Connecting them was a consistent philosophy. Complexity is unavoidable behind the scenes, but it should produce clarity for the audience.

    Stateman began with the unusual nature of sound itself. A visual composition can be stopped and examined. People can point towards a particular area of an image, discuss its composition and compare alternatives while the material remains stationary. Sound exists through time. A mixer listens, makes an adjustment, returns to an earlier point and listens again. Creative refinement becomes a repeated movement backwards and forwards through the material, making the number of meaningful passes completed within a working day directly relevant to the sophistication of the result.

    Such a process makes both technology and collaboration important, but neither can replace judgement. Stateman described audio as a discipline in which scientific knowledge and subjective interpretation continually meet. Vibration can be measured objectively, while a listener’s response to it cannot be reduced so easily. Designers therefore work simultaneously with physical systems and human perception. Loudspeakers, auditoria, codecs and playback levels matter, but so do memory, expectation, emotion and attention. Even the most technically controlled production process eventually encounters a listener whose response cannot be engineered with the same certainty as the system delivering the sound.

    Perhaps this uncertainty helps explain why Stateman spoke so strongly about collaboration. Rather than presenting professional success as the achievement of a solitary creative individual, he described a career built through relationships with people whose abilities complemented his own. Soundelux, the company he established with fellow sound designer Lon Bender, grew from a small operation into an organisation employing hundreds of people across several cities. Its development depended upon far more than creative talent. Business management, accounting, engineering, technology, sales and production all had to support the work of designers, editors and mixers.

    The lesson for students was not that everyone should attempt to build a large company. Stateman’s broader point concerned complementary ability. Nobody needs to become equally skilled at every aspect of creative and professional life. Someone with little interest in finance needs a trustworthy person who understands it. A creative specialist working on complex productions benefits from engineers and technologists capable of turning ideas into reliable systems. Partnerships become valuable when they extend what a group can imagine and accomplish rather than merely reproducing the same abilities several times.

    Professional value also emerges through reliability. Stateman described a valuable colleague in strikingly simple terms: somebody who can understand a problem, take responsibility for it and allow everyone else to stop worrying about it. Creative ability matters enormously, but large productions depend upon trust. A person who solves a problem without creating several new ones becomes increasingly valuable to the people around them. Careers are built not only through the quality of isolated work, but through the confidence that others can place in somebody when a difficult problem arrives.

    Filmmaking itself operates through the same interdependence. Every specialist inevitably perceives the project through a particular discipline. Sound designers think about sound, composers about music, cinematographers about light and composition, costume designers about clothing and production designers about the physical world. Stateman compared this to a collection of unreliable narrators, each understanding the film from a particular perspective. The director’s responsibility is to bring those partial perspectives together into a coherent experience.

    Sound professionals therefore need both commitment to their own discipline and awareness of the larger work. Stateman’s long relationships with individual filmmakers allowed such understanding to develop across several projects. By the time of Once Upon a Time in Hollywood, he had worked with Quentin Tarantino on seven films. Repeated collaboration created trust and a shared shorthand, allowing Stateman to take substantial creative ownership of the soundtrack while remaining clear that every decision ultimately served the director’s film.

    That balance between ownership and service runs through much of professional sound design. Creative contribution requires conviction. A designer who merely waits for instructions cannot provide the full value of specialist expertise. Yet conviction must remain connected to the filmmaker’s intention rather than becoming an opportunity to demonstrate technique. Stateman’s account suggested a form of authorship that is confident without becoming possessive. Sound teams need enough ownership to make strong decisions, while recognising that those decisions belong within a larger work whose purpose they must understand.

    Understanding that purpose becomes increasingly important as the available technology grows more powerful. More capability does not automatically justify more activity, and one of Stateman’s clearest demonstrations came from work beyond conventional cinema. Soundelux developed audio and show-control systems for large theme-park attractions, including a Terminator 2 experience designed to move hundreds of visitors through repeated performances every day. Audiences passed through a pre-show before entering a theatre containing three 3D IMAX screens and a motion base. Reliability was essential, but technical reliability alone could not create an engaging experience.

    Despite having access to the motion platform throughout the attraction, the production reserved its major movement for a single moment. Audiences were allowed to become comfortable before the platform suddenly dropped. Their physical reaction was powerful precisely through its rarity. Continual movement would have made the mechanism familiar and reduced its dramatic value. Behind one carefully timed surprise sat engineering, amplification, show control, multiple screens, motion systems and the work of a large team. None of that complexity needed to become the audience’s concern. They experienced the result.

    The same principle shaped Stateman’s approach to film sound. A system may offer hundreds of channels, extensive spatial control and enormous dynamic range, but the designer still needs to decide which possibilities deserve to be used. Creative sophistication can appear through restraint, and technical capability becomes valuable when it helps direct attention rather than continually demanding it.

    When asked how he decides what an audience should attend to, Stateman returned to simplicity. Film combines visual and auditory information, but listeners cannot process every available element with equal attention. A dense soundtrack may contain extraordinary detail while communicating very little. The designer’s task is therefore not to make everything audible at once, but to create a clear path through the scene.

    Dialogue often provides the starting point. Human listeners extract extraordinary amounts of information from voices. Words communicate explicit meaning, while rhythm, pitch, timbre, hesitation and vocal effort reveal character and emotion. If audiences struggle to understand what somebody is saying, an essential layer of narrative and performance has been weakened. Music offers another route through the experience, leading or following emotional movement, preparing audiences for change or allowing feeling to emerge after an event. Environments and sound effects establish space, physicality, scale, tension and perspective. None of these categories possesses permanent priority. What matters is understanding what the audience should receive from a particular moment and organising the soundtrack accordingly.

    Stateman’s preferred method was additive rather than deductive. Instead of filling a soundtrack with every plausible sound and gradually removing whatever causes problems, begin with what is essential. Establish the dialogue. Introduce music when the scene requires it. Add environment to create space and context. Bring in effects deliberately, allowing each contribution to justify its presence. This approach makes purpose part of the design process before complexity accumulates.

    Spatial sound presents the same challenge on another scale. Contemporary systems allow designers to place and move sounds through increasingly elaborate speaker arrangements, but movement acquires meaning only through its relationship with the story. Surrounding listeners with constant activity can make spatial information less expressive rather than more. If every sound moves, movement itself loses significance. Space becomes useful when its behaviour supports attention, perspective or dramatic intention.

    For Once Upon a Time in Hollywood, the absence of a conventional composed score created a distinctive set of possibilities. Music arrived through records, radio and material connected closely with the period represented by the film. Sound design could occupy areas that might otherwise have belonged to score, contributing low-frequency energy, changes in texture and transitions that influenced mood without announcing themselves as musical cues. The team did not begin by asking how every capability of contemporary soundtrack production could be demonstrated. They considered what would make the film feel connected to 1969.

    That thinking also informed Stateman’s use of Dolby Atmos. He regarded the format as an impressive creative environment, but not as a reason to send objects continually moving through the auditorium. Early experimentation with expanded spatial systems could become cluttered when additional speakers were treated as spaces waiting to be filled. Stateman instead described a stable foundation with carefully selected individual elements used when spatial movement genuinely contributed to the experience.

    For a film drawing heavily upon period recordings, radio and two-channel music sources, aggressive object movement could have conflicted with the aesthetic world being created. A technically advanced format can sometimes serve a film most effectively by concealing its sophistication. As with the single movement of the Terminator 2 platform, possibility acquires value through selection.

    Working with very different directors reinforced Stateman’s resistance to universal solutions. The sonic worlds of John Hughes and Oliver Stone required radically different forms of expression. Moving between comedy and films such as JFK, Born on the Fourth of July and Natural Born Killers prevented one successful method from becoming a formula applied repeatedly. An early mentor had encouraged Stateman to approach every project as a new problem and continue trying different methods rather than relying upon established answers.

    Experience, from this perspective, should expand the vocabulary available to a practitioner rather than narrow the range of possible responses. A Quentin Tarantino film does not require the same sonic logic as an Oliver Stone film, and neither can be approached as a variation of a John Hughes comedy. Even repeated work with the same director changes as the project changes. Trust allows a shorthand to develop, but familiarity should not turn into repetition.

    Comedy offered a useful example of this contextual thinking. Familiarity can establish a pattern, while an unexpected interruption creates surprise. Yet the same broad relationship between expectation and disruption can also support drama, suspense and horror. A technique has no fixed emotional meaning outside its context. What matters is the audience’s developing expectation and the moment at which the soundtrack confirms, delays or overturns it.

    Such decisions require somebody to understand more than individual sounds, which led towards the most distinctive professional idea in Stateman’s lecture. He argued that contemporary practitioners should think of themselves not only as sound designers, but as sound directors, sound producers and sound designers. These are not three disconnected jobs. They represent different perspectives on a shared responsibility for the complete sonic experience.

    The sound director understands intention. This person can discuss the desired experience with the filmmaker, interpret creative needs and establish an overall direction for the soundtrack. The sound producer understands resources. Schedules, budgets, staffing and priorities determine what can be achieved and where effort should be concentrated. The sound designer turns those intentions and resources into creative work. Stateman regarded the combination of these abilities as a professional ideal.

    Directors themselves often move through a similar expansion of responsibility. A writer may become a director to shape the interpretation of the material, then become a producer to gain greater influence over resources and priorities. Stateman argued that sound professionals can develop in a comparable way. Technical expertise remains essential, but understanding intention, organising people and resources and making creative decisions gives the practitioner a more substantial role in shaping the work.

    His career offers several examples of this expanded responsibility. Building large creative teams required production thinking. Designing complex theme-park attractions demanded an understanding of playback systems and engineering. Long-term relationships with directors depended upon interpreting intention rather than waiting for technical instructions. Global distribution and localisation required the sound team to consider what happens after the supposedly finished soundtrack leaves the mixing stage.

    At that point, the problem of control becomes larger. Stateman may work in a highly controlled Atmos environment containing an extensive loudspeaker system, while an audience member eventually hears the result through earbuds on a train. Another listener watches through a television with speakers facing towards a wall. Somebody else lies in bed listening quietly while another person sleeps nearby. A theatrical soundtrack mixed at a high reference level may later be heard at a dramatically lower domestic level.

    No single master can guarantee an identical perceptual experience across all of those conditions. Low-level detail that remains clear in a cinema may disappear when the entire soundtrack is turned down. Wide dynamics that feel exciting in a theatre can become impractical late at night. Dialogue that is intelligible through one system may become difficult to follow through another. Translation therefore concerns more than whether a format can technically fold down from one speaker configuration to another. It concerns what survives perceptually when the listener, environment and playback level change.

    Even cinemas introduce uncertainty. Stateman described discussions with theatre owners about films being played below their intended reference levels. Their explanation was practical. A film begins at the specified level, somebody complains, and the level is reduced. Further complaints lead to further reductions until complaints stop. From the exhibitor’s perspective, this is a rational response to the audience in the room.

    Production can create pressure in the opposite direction. Filmmakers listening in a controlled mixing environment may repeatedly ask for greater excitement, prompting levels to rise until the result satisfies the room. One part of the system pushes upwards while another turns the finished film down. The problem cannot be solved by allowing dialogue, music and effects departments to maximise their own material independently. Somebody needs to maintain responsibility for the dynamic shape of the complete experience.

    Dynamics, in this sense, organise attention across time. Genuine intensity requires quieter material around it. Low-level information needs to remain meaningful under realistic playback conditions, while moments of scale need enough contrast to feel exceptional. The same principle that made one movement of the Terminator 2 platform effective applies to the soundtrack more broadly. Constant intensity reduces the expressive power of intensity itself.

    Global distribution introduces another kind of variation. A soundtrack may need to function across many languages, each with different rhythms, durations and vocal characteristics. Preserving quality across those versions is not merely an administrative task completed after the creative work. It is an audio-design problem involving performance, mixing, workflow and technology.

    Stateman’s approach was to break an apparently overwhelming challenge into smaller problems without losing sight of the whole. A modern feature soundtrack can require enormous quantities of editorial and mixing time, making meaningful control by one person impossible. Large teams become necessary, yet specialisation creates the risk that individuals see only their own component. The sound director-producer-designer model provides a means of maintaining overall intention while complex work is distributed among specialists.

    Localisation makes the limitations of a single fixed soundtrack particularly visible. Different languages alter timing and vocal behaviour, while different audiences and distribution systems introduce further variables. Streaming services now possess forms of information and control unavailable to earlier theatrical systems. Different languages and formats can already be delivered to different users. Stateman’s discussion suggested a logical extension of that capability: audio systems that respond more directly to individual listeners and the circumstances in which they are listening.

    Rather than treating a soundtrack as one fixed master expected to survive every possible situation, elements could remain available for intelligent recombination. Dialogue might receive greater prominence where needed. Dynamic behaviour could change for quiet listening. Different relationships between elements might be created according to the user’s environment, hearing ability or preferences.

    Such possibilities do not necessarily weaken creative intention. They raise a more fundamental question about what preserving intention actually means. A listener who cannot understand the dialogue is not receiving the intended experience merely through being sent the same electrical signal as everyone else. Someone listening quietly in bed occupies a different perceptual situation from an audience inside a calibrated cinema. Identical delivery does not guarantee equivalent experience.

    Adaptive audio therefore extends rather than abandons Stateman’s earlier principles. If the purpose of sound design is to guide attention, communicate story and create emotional relationships, then changes in listening circumstances matter. The challenge is to determine which aspects of an experience must remain stable and which can change to preserve them. A theme-park attraction, theatrical soundtrack, localised release and adaptive audio system appear very different, yet each requires designers to think about the complete journey between creative intention and audience experience.

    Across Stateman’s lecture, simplicity emerged not as an absence of complexity, but as its successful organisation. Hundreds of people may contribute to a project. Thousands of hours may be spent editing and mixing. Sophisticated engineering may sit behind the delivery system. The audience does not need to experience any of that as complication. They need to understand a voice, feel a change in atmosphere, anticipate an event or be surprised by a sudden movement.

    A film soundtrack can contain vast numbers of tracks, edits, recordings and processing decisions while resolving into a clear moment of attention. Achieving that clarity requires more than technical skill. Designers need to understand what matters within the scene. Producers need to organise the resources that make the work possible. Sound directors need to maintain the relationship between individual decisions and the complete experience. Teams need complementary expertise and enough trust for difficult problems to be handed to people capable of solving them.

    Stateman’s lecture ultimately presented sound design as the organisation of attention through time. The designer decides what matters now, what can wait, what should disappear and what should arrive unexpectedly. The producer makes those decisions achievable. The director maintains an understanding of why they matter. Technology expands the available possibilities, but cannot decide which ones belong in the experience.

    That distinction becomes more important as audio technology develops. Spatial formats offer increasingly detailed control inside theatres and listening rooms. Streaming services distribute content globally and create new demands for localisation. Object-based systems can preserve elements beyond a fixed master, while adaptive delivery may eventually allow soundtracks to respond more intelligently to listeners and environments. Each development increases possibility, but also increases the need for judgement.

    Somewhere beyond all of those systems, a listener is trying to follow a story. They may be sitting inside an elaborate cinema, travelling on a train, watching a laptop or lying quietly in bed. The sound professional cannot control every condition surrounding that experience, but can understand the relationships that matter. Dialogue can remain clear. Attention can be guided. Complexity can be organised. Dynamics can create contrast rather than exhaustion. Technology can serve intention rather than advertise itself.

    The response to uncertainty is not complete control. It is to design intelligently for variation while remaining clear about intention. Build teams capable of solving problems that no individual could solve alone. Understand the whole experience while taking responsibility for one part of it. Use complexity behind the scenes to create clarity for the listener. Know the story well enough to recognise when a convention should be followed, when it should be overturned and when the most powerful use of a technology is to leave it silent.

    A soundtrack may be shaped inside a carefully controlled room, but it does not remain there. It travels through cinemas, languages, formats, devices, environments and listeners. Every stage changes the circumstances in which the work will be experienced. Stateman’s lecture suggested that the future of sound design lies not in pretending those differences can be eliminated, but in understanding them well enough to preserve what matters. The mix may be finished when it leaves the studio. The experience begins when somebody, somewhere, presses play.

  • How Do You Design the Sound of Reality? Sefi Carmel on Documentary Sound, Perception, and the Ethics of Construction

    Sefi Carmel

    How do you design the sound of reality?

    A whale dives beneath the surface of the Atlantic Ocean. Filmed from a distant boat, its tail disappears into the water and a splash is heard. The moment appears entirely natural, yet the camera crew may have been hundreds of metres away, surrounded by engine noise and incapable of recording anything resembling the sound presented in the finished film. Perhaps the splash came from a sound-effects library. Perhaps somebody dropped an object into water. Perhaps a Foley artist moved a scuba flipper through a bathtub. Does adding that sound make the documentary less truthful, or does it help audiences experience an event that genuinely occurred but could not be captured adequately during filming?

    During his online guest lecture for Edinburgh Napier University, London-based sound designer, composer and dubbing mixer Sefi Carmel explored the creative and ethical questions surrounding soundtrack creation for documentaries. Drawing upon experience of mixing more than one hundred documentaries, he challenged the assumption that factual filmmaking requires a fundamentally different sonic vocabulary from drama. Dialogue, music, atmospheres, spot effects, Foley and abstract sound design can all contribute towards documentary storytelling. The crucial issue is not whether a sound was recorded at the moment shown on screen, but what its addition asks the audience to believe.

    Throughout the lecture, Carmel returned to a broad understanding of the soundtrack. Everything emerging from the speakers belongs to a single composition created in relationship with the image. Dialogue, music and effects may be separated technically, but audiences experience their combined movement through time and space. Like music, a soundtrack arranges rhythm, dynamics and timbre to communicate emotions and ideas. Documentary sound therefore involves much more than cleaning interviews and placing music underneath them. It is the construction of an audiovisual experience whose materials may come from reality while their organisation remains an act of filmmaking.

    Carmel began by questioning a belief he had held as a young sound designer. News reporting and documentary filmmaking had once appeared to occupy similar positions on a spectrum between reality and fiction. At one extreme, news aspires towards an account of events with minimal manipulation. At the other, drama openly asks audiences to accept scripted performances, constructed edits, designed sound effects and music intended to influence emotion. Documentary can initially appear closer to the first model, yet Carmel’s experience of the form led him towards a different conclusion. A feature-length documentary still has to hold an audience for sixty, ninety or more minutes. It communicates ideas, develops relationships, establishes places, controls pace and creates emotional movement. Those demands make it filmmaking rather than an extended news report.

    Understanding documentary in those terms considerably widens its creative possibilities. Music can shape emotional interpretation, sound effects can strengthen actions and atmospheres can establish locations that production recordings fail to communicate. Archive footage can be reconstructed into a convincing audiovisual world, Foley can restore physical detail and abstract sound design can emphasise an important transition or idea. Documentary makers have access to almost the complete filmmaking toolbox. With that freedom comes an ethical problem. If documentary claims a relationship with reality, how far can its soundtrack depart from literal recording before enhancement becomes deception?

    The whale provides a useful test. The animal really entered the water and its tail really created a splash, even though the filmmakers could not capture that sound from their position. Adding a plausible splash does not invent the event. It reconstructs an acoustic consequence already implied by the image, allowing audiences to experience the animal’s movement and scale more immediately without asking them to believe that something happened when it did not. Matters become more difficult when sound influences interpretation rather than restoring an unheard event. Carmel drew his clearest ethical line around speech. Reconstructing the likely sound of an action differs substantially from editing somebody’s words in a manner that changes their meaning or context. The latter can alter the evidence from which audiences understand an event. For Carmel, creative sound design remains legitimate when it serves the film without becoming a gross lie.

    A much less serious encounter with perceptual truth came during his work on a documentary about the Winter Olympics. Faced with a distant shot of a skier slaloming down a snowy slope, the director wanted the movement to have greater sonic presence. Library searches failed to produce a suitable recording, so Carmel created one himself. The director liked the result and repeatedly asked what had produced it. Carmel initially refused to answer, but eventually revealed the source: a knife scraping across toast. The sound had worked perfectly while the director perceived it as skiing. Once its origin became known, however, the illusion collapsed. The director could hear only toast and eventually asked for it to be removed. Nothing in the waveform had changed. Image, expectation and context had previously allowed one event to become another, while knowledge of the source created a new and apparently irreversible interpretation.

    Listeners do not identify sounds solely through their acoustic properties. Visual information, expectation, context and prior knowledge shape perception. A recording does not acquire credibility merely through sharing a physical origin with the object shown on screen, and a constructed sound does not automatically become deceptive through coming from somewhere else. Credibility emerges from the relationship between sound, image and meaning. Documentary sound design occupies a space between physical truth and perceptual credibility, requiring the designer to consider where a sound came from alongside what audiences understand it to represent.

    Reality presents another difficulty long before questions of creative enhancement arise. Documentary crews work in environments that cannot be controlled in the manner of a drama production. Interviews happen near roads, beneath aircraft routes, inside noisy buildings and beside machinery. Important moments may occur only once, leaving post-production dependent upon whatever the location recordist managed to capture. Documentary also has limited access to ADR. Replacing a contributor’s voice with a later studio performance can compromise spontaneity, authenticity and practicality. An imperfect recording of an essential contribution therefore has to be made intelligible and aesthetically acceptable even when the original conditions were hostile.

    Some problems are comparatively manageable. Low-frequency rumble can be reduced, while constant noises such as air conditioning, electrical hum or an aircraft cabin may respond well to noise reduction. Variable interference presents a much harder problem. An accelerating motorcycle can move through the same frequency regions as speech while continually changing its spectral character, making removal difficult without damaging the voice. Aggressive processing introduces further compromises. Equalisation can isolate the most intelligible part of a voice while leaving it thin and unnatural. Noise reduction can remove interference at the cost of audible artefacts. A line may become technically clearer but aesthetically less convincing. Restoration is not a contest to remove the greatest possible quantity of unwanted sound. Every intervention changes the material audiences hear.

    Carmel connected this problem to the idea of aesthetic disturbance. Drawing upon Ludwig Wittgenstein, he described aesthetics partly through the recognition of something being wrong: a picture hanging at an angle, for example, produces a disturbance that disappears when the relationship is corrected. Documentary sound can create similar disturbances. A close image accompanied by an unexpectedly distant, reverberant voice may produce audiovisual dissonance. Distortion attracts attention, while harshness, brittleness or excessive boxiness can make listening uncomfortable even when every word remains understandable. Intelligibility is necessary but insufficient. Dialogue also needs to feel congruent with the image and surrounding soundtrack, and a technically rescued voice can still undermine a scene if its perspective, spectrum or acoustic character appears disconnected from what viewers see.

    Modern restoration tools have expanded what can be recovered. Carmel discussed noise reduction, de-clicking, de-crackling, de-clipping, dereverberation and spectral repair as processes capable of rescuing recordings that might once have been considered unusable. Constant noise can sometimes be reduced substantially, while sophisticated interpolation may reconstruct clipped or distorted speech with surprising effectiveness. Greater capability does not remove the need for judgement. Restoration can erase information that belongs to the story. Crackle on an old recording may be technically undesirable while simultaneously communicating historical distance. The noise of an old shellac disc or archive recording can help audiences understand that they are hearing material from another period. Removing every imperfection may weaken the narrative. Technical possibility matters less than understanding what the existing sound already communicates.

    This tension between repair and preservation led to one of Carmel’s central principles: ideally, every layer of the soundtrack should be capable of telling the story. Dialogue carries narrative through words. An atmosphere can communicate location, weather, time of day and activity without spoken explanation. Music can reveal emotional direction. A forest atmosphere containing birds, wind and running water immediately places listeners within a particular kind of environment. Each layer contributes different information, and the complete soundtrack emerges from their interaction rather than from one dominant element surrounded by decoration.

    Documentary production rarely provides ideal materials, so those layers often have to support one another. Aggressively cleaned dialogue recorded on a windy location may no longer contain enough environmental information to make its setting believable. Carefully constructed atmospheres can return that context. Sound effects can reinforce actions the original recording failed to capture, while music can support emotional movement that damaged or fragmented production sound cannot carry alone. None of these layers needs to become conspicuous. A weak recording may become convincing once placed inside an appropriately designed environment, since audiences hear relationships between elements rather than evaluating every track in isolation.

    Voiceover introduces another distinctive element within those relationships. Casting is often a directorial decision, though Carmel argued that sound professionals can contribute valuable thinking when invited into the process. The appropriate voice depends upon the subject, intended audience and emotional character of the film. A documentary about the rise of a young pop group requires a different vocal identity from one exploring humpback whales in the North Atlantic. Performance matters as much as casting. Tone, energy, pacing and character shape the audience’s relationship with the film, while the chosen voice influences the space available for music, effects and atmosphere. Voiceover is part of the soundtrack’s composition, not information simply placed above it.

    Music performs an equally integrated role. Carmel distinguished between music existing within the world shown on screen and music functioning as score. Source music might come from a visible performer, radio, jukebox or other plausible location within the scene, while score operates outside that visible world and shapes emotion from another level. Documentary makers can also blur the distinction deliberately. An old rock-and-roll song treated with band-limiting and room reverberation might appear to come from a jukebox in a diner, allowing music to reinforce the setting, contribute historical or cultural information and perhaps help mask weaknesses in the location sound.

    Selecting or composing music requires an understanding of everything else occupying the scene. Dialogue-heavy sequences need music capable of supporting speech without continually competing for attention. Dense lead instruments or vocals can occupy perceptual and spectral territory similar to the human voice, forcing the music much lower in the mix. More restrained arrangements can create emotional colour while leaving space for narration and interviews. Other sequences allow music to carry more of the narrative. An expansive aerial view with little dialogue may support a large thematic statement that would overwhelm an intimate interview. Musical effectiveness in isolation matters less than the role a piece needs to perform at a particular moment and the relationships it forms with the rest of the soundtrack.

    Location dialogue complicates those relationships further. A controlled voiceover recording tends to maintain comparatively stable level and performance, while spontaneous speech can vary considerably. Contributors may begin sentences with energy and trail away as thoughts conclude. Music automation must respond to those changing patterns. A static reduction may leave quieter words obscured or make stronger phrases feel unnecessarily exposed. Mixing becomes a continual negotiation between intelligibility and musical continuity, with the soundtrack moving around the natural behaviour of voices that were never performed for the convenience of the mixer.

    Atmospheres perform several roles simultaneously. Most obviously, they tell audiences where they are. Traffic, birds, wind, room tone, distant machinery or human activity can define an environment before viewers consciously analyse the image. They also smooth editorial transitions. Documentary scenes are frequently assembled from material recorded at different moments, positions or even days, and a continuous environmental bed can help separate pieces of location sound feel as though they belong to a coherent space. Carmel identified another, less obvious function: atmospheres can contribute spectral balance. If a scene feels sonically empty within a particular frequency region, an appropriate environmental layer can help create a more aesthetically satisfying whole. The choice still needs to make narrative sense, but storytelling and sonic composition overlap here. Atmosphere can provide information, continuity and texture at the same time.

    Spot effects operate on a more local scale. A car door closes, a telephone is placed down or a gun fires. Brief synchronised events can reinforce visible actions and restore details absent from production recordings. Archive footage provides particularly rich opportunities for this kind of reconstruction, especially when historical images arrive without usable synchronous sound. Old footage may be silent or accompanied by narration and music unsuitable for the contemporary documentary. A battlefield sequence showing tanks, artillery and soldiers therefore presents the sound designer with an empty world that needs to be rebuilt.

    For Carmel, archive reconstruction can be approached with the same dramatic ambition used in fiction. Tanks can advance through the frame, gunfire can occupy different distances and artillery can establish scale, while wind across an exposed landscape gives the environment continuity. The designer might process the soundtrack to suggest historical recording technology or create a vivid modern sound world that places the audience imaginatively inside the event. A deliberately aged soundtrack reminds viewers that they are encountering archive material, while a contemporary reconstruction can reduce historical distance and make an event feel immediate. Neither choice is neutral. Both interpret history, leaving the designer to consider the relationship the documentary seeks to create between the audience and the past.

    Foley can contribute in much the same way, although documentary schedules and budgets rarely permit complete coverage. Selective use can still transform significant moments. A historical reconstruction showing chainmail being worn, armour handled or a sword drawn may deserve detailed physical sound even when none was captured during filming. If the moment carries narrative importance, there is no reason to reject Foley merely through an assumption that documentary sound must remain limited to location recordings. Abstract sound design extends the same principle further. Drones, impacts and heavily processed transitions can strengthen important ideas or structural moments, creating unease, giving a cut greater dramatic force or helping an audience experience a transition emotionally as well as intellectually.

    For Carmel, the legitimacy of these devices depends upon purpose. Dramatic sound should strengthen the film rather than substitute manipulation for argument. Documentary editing, cinematography and music already influence how audiences understand material, and sound design participates in the same process. Ethical responsibility lies in recognising what each construction communicates rather than pretending that construction does not occur. Documentaries are built from real people, events, evidence and places, yet films do not emerge automatically from those materials. Someone selects shots, orders sequences, chooses where music begins, decides when silence matters and determines which details audiences hear. Soundtrack creation is part of that authorship.

    Creative decisions are only part of the work. The documentary must also survive the circumstances in which it will be heard. Carmel emphasised that mixing begins with the destination. A theatrical documentary, television broadcast, online film and festival screening present different playback conditions and technical expectations. Format, dynamics, equalisation and level decisions need to reflect those contexts. A large theatrical environment can support substantial low-frequency extension and wider dynamics, while television playback may occur through much smaller speakers and in less controlled surroundings. Processing that creates clarity in one context can become harsh or excessive in another.

    Platform awareness connects technical delivery directly with audience experience. Spectral energy inaudible on small television speakers can still consume headroom. Extreme dynamics may work beautifully in a cinema while causing viewers at home to continually adjust volume. Loudness standards formalise part of that relationship, particularly for broadcast delivery, but Carmel’s broader point concerned the distribution of intensity across the film. Loudness can be understood as a budget. If every moment is treated as maximally intense, little room remains for genuine peaks. Dynamic planning becomes another form of storytelling, allowing contrast to carry dramatic meaning rather than treating level merely as a compliance problem.

    Compression and limiting require similar contextual judgement. A theatrical mix may use comparatively subtle master processing, preserving headroom and contrast, while television and online material may tolerate or require greater control. One version is not inherently superior to another. Each mix needs to function within the medium for which it is intended, preserving the film’s intentions under different listening conditions.

    Deliverables extend that responsibility beyond the primary audience. Documentary films may travel between territories and require new narration or dubbed dialogue. Music and effects tracks need to support localisation rather than reproduce automation created around the timing of the original language. Carmel explained the value of providing undipped music for this purpose. In an English version, music may be reduced beneath a particular phrase and raised again when the speaker stops. A translated version may take longer or shorter to communicate the same idea. If the music stem already contains automation tied to the English timing, the foreign-language mixer inherits a structure that no longer fits. Providing material without those dialogue-specific reductions allows the new mix to respond properly to the translated performance.

    A soundtrack therefore exists as more than a finished mix. It may need to survive new languages, platforms and contexts, and good delivery anticipates the work of people who will encounter the material later. Carmel ended with an even simpler responsibility: check the work. Quality control may appear less intellectually exciting than documentary ethics or perceptual sound design, yet it protects every creative decision made before delivery. A mixer can spend days constructing a sophisticated soundtrack and still send an unusable file through a routing mistake or export error. Recording or exporting something does not prove that the expected material exists in the resulting file. Listen to it. Watch it. Check it.

    That practical instruction sits neatly alongside the lecture’s larger argument. Documentary soundtrack creation moves constantly between interpretation and responsibility. Designers can construct sounds that were never recorded, rebuild silent archives, use Foley, shape emotion through music and introduce dramatic sonic devices. Those freedoms demand judgement. Does a sound restore an experience, clarify it, interpret it or falsify it? What does it ask the audience to believe? Does it support the film’s argument without altering the meaning of its evidence?

    Carmel’s lecture presented documentary sound as an art of constructing experience from incomplete reality. Location recordings arrive damaged. Cameras capture actions from distances at which their sounds cannot be heard. Archive images survive after their original acoustic worlds have disappeared. Interviews need to coexist with music, while fragmented scenes require atmospheres capable of making them feel continuous. The documentary soundtrack is built through responses to these absences. Sometimes the response is technological: remove a constant noise, repair distortion or restore intelligibility. Sometimes it is editorial: create a continuous atmosphere around fragmented material. Sometimes it is performative: add Foley to a significant physical action. Sometimes it is musical, shaping the emotional direction of a sequence. At other moments, the solution may be a completely unrelated object whose acoustic behaviour happens to make an image believable.

    A knife scraping toast can become a skier moving across snow, at least until somebody learns the secret. Reality does not arrive in post-production as a complete audiovisual object waiting to be preserved. It arrives as recordings, images, testimony, fragments and absences. Filmmakers decide how those materials should be organised into an experience that audiences can follow and understand. Sound designers participate in that process by reconstructing relationships between actions and consequences, voices and spaces, images and expectations.

    Construction and dishonesty are not the same thing. A documentary soundtrack can be richly designed while remaining faithful to the people and events it represents. Literal accuracy may sometimes be essential. At other moments, perceptual credibility communicates an experience more effectively than an unusable or absent location recording ever could. A splash can give weight to a whale entering the ocean. An atmosphere can return a damaged interview to its environment. Designed sound can give silent archive footage physical immediacy. Music can reveal emotional relationships without changing the evidence shown on screen.

    Documentary sound occupies the space between what happened, what could be recorded and what audiences need in order to understand and feel the film. Carmel’s lecture showed that this space is not a technical inconvenience to be hidden. Much of the creative work begins there. The documentary sound designer cannot preserve every sound of reality, since many were never captured in the first place. The responsibility is to decide what should be repaired, what can be reconstructed, what needs to remain imperfect and what must never be changed.

  • How Do You Make a Game Feel Dangerous? Will Morton on Emotion, Attention, and Designing Sound for the Player

    Will Morton

    How do you make a game feel dangerous?

    A gun can sound enormous and still fail to make a gunfight frightening. Every weapon may have a powerful attack, convincing mechanical detail and an impressive environmental tail, yet the player can remain strangely detached from the danger. Solving that problem may have little to do with redesigning the weapon itself. Bullets pass close to the head. Impacts strike nearby walls with exaggerated force. Brickwork breaks apart, fragments scatter and the environment appears to react violently to the threat. During his online guest lecture for Edinburgh Napier University, game audio designer Will Morton explored how sound can shape emotion, focus attention and guide players through complex interactive experiences. Drawing upon twelve years at Rockstar North, where his work included the Grand Theft Auto series, Red Dead Redemption and L.A. Noire, followed by the establishment of Solid Audio Works with fellow former Rockstar audio specialist Craig Connor, Morton presented game sound design as a discipline of selection. Thousands of sounds may exist within a game, but their value depends upon knowing which ones matter at any particular moment. Throughout the lecture, one principle repeatedly emerged. A designer must ask not only what a game world should sound like, but what the player needs to hear in order to feel what the game intends them to feel.

    Morton began by placing creative decisions within the realities of AAA game production. Large games are expensive, technically constrained and continually changing. Development rarely follows a fixed design from beginning to end. Features evolve, producers reconsider decisions and new requests arrive late in production. Platform restrictions impose further limits through storage, memory, streaming performance and processing capability. Scale introduces organisational complexity as well. Larger games require larger teams, while experienced creative specialists can find increasing amounts of their time absorbed by scheduling, administration and coordination. Sound design develops inside a moving system of technical, financial and organisational constraints. Success requires more than imagining an ideal soundtrack. Designers must create one capable of surviving years of changing requirements.

    Historical changes in technology have altered the scale of those constraints without eliminating them. Morton recalled Commodore 64 composer Martin Galway fitting numerous sound effects into approximately one kilobyte of memory. Restrictions of that magnitude appear almost comic from the perspective of contemporary production, yet modern open-world games can still leave audio teams fighting for storage and memory. Vastly greater resources are now available, while games simultaneously attempt to represent entire cities, landscapes and fictional worlds. Technological abundance creates new possibilities, but ambition expands alongside it. Designers still have to decide where limited resources will make the greatest contribution.

    Money introduces another set of choices. Sound effects require designers, recording equipment, locations, editing time, libraries, software and continually changing computer systems. Dialogue adds writers, actors, directors, studios, recording staff and extensive editing. Music may involve composition, licensing, performers, orchestras, specialist recording facilities and interactive implementation. Morton’s overview exposed the consequences hidden behind apparently simple creative ambitions. Another recording variation, character voice or interactive music feature consumes time, money, memory and attention that cannot be spent elsewhere.

    Dialogue provided one of the clearest examples of complexity hiding behind familiar production tasks. Morton spent much of his Rockstar career dividing his time between sound design and dialogue before the scale of Grand Theft Auto V led him to work entirely as dialogue supervisor. For story-heavy games, dialogue cannot be treated as a sequence of lines requested by designers and recorded by actors. Repetition needs consideration. Lines must make sense across changing gameplay situations. Story information may unfold over many hours or days of play, while players can interrupt, delay or alter the circumstances in which dialogue occurs. A dialogue designer therefore needs to understand the interactive structure of the game as deeply as the individual performances being recorded.

    Direction presents a related challenge. Experience in film and television can produce excellent performances, yet game dialogue introduces problems absent from linear media. Cutscenes may operate much like conventional scenes, while in-game dialogue can occur within changing gameplay circumstances and across story structures experienced differently by individual players. Directors need more than an ability to elicit compelling performances. Detailed knowledge of the script, game and eventual context of each line becomes essential. A performance recorded in isolation must later remain convincing within circumstances that may not even be visible inside the recording studio.

    Morton also challenged assumptions about professional recording. Technically excellent dialogue can be captured with comparatively modest equipment in a carefully treated space. A suitable microphone, simple accessories and an acoustically controlled room may produce results approaching those of a far more expensive studio. Recording quality, however, forms only one part of a professional session. High-profile performers need confidence that they are participating in a serious production. Environment, organisation and treatment of the actor can influence trust in the project and relationships with agents and management. Professionalism encompasses the experience surrounding a recording alongside the technical quality of the resulting file.

    From these production realities, Morton moved towards the central creative argument of the lecture. Designers can easily assume that everything visible in a game should automatically produce a sound. Faced with an unsounded world, teams begin filling every action with detail. Cars receive engines and collisions. Characters acquire footsteps. Objects gain interactions. Environments fill with ambiences. A technically comprehensive and entirely plausible soundtrack gradually emerges, yet plausibility alone cannot guarantee clarity, excitement or emotional effect.

    Practical concerns provide one reason for restraint. Every additional sound requires creation, editing, implementation, memory and testing. Artistic considerations are even more important. If everything demands attention simultaneously, nothing receives focus. Morton encouraged designers to identify what deserves to be heard and what can remain absent. Important sounds require room to breathe. Dynamics rely upon contrast, while emotional emphasis depends upon moving attention between elements. Silence and omission become active design decisions.

    Player experience consequently takes priority over literal acoustic reconstruction. Real events often sound less dramatic than audiences expect. A gunshot captured from a particular position may seem surprisingly small. A suppressed weapon does not necessarily produce the familiar cinematic whisper audiences have learned to associate with it. A real minigun may collapse into an almost continuous mechanical roar instead of revealing every stage of its operation. Decades of film, television and games have established sonic conventions that now influence how audiences expect objects and events to behave. Designers work within that accumulated perceptual history.

    Morton demonstrated the point by comparing cinematic and real recordings of suppressed firearms. A familiar designed version sounded short, controlled and immediately recognisable, carrying characteristics audiences strongly associate with a silenced weapon. Real recordings behaved quite differently. Neither approach offered a universal answer. Documentary representation might favour acoustic accuracy, while an action game may need immediate recognition, excitement and dramatic satisfaction. Before designing a sound, Morton considers the role accuracy should play, whether an event needs to feel larger than reality and how strongly audience expectations should influence the result.

    Miniguns provided an even clearer illustration. Morton discussed the famous weapon sequence in Predator, where mechanical movement, spinning and other details reinforce the spectacle shown on screen. Real recordings present a very different impression, dominated by the extraordinary density of rapid gunfire. Grand Theft Auto: Vice City and Grand Theft Auto V offered further interpretations, each constructing the weapon differently, while Terminator 2 adopted another cinematic approach that remained closer to aspects of the real sound. Comparing them did not reveal a correct minigun. Instead, their differences showed how each design serves the experience of a particular production.

    Reality and relevance are therefore separate considerations. A game designed as escapism may gain little from reproducing everyday acoustic experience with complete fidelity. Players do not always need to hear events from the perspective of ordinary observers. They need a version that communicates the event’s importance within the game. Sound design selects, enlarges, simplifies and reshapes reality according to dramatic purpose.

    Attention can be directed just as effectively through subtraction. Morton used a remotely detonated explosive from Grand Theft Auto IV: The Ballad of Gay Tony to demonstrate how removing surrounding sound can make a single event dominate awareness. Immediately before the explosion, the wider soundtrack recedes and the bomb’s warning becomes the focus. He compared the moment with the seismic charge sequence from Star Wars: Episode II, where a brief interruption in the expected sound field creates anticipation and gives the subsequent event greater impact. Both sequences derive part of their power from absence.

    Focus involves more than increasing the level of an important sound. Every event competes with its surroundings. Removing distractions can achieve more than additional layers, greater loudness or further spectral exaggeration. Presence acquires meaning through contrast with absence. A fraction of a second of reduced activity can prepare an event more effectively than a continuously dense soundtrack.

    Morton’s most revealing example concerned a producer asking for guns to sound more dangerous. Existing weapon sounds already seemed successful to the audio team. They contained convincing mechanical detail, a strong initial attack, substantial body and satisfying environmental tails. Reworking those qualities did not solve the problem. Although the request appeared to concern the sound of the guns, the underlying dissatisfaction was emotional.

    An unexpected experience away from the studio suggested another approach. During a paintball game, Morton found himself sheltering behind structures made from metal oil drums. The paintball markers themselves produced relatively insignificant sounds. Fear came from sudden, violent impacts striking the metal around him. His immediate environment appeared to contract as attention focused upon nearby impacts, resonances and the sense that projectiles were arriving from directions he could not fully control. The weapon itself was not frightening. Being under fire was.

    Recognising that distinction transformed the design problem. Bullet passes became more prominent, with their levels responding more strongly to proximity. Impacts against brick and concrete gained force. Low-frequency energy added physical weight, while debris and crumbling material made the environment appear to react to nearby gunfire. Details that might be acoustically subordinate to a real gunshot were deliberately brought forward. Stronger weapon sounds had never been the answer. Gunfights needed to communicate vulnerability and danger.

    Here Morton identified a broader professional skill. Directors and producers rarely describe every sound problem in acoustic terms. They may ask for something to be louder, bigger, darker, faster or more dangerous while expressing dissatisfaction with an emotional result. Following the literal wording can send a designer towards the wrong solution. Morton argued for interpreting the intention behind the request. Someone asking for a more dangerous gun may actually be asking to feel vulnerable. Once the desired experience becomes clear, the designer can decide which part of the sound world needs to change.

    Creative collaboration therefore requires translation between intention and acoustic action. Sound professionals develop specialised vocabularies for frequency, dynamics, spatial behaviour, envelope and processing. Producers may describe experiences through emotion, imagery or metaphor. Useful information can exist within both forms of language. Professional expertise includes moving between them and identifying the experience concealed inside an apparently vague request.

    Gunfire also demonstrates how game audio operates through systems of cause and consequence. A weapon consists of more than the sound emitted when a trigger is pulled. Projectiles move through space. Near misses pass the player. Bullets strike walls. Debris falls afterwards. Environments respond differently according to material and distance. Danger emerges from relationships between these events. Concentrating exclusively upon the source can leave the wider experience emotionally incomplete.

    Similar systems govern the game mix. Morton described mixing a large game as a potentially overwhelming process involving thousands of assets and potentially hundreds of simultaneous channels, all changing according to gameplay. Exact combinations cannot be predicted in advance as they can within a linear soundtrack. Dialogue, music, ambience, vehicles, weapons, footsteps and environmental interactions combine differently according to player behaviour. Even excellent individual assets can produce an exhausting result when their relationships are poorly controlled.

    Morton recalled playing games whose soundtracks felt like continuous acoustic assault. Constant density can be as tiring as excessive level. When every category remains active and prominent, listeners receive no opportunity for recovery and little indication of where attention belongs. More detail can therefore produce less communication.

    Working with Craig Connor, Morton developed a useful analogy for managing this complexity: approach the game as a music producer approaches a song. Every element needs an appropriate place and enough space around it. The comparison does not imply imposing a static musical mix upon an interactive system. Its value lies in encouraging relational thinking. Sounds acquire meaning through their positions alongside other elements rather than through isolated perfection.

    Arrangement offers another useful parallel. Music producers rarely expect every instrument to occupy the foreground continuously. Parts enter and leave, density changes and contrast creates structure. Individual elements occupy different spectral, spatial and dynamic roles. Interactive audio can apply similar principles while responding continuously to player action, with priority changing according to what the player is doing and what the game needs to communicate.

    Waiting until asset production is complete before attempting a final mix creates serious problems. By the time thousands of sounds have accumulated, assumptions about level, density and priority may already be embedded throughout the project. Assets designed without a meaningful reference mix can prove difficult to reconcile, particularly when each has been created to sound impressive in isolation.

    Morton recommended an iterative alternative. Early in development, designers can choose a small but representative collection of sounds and make them work together convincingly. A useful reference might include something loud, such as a gunshot or explosion, alongside quieter material such as footsteps. Approximate ambience establishes the environmental bed, while dialogue examples can cover a range from quiet speech through ordinary conversation to shouting. Mixing these elements early creates a working scale for everything that follows.

    New assets can then be designed in relation to an existing sonic framework. Designers know roughly where a sound needs to sit and how much space surrounds it. Original reference assets may eventually be replaced, but their early role remains valuable. They establish relationships before growing complexity makes fundamental decisions harder to change.

    Mixing, from this perspective, becomes part of sound design rather than a finishing process. Working without context encourages every gun, vehicle, impact and interaction to become enormous, detailed and impressive. Once combined, they compete. An evolving reference mix permits more varied decisions. Some sounds can remain small. Others can be narrow, distant or restrained. Detail can be reserved for moments when players have enough space to perceive it.

    Focus also connects sound directly to gameplay. Players continually decide where to look, where to move and what to do next. Audio can support those decisions by drawing attention towards useful information or allowing distractions to recede. An approaching threat, important character, changing environment or imminent event may receive temporary priority. Elements contributing little to the current experience can move into the background.

    Informational and emotional focus can coexist. A soundtrack can communicate danger without identifying an enemy’s exact position, or create suspense without explaining precisely what will happen next. Morton’s examples showed how game audio shapes a player’s state of mind while remaining part of the fictional world. Excitement, suspense, drama and humour are designed responses. A weapon feels powerful partly through its sound. An approaching explosion gains anticipation from the quiet preceding it. A gunfight becomes dangerous when incoming fire appears to tear apart the world immediately around the player.

    Realism consequently remains flexible. Effective sounds may preserve recognisable aspects of reality while exaggerating qualities useful to the experience. A weapon can retain enough mechanical identity to remain believable while gaining additional weight. An impact can exceed its real counterpart without feeling inappropriate to the image. A suppressed firearm can satisfy an established cultural expectation even when that expectation differs from literal acoustic reality.

    Morton resisted turning these observations into universal rules. Predator and Terminator 2 can present radically different miniguns while both succeeding on their own terms. A game pursuing realism may demand a different balance from an exaggerated action title. Intimate narrative experiences may use restraint where large-scale spectacle needs greater sonic scale. Designers need to understand what their particular project is asking players to experience.

    Selectivity also shapes resource allocation. No production can pursue every possible recording session, dialogue variation or interactive feature. Large games generate an almost unlimited number of potential tasks while budgets and schedules remain finite. Designers select which events deserve sound, which sounds deserve prominence, which systems justify development time and which details will genuinely improve the experience.

    Morton’s discussion of sound libraries introduced a longer view of professional practice. Building a useful collection of original recordings is expensive, but opportunities to capture interesting material should be taken when possible. A sound recorded today may remain unused for years before finding its purpose. Game audio professionals develop habits of listening beyond individual projects, continually noticing potential material in the world around them.

    His paintball experience represents an even deeper form of professional listening. Morton did not return merely with a useful recording. He had experienced a relationship between threat, proximity and environmental impact that changed how he understood a design problem. Everyday experiences can reveal how attention and emotion respond to sound. Listening professionally involves recognising such relationships as well as collecting interesting timbres.

    Technical constraints and psychological effects remain closely connected throughout Morton’s practice. Memory budgets, recording costs, dialogue systems, asset creation and mixing strategies eventually converge upon one question: what experience is the player having? Sophisticated technology has limited value when it does not support that experience. Conversely, a simple decision such as briefly removing surrounding sound can transform a moment when it directs attention effectively.

    By the end of the lecture, Morton had presented AAA game audio as a discipline balanced between enormous complexity and deliberate restraint. Contemporary designers have access to more memory, processing, channels and real-time synthesis than earlier generations could have imagined. Additional technical capacity, however, does not remove the need to choose. More available sounds do not make more audible sounds desirable. Processing power cannot determine where attention belongs. Larger worlds make focus increasingly important.

    His account also challenged the familiar suggestion that audiences notice game audio only when it fails. Players respond to powerful sound design even when they cannot describe every mechanism behind it. They feel the danger of a gunfight, anticipate an explosion during a sudden moment of quiet and recognise when a world has enough space to breathe. Good audio can elevate games whose code, visuals and other systems already represent enormous investment. Its contribution reaches far beyond correcting problems. Sound helps determine the emotional character of the experience.

    Morton’s lecture ultimately revealed game sound design as the management of attention through an interactive world. Designers decide when reality should be preserved and when expectation should take priority. Silence can become more powerful than another layer. A producer’s request needs to be interpreted through the emotion it seeks to achieve. A gun may already sound excellent while the gunfight surrounding it remains ineffective. Thousands of assets can exist within a game, yet the success of the soundtrack depends upon knowing which ones deserve attention at a particular moment.

    Perhaps the most important question is simply what the player needs to hear now. Sometimes the answer is a spectacular weapon. At another moment, it is the violent impact of a projectile against the wall beside them. Elsewhere, a quiet footstep, distant ambience or line of dialogue needs enough space to be understood. Occasionally, almost everything should disappear. Game audio becomes powerful when sound shapes experience rather than catalogues events. A virtual world may contain thousands of possible sounds. The art lies in choosing the ones that make the player feel something.

  • How Do You Perform a World Through Sound? Jason Swanscott on Foley, Improvisation, and Making Movement Believable

    Jason Swanscott

    How do you perform a world through sound?

    A character crosses a room, adjusts a coat, places a glass on a table and sits down. Nothing about the sequence appears extraordinary, yet its soundtrack may contain dozens of carefully performed details. Every footstep must carry the correct weight. Fabric must move with the body rather than merely rustle somewhere in the background. The glass must sound appropriate for its size and surface, while the chair should respond convincingly to the character’s movement. None of these sounds is likely to attract conscious attention, yet without them the scene can feel strangely empty. During his online guest lecture for Edinburgh Napier University, veteran Foley artist Jason Swanscott explored the extraordinary craft behind these apparently ordinary sounds. Drawing upon almost three decades of work across film, television and games, he revealed Foley as something far richer than synchronising footsteps and props to pictures. It is a form of performance in which movement, character, materials, objects and physical spaces are interpreted through sound. Throughout the session, one idea repeatedly emerged. Foley does not simply reproduce what can be seen on screen. It performs the physical behaviour of an entire world.

    Swanscott began by establishing the three principal elements of Foley: movement, footsteps and spot effects. This distinction provides a useful practical framework, though each category immediately reveals the interpretive nature of the work. Movement concerns the subtle sounds produced as characters shift their bodies and clothing. Footsteps recreate their interaction with the ground. Spot effects encompass the countless objects they touch, lift, open, close, carry and manipulate. Together, these tracks restore a physical presence that may be missing from production recordings or require greater definition within the finished soundtrack. Their purpose is not to make the soundtrack busier. They make bodies, materials and objects feel as though they occupy the world shown on screen.

    Movement often begins with fabric, but choosing a material is only the first decision. A suit jacket does not behave like a tweed coat, while cotton, silk and heavier period fabrics each possess different textures and patterns of movement. Swanscott watches how the character moves and performs those movements through the selected material. A slight adjustment in a chair requires a different gesture from somebody running, fighting or struggling into a coat. Sound follows action rather than existing as a continuous layer of generic clothing noise. The artist watches the body, interprets its movement and recreates its acoustic consequences.

    Footsteps make the relationship between sound and performance even clearer. Correct shoes and surfaces matter enormously. Foley stages therefore contain collections of footwear and recording surfaces capable of representing different periods, occupations, characters and environments. Leather-soled shoes suggest a different person from combat boots or high heels, while wood, gravel, tile and cobblestone each change the relationship between the performer and the ground. Owning the appropriate shoe and standing on the appropriate surface, however, does not guarantee a convincing footstep. The Foley artist must perform the character. Weight, pace, hesitation, confidence, exhaustion and emotional state all influence the rhythm and physical force of movement. A footstep can communicate who the character is and how they are moving through the scene.

    Spot effects extend this performance across the material world. A hand lifts a glass, a door opens, a weapon is drawn or an object falls onto a surface. Some actions can be recreated using objects similar to those shown in the image. Others require considerable invention. A sword may acquire additional metallic resonance to give its movement greater presence. A body impact may emerge from materials whose physical origins bear little resemblance to the action on screen. Swanscott’s task is to produce the sound that makes the visible action believable, regardless of whether the studio prop resembles its apparent source.

    One of the most interesting themes running through the lecture concerned the difference between physical truth and perceptual truth. Real events do not always produce the sounds audiences expect from them. A literal recording may appear weak, ambiguous or dramatically inappropriate once placed against an image. Foley therefore occupies a curious territory between reality and expectation. The artist frequently uses an object that is physically wrong to create a sound that feels perceptually right.

    Familiar examples are entertaining precisely through the unlikely relationship between source and result. Celery and cabbage can provide organic fractures and breaks. Bananas and watermelons can contribute the resistance and wetness required for sounds involving cutting flesh. Water, washing-up liquid and wallpaper paste can create viscous materials for blood and other bodily effects. Wet newspaper and other organic materials may provide textures required for medical and forensic scenes. Real offal can also be used, yet Swanscott’s examples demonstrate that literal materials hold no automatic claim to greater credibility. Screen reality is constructed through judgement rather than fidelity to the original source.

    This becomes especially apparent in genres where audiences possess strong expectations about events that few people have experienced directly. A cinematic punch must communicate weight and violence immediately. Different materials and performances can be combined to create that impression of force. Leather, boxing gloves, weighted impacts and other layers contribute qualities that the image appears to demand. Listeners perceive a body being struck even though the sound may have been constructed from objects with entirely different identities. Foley succeeds when attention remains on the event rather than the materials used to create it.

    Different productions also demand distinct styles of performance. A period drama such as Emma requires attention to the weight and behaviour of long dresses, delicate gloves, heeled footwear, polished floors and carefully handled objects. Its physical world is controlled and refined, so the Foley performance must support that quality. An action film such as Kick-Ass demands something entirely different. Fights require aggressive bodily performance, impacts, falls, furniture movement and layers of activity that communicate speed and force. Horror presents another challenge, where small sounds can become disproportionately important. A creaking surface, distant movement or restrained scrape may contribute more tension than an obviously dramatic effect. Foley changes with the storytelling language of the production.

    Across these genres, Swanscott must interpret images remarkably quickly. Foley artists frequently begin work without having watched the production in advance. They may receive limited notes and occasional guidance about particular details such as footwear or costume, while much of the decision-making happens as the material appears before them. A contemporary drama, period production, action sequence or fantasy world immediately suggests different requirements, yet the artist must continually decide what deserves performance, which materials might work and how much detail the scene can support.

    Such a working method reveals a form of expertise that is difficult to reduce to written instructions. Swanscott’s practice depends upon an accumulated relationship between sight, movement and sound. Watching a character walk prompts immediate decisions about shoe, surface, rhythm and weight. While performing those footsteps, he may already be noticing the objects that will require attention during the later spot-effects pass. Several forms of attention overlap. He is performing the current sound, watching synchronisation, interpreting character and preparing mentally for sounds that have yet to be recorded.

    Pressure makes this embodied expertise particularly visible. Swanscott contrasted the schedules available to different productions. A forty-five-minute ITV drama episode might receive only two days of Foley work, while a longer BBC production such as Silent Witness may allow around five. Feature films can provide considerably longer schedules, with major productions sometimes allowing several weeks. More time allows experimentation, comparison and repeated refinement. Fast television requires strong decisions almost immediately.

    On a two-day television schedule, movement and footsteps may occupy the first day, with spot effects following on the second. Every character still needs to move through the programme. Shoes still need to match, surfaces still need to change and objects still need to acquire physical presence. The schedule compresses the time available without removing the detail. Speed in this context means more than moving quickly. It depends upon recognising patterns, anticipating requirements and drawing upon a sufficiently developed physical and sonic vocabulary that experimentation can happen almost instantaneously.

    Years of this practice turn Foley stages into unusual archives of material culture. Shoes, fabrics, doors, glasses, tools, weapons, furniture and fragments of apparently unremarkable objects accumulate over time. Their value bears little relationship to monetary worth or original purpose. An object becomes valuable when the artist recognises useful sonic behaviour within it. Swanscott recalled finding a discarded sink outside a studio and bringing it inside. Its metallic resonance later proved ideal for a sound effect. A small paving slab could suggest the movement of a vast stone door. Neither object is limited by its physical scale or intended function once the artist begins to imagine what its sound might become.

    Foley therefore requires a peculiar form of auditory imagination. Swanscott has developed knowledge of how materials behave, then uses that knowledge to recognise relationships between sounds and images that may initially appear unrelated. He hears what an object can become when performed differently, recorded from another perspective or combined with another layer.

    One of the clearest examples came from The Legend of Tarzan. Swanscott recreated the heavy, padded movement of gorillas using boxing gloves. The relationship becomes logical when considered through physical qualities rather than visual resemblance: both can produce broad, soft impacts carrying considerable apparent weight. Discovering that relationship requires an artist to think in terms of behaviour. The boxing glove becomes useful through what it can perform.

    Other problems require materials that can be made to behave in particular ways. A snake moving across a surface can be suggested through materials such as cornflour-filled pillowcases, where the shifting internal texture creates the required movement. Magical or supernatural objects may combine recognisable physical impacts with resonant elements that suggest something beyond ordinary reality. In The Witcher, Swanscott described combining physical sounds with the resonance of a singing wine glass to give an enchanted object both material weight and a supernatural quality. Listeners need to believe that the object occupies physical space while also sensing that it belongs to a world governed by unfamiliar forces. Foley can communicate both qualities at once.

    Science fiction extends this challenge further. Productions such as Intergalactic require sounds for technologies without direct real-world equivalents. Spaceship doors, airlocks, damaged suits and magnetic footwear still need to communicate material, force and function. Heavy work boots combined with a rubbery surface can suggest magnetic contact, while compressed air and vocal techniques can contribute to the impression of air escaping from a damaged suit. Even within imaginary worlds, invention remains grounded in physical performance.

    Visually synthetic productions make this physical grounding especially valuable. Computer-generated creatures, magical objects and futuristic environments remove any possibility of simply recording what happened during filming. Audiences nevertheless expect those worlds to possess internal physical coherence. A creature needs weight. A machine needs resistance. Clothing must move with bodies, and objects must appear to interact with surfaces. Familiar materials give impossible worlds a physical logic that listeners recognise instinctively.

    Sometimes the object alone cannot create that credibility. Acoustic space can become part of the performance. Swanscott discussed work on Captain Phillips, where scenes inside the enclosed lifeboat required a claustrophobic acoustic quality. Performing the Foley conventionally within an open recording space would not reproduce the confined reflections demanded by the scene. The team instead worked inside an overturned water tank, using the enclosure itself to create the required resonance.

    That example broadens the definition of the Foley instrument. The performer and prop exist within an acoustic relationship shaped by surfaces, boundaries, reflections and microphone perspective. For the lifeboat, claustrophobia emerged through the physical conditions of recording. The space itself became part of the performance.

    Large scenes demand another kind of construction. A soldier running through a battlefield is represented by far more than footsteps on dirt. Clothing moves, equipment shifts, buckles rattle and weapons interact with the body. Each component may be relatively simple in isolation, yet together they create the physical complexity of a person carrying weight through an environment. Foley gives the body layers. Listeners may never consciously identify each buckle or movement, but their combined behaviour makes the character feel physically present.

    Close collaboration with sound editors allows this layered physicality to remain flexible. The artist performs movement, footsteps and spot effects, while editors organise, refine and prepare those recordings for the wider soundtrack. Individual elements can be adjusted independently, allowing movement, footsteps and object interactions to find appropriate relationships within the final mix. Additional sound effects or processing may supplement the performance, while the Foley recording provides an organic foundation tied directly to the rhythm of the image.

    That connection to performance becomes especially important when production dialogue has been replaced. ADR can separate dialogue from the incidental sounds originally captured on set. Footsteps, clothing and object interaction may then need to be reconstructed around the replacement dialogue so that the scene regains its physical continuity. Foley restores relationships between bodies and environments that post-production processes may have separated.

    Swanscott’s examples also challenged any assumption that every visible action requires equal sonic emphasis. Foley serves storytelling, and different scenes demand different levels of detail and exaggeration. Restrained handling of porcelain in a period drama communicates something very different from the aggressive materiality of an action sequence. Horror may depend upon an isolated movement sound emerging from otherwise restrained surroundings. Fantasy may require a familiar impact combined with an unfamiliar resonance. Every scene asks the artist to decide how an event should sound and how strongly the audience should experience it.

    Improvisation is central to these decisions, though professional improvisation is far from random. Years spent learning the behaviour of materials, surfaces and objects allow relationships to be tested quickly when an unexpected problem appears. A discarded sink becomes useful through recognition of its resonant potential. Boxing gloves become gorilla feet through their combination of softness and apparent weight. What looks spontaneous from outside the studio rests upon a deeply developed vocabulary of physical sound.

    Swanscott’s own body is equally important. Fight sequences may require vigorous physical movement. Footsteps must be performed with the rhythm and weight of somebody whose body may differ considerably from that of the artist. A delicate character, exhausted soldier and threatening pursuer cannot all be represented through the same neutral walking pattern. The Foley artist inhabits movement without being visible. His body becomes an interpretive instrument.

    Precision and expressiveness therefore exist together. Synchronisation matters. A footstep landing visibly out of time can break the connection between sound and image. Accurate timing, however, cannot rescue a performance with the wrong weight, rhythm or intention. Like other forms of performance, Foley uses technical control to enable expression. The artist must arrive at the correct moment while making that moment belong to the character.

    Interactive media changes the structure within which these performances operate. Swanscott’s work extends into video games, where sounds cannot always be performed against a fixed sequence experienced identically by every player. Games such as Alien: Isolation require movement and environmental sounds that respond dynamically to player behaviour. Footsteps and interactions need variation, while sounds must remain convincing when triggered in different sequences and contexts. The direct relationship between performer and fixed picture becomes a system of potential relationships between recorded material and player action.

    Embodied performance remains central even within that nonlinear structure. Weight, texture and physical credibility still matter. Repetition must be controlled so that repeated actions do not expose the limited number of recordings behind them. Direction and distance gain importance as sounds operate within spatial environments. The player may determine when a footstep occurs, while the Foley artist still determines what that step communicates.

    Changes in production systems place pressure on this kind of embodied practice. Swanscott described compressed schedules alongside continuing expectations for highly detailed soundtracks. Productions that once allowed longer periods for Foley may now expect comparable quality in fewer days. Outsourcing and remote working can reduce the close feedback once shared between Foley stages and other post-production departments, while sound libraries provide increasingly convenient alternatives for some categories of material.

    A library effect and a Foley performance differ most fundamentally in their relationship to a particular moment. Foley is created for a specific body, action and dramatic context. Pace, force, rhythm and texture can change in direct response to the image. Libraries can contain exceptional recordings, but the relationship between a pre-existing sound and a new movement must be constructed afterwards. Foley creates that relationship through performance.

    Technology continues to change how this embodied knowledge is captured and used. Digital workflows have transformed recording and editing. New microphones can reveal forms of vibration and underwater sound that conventional recording methods may miss. Synthetic imagery creates opportunities for invention, while games require nonlinear systems rather than fixed sequences. Swanscott’s practice has adapted continually to these changes. At its centre remains a physical act: watching movement, understanding its dramatic purpose and finding a way to perform its sound.

    Preserving the profession therefore involves more than documenting a catalogue of techniques. A list can identify suitable shoes, useful surfaces and familiar prop substitutions, but it cannot easily communicate the embodied timing required to perform another person’s movement or the instinct that allows an artist to hear a discarded object and recognise its potential. Swanscott emphasised the importance of mentorship and training within a craft learned through watching, listening, trying, failing and gradually developing a relationship with materials.

    Perhaps this is the most revealing way to understand Foley. Spectacular examples naturally attract attention: vegetables become broken bones, boxing gloves become gorillas and an overturned water tank becomes a lifeboat. Yet the deeper skill lies in the decisions connecting those materials to the image. A watermelon has no inherent cinematic meaning. Its value emerges through how it is performed, where the microphone is placed, which qualities of its sound are useful and whether the resulting texture belongs within the world of the scene.

    Swanscott’s lecture consequently revealed a discipline of transformation. Fabric becomes bodily movement. Shoes become character. Small objects become enormous mechanisms. Familiar materials give physical credibility to creatures and technologies that have never existed. A recording space becomes part of a fictional environment, while a performer standing on a Foley stage inhabits the movement of somebody in another place, another period or an entirely imaginary world.

    Finished Foley conceals this constant process of interpretation. Audiences see a character walk and hear footsteps. They see clothing move and hear fabric. They watch an object fall and accept its weight. The connection appears inevitable, even though every sound may have required decisions about material, performance, surface, timing, perspective and dramatic emphasis. As the performance becomes more convincing, the labour required to create that relationship becomes less visible.

    By the end of the lecture, Swanscott had transformed the Foley stage from a room full of shoes, fabrics, surfaces and strange objects into something closer to a physical imagination of the screen. Nothing in the room has only one identity. A boxing glove can become a gorilla. A paving slab can become monumental architecture. A discarded sink can acquire a new purpose through its resonance. The artist’s task is to watch movement, understand character and hear possibilities hidden inside ordinary materials. Foley gives bodies weight, objects substance and imaginary places a physical life. When the work succeeds, audiences do not hear somebody performing a world in a studio. They simply believe that the world was always there.