Author: iainmcgregor

  • How Do You Make a Game Feel Dangerous? Will Morton on Emotion, Attention, and Designing Sound for the Player

    Will Morton

    How do you make a game feel dangerous?

    A gun can sound enormous and still fail to make a gunfight frightening. Every weapon may have a powerful attack, convincing mechanical detail and an impressive environmental tail, yet the player can remain strangely detached from the danger. Solving that problem may have little to do with redesigning the weapon itself. Bullets pass close to the head. Impacts strike nearby walls with exaggerated force. Brickwork breaks apart, fragments scatter and the environment appears to react violently to the threat. During his online guest lecture for Edinburgh Napier University, game audio designer Will Morton explored how sound can shape emotion, focus attention and guide players through complex interactive experiences. Drawing upon twelve years at Rockstar North, where his work included the Grand Theft Auto series, Red Dead Redemption and L.A. Noire, followed by the establishment of Solid Audio Works with fellow former Rockstar audio specialist Craig Connor, Morton presented game sound design as a discipline of selection. Thousands of sounds may exist within a game, but their value depends upon knowing which ones matter at any particular moment. Throughout the lecture, one principle repeatedly emerged. A designer must ask not only what a game world should sound like, but what the player needs to hear in order to feel what the game intends them to feel.

    Morton began by placing creative decisions within the realities of AAA game production. Large games are expensive, technically constrained and continually changing. Development rarely follows a fixed design from beginning to end. Features evolve, producers reconsider decisions and new requests arrive late in production. Platform restrictions impose further limits through storage, memory, streaming performance and processing capability. Scale introduces organisational complexity as well. Larger games require larger teams, while experienced creative specialists can find increasing amounts of their time absorbed by scheduling, administration and coordination. Sound design develops inside a moving system of technical, financial and organisational constraints. Success requires more than imagining an ideal soundtrack. Designers must create one capable of surviving years of changing requirements.

    Historical changes in technology have altered the scale of those constraints without eliminating them. Morton recalled Commodore 64 composer Martin Galway fitting numerous sound effects into approximately one kilobyte of memory. Restrictions of that magnitude appear almost comic from the perspective of contemporary production, yet modern open-world games can still leave audio teams fighting for storage and memory. Vastly greater resources are now available, while games simultaneously attempt to represent entire cities, landscapes and fictional worlds. Technological abundance creates new possibilities, but ambition expands alongside it. Designers still have to decide where limited resources will make the greatest contribution.

    Money introduces another set of choices. Sound effects require designers, recording equipment, locations, editing time, libraries, software and continually changing computer systems. Dialogue adds writers, actors, directors, studios, recording staff and extensive editing. Music may involve composition, licensing, performers, orchestras, specialist recording facilities and interactive implementation. Morton’s overview exposed the consequences hidden behind apparently simple creative ambitions. Another recording variation, character voice or interactive music feature consumes time, money, memory and attention that cannot be spent elsewhere.

    Dialogue provided one of the clearest examples of complexity hiding behind familiar production tasks. Morton spent much of his Rockstar career dividing his time between sound design and dialogue before the scale of Grand Theft Auto V led him to work entirely as dialogue supervisor. For story-heavy games, dialogue cannot be treated as a sequence of lines requested by designers and recorded by actors. Repetition needs consideration. Lines must make sense across changing gameplay situations. Story information may unfold over many hours or days of play, while players can interrupt, delay or alter the circumstances in which dialogue occurs. A dialogue designer therefore needs to understand the interactive structure of the game as deeply as the individual performances being recorded.

    Direction presents a related challenge. Experience in film and television can produce excellent performances, yet game dialogue introduces problems absent from linear media. Cutscenes may operate much like conventional scenes, while in-game dialogue can occur within changing gameplay circumstances and across story structures experienced differently by individual players. Directors need more than an ability to elicit compelling performances. Detailed knowledge of the script, game and eventual context of each line becomes essential. A performance recorded in isolation must later remain convincing within circumstances that may not even be visible inside the recording studio.

    Morton also challenged assumptions about professional recording. Technically excellent dialogue can be captured with comparatively modest equipment in a carefully treated space. A suitable microphone, simple accessories and an acoustically controlled room may produce results approaching those of a far more expensive studio. Recording quality, however, forms only one part of a professional session. High-profile performers need confidence that they are participating in a serious production. Environment, organisation and treatment of the actor can influence trust in the project and relationships with agents and management. Professionalism encompasses the experience surrounding a recording alongside the technical quality of the resulting file.

    From these production realities, Morton moved towards the central creative argument of the lecture. Designers can easily assume that everything visible in a game should automatically produce a sound. Faced with an unsounded world, teams begin filling every action with detail. Cars receive engines and collisions. Characters acquire footsteps. Objects gain interactions. Environments fill with ambiences. A technically comprehensive and entirely plausible soundtrack gradually emerges, yet plausibility alone cannot guarantee clarity, excitement or emotional effect.

    Practical concerns provide one reason for restraint. Every additional sound requires creation, editing, implementation, memory and testing. Artistic considerations are even more important. If everything demands attention simultaneously, nothing receives focus. Morton encouraged designers to identify what deserves to be heard and what can remain absent. Important sounds require room to breathe. Dynamics rely upon contrast, while emotional emphasis depends upon moving attention between elements. Silence and omission become active design decisions.

    Player experience consequently takes priority over literal acoustic reconstruction. Real events often sound less dramatic than audiences expect. A gunshot captured from a particular position may seem surprisingly small. A suppressed weapon does not necessarily produce the familiar cinematic whisper audiences have learned to associate with it. A real minigun may collapse into an almost continuous mechanical roar instead of revealing every stage of its operation. Decades of film, television and games have established sonic conventions that now influence how audiences expect objects and events to behave. Designers work within that accumulated perceptual history.

    Morton demonstrated the point by comparing cinematic and real recordings of suppressed firearms. A familiar designed version sounded short, controlled and immediately recognisable, carrying characteristics audiences strongly associate with a silenced weapon. Real recordings behaved quite differently. Neither approach offered a universal answer. Documentary representation might favour acoustic accuracy, while an action game may need immediate recognition, excitement and dramatic satisfaction. Before designing a sound, Morton considers the role accuracy should play, whether an event needs to feel larger than reality and how strongly audience expectations should influence the result.

    Miniguns provided an even clearer illustration. Morton discussed the famous weapon sequence in Predator, where mechanical movement, spinning and other details reinforce the spectacle shown on screen. Real recordings present a very different impression, dominated by the extraordinary density of rapid gunfire. Grand Theft Auto: Vice City and Grand Theft Auto V offered further interpretations, each constructing the weapon differently, while Terminator 2 adopted another cinematic approach that remained closer to aspects of the real sound. Comparing them did not reveal a correct minigun. Instead, their differences showed how each design serves the experience of a particular production.

    Reality and relevance are therefore separate considerations. A game designed as escapism may gain little from reproducing everyday acoustic experience with complete fidelity. Players do not always need to hear events from the perspective of ordinary observers. They need a version that communicates the event’s importance within the game. Sound design selects, enlarges, simplifies and reshapes reality according to dramatic purpose.

    Attention can be directed just as effectively through subtraction. Morton used a remotely detonated explosive from Grand Theft Auto IV: The Ballad of Gay Tony to demonstrate how removing surrounding sound can make a single event dominate awareness. Immediately before the explosion, the wider soundtrack recedes and the bomb’s warning becomes the focus. He compared the moment with the seismic charge sequence from Star Wars: Episode II, where a brief interruption in the expected sound field creates anticipation and gives the subsequent event greater impact. Both sequences derive part of their power from absence.

    Focus involves more than increasing the level of an important sound. Every event competes with its surroundings. Removing distractions can achieve more than additional layers, greater loudness or further spectral exaggeration. Presence acquires meaning through contrast with absence. A fraction of a second of reduced activity can prepare an event more effectively than a continuously dense soundtrack.

    Morton’s most revealing example concerned a producer asking for guns to sound more dangerous. Existing weapon sounds already seemed successful to the audio team. They contained convincing mechanical detail, a strong initial attack, substantial body and satisfying environmental tails. Reworking those qualities did not solve the problem. Although the request appeared to concern the sound of the guns, the underlying dissatisfaction was emotional.

    An unexpected experience away from the studio suggested another approach. During a paintball game, Morton found himself sheltering behind structures made from metal oil drums. The paintball markers themselves produced relatively insignificant sounds. Fear came from sudden, violent impacts striking the metal around him. His immediate environment appeared to contract as attention focused upon nearby impacts, resonances and the sense that projectiles were arriving from directions he could not fully control. The weapon itself was not frightening. Being under fire was.

    Recognising that distinction transformed the design problem. Bullet passes became more prominent, with their levels responding more strongly to proximity. Impacts against brick and concrete gained force. Low-frequency energy added physical weight, while debris and crumbling material made the environment appear to react to nearby gunfire. Details that might be acoustically subordinate to a real gunshot were deliberately brought forward. Stronger weapon sounds had never been the answer. Gunfights needed to communicate vulnerability and danger.

    Here Morton identified a broader professional skill. Directors and producers rarely describe every sound problem in acoustic terms. They may ask for something to be louder, bigger, darker, faster or more dangerous while expressing dissatisfaction with an emotional result. Following the literal wording can send a designer towards the wrong solution. Morton argued for interpreting the intention behind the request. Someone asking for a more dangerous gun may actually be asking to feel vulnerable. Once the desired experience becomes clear, the designer can decide which part of the sound world needs to change.

    Creative collaboration therefore requires translation between intention and acoustic action. Sound professionals develop specialised vocabularies for frequency, dynamics, spatial behaviour, envelope and processing. Producers may describe experiences through emotion, imagery or metaphor. Useful information can exist within both forms of language. Professional expertise includes moving between them and identifying the experience concealed inside an apparently vague request.

    Gunfire also demonstrates how game audio operates through systems of cause and consequence. A weapon consists of more than the sound emitted when a trigger is pulled. Projectiles move through space. Near misses pass the player. Bullets strike walls. Debris falls afterwards. Environments respond differently according to material and distance. Danger emerges from relationships between these events. Concentrating exclusively upon the source can leave the wider experience emotionally incomplete.

    Similar systems govern the game mix. Morton described mixing a large game as a potentially overwhelming process involving thousands of assets and potentially hundreds of simultaneous channels, all changing according to gameplay. Exact combinations cannot be predicted in advance as they can within a linear soundtrack. Dialogue, music, ambience, vehicles, weapons, footsteps and environmental interactions combine differently according to player behaviour. Even excellent individual assets can produce an exhausting result when their relationships are poorly controlled.

    Morton recalled playing games whose soundtracks felt like continuous acoustic assault. Constant density can be as tiring as excessive level. When every category remains active and prominent, listeners receive no opportunity for recovery and little indication of where attention belongs. More detail can therefore produce less communication.

    Working with Craig Connor, Morton developed a useful analogy for managing this complexity: approach the game as a music producer approaches a song. Every element needs an appropriate place and enough space around it. The comparison does not imply imposing a static musical mix upon an interactive system. Its value lies in encouraging relational thinking. Sounds acquire meaning through their positions alongside other elements rather than through isolated perfection.

    Arrangement offers another useful parallel. Music producers rarely expect every instrument to occupy the foreground continuously. Parts enter and leave, density changes and contrast creates structure. Individual elements occupy different spectral, spatial and dynamic roles. Interactive audio can apply similar principles while responding continuously to player action, with priority changing according to what the player is doing and what the game needs to communicate.

    Waiting until asset production is complete before attempting a final mix creates serious problems. By the time thousands of sounds have accumulated, assumptions about level, density and priority may already be embedded throughout the project. Assets designed without a meaningful reference mix can prove difficult to reconcile, particularly when each has been created to sound impressive in isolation.

    Morton recommended an iterative alternative. Early in development, designers can choose a small but representative collection of sounds and make them work together convincingly. A useful reference might include something loud, such as a gunshot or explosion, alongside quieter material such as footsteps. Approximate ambience establishes the environmental bed, while dialogue examples can cover a range from quiet speech through ordinary conversation to shouting. Mixing these elements early creates a working scale for everything that follows.

    New assets can then be designed in relation to an existing sonic framework. Designers know roughly where a sound needs to sit and how much space surrounds it. Original reference assets may eventually be replaced, but their early role remains valuable. They establish relationships before growing complexity makes fundamental decisions harder to change.

    Mixing, from this perspective, becomes part of sound design rather than a finishing process. Working without context encourages every gun, vehicle, impact and interaction to become enormous, detailed and impressive. Once combined, they compete. An evolving reference mix permits more varied decisions. Some sounds can remain small. Others can be narrow, distant or restrained. Detail can be reserved for moments when players have enough space to perceive it.

    Focus also connects sound directly to gameplay. Players continually decide where to look, where to move and what to do next. Audio can support those decisions by drawing attention towards useful information or allowing distractions to recede. An approaching threat, important character, changing environment or imminent event may receive temporary priority. Elements contributing little to the current experience can move into the background.

    Informational and emotional focus can coexist. A soundtrack can communicate danger without identifying an enemy’s exact position, or create suspense without explaining precisely what will happen next. Morton’s examples showed how game audio shapes a player’s state of mind while remaining part of the fictional world. Excitement, suspense, drama and humour are designed responses. A weapon feels powerful partly through its sound. An approaching explosion gains anticipation from the quiet preceding it. A gunfight becomes dangerous when incoming fire appears to tear apart the world immediately around the player.

    Realism consequently remains flexible. Effective sounds may preserve recognisable aspects of reality while exaggerating qualities useful to the experience. A weapon can retain enough mechanical identity to remain believable while gaining additional weight. An impact can exceed its real counterpart without feeling inappropriate to the image. A suppressed firearm can satisfy an established cultural expectation even when that expectation differs from literal acoustic reality.

    Morton resisted turning these observations into universal rules. Predator and Terminator 2 can present radically different miniguns while both succeeding on their own terms. A game pursuing realism may demand a different balance from an exaggerated action title. Intimate narrative experiences may use restraint where large-scale spectacle needs greater sonic scale. Designers need to understand what their particular project is asking players to experience.

    Selectivity also shapes resource allocation. No production can pursue every possible recording session, dialogue variation or interactive feature. Large games generate an almost unlimited number of potential tasks while budgets and schedules remain finite. Designers select which events deserve sound, which sounds deserve prominence, which systems justify development time and which details will genuinely improve the experience.

    Morton’s discussion of sound libraries introduced a longer view of professional practice. Building a useful collection of original recordings is expensive, but opportunities to capture interesting material should be taken when possible. A sound recorded today may remain unused for years before finding its purpose. Game audio professionals develop habits of listening beyond individual projects, continually noticing potential material in the world around them.

    His paintball experience represents an even deeper form of professional listening. Morton did not return merely with a useful recording. He had experienced a relationship between threat, proximity and environmental impact that changed how he understood a design problem. Everyday experiences can reveal how attention and emotion respond to sound. Listening professionally involves recognising such relationships as well as collecting interesting timbres.

    Technical constraints and psychological effects remain closely connected throughout Morton’s practice. Memory budgets, recording costs, dialogue systems, asset creation and mixing strategies eventually converge upon one question: what experience is the player having? Sophisticated technology has limited value when it does not support that experience. Conversely, a simple decision such as briefly removing surrounding sound can transform a moment when it directs attention effectively.

    By the end of the lecture, Morton had presented AAA game audio as a discipline balanced between enormous complexity and deliberate restraint. Contemporary designers have access to more memory, processing, channels and real-time synthesis than earlier generations could have imagined. Additional technical capacity, however, does not remove the need to choose. More available sounds do not make more audible sounds desirable. Processing power cannot determine where attention belongs. Larger worlds make focus increasingly important.

    His account also challenged the familiar suggestion that audiences notice game audio only when it fails. Players respond to powerful sound design even when they cannot describe every mechanism behind it. They feel the danger of a gunfight, anticipate an explosion during a sudden moment of quiet and recognise when a world has enough space to breathe. Good audio can elevate games whose code, visuals and other systems already represent enormous investment. Its contribution reaches far beyond correcting problems. Sound helps determine the emotional character of the experience.

    Morton’s lecture ultimately revealed game sound design as the management of attention through an interactive world. Designers decide when reality should be preserved and when expectation should take priority. Silence can become more powerful than another layer. A producer’s request needs to be interpreted through the emotion it seeks to achieve. A gun may already sound excellent while the gunfight surrounding it remains ineffective. Thousands of assets can exist within a game, yet the success of the soundtrack depends upon knowing which ones deserve attention at a particular moment.

    Perhaps the most important question is simply what the player needs to hear now. Sometimes the answer is a spectacular weapon. At another moment, it is the violent impact of a projectile against the wall beside them. Elsewhere, a quiet footstep, distant ambience or line of dialogue needs enough space to be understood. Occasionally, almost everything should disappear. Game audio becomes powerful when sound shapes experience rather than catalogues events. A virtual world may contain thousands of possible sounds. The art lies in choosing the ones that make the player feel something.

  • How Do You Perform a World Through Sound? Jason Swanscott on Foley, Improvisation, and Making Movement Believable

    Jason Swanscott

    How do you perform a world through sound?

    A character crosses a room, adjusts a coat, places a glass on a table and sits down. Nothing about the sequence appears extraordinary, yet its soundtrack may contain dozens of carefully performed details. Every footstep must carry the correct weight. Fabric must move with the body rather than merely rustle somewhere in the background. The glass must sound appropriate for its size and surface, while the chair should respond convincingly to the character’s movement. None of these sounds is likely to attract conscious attention, yet without them the scene can feel strangely empty. During his online guest lecture for Edinburgh Napier University, veteran Foley artist Jason Swanscott explored the extraordinary craft behind these apparently ordinary sounds. Drawing upon almost three decades of work across film, television and games, he revealed Foley as something far richer than synchronising footsteps and props to pictures. It is a form of performance in which movement, character, materials, objects and physical spaces are interpreted through sound. Throughout the session, one idea repeatedly emerged. Foley does not simply reproduce what can be seen on screen. It performs the physical behaviour of an entire world.

    Swanscott began by establishing the three principal elements of Foley: movement, footsteps and spot effects. This distinction provides a useful practical framework, though each category immediately reveals the interpretive nature of the work. Movement concerns the subtle sounds produced as characters shift their bodies and clothing. Footsteps recreate their interaction with the ground. Spot effects encompass the countless objects they touch, lift, open, close, carry and manipulate. Together, these tracks restore a physical presence that may be missing from production recordings or require greater definition within the finished soundtrack. Their purpose is not to make the soundtrack busier. They make bodies, materials and objects feel as though they occupy the world shown on screen.

    Movement often begins with fabric, but choosing a material is only the first decision. A suit jacket does not behave like a tweed coat, while cotton, silk and heavier period fabrics each possess different textures and patterns of movement. Swanscott watches how the character moves and performs those movements through the selected material. A slight adjustment in a chair requires a different gesture from somebody running, fighting or struggling into a coat. Sound follows action rather than existing as a continuous layer of generic clothing noise. The artist watches the body, interprets its movement and recreates its acoustic consequences.

    Footsteps make the relationship between sound and performance even clearer. Correct shoes and surfaces matter enormously. Foley stages therefore contain collections of footwear and recording surfaces capable of representing different periods, occupations, characters and environments. Leather-soled shoes suggest a different person from combat boots or high heels, while wood, gravel, tile and cobblestone each change the relationship between the performer and the ground. Owning the appropriate shoe and standing on the appropriate surface, however, does not guarantee a convincing footstep. The Foley artist must perform the character. Weight, pace, hesitation, confidence, exhaustion and emotional state all influence the rhythm and physical force of movement. A footstep can communicate who the character is and how they are moving through the scene.

    Spot effects extend this performance across the material world. A hand lifts a glass, a door opens, a weapon is drawn or an object falls onto a surface. Some actions can be recreated using objects similar to those shown in the image. Others require considerable invention. A sword may acquire additional metallic resonance to give its movement greater presence. A body impact may emerge from materials whose physical origins bear little resemblance to the action on screen. Swanscott’s task is to produce the sound that makes the visible action believable, regardless of whether the studio prop resembles its apparent source.

    One of the most interesting themes running through the lecture concerned the difference between physical truth and perceptual truth. Real events do not always produce the sounds audiences expect from them. A literal recording may appear weak, ambiguous or dramatically inappropriate once placed against an image. Foley therefore occupies a curious territory between reality and expectation. The artist frequently uses an object that is physically wrong to create a sound that feels perceptually right.

    Familiar examples are entertaining precisely through the unlikely relationship between source and result. Celery and cabbage can provide organic fractures and breaks. Bananas and watermelons can contribute the resistance and wetness required for sounds involving cutting flesh. Water, washing-up liquid and wallpaper paste can create viscous materials for blood and other bodily effects. Wet newspaper and other organic materials may provide textures required for medical and forensic scenes. Real offal can also be used, yet Swanscott’s examples demonstrate that literal materials hold no automatic claim to greater credibility. Screen reality is constructed through judgement rather than fidelity to the original source.

    This becomes especially apparent in genres where audiences possess strong expectations about events that few people have experienced directly. A cinematic punch must communicate weight and violence immediately. Different materials and performances can be combined to create that impression of force. Leather, boxing gloves, weighted impacts and other layers contribute qualities that the image appears to demand. Listeners perceive a body being struck even though the sound may have been constructed from objects with entirely different identities. Foley succeeds when attention remains on the event rather than the materials used to create it.

    Different productions also demand distinct styles of performance. A period drama such as Emma requires attention to the weight and behaviour of long dresses, delicate gloves, heeled footwear, polished floors and carefully handled objects. Its physical world is controlled and refined, so the Foley performance must support that quality. An action film such as Kick-Ass demands something entirely different. Fights require aggressive bodily performance, impacts, falls, furniture movement and layers of activity that communicate speed and force. Horror presents another challenge, where small sounds can become disproportionately important. A creaking surface, distant movement or restrained scrape may contribute more tension than an obviously dramatic effect. Foley changes with the storytelling language of the production.

    Across these genres, Swanscott must interpret images remarkably quickly. Foley artists frequently begin work without having watched the production in advance. They may receive limited notes and occasional guidance about particular details such as footwear or costume, while much of the decision-making happens as the material appears before them. A contemporary drama, period production, action sequence or fantasy world immediately suggests different requirements, yet the artist must continually decide what deserves performance, which materials might work and how much detail the scene can support.

    Such a working method reveals a form of expertise that is difficult to reduce to written instructions. Swanscott’s practice depends upon an accumulated relationship between sight, movement and sound. Watching a character walk prompts immediate decisions about shoe, surface, rhythm and weight. While performing those footsteps, he may already be noticing the objects that will require attention during the later spot-effects pass. Several forms of attention overlap. He is performing the current sound, watching synchronisation, interpreting character and preparing mentally for sounds that have yet to be recorded.

    Pressure makes this embodied expertise particularly visible. Swanscott contrasted the schedules available to different productions. A forty-five-minute ITV drama episode might receive only two days of Foley work, while a longer BBC production such as Silent Witness may allow around five. Feature films can provide considerably longer schedules, with major productions sometimes allowing several weeks. More time allows experimentation, comparison and repeated refinement. Fast television requires strong decisions almost immediately.

    On a two-day television schedule, movement and footsteps may occupy the first day, with spot effects following on the second. Every character still needs to move through the programme. Shoes still need to match, surfaces still need to change and objects still need to acquire physical presence. The schedule compresses the time available without removing the detail. Speed in this context means more than moving quickly. It depends upon recognising patterns, anticipating requirements and drawing upon a sufficiently developed physical and sonic vocabulary that experimentation can happen almost instantaneously.

    Years of this practice turn Foley stages into unusual archives of material culture. Shoes, fabrics, doors, glasses, tools, weapons, furniture and fragments of apparently unremarkable objects accumulate over time. Their value bears little relationship to monetary worth or original purpose. An object becomes valuable when the artist recognises useful sonic behaviour within it. Swanscott recalled finding a discarded sink outside a studio and bringing it inside. Its metallic resonance later proved ideal for a sound effect. A small paving slab could suggest the movement of a vast stone door. Neither object is limited by its physical scale or intended function once the artist begins to imagine what its sound might become.

    Foley therefore requires a peculiar form of auditory imagination. Swanscott has developed knowledge of how materials behave, then uses that knowledge to recognise relationships between sounds and images that may initially appear unrelated. He hears what an object can become when performed differently, recorded from another perspective or combined with another layer.

    One of the clearest examples came from The Legend of Tarzan. Swanscott recreated the heavy, padded movement of gorillas using boxing gloves. The relationship becomes logical when considered through physical qualities rather than visual resemblance: both can produce broad, soft impacts carrying considerable apparent weight. Discovering that relationship requires an artist to think in terms of behaviour. The boxing glove becomes useful through what it can perform.

    Other problems require materials that can be made to behave in particular ways. A snake moving across a surface can be suggested through materials such as cornflour-filled pillowcases, where the shifting internal texture creates the required movement. Magical or supernatural objects may combine recognisable physical impacts with resonant elements that suggest something beyond ordinary reality. In The Witcher, Swanscott described combining physical sounds with the resonance of a singing wine glass to give an enchanted object both material weight and a supernatural quality. Listeners need to believe that the object occupies physical space while also sensing that it belongs to a world governed by unfamiliar forces. Foley can communicate both qualities at once.

    Science fiction extends this challenge further. Productions such as Intergalactic require sounds for technologies without direct real-world equivalents. Spaceship doors, airlocks, damaged suits and magnetic footwear still need to communicate material, force and function. Heavy work boots combined with a rubbery surface can suggest magnetic contact, while compressed air and vocal techniques can contribute to the impression of air escaping from a damaged suit. Even within imaginary worlds, invention remains grounded in physical performance.

    Visually synthetic productions make this physical grounding especially valuable. Computer-generated creatures, magical objects and futuristic environments remove any possibility of simply recording what happened during filming. Audiences nevertheless expect those worlds to possess internal physical coherence. A creature needs weight. A machine needs resistance. Clothing must move with bodies, and objects must appear to interact with surfaces. Familiar materials give impossible worlds a physical logic that listeners recognise instinctively.

    Sometimes the object alone cannot create that credibility. Acoustic space can become part of the performance. Swanscott discussed work on Captain Phillips, where scenes inside the enclosed lifeboat required a claustrophobic acoustic quality. Performing the Foley conventionally within an open recording space would not reproduce the confined reflections demanded by the scene. The team instead worked inside an overturned water tank, using the enclosure itself to create the required resonance.

    That example broadens the definition of the Foley instrument. The performer and prop exist within an acoustic relationship shaped by surfaces, boundaries, reflections and microphone perspective. For the lifeboat, claustrophobia emerged through the physical conditions of recording. The space itself became part of the performance.

    Large scenes demand another kind of construction. A soldier running through a battlefield is represented by far more than footsteps on dirt. Clothing moves, equipment shifts, buckles rattle and weapons interact with the body. Each component may be relatively simple in isolation, yet together they create the physical complexity of a person carrying weight through an environment. Foley gives the body layers. Listeners may never consciously identify each buckle or movement, but their combined behaviour makes the character feel physically present.

    Close collaboration with sound editors allows this layered physicality to remain flexible. The artist performs movement, footsteps and spot effects, while editors organise, refine and prepare those recordings for the wider soundtrack. Individual elements can be adjusted independently, allowing movement, footsteps and object interactions to find appropriate relationships within the final mix. Additional sound effects or processing may supplement the performance, while the Foley recording provides an organic foundation tied directly to the rhythm of the image.

    That connection to performance becomes especially important when production dialogue has been replaced. ADR can separate dialogue from the incidental sounds originally captured on set. Footsteps, clothing and object interaction may then need to be reconstructed around the replacement dialogue so that the scene regains its physical continuity. Foley restores relationships between bodies and environments that post-production processes may have separated.

    Swanscott’s examples also challenged any assumption that every visible action requires equal sonic emphasis. Foley serves storytelling, and different scenes demand different levels of detail and exaggeration. Restrained handling of porcelain in a period drama communicates something very different from the aggressive materiality of an action sequence. Horror may depend upon an isolated movement sound emerging from otherwise restrained surroundings. Fantasy may require a familiar impact combined with an unfamiliar resonance. Every scene asks the artist to decide how an event should sound and how strongly the audience should experience it.

    Improvisation is central to these decisions, though professional improvisation is far from random. Years spent learning the behaviour of materials, surfaces and objects allow relationships to be tested quickly when an unexpected problem appears. A discarded sink becomes useful through recognition of its resonant potential. Boxing gloves become gorilla feet through their combination of softness and apparent weight. What looks spontaneous from outside the studio rests upon a deeply developed vocabulary of physical sound.

    Swanscott’s own body is equally important. Fight sequences may require vigorous physical movement. Footsteps must be performed with the rhythm and weight of somebody whose body may differ considerably from that of the artist. A delicate character, exhausted soldier and threatening pursuer cannot all be represented through the same neutral walking pattern. The Foley artist inhabits movement without being visible. His body becomes an interpretive instrument.

    Precision and expressiveness therefore exist together. Synchronisation matters. A footstep landing visibly out of time can break the connection between sound and image. Accurate timing, however, cannot rescue a performance with the wrong weight, rhythm or intention. Like other forms of performance, Foley uses technical control to enable expression. The artist must arrive at the correct moment while making that moment belong to the character.

    Interactive media changes the structure within which these performances operate. Swanscott’s work extends into video games, where sounds cannot always be performed against a fixed sequence experienced identically by every player. Games such as Alien: Isolation require movement and environmental sounds that respond dynamically to player behaviour. Footsteps and interactions need variation, while sounds must remain convincing when triggered in different sequences and contexts. The direct relationship between performer and fixed picture becomes a system of potential relationships between recorded material and player action.

    Embodied performance remains central even within that nonlinear structure. Weight, texture and physical credibility still matter. Repetition must be controlled so that repeated actions do not expose the limited number of recordings behind them. Direction and distance gain importance as sounds operate within spatial environments. The player may determine when a footstep occurs, while the Foley artist still determines what that step communicates.

    Changes in production systems place pressure on this kind of embodied practice. Swanscott described compressed schedules alongside continuing expectations for highly detailed soundtracks. Productions that once allowed longer periods for Foley may now expect comparable quality in fewer days. Outsourcing and remote working can reduce the close feedback once shared between Foley stages and other post-production departments, while sound libraries provide increasingly convenient alternatives for some categories of material.

    A library effect and a Foley performance differ most fundamentally in their relationship to a particular moment. Foley is created for a specific body, action and dramatic context. Pace, force, rhythm and texture can change in direct response to the image. Libraries can contain exceptional recordings, but the relationship between a pre-existing sound and a new movement must be constructed afterwards. Foley creates that relationship through performance.

    Technology continues to change how this embodied knowledge is captured and used. Digital workflows have transformed recording and editing. New microphones can reveal forms of vibration and underwater sound that conventional recording methods may miss. Synthetic imagery creates opportunities for invention, while games require nonlinear systems rather than fixed sequences. Swanscott’s practice has adapted continually to these changes. At its centre remains a physical act: watching movement, understanding its dramatic purpose and finding a way to perform its sound.

    Preserving the profession therefore involves more than documenting a catalogue of techniques. A list can identify suitable shoes, useful surfaces and familiar prop substitutions, but it cannot easily communicate the embodied timing required to perform another person’s movement or the instinct that allows an artist to hear a discarded object and recognise its potential. Swanscott emphasised the importance of mentorship and training within a craft learned through watching, listening, trying, failing and gradually developing a relationship with materials.

    Perhaps this is the most revealing way to understand Foley. Spectacular examples naturally attract attention: vegetables become broken bones, boxing gloves become gorillas and an overturned water tank becomes a lifeboat. Yet the deeper skill lies in the decisions connecting those materials to the image. A watermelon has no inherent cinematic meaning. Its value emerges through how it is performed, where the microphone is placed, which qualities of its sound are useful and whether the resulting texture belongs within the world of the scene.

    Swanscott’s lecture consequently revealed a discipline of transformation. Fabric becomes bodily movement. Shoes become character. Small objects become enormous mechanisms. Familiar materials give physical credibility to creatures and technologies that have never existed. A recording space becomes part of a fictional environment, while a performer standing on a Foley stage inhabits the movement of somebody in another place, another period or an entirely imaginary world.

    Finished Foley conceals this constant process of interpretation. Audiences see a character walk and hear footsteps. They see clothing move and hear fabric. They watch an object fall and accept its weight. The connection appears inevitable, even though every sound may have required decisions about material, performance, surface, timing, perspective and dramatic emphasis. As the performance becomes more convincing, the labour required to create that relationship becomes less visible.

    By the end of the lecture, Swanscott had transformed the Foley stage from a room full of shoes, fabrics, surfaces and strange objects into something closer to a physical imagination of the screen. Nothing in the room has only one identity. A boxing glove can become a gorilla. A paving slab can become monumental architecture. A discarded sink can acquire a new purpose through its resonance. The artist’s task is to watch movement, understand character and hear possibilities hidden inside ordinary materials. Foley gives bodies weight, objects substance and imaginary places a physical life. When the work succeeds, audiences do not hear somebody performing a world in a studio. They simply believe that the world was always there.

  • How Do You Make ADR Sound Like It Was Never Replaced? Paul Carden and Chris Navarro on Performance, Technology, and the Art of Dialogue Replacement

    Paul Carden and Chris Navarro

    How do you make ADR sound like it was never replaced?

    A line of dialogue may last only a few seconds, yet replacing it convincingly can require an extraordinary combination of preparation, performance, technology and judgement. The original production recording must first be identified as unusable, the replacement carefully documented and prepared, the actor returned to the emotional and physical circumstances of a performance recorded months earlier, and the new dialogue captured so that it matches the timing, vocal quality, microphone perspective and acoustic character of the original scene. If the process succeeds, the audience should never know that any of this work happened. During their joint online guest lecture for Edinburgh Napier University, ADR supervisor Paul Carden and ADR mixer Chris Navarro demonstrated this complete process from beginning to end. Rather than discussing Automated Dialogue Replacement only in theory, they created a deliberately compromised line of production dialogue, prepared it for replacement, recorded it on an ADR stage and evaluated the result. Their demonstration revealed a process in which meticulous preparation and technical fluency serve a deceptively simple objective: allowing everyone involved to concentrate upon the performance.

    Carden began not in a recording studio, but outside beside a car. His objective was to demonstrate how an apparently simple piece of dialogue could become unusable during production. Wearing a lavalier microphone, he performed the five-word line “I’m late for work” while getting into the vehicle. The exercise immediately revealed how vulnerable production dialogue can be. Clothing could obscure the microphone, a zipped jacket could change the sound, movement could complicate recording and the closing car door could mask the line itself. The example was deliberately simple. Carden was standing almost still, knew exactly what he was going to say and had constructed the situation specifically for the demonstration. On a real production, actors may be walking, running, interacting with objects and performing emotionally demanding scenes while the production sound team manages multiple radio microphones and one or more booms. From this perspective, the surprise is often not that some lines require replacement, but that so much production dialogue survives at all.

    The demonstration also established one of the central tensions within ADR. Production dialogue is not simply speech recorded on location. It contains a performance created at a particular moment, within a particular physical environment, as part of an interaction with other performers. Months later, an actor may arrive on an ADR stage while already immersed in an entirely different project and be asked to recreate a few seconds from that earlier performance. The recording environment offers little of the original context. The actor stands in a comparatively neutral room, watches an image on a screen, listens for cues and attempts to reproduce not merely the words, but the emotional and physical conditions under which those words were originally spoken. Carden emphasised that this is one reason actors can find ADR difficult. Matching a line involves returning to a performance that may no longer feel immediate or familiar.

    Before any actor reaches the stage, however, the line must become an ADR cue. Carden demonstrated this preparation process by locating the damaged dialogue within the picture, defining the cue and documenting the reason for replacement. The cue was assigned a unique identifier, linked to its scene information and accompanied by notes explaining that the car door had closed over the line. He also demonstrated the value of professional flexibility. Where a production line was imperfect but potentially usable, he might identify the replacement as optional rather than forcing an unnecessary argument over whether it must be replaced. The cue sheet then carried the information required by the recording stage, including project details, version information, character and actor names, cue number, timecode and recording information. A five-word performance therefore arrived at the stage supported by an extensive information system designed to ensure that everybody was working on the correct material.

    Version control was particularly important. Carden described production teams as living and dying by version dates, reflecting the practical reality that picture changes can quickly make carefully prepared cues inaccurate. Cue numbering performs a similarly essential role. Each take is voice-slated so that the recording itself retains its identity even if paperwork becomes separated from the audio. These details may appear administrative when viewed from outside professional production, yet they protect the continuity of the entire process. An ADR session may contain hundreds of replacement lines recorded across multiple days, actors and facilities. The actor should not have to think about whether the cue is correctly identified or whether the stage is working from the right picture version. Preparation creates the stability within which performance can happen.

    Carden also offered one deceptively simple instruction for aspiring ADR mixers: keep recording until somebody clearly indicates that the take has ended. An actor may continue the performance, a director may decide to record another version immediately or the session may move into wild recording without a formal interruption. Stopping too early can lose useful material for no meaningful benefit. Storage is cheap; an unrecoverable performance is not. The advice reflected a wider principle that would continue through the lecture. Good ADR practice depends upon remaining attentive to what is happening in the room rather than allowing the machinery of recording to dictate the session.

    The second stage of the demonstration moved into Navarro’s ADR facility. Even though the project consisted of a single line created for the lecture, he established the session as though it were a conventional production. Folder structures, documents, Pro Tools sessions and media were organised according to the same principles he would use for a project involving one session or hundreds. His ADR template was already configured for the ordinary demands of recording, while remaining adaptable when unusual situations arose, such as multiple actors performing together with several microphones each. A good template does not prescribe every session in advance. It removes predictable technical work so that attention remains available for situations that cannot be predicted.

    Even the imported picture introduced a practical lesson. Navarro noticed that the system was responding sluggishly and identified the compressed H.264 picture as a likely cause, since the computer had to decode it continuously during playback. The moment was minor, but revealing. Professional technical fluency often appears not through dramatic troubleshooting, but through the ability to recognise quickly why a system is behaving unexpectedly and continue without allowing the problem to dominate the room. Earlier, Carden had offered similarly pragmatic advice about computer failure: save the work, restart the system and continue rather than allowing panic to consume valuable time. Both speakers treated technical expertise as calm familiarity rather than technological display.

    Microphone selection then revealed how ADR matching differs from conventional voice recording. The objective is not to capture the most beautiful possible version of the actor’s voice. It is to create a recording that can inhabit the existing production soundtrack without attracting attention. Navarro showed that the stage had previously been configured for voice-over work, with microphones positioned closely and directly to produce the clear, present sound required for narration. ADR demanded a different approach. Carden had recorded the original line using a lavalier, so the replacement needed to reproduce the qualities of that production perspective rather than simply offering a technically superior recording.

    Carden explained that ADR sessions commonly record both boom and lavalier microphones, ideally using the same models employed during production. Even when the original line appears to come predominantly from one microphone, alternatives can prove unexpectedly useful. Navarro demonstrated this by positioning a shotgun microphone close to Carden but substantially off-axis. Pointing it directly at the performer would have produced a cleaner and more extended sound, but that was not necessarily desirable. The off-axis position rejected particular frequencies and created a tonal character that could potentially sit closer to the lavalier recording. The result was counterintuitive: the boom microphone could, under those conditions, sound more like the production lavalier than the replacement lavalier itself. Matching therefore began not with assumptions about microphone categories, but with listening.

    This distinction between technical quality and contextual suitability runs throughout professional ADR. A beautifully recorded line can fail if it sounds too clean, too close, too rich or too controlled for the image and surrounding production dialogue. Conversely, a microphone position that would appear unconventional in another recording context may provide exactly the spectral character required for a convincing match. The mixer must understand microphones well enough to use their imperfections deliberately. The question is never simply which microphone sounds best. It is which recording can become part of the scene without revealing the process that created it.

    Once Carden stepped in front of the microphone, the demonstration moved from equipment towards performance. The familiar system of three audible beeps established the timing, with the actor beginning the line where an imaginary fourth beep would occur. In principle, the task appeared straightforward. In practice, Carden repeatedly entered early, became self-conscious about the timing and discovered how quickly an apparently trivial five-word line could become difficult once performance, synchronisation and technical awareness competed for attention. His reaction gave the students an unusually useful demonstration of the psychological demands placed upon actors during ADR. Knowing the line is not enough. Understanding the timing is not enough. Even achieving synchronisation is not enough. The replacement still has to sound as though it belongs to the original performance.

    At one point, Carden produced a take that was synchronised correctly but immediately recognised that the performance itself was wrong. Navarro’s response was subtle. He did not ask for shouting or a dramatically louder delivery. He suggested only slightly more projection. The difference was not simply level. A small change in the physical production of the voice altered its tone and gave the line the additional presence heard in the original recording. When the replacement was compared directly with production dialogue, the improvement became obvious. The exercise demonstrated that ADR matching cannot be reduced to waveform alignment, pitch correction or equalisation. Performance changes the spectrum of the voice before any microphone or processor becomes involved.

    Navarro later developed this point in greater depth. Two performances can have similar apparent loudness and pitch while differing substantially in vocal quality. An actor performing naturally on set may have a relaxed throat and a particular physical relationship with the surrounding scene. On the ADR stage, tension, self-consciousness or the effort to satisfy technical instructions can change the voice. A performer may become tighter, brighter or less natural even while reproducing the words and timing accurately. Matching therefore requires attention to qualities that are difficult to describe numerically. Projection, pitch, volume and rhythm matter, but so do muscular relaxation, breath and the physical origin of the voice.

    This presents the ADR mixer with a delicate problem. The mixer may hear precisely what is preventing a line from matching, yet communicating every technical observation to the actor may make the performance worse. Navarro warned that performers can absorb only so many notes before they begin thinking about the mechanics of speech rather than the character. A useful intervention must therefore translate technical listening into language that supports performance. Sometimes the correct decision is to offer a suggestion. Sometimes it is to communicate through the director. Sometimes it is to recognise that an imperfection is not important enough to justify disturbing the creative balance of the room. Hearing a problem and knowing whether to mention it are separate professional skills.

    The lecture also demonstrated an older approach to ADR recording that remains remarkably useful. Navarro sampled the original production line and played it repeatedly into Carden’s headphones. Instead of concentrating simultaneously upon picture, beeps and performance, Carden could hear the original line and immediately reproduce it, repeating the process several times. Navarro then edited the resulting takes into position and compared them with the production recording. The method offered a direct reference for timing, intonation, volume and vocal character, allowing the performer to respond to the sound of the original performance rather than attempting to reconstruct every detail intellectually.

    Navarro’s explanation of this technique revealed an important insight into perception. Picture can sometimes become a distraction. An actor watching for a mouth movement may wait until it becomes visually apparent, by which time the correct moment to begin has already passed. If the original audio is correctly synchronised to picture, matching its rhythm and timing can naturally reproduce picture sync. The performer can therefore concentrate upon hearing and responding rather than continually monitoring several streams of information at once. Despite the age of the sampling technique, Navarro regarded it as one of the most useful tools available on the stage. Its continued value comes not from technological sophistication, but from the way it simplifies the performer’s task.

    The same objective shaped Navarro’s control room. His system contained extensive routing, multiple microphone inputs, separate monitoring paths and a heavily customised control surface, yet the purpose of this complexity was to make the session itself feel simple. He might need to manage separate mixes for the control room, recording stage, actor headphones, supervisor headphones and remote participants, with each requiring different material during rehearsal, recording and playback. Attempting to reconfigure every route manually between passes would slow the session and create repeated opportunities for error. Navarro therefore designed systems that allowed complex changes to happen immediately.

    His customised keypad provided a particularly revealing example. Individual buttons could trigger sequences of macros that armed tracks, initiated recording and performed repetitive editing operations once a take had finished. Navarro had spent months developing the core system and continued refining it whenever he noticed himself repeating unnecessary keyboard operations. He had begun his career as an ADR recordist and understood the division of labour on a two-person stage, where one person could manage recordings and files while the mixer concentrated upon the performers. Working alone, he used automation to reproduce some of that support, delegating repetitive technical operations to macros so that his attention could remain directed towards the stage.

    This led to one of Navarro’s most important ideas: the ADR mixer must work at the speed of creativity. An actor or director may suddenly discover a new approach to a line and want to record it immediately. If the mixer responds by asking everyone to wait while tracks are configured, routes changed or files prepared, the idea may lose its immediacy. A delay of only a few seconds can alter the atmosphere of a performance. Technical speed therefore has a human purpose. The mixer learns the system thoroughly enough that machinery does not interrupt thought.

    Navarro connected this principle to advice he had received about respected ADR mixer Tommy O’Connell. Asked what distinguished O’Connell’s work, a sound editor gave a simple answer: he anticipates. Useful actions are completed before anyone needs to request them. Navarro interpreted this not as a mysterious talent, but as the result of attention. If the mixer is watching the stage, listening to conversations and understanding the direction in which a session is moving, preparation for the next action can begin before a formal instruction arrives. Anticipation therefore depends upon both technical readiness and social awareness. A mixer whose attention is buried in the workstation may complete every requested task correctly while still remaining one step behind the session.

    Carden’s preparation of the cue and Navarro’s management of the stage reveal two sides of the same professional process. Carden reduces uncertainty before the session begins through accurate cueing, documentation, version control and communication. Navarro reduces friction during the session through templates, routing, automation and anticipation. Neither form of preparation is intended to make the process more rigid. Both preserve the possibility that an actor, director or mixer can respond immediately when something unexpected and valuable occurs.

    Microphone signal flow offered another example of this balance. Navarro described a deliberately clean recording path from microphone through preamplifier into Pro Tools, avoiding unnecessary outboard processing. Within the workstation, compression and equalisation could be used sparingly, but he warned against making irreversible decisions without good reason. Recording completely flat preserves maximum flexibility for the re-recording mixer, while careful decisions made during the session can still produce useful, committed tracks. Heavy processing may be difficult to undo later. A technically impressive recording decision is not valuable if it reduces the options available to the production.

    Navarro’s session was configured to make microphone comparison immediate. Several microphone inputs could remain available, feeding dedicated record tracks at appropriate levels. A boom and lavalier could be recorded simultaneously as discrete channels, preserving both perspectives for later evaluation. The workflow reflected the uncertainty inherent in matching. The microphone expected to work best may not produce the most convincing result once the line is placed against production. Recording alternatives gives later editors and mixers material from which to construct the most believable transition.

    His discussion of recording format was equally pragmatic. For conventional ADR, he worked at the established production standard of 48 kHz, 24-bit audio. Higher sample rates could be valuable for sound effects intended for extensive manipulation, where recordings might later be slowed or processed heavily. ADR serves a different purpose. Recording at unnecessarily high rates would increase storage requirements before the material was eventually converted to the format used by the production. The appropriate technical choice depends upon what will happen to the sound. More data is not automatically more useful.

    As the lecture progressed, it became increasingly clear that the most difficult parts of ADR were not contained within any equipment specification. Navarro estimated that learning the technical fundamentals and becoming comfortable with Pro Tools took years, yet he distinguished that competence from the broader ability to run a session. Recording dialogue in sync with picture is ultimately a technical process that can be learned. Managing actors, directors, supervisors, performance, uncertainty and the emotional energy of the room represents a different level of expertise.

    Actors may arrive nervous. Directors may be highly collaborative, completely self-sufficient or resistant to suggestions. A performer may struggle with a line while becoming increasingly aware of the difficulty. The mixer must understand how much intervention the situation can support. Navarro described the importance of creating a calm and comfortable environment, particularly when clients are unfamiliar with the process. Confidence can be communicated without dominance. If someone is uncertain, the mixer can explain what will happen, answer questions and demonstrate that the session is under control. Technical authority becomes useful when it reduces anxiety rather than displaying superiority.

    Carden and Navarro also acknowledged the professional judgement involved in deciding when not to intervene. A mixer may recognise an aspect of a performance that could be improved, yet nobody has asked for technical input and the issue may not be important enough to disrupt the session. Navarro framed this as a balancing act. Would the intervention materially improve the line? Is the director open to suggestions? Will another note distract the actor from a performance that is already working? Expertise includes recognising problems, while professional maturity requires distinguishing consequential problems from imperfections that do not matter.

    The ADR stage therefore brings technical precision and emotional sensitivity into the same process. The mixer must hear minute changes in vocal quality, understand microphone behaviour, maintain synchronisation, manage several monitoring environments and operate the recording system almost instinctively. At the same time, attention must remain on body language, conversation, confidence, frustration and creative momentum. Mastery of the technology is essential precisely so that the technology no longer consumes the attention needed elsewhere.

    Carden’s demonstration also placed the scale of professional ADR into perspective. Their five-word line required production recording, cue preparation, documentation, stage setup, microphone selection, rehearsal, multiple takes, alternative recording methods, editing and comparison. A skilled actor might complete around ten or twelve conventional cues in an hour, while a feature film may contain 150 or 200 ADR lines. Complicated performances naturally take longer. What appears to an audience as a few moments of seamless dialogue can therefore represent days of concentrated work.

    Yet the success of that work is measured largely through its invisibility. When Navarro compared Carden’s replacement line with the production recording, the evaluation concerned relationships: volume, tone, vocal quality, perspective and the way the replacement interacted with the surrounding scene. The door slam that had originally damaged the line was restored around the new performance, returning the replacement to the event from which it had temporarily been separated. Heard alone, the ADR recording was merely a voice on a stage. Placed back into the scene, it became part of an action.

    This transformation captures the philosophy shared by both speakers. ADR is often described as the replacement of unusable dialogue, but their demonstration showed that replacement is only the beginning of the problem. The objective is reconstruction. The actor reconstructs a performance. The mixer reconstructs a microphone perspective. Editors reconstruct timing and continuity. The final soundtrack reconstructs the relationship between voice, action and environment so successfully that the audience experiences a single uninterrupted moment.

    The joint nature of the lecture made this especially clear. Carden approached ADR through the complete production process, showing how a line travels from location problem to documented cue and finally to the recording stage. Navarro approached the same process from inside the control room, revealing the technical systems, listening skills and interpersonal judgement required to capture a convincing replacement. Their perspectives met at the point where professional preparation serves human performance.

    By the end of the session, the deliberately damaged line had become something much larger than a technical demonstration. It revealed why ADR demands more than synchronisation, why the cleanest microphone is not always the correct microphone, why a performer can match timing while missing the voice, why an old sampler can remain useful in a modern digital workflow and why a highly automated control room can make a session feel more human rather than less. Above all, Carden and Navarro showed that successful ADR depends upon where attention is directed. The actor should be thinking about performance rather than machinery. The director should be thinking about the scene rather than routing. The mixer should be watching and listening to the room rather than fighting the workstation. The audience, finally, should be thinking about none of these things. When every part of the process works together, they simply hear a character speak.

  • How Do You Mix Television Sound Under Pressure? Frank Morrone on Dialogue, Workflow, and the Art of Re-Recording

    Frank Morrone

    How do you mix television sound under pressure?

    A television soundtrack may contain hundreds of dialogue recordings, sound effects, Foley performances, backgrounds, ADR takes and music stems, all competing for space within a mix that must remain clear, emotionally convincing and technically suitable for broadcast. The audience should never become aware of that complexity. They should simply understand every line, believe every environment and remain absorbed in the story. During his online guest lecture for Edinburgh Napier University, re-recording mixer Frank Morrone explored the craft behind achieving that apparent simplicity. Drawing upon a career in film and television that began in 1979, and projects including Lost, The Strain, Sleepy Hollow and Criminal Minds: Beyond Borders, he revealed a discipline shaped equally by technology, organisation, collaboration and judgement. Throughout the session, one principle emerged repeatedly. The most effective mixing workflows allow enormous technical complexity to disappear behind the story.

    Morrone began by tracing a career that developed across several different areas of professional audio. His earliest work took place in music studios, recording jazz and orchestral film scores before following those recordings into the dubbing theatre and becoming increasingly interested in the way complete soundtracks were assembled. A move into post-production allowed him to work across dialogue editing, music editing and Foley recording before concentrating upon re-recording mixing. That breadth of experience shaped the collaborative philosophy running throughout the lecture. Unlike a music recording session, where one engineer may remain closely involved from recording through to the final mix, film and television sound brings together work created by many different specialists. The dub stage is where those contributions finally meet. Successful mixing therefore depends upon understanding not only the material itself, but the people, processes and decisions that produced it.

    Television makes this collaboration particularly demanding. Morrone described an industry in which track counts have continued to increase while schedules and budgets have become progressively tighter. Sophisticated surround mixes must be created rapidly, and emerging formats add further complexity without removing the need to support conventional playback systems. His response is not simply to work faster. It is to design workflows that remove unnecessary decisions from the mixing stage. As soon as he joins a project, he communicates with the supervising sound editor about track requirements and provides a starting template so that incoming material already fits an established structure. Organisation begins before the mixer enters the room. Under severe time pressure, the ability to find, control and compare material immediately becomes part of the creative process itself.

    This reveals something deeper about Morrone’s understanding of expertise. His templates, early conversations with sound editors, knowledge of production microphones, preparation of alternative takes and habit of printing completed passes all anticipate problems before they are allowed to interrupt the mix. The same thinking extends beyond the dubbing theatre. He considers how broadcast processing will react to dynamics, how a mix will translate to domestic systems and how material created for one format will behave when heard through another. Professional experience, in this sense, is not simply the ability to solve problems quickly. It is the ability to recognise where problems are likely to emerge and construct a workflow in which many of them have already been addressed before they become urgent.

    The scale of Lost provided a striking illustration. Morrone showed the students sessions containing extraordinary numbers of elements, including dozens of tracks dedicated to the Smoke Monster alone. Its identity emerged from a deliberately ambiguous combination of animal voices, pneumatic machinery, roller-coaster wheels and other contrasting sources, creating something that resisted being understood as either entirely organic or entirely mechanical. Those effects existed alongside hard effects, backgrounds, Foley, production dialogue, ADR, group recordings and substantial music deliveries. The technology available at the time imposed strict limits upon voices and processing, requiring careful decisions about resource allocation as well as creative balance. Complexity could not simply be solved by adding more processing. The session itself had to be organised so that the mixers could navigate it instinctively.

    Custom fader layouts and VCA groups became essential to that process. Morrone described arranging controls so that principal dialogue could immediately be balanced against ADR, group recordings and music, while different categories of material remained independently accessible. Music sources could be separated from score, while dialogue in different languages could be isolated for international deliverables. Every layer of organisation reduced the time between hearing a problem and solving it. This became especially important during pilot season, when mixers might receive unusually elaborate material without first having time to develop a workflow around the programme. A sufficiently flexible template must already be capable of accommodating whatever arrives. Preparation, in this context, creates the conditions in which creative decisions can still be made under pressure.

    The production of Lost also demonstrated the tension between creative ambition and delivery requirements. Morrone recalled the exceptional resources devoted to the programme, including a large pilot budget and Michael Giacchino’s insistence upon recording a live orchestra for each episode. Yet the soundtrack still had to survive the restrictions of television broadcast. The team therefore created a more dynamic version for DVD before producing a contained broadcast mix designed to survive transmission processing. A similar approach was later adopted on The Strain. The distinction mattered. A mix can remain technically within specification and still behave poorly when subsequent broadcast processing responds to excessive dynamics. Re-recording therefore requires mixers to think beyond the dubbing theatre. They are mixing for every system through which the programme will eventually reach its audience.

    Dialogue occupied the centre of Morrone’s approach. His first objective is always to preserve the production performance wherever possible. When ADR has been recorded, he wants to know why. A line replaced for performance reasons presents a different problem from one replaced to solve a technical fault. If the director wanted a different performance, Morrone respects that decision while keeping the original available as an alternative. If the problem was technical, he first explores whether the production recording can be repaired. Modern restoration tools have dramatically expanded what can be rescued, reducing the need to replace performances that may possess subtleties difficult to recreate months later in an ADR studio.

    When ADR is necessary, matching involves far more than applying equalisation and reverb. Morrone keeps production dialogue available alongside the replacement so that he can compare transitions directly, matching tone, acoustic environment, pacing and performance. He prefers to receive several strong takes and recordings from both boom and lavalier microphones, recognising that the microphone apparently closest to the production perspective is not always the easiest to integrate. Understanding which microphones and wireless systems were used during production can also provide valuable clues, particularly where transmission systems have imparted their own sonic characteristics. ADR matching consequently becomes a process of reconstructing relationships rather than searching for a single corrective setting.

    Technology has transformed that work, though Morrone repeatedly warned against allowing powerful restoration tools to encourage excessive processing. Noise reduction, spectral repair, ambience matching, EQ matching and dereverberation can rescue material that would previously have required replacement. Yet he deliberately removes less noise than might appear necessary when dialogue is heard in isolation. Once backgrounds, effects and music return, much of the remaining noise may be perceptually masked. Processing that sounds impressively clean in solo can leave dialogue lifeless and constricted in the finished mix. Morrone therefore keeps copies of original material and sometimes returns to less processed versions during the final mix. The objective is not the cleanest possible dialogue track. It is dialogue that remains natural and convincing within the complete soundtrack.

    His attitude towards restoration reveals a broader philosophy of technology. Morrone began working with magnetic tape, a Cat 43 noise reduction unit and a notch filter, and he clearly values the extraordinary capabilities available to contemporary mixers. Yet greater technical power has not removed the need for judgement. In many respects, it has increased it. The ability to remove more noise does not mean that more noise should be removed, just as the ability to place sound almost anywhere within an immersive field does not mean that every available position should be used. New tools expand the range of possible decisions. They do not determine which decisions are appropriate. Throughout Morrone’s lecture, technical capability remained subordinate to perception.

    This distinction between isolation and context became one of the lecture’s most important ideas. Morrone described television mixing as beginning with dialogue and music, establishing the foundation against which the effects mixer can develop backgrounds and action. Once those elements come together, the mixers decide what should drive the scene. Some moments belong to music, others to effects, while still others require a more subjective perspective or a deliberate reduction of material. Experienced mixing partners develop an almost instinctive understanding of one another’s decisions. If an effect obscures a line during an early pass, Morrone knows that a trusted colleague will create space for it during refinement. Collaboration is therefore more than the division of tracks between two people. It is a shared process of deciding where the audience should listen.

    The idea that sounds must be judged in context reaches far beyond balancing dialogue against effects. Dialogue that appears too noisy when soloed may become entirely convincing once the world of the scene surrounds it. ADR that attracts attention after twenty repeated comparisons may pass unnoticed when the audience encounters it once within a continuous performance. Group recordings that sound absurd under isolated scrutiny may perform their role perfectly when placed at the correct distance behind principal dialogue. Morrone’s examples repeatedly challenged the assumption that individual elements should be perfected independently. The meaningful unit of evaluation is ultimately the audience’s experience of the scene. A soundtrack succeeds through relationships between sounds rather than the isolated perfection of its components.

    Normally, the priority within those relationships remains dialogue. Morrone described it as both the foundation of the mix and the principal carrier of storytelling. His concern with intelligibility, however, extends beyond maintaining a technical hierarchy between dialogue, music and effects. It is fundamentally about preserving attention. His most revealing test is not a meter reading but a listener asking what somebody has just said. At that moment, the audience member has been pulled out of the story and required to think about the failure of the soundtrack. Clear dialogue therefore supports immersion precisely by avoiding attention to itself. The better the mix communicates, the less the audience needs to think about the process of communication.

    Yet the rule is not absolute. For one sequence in The Family, a child was being used as bait to attract a kidnapper in a crowded shopping centre. The environment needed to feel genuinely busy, yet the available collection of separate crowd recordings and reverberant elements never created a convincing whole. Morrone took a six-channel recorder into a real shopping centre, captured the food court from different perspectives and brought the recordings back into the mix. The result allowed the environment to crowd the dialogue slightly, which was precisely what the scene required. Clarity remains fundamental, but realism sometimes depends upon controlled difficulty.

    The shopping-centre sequence illustrates an important tension within Morrone’s approach. His practice is built upon strong principles, but those principles do not become inflexible rules. Dialogue normally takes priority, except when allowing the environment to interfere with it makes the dramatic situation more believable. Restoration should preserve intelligibility, except when excessive cleaning destroys naturalness. Acoustic simulation is valuable, except when recording the real physical relationship between people and space produces a more convincing result. Expertise therefore involves knowing the rules well enough to understand what they protect, then recognising the moments when the needs of the scene justify bending them.

    This willingness to leave the dubbing theatre and record real spaces appeared repeatedly throughout the session. Morrone discussed impulse responses, convolution reverbs and carefully developed presets for rooms, vehicles and other environments, all of which provide useful starting points for worldizing sound. Yet he remained pragmatic about their limitations. A real location does not necessarily sound like the space suggested by the finished image, and even a carefully captured impulse response may require additional reflection, delay or reverberation before it feels convincing. Cars present an especially difficult problem, combining strong early reflections from glass with highly absorptive surfaces elsewhere. Experience gradually produces a library of useful starting points, though listening still determines the final result.

    Sometimes no simulation is as convincing as returning to the physical situation itself. Morrone described a scene in which a group of foster children were supposed to be creating chaos upstairs while a conversation took place below. Studio-recorded group voices did not reproduce the peculiar combination of footfalls, structural transmission and reflections travelling down a staircase into another room. His solution was direct. He gathered children on the upper floor of a house, recorded from downstairs and captured the complete acoustic event as it occurred. No increasingly elaborate chain of processing was required. The physical relationship between performers, building and microphone provided what the scene needed.

    The pressures of television production make such judgement particularly important. Morrone compared television mixing to boot camp. Schedules leave little room for hesitation, and mixers must develop workflows capable of producing strong results quickly. Once a scene has been successfully mixed, he prefers to print it rather than trusting that automation will remain untouched throughout later work. Accidental writes, changed sends and technical errors can occur even in experienced hands. Printing completed work provides security and allows later changes to be punched into established stems. Efficiency does not mean rushing blindly. It means reducing the opportunities for avoidable problems to consume the limited time available.

    Long sessions also introduce a more human limitation: hearing fatigue. Morrone described working on the highly dynamic soundtrack of The Strain, where sustained exposure to loud material forced him to think deliberately about auditory recovery. His solution was simple but important. He left the room for short periods, walked and allowed his ears to recover while his mixing partner continued working. The two mixers could alternate demanding passes, giving each other opportunities for rest without stopping the session. After decades of mixing, one of the useful discoveries was not another plug-in or processor, but the value of leaving the chair for ten minutes. Professional listening depends upon recognising the limits of the listener.

    Morrone’s discussion of client relationships revealed another dimension of the re-recording mix that students may rarely encounter in technical demonstrations. Mixers are working with directors, producers and other clients who may have strong preferences that differ from their own. Morrone described situations in which clients wanted music loud enough to compete with dialogue. His responsibility was to explain the likely consequences, demonstrate how the mix translated at lower levels and on smaller monitors, and search for a compromise that preserved the client’s intention while protecting intelligibility as far as possible. The mixer offers expertise, but does not own the programme. Knowing which decisions are worth challenging and which require accommodation is part of the craft.

    His description of deciding whether an issue represents a “hill to die on” reveals a sophisticated understanding of professional authority. Expertise does not give the mixer unlimited control over the work, nor does collaboration require the abandonment of professional judgement. The mixer must advocate for the audience, explain likely consequences and make alternatives audible, while recognising that the final creative intention belongs to the client. Professional judgement therefore includes negotiation. Sometimes expertise means defending a decision. Sometimes it means finding a compromise that neither side initially imagined. Occasionally, it means implementing a choice that remains contrary to personal taste while ensuring that it works as successfully as possible.

    Technical fluency plays an important role in maintaining those relationships. When a client requests a change, Morrone wants to make it immediately, play it once and continue. Searching through tracks or repeatedly troubleshooting a familiar process changes the atmosphere of the room and interrupts attention to the programme itself. A well-designed session keeps the conversation focused upon storytelling and communication. The deeper the mixer’s command of the tools, routing and session layout, the less those systems intrude upon the creative discussion.

    This is another form of transparency. Morrone repeatedly returned to the idea that technology should disappear from the client’s experience, yet transparency does not mean that technology has become unimportant. The opposite is closer to the truth. Considerable technical knowledge is required to make complex systems feel immediate. Templates, custom layouts, routing, monitoring, printed stems and intimate familiarity with the workstation create an environment in which a creative request can become an audible result without breaking the flow of the session. Mastery becomes visible through the absence of friction.

    The growth of immersive sound has expanded this challenge further. Morrone described mixers as working between extremes, from Dolby Atmos and sophisticated home theatres to stereo playback and mobile devices with earbuds. His philosophy was to begin with the best mix possible in the most capable format, then ensure that it translates successfully into simpler ones. Immersive technology may offer extraordinary spatial possibilities, but Morrone remained cautious about using novelty without considering perception. In particular, he defended dialogue remaining anchored to the centre. Moving voices between speakers can alter timbre and create distracting changes that audiences may notice without understanding their source. New formats create possibilities, though they do not invalidate principles developed through decades of listening.

    His argument about dialogue placement is particularly revealing. Audiences do not need to identify the technical source of a problem in order to experience discomfort. Morrone described viewers sensing that something was wrong when dialogue moved around an immersive field, even when they could not explain precisely what disturbed them. This places an unusual responsibility upon the mixer. Professional listening must sometimes diagnose experiences that ordinary listeners can feel but cannot name. The purpose of expertise is not to dismiss those responses as technically uninformed, but to understand the perceptual conditions that produced them.

    The same concern for translation shaped his approach to low frequencies. Subwoofers vary enormously between listening environments, and domestic listeners frequently adjust them far beyond calibrated levels. Morrone therefore warned against depending entirely upon the LFE channel for the weight of a soundtrack. Low-frequency energy can also be carried through the main channels, creating a result that remains powerful across a wider range of playback systems. His experience of hearing a domestic subwoofer struggle with low-frequency material from Sleepy Hollow reinforced the point. A soundtrack must survive real listening environments, not merely sound impressive on a perfectly calibrated dubbing stage.

    As the discussion widened, Morrone considered the future possibility of mixes that adapt more intelligently to different devices and contexts. Streaming, immersive audio, virtual reality and personalised playback were already creating pressure for soundtracks to function across radically different systems. Yet his underlying philosophy remained remarkably consistent. Whatever the format, begin with the strongest possible mix, preserve the storytelling hierarchy and understand how human perception responds to the result. Technology changes quickly. The responsibility to guide attention and communicate narrative does not.

    Questions from the students returned the discussion to practical preparation. Morrone strongly supported editors delivering material that is already sensibly balanced before it reaches the stage. Dialogue editors who use clip gain to create consistent levels save valuable mixing time, while backgrounds and effects arriving close to useful operating levels allow the mixer to begin creatively rather than first correcting avoidable problems. He described requesting particular editors for demanding productions precisely for this reason. Good preparation is noticed. In a professional environment where a pilot may need to be mixed in only a few days, the person who consistently delivers well-organised, intelligently balanced material becomes someone mixers actively want on the next project.

    His discussion of ADR offered another deceptively simple lesson about perception. Morrone sometimes works on difficult replacement lines privately through headphones while the effects mixer is making a pass. The client then hears the finished line only once in context rather than listening to it repeated dozens of times during adjustment. Repetition directs attention towards the repair and teaches the listener exactly where to expect it. The same awareness shaped his humorous rule about never soloing loop group in front of a client. Background conversations that work perfectly as part of a scene may sound absurd when isolated and scrutinised. What matters is not whether every element survives examination on its own, but whether it performs its intended role in the scene.

    These examples reveal that Morrone is not simply mixing sound. He is managing attention, expectation and knowledge. Once listeners have been taught where an edit exists, they may hear it differently. Once an element has been isolated, they may judge it according to criteria that have little relevance to its actual purpose. Perception is shaped not only by acoustic information, but by what listeners have been encouraged to notice. Part of the mixer’s craft therefore lies in protecting the audience’s experience from unnecessary awareness of the mechanisms used to construct it.

    Yet group recording could also become a powerful storytelling tool. On Criminal Minds: Beyond Borders, episodes moved between international locations while much of the production remained based in Los Angeles. Carefully performed local-language group recordings, combined with music and other environmental elements, became essential to establishing each location convincingly. Here, group material could be brought forward rather than hidden. There was no universal rule governing how loudly an element should be mixed. Its appropriate level depended upon what the scene needed to communicate.

    By the end of the session, Morrone’s account of re-recording mixing had moved far beyond faders, plug-ins and delivery specifications. The technology matters enormously, as do templates, routing, restoration tools, monitoring and control surfaces, but those things serve a larger process. A mixer must understand performance, storytelling, perception, collaboration, translation and the subtle politics of working with clients under pressure. The session may contain hundreds of tracks, yet the audience should hear a coherent world rather than the complexity required to construct it. Morrone’s lecture revealed a craft built upon anticipation, contextual judgement and the careful management of attention. Preparation preserves the possibility of creativity under pressure. Technical knowledge allows technology to disappear from the conversation. Rules provide essential foundations, while experience reveals when the needs of a scene require them to bend. Great television sound is not created by making every element impressive or every recording perfect in isolation. It emerges from understanding what the audience needs to hear, recognising what they should never need to notice, and making hundreds of individual decisions feel like one continuous experience.

  • How Do You Build a Sonic Brand? Sean Thornton on Audio Branding, User Experience, and Designing with Sound

    Sean Thornton

    How do you build a sonic brand?

    Most people can recognise a familiar logo within a fraction of a second. Distinctive colours, typography and visual symbols allow organisations to establish an identity that remains remarkably consistent across products, advertising and digital services. Sound, however, has often been treated very differently. Many organisations commission an audio logo, perhaps a short mnemonic played at the end of an advertisement, and regard the task as complete. During his online guest lecture for Edinburgh Napier University, Sean Thornton challenged this assumption directly. Drawing upon more than a decade working in audio branding and as co-founder of Audio UX, he argued that effective sonic branding is not created through isolated assets but through coherent systems. Throughout the lecture, one idea repeatedly emerged. Sound should be treated as a complete design language rather than a collection of individual sounds.

    Thornton began by reflecting upon his own professional journey. Although he initially studied music with ambitions of becoming a film composer, early experiences within game audio and commercial music gradually led him towards a field that scarcely existed as a recognised discipline when he was a student. That progression illustrated one of the lecture’s broader themes. Creative careers rarely follow predictable paths. Formal education provides an important foundation, though many professional specialisms emerge only through experimentation, curiosity and the willingness to explore opportunities beyond conventional career routes. Audio branding, he suggested, represents precisely this kind of evolving discipline, combining composition, sound design, psychology, branding and user experience into a field that continues to redefine itself.

    Before discussing design methods, Thornton carefully distinguished between several closely related concepts. Audio branding, in its simplest form, involves the strategic use of sound to establish or reinforce identity. Yet he argued that this definition captures only part of the picture. Organisations increasingly communicate across websites, mobile applications, physical products, voice assistants, advertising, podcasts and public spaces. Users encounter brands through countless interactions rather than through a single advertisement or logo. If every one of those interactions produces unrelated sounds, opportunities to build familiarity and trust are quickly lost. Holistic audio branding therefore asks a different question. Instead of considering how an individual sound performs in isolation, it asks how every auditory experience contributes towards a coherent and recognisable identity. The objective is not consistency through repetition, but consistency through design.

    This broader perspective naturally led to Thornton’s discussion of Audio User Experience, or Audio UX, which underpins his company’s approach. The distinction is subtle yet important. Audio branding concerns the sounds themselves. Audio UX concerns how people experience those sounds. A perfectly crafted sonic identity still fails if it frustrates, distracts or overwhelms the people who encounter it. Throughout the lecture, Thornton repeatedly encouraged students to approach sound through empathy before aesthetics. Designers should not begin by asking whether a sound expresses the personality of a brand. They should first ask whether it genuinely improves the experience of the person hearing it. Successful sonic branding therefore begins with people rather than brands. When the user experience has been designed well, brand recognition becomes a natural consequence rather than the primary objective.

    This philosophy becomes much clearer when viewed alongside visual identity design. Few organisations would expect a designer to create only a logo while ignoring colours, typography, photography and layout. Successful visual identities are built from systems of related elements that can be applied consistently across many different contexts. Thornton argued that sound should be approached in exactly the same way. Instead of delivering a handful of isolated assets, audio designers should create flexible frameworks capable of supporting products, advertising, digital interfaces, physical environments and future developments that may not yet exist. Individual sounds remain important, though their real value lies in how they work together. Sonic branding therefore becomes less about composing memorable cues and more about designing a language that other designers can continue to use.

    One of the lecture’s most distinctive ideas concerned the difference between creating assets and creating design systems. Traditional projects often revolve around a short audio logo, a notification sound or a piece of advertising music delivered as a finished product. Thornton argued that this approach solves only today’s problem. Instead, he proposed building reusable sonic building blocks from which many future experiences could emerge. Characteristic instrumental colours, recurring melodic gestures, harmonic language, rhythmic behaviour and carefully selected textures become components within a wider design framework rather than fixed compositions. The designer is no longer delivering a finished collection of sounds. They are creating the vocabulary from which an organisation’s sonic identity can continue to grow.

    Thornton illustrated this philosophy through the development of a comprehensive sonic identity for USA Today. Rather than beginning with musical preferences or fashionable production styles, the project started by asking a more fundamental question. What should one of America’s largest news organisations sound like? The resulting concept centred upon the idea of a sonic mosaic that reflected the diversity of voices, cultures and experiences represented within the publication. Recordings of instruments associated with different regions of the United States were combined with contributions from a wide variety of people, gradually creating a distinctive palette of sonic materials. These recordings were never intended simply to become finished pieces of music. They formed the foundation of a much larger design system from which future compositions, interface sounds, broadcast material and branded experiences could all develop while remaining recognisably connected. The case study demonstrated that successful audio branding begins long before composition. It begins by defining the ideas, relationships and design principles from which every subsequent sound will emerge.

    Perhaps the most forward-looking aspect of Thornton’s presentation concerned what happens after a sonic identity has been created. Visual brands are supported by style guides, governance and clear documentation that help maintain consistency across years of future development. Thornton argued that sonic identities require the same level of stewardship. Designers should not simply hand over a collection of audio files. They should provide the principles, frameworks and guidance that allow other teams to implement those sounds consistently across new products, services and technologies. Designing a sonic brand therefore involves far more than composing memorable sounds. It means creating a living design language capable of evolving without losing its identity. That broader challenge would become the focus of the remainder of the lecture, as Thornton explored how these systems are implemented, managed and refined over time.

    Having established the principles behind holistic audio branding, Thornton turned to the practical question that ultimately determines whether those ideas succeed. Designing a sonic identity is only the beginning. The greater challenge lies in ensuring that it remains useful, recognisable and adaptable as organisations evolve. Throughout the second half of the lecture, he argued that the true value of an audio brand emerges not from individual assets, but from the systems that allow those assets to be deployed intelligently across countless future experiences.

    The USA Today project provided a detailed illustration of this philosophy in practice. Rather than delivering a fixed collection of completed sounds, Thornton’s team developed a series of reusable audio components that could be assembled in different ways depending upon the context. This modular approach reflected the same design thinking used throughout modern visual branding. Designers rarely create a new colour palette or typeface for every campaign. Instead, they draw from a shared system that allows individual pieces of communication to remain distinctive while still feeling unmistakably connected. Thornton argued that sound should operate according to exactly the same principle. By working with carefully designed components instead of inflexible assets, organisations gain the freedom to create new experiences without continually reinventing their sonic identity. The result is a brand that remains recognisable while continuing to evolve alongside new products, technologies and audiences.

    One of the most revealing aspects of this system involved thinking about sound through the language of design rather than music. Thornton described key signatures as global design patterns capable of giving different pieces of audio a shared tonal identity. Characteristic textures functioned in much the same way as colours within a visual brand, while processed recordings became recognisable building blocks that could be reused in multiple contexts. Even small musical gestures could perform the same role as graphic motifs or visual icons, quietly reinforcing identity without demanding attention. This perspective shifted the discussion away from individual compositions towards relationships between sonic elements. Listeners may never consciously recognise these patterns, yet together they help establish familiarity across an increasingly diverse collection of products and services.

    The USA Today identity demonstrated how this philosophy could be implemented across very different media. A podcast introduction, a smart speaker news briefing and a broadcast sequence all required different durations, different pacing and different levels of branding. Treating each as an independent project would inevitably weaken consistency. Instead, every experience drew from the same underlying design language while adapting its presentation to suit the situation. Thornton described, for example, how a brief branded audio fingerprint proved more appropriate for voice assistants than a longer musical introduction. Listeners asking for the latest headlines wanted information quickly. A concise sonic reminder of the brand communicated identity without delaying access to the content. User experience and branding were therefore not competing priorities. Each strengthened the other.

    This naturally led to one of Thornton’s strongest practical messages. Designing an effective sonic identity is only half of the task. The identity must also be implemented consistently by the many people who will use it in the future. A beautifully designed audio logo achieves very little if nobody understands when it should be used, when silence would be more appropriate or how new material should relate to the existing system. Implementation therefore becomes a creative activity rather than an administrative one. Clear principles, documentation and governance allow organisations to grow their sonic identities confidently without gradually losing the coherence that made them distinctive in the first place.

    Voice itself formed another important component of this broader system. Thornton argued that organisations should think beyond selecting an appropriate voice actor. Voice also encompasses conversational style, language, pacing and the structure of spoken interactions. These decisions become particularly significant as brands increasingly communicate through voice assistants, automated services and accessibility technologies. Designing effective conversations therefore requires the same careful planning applied to music and sound design. Every spoken interaction contributes towards the wider experience of the brand, making conversation design an integral part of holistic audio branding rather than a separate discipline.

    Accessibility emerged as a recurring theme throughout this discussion. Thornton repeatedly returned to the idea that good sonic design should improve experiences for everyone rather than simply strengthening recognition. Voice interfaces, carefully considered auditory feedback and appropriately designed sonic cues can all help create more inclusive interactions for people with different sensory abilities. Importantly, these improvements should not be regarded as specialist features added for a small minority of users. Designing with accessibility in mind frequently produces better experiences for everybody. Rather than seeing accessibility as a constraint, Thornton encouraged students to view it as an opportunity to create richer, more intuitive and more human-centred experiences. Once again, the lecture returned to its central principle. Successful sonic branding begins with people rather than brands.

    Towards the end of the lecture, Thornton reflected upon the future of the discipline with considerable optimism. Advances in spatial audio, wearable technology and increasingly precise location-aware devices present exciting opportunities to rethink how organisations communicate through sound. Yet these possibilities also introduce important ethical questions. Thornton imagined a future in which highly personalised spatial advertisements might follow listeners through everyday environments, responding continuously to their location and behaviour. Although technically possible, he questioned whether such applications would genuinely improve people’s lives. New technologies, he argued, should first be viewed as opportunities to create more meaningful experiences rather than simply more opportunities for marketing. The question is never merely what audio branding can do, but what it should do.

    This concern led naturally to a broader critique of current practice. Thornton observed that many organisations continue to treat sonic branding as a superficial exercise, commissioning generic three-note logos or brief musical signatures without considering the wider user experience. As more companies embrace sound, he believes the greatest opportunity lies not in creating more branded audio, but in creating better branded audio. Holistic systems, careful implementation and genuine concern for the listener allow organisations to stand out far more effectively than louder or more intrusive branding ever could. Thoughtful design therefore becomes both a competitive advantage and an ethical responsibility.

    The lecture concluded by returning to one of its earliest themes. Brands are not static objects. They evolve continually alongside the people who use them. Thornton argued that sonic identities should therefore be managed in much the same way as living organisms. They require ongoing stewardship, periodic refinement and careful adaptation as technologies, expectations and cultural contexts change. An audio brand is never truly finished. It develops through continual listening, evaluation and improvement. Ultimately, Thornton challenged students to stop thinking about audio branding as the design of sounds and instead see it as the design of relationships. Every carefully considered interaction strengthens the connection between people and the organisations they encounter. When approached in this way, sonic branding becomes less about recognition and more about creating experiences that people genuinely value.

  • How Do You Design the Sound of a Blockbuster Game? Michael Caisley on Creativity, Recording, and Crafting the Sound of Call of Duty

    Michael Caisley

    How do you design the sound of a blockbuster game?

    Modern video games are built from extraordinarily complex systems. Artificial intelligence, physics, animation, graphics and networking all operate simultaneously to create worlds that respond continuously to the player’s decisions. Sound design must function within that same complexity. Unlike film, where every frame is predetermined, game audio unfolds differently every time someone plays. Thousands of individual sounds interact dynamically, responding to changing environments, player behaviour and gameplay events without losing clarity or dramatic impact. During his online guest lecture for Edinburgh Napier University, Michael Caisley drew upon his experience as Senior Sound Designer on Call of Duty: Advanced Warfare to explore how one of the industry’s largest productions approached this challenge. Throughout the session, one principle emerged repeatedly. Great game audio is designed as a complete system rather than a collection of individual sound effects.

    This philosophy shaped every stage of the project’s development. Rather than asking how individual weapons, footsteps or explosions should sound, the audio team began with a broader question. How should the player experience the world? Every recording, editing decision and implementation technique ultimately served that objective. Sound design therefore became an exercise in shaping perception rather than simply producing assets. Individual recordings remained important, though their true value emerged only through the relationships they formed with every other element of the soundtrack. The player never experiences sounds in isolation. They experience an acoustic world.

    Caisley explained that this perspective influenced one of the team’s earliest decisions. Although Call of Duty already possessed an established sonic identity developed across multiple successful titles, the audio team resisted the temptation simply to inherit those conventions. Instead, they treated Advanced Warfare as an opportunity to rethink the game’s entire sound philosophy from first principles. Existing assets, familiar production techniques and long-standing implementation methods were all reconsidered. Their ambition was not to reject the past, but to ensure that every creative decision continued to serve the experience they wanted players to have. Innovation therefore emerged through careful questioning rather than change for its own sake.

    That philosophy also transformed the relationship between sound design and implementation. In many production pipelines, sound designers create assets that are later integrated into the game by other specialists. Caisley described a markedly different approach. Sound designers remained responsible for implementation inside the game itself, allowing them to shape how recordings behaved once they became part of the interactive experience. The timing of a sound, the circumstances under which it played, the way it interacted with other events and its contribution to the overall mix all became part of the design process. Creating an excellent recording represented only the beginning. The player’s experience ultimately depended upon how successfully that recording functioned within the wider system. Implementation was therefore not separate from sound design. It was an essential part of it.

    The same systems-oriented thinking naturally extended to recording. Rather than relying primarily upon commercial sound libraries, the team invested heavily in producing original recordings specifically for the game. Specialist libraries remained valuable resources, particularly carefully curated collections produced by experienced field recordists, though Caisley consistently argued that original recording provides opportunities to discover sounds that nobody else possesses. More importantly, recording becomes a creative process rather than simply a method of gathering raw material. Unexpected textures, unusual perspectives and subtle acoustic details often emerge only when designers capture sounds for themselves. Distinctive game audio begins long before editing or implementation. It begins with listening carefully to the world.

    One particularly revealing example involved footsteps. Traditional Foley often records isolated footsteps on carefully prepared surfaces inside controlled studio environments. Caisley questioned whether this approach remained appropriate for a first-person game in which movement is experienced continuously through the player rather than observed from an external viewpoint. Instead, the team carried lightweight portable recorders into forests, hillsides and outdoor locations, capturing complete performances that naturally progressed from walking to running and sprinting. Rather than constructing movement artificially from disconnected recordings, they captured the changing rhythm, effort and momentum that emerge naturally when people move through real environments. The resulting recordings felt noticeably more convincing, illustrating that authenticity sometimes depends less upon technical precision than upon preserving the natural behaviour of the performer.

    The recording equipment itself reflected the same practical philosophy. Caisley encouraged students not to become preoccupied with expensive technology at the expense of creative opportunity. Much of the team’s field recording relied upon compact portable recorders that could be deployed quickly whenever an interesting sound presented itself. Mounted directly onto lightweight boom poles, these systems reduced handling noise while allowing recording sessions to remain flexible and spontaneous. The lesson extended far beyond the specific equipment being used. Interesting sounds rarely arrive when it is convenient to record them. Designers therefore benefit from tools that allow them to respond immediately rather than waiting for ideal conditions or elaborate recording setups. Creativity, he suggested, often rewards preparedness more than perfection.

    The same willingness to question established practice shaped the recording of weapons. Rather than organising one large recording session intended to capture every firearm in a single location, the team divided the work across numerous smaller sessions. This approach simplified logistics, though its greatest benefit proved creative rather than organisational. Each session could be reviewed afterwards, allowing the team to identify opportunities for improvement before returning to record additional material. Different environments also introduced naturally varying acoustic characteristics, providing a richer collection of perspectives than a single location could have offered. Recording therefore became an iterative process in which every session informed the next. The objective was not simply to accumulate material, but to refine the sonic identity of the game through continual experimentation.

    Perhaps the most important lesson from this stage of the lecture concerned the relationship between individual sounds and the finished player experience. Caisley observed that players rarely remember isolated recordings. They remember moments. The impact of those moments depends upon countless design decisions working together, from recording and editing through implementation, mixing and gameplay design. The audio team’s objective was therefore never to create the loudest explosion or the most detailed weapon recording. It was to build a soundtrack in which every element supported the player’s understanding of the world. Call of Duty: Advanced Warfare consequently adopted a more dynamic approach to mixing, allowing important sounds to occupy the foreground while leaving space for the rest of the soundtrack to breathe. Restraint became every bit as valuable as spectacle. The most memorable moments did not emerge from individual sound effects alone. They emerged from a coherent acoustic world in which every element strengthened the player’s belief that the environment around them was responsive, believable and alive.

    Having established the technical foundations of the project, Caisley turned towards the creative decisions that ultimately give a game its identity. Recording and implementation provide the raw materials, though they do not determine how a player experiences a moment. That depends upon judgement. Throughout the remainder of the session, he returned repeatedly to an idea that sounds deceptively simple but lies at the heart of professional sound design. Every sound reflects a design decision. The role of the sound designer is not merely to create convincing audio, but to decide what deserves to be heard, when it should be heard and, just as importantly, what should remain absent.

    This philosophy shaped the way Caisley approached almost every design problem. Instead of searching immediately for the perfect recording, he preferred to build what he described as palettes of possibilities. Families of related sounds sharing particular textures, movements and tonal characteristics were assembled through recording, processing and experimentation. Organic recordings of motors, impacts, machinery and environmental sounds were manipulated repeatedly, gradually forming a collection of materials from which the final design could emerge. Creativity therefore developed through exploration instead of beginning with a predetermined solution. Designers rarely know exactly what they are searching for at the start of a project. They discover it by experimenting until unexpected relationships begin to reveal themselves.

    His workflow reflected the same exploratory mindset. Projects often began in apparent disorder, with sounds accumulating rapidly as multiple ideas were investigated simultaneously. Immediate organisation was deliberately given lower priority than experimentation. Once a broad range of possibilities had been created, the process shifted towards careful refinement. Caisley compared this approach to sculpting. A sculptor begins with a block of material and gradually removes everything that does not belong until the final form becomes visible. Sound design, he suggested, often develops in exactly the same way. Instead of continually asking what should be added, designers should also ask what can be removed.

    This idea challenges one of the most common assumptions made by new sound designers. Richer sound does not necessarily result from adding more layers. As recordings accumulate, frequency masking increases, textures become crowded and important details begin to disappear. Caisley described repeatedly muting, removing and simplifying elements until only those making a genuine contribution remained. Equalisation, dynamics processing, timing adjustments and careful layering all supported this process, though none represented the objective in itself. Their purpose was to improve clarity, strengthen communication and ensure that every remaining sound justified its place within the mix. Professional sound design therefore depends less upon the quantity of material than upon the quality of the decisions shaping it.

    A particularly memorable example came from a sequence in which the player escapes across a glass roof before an ally destroys the structure beneath pursuing enemies. The obvious solution might appear to involve recording increasingly dramatic glass impacts before combining them into one spectacular crash. Caisley approached the problem very differently. The event was divided into a sequence of distinct dramatic stages. Initial bullet impacts, subtle structural weakening, growing instability and the final collapse each received their own carefully judged sonic treatment. Texture, pacing and silence changed gradually as the scene unfolded, allowing players to follow the progression of the collapse as a connected series of events rather than experiencing a single overwhelming burst of noise. The sequence derived its dramatic impact from the way the sound evolved over time, allowing the narrative of the scene to unfold naturally through listening as well as through the visuals.

    The same attention to dramatic pacing shaped Caisley’s approach to synchronisation. Students often assume that every visible action should be matched precisely by an accompanying sound. Professional practice, he suggested, is considerably more nuanced. Delaying one sound slightly, allowing another to emerge first or simplifying an otherwise crowded moment can produce a stronger dramatic effect than strict synchronisation alone. Rhythm, pacing, expectation and contrast all become compositional tools that guide the player’s attention. Instead of following every visual event mechanically, sound design helps determine what players notice, what they anticipate and how they interpret the unfolding action. Games therefore rely upon many of the same principles of dramatic storytelling found in music and cinema, while remaining responsive to player interaction.

    Equally revealing was Caisley’s discussion of realism. Throughout the lecture, he challenged the assumption that authentic sound must originate from authentic sources. Recording larger explosions does not necessarily produce better explosions, nor does striking more metal automatically create more convincing mechanical impacts. Professional sound designers routinely combine recordings whose original sources bear little resemblance to the finished result. Environmental ambiences, machinery, organic textures and countless unexpected recordings may all contribute qualities that literal recording alone cannot provide. What ultimately matters is not the origin of the sound, but whether it supports the player’s perception of the world. Believability depends upon the finished experience rather than literal accuracy.

    Technical processing formed part of this broader creative process rather than existing as an end in itself. Equalisation, compression, distortion and other processing tools undoubtedly shape the final soundtrack, though Caisley resisted presenting them as universal recipes. Every adjustment served a specific purpose within the wider composition. Heavy compression might transform an otherwise unremarkable recording into the perfect supporting layer. Subtle timing adjustments could reveal details previously hidden within the mix. Equalisation often preserved recordings that might otherwise have been discarded. Considered individually, many processed sounds appeared incomplete or even unattractive. Their value emerged only through their relationship with every other element. As throughout the lecture, the emphasis remained firmly upon systems rather than isolated sounds.

    Towards the end of the session, Caisley reflected upon the qualities that distinguish successful sound designers from merely competent technicians. Technical expertise undoubtedly matters, though he argued that curiosity, collaboration and the willingness to accept constructive criticism exert a far greater influence over long-term professional development. Working alongside experienced colleagues continually challenges assumptions and exposes designers to alternative ways of thinking. Equally valuable is the habit of listening analytically to other people’s work. Rather than deciding whether an entire game succeeds or fails, Caisley encouraged students to identify individual moments that demonstrate particularly thoughtful creative decisions. Examining one successful interaction in depth often teaches far more than making broad judgements about an entire soundtrack. Developing as a sound designer therefore depends as much upon careful listening as upon creating new sounds.

    Taken together, Caisley’s presentation revealed that blockbuster game audio is built as much through judgement as through technology. Recording, editing, implementation and mixing undoubtedly provide the necessary tools, though those tools acquire meaning only through the decisions that shape them. Every sound exists in relation to every other sound, every moment contributes to a larger dramatic experience and every creative choice influences how players understand the world around them. Sound design is not the art of creating more sound, but of making better decisions. Technology provides the tools. Careful listening, thoughtful judgement and an understanding of human perception transform those tools into interactive experiences that players instinctively accept as real.

  • How Do Robots Communicate Through Sound? Connor Moore on Audio UX, Robotics, and Designing Meaningful Interactions

    Connor Moore

    How do robots communicate through sound?

    People increasingly interact with technology through sound. Smartphones acknowledge completed payments, electric vehicles alert drivers to potential hazards, wearable devices provide subtle notifications and intelligent products communicate through a growing vocabulary of tones, chimes and alerts. Yet these sounds rarely receive the same attention as visual design. During his online guest lecture for Edinburgh Napier University, Connor Moore explored the growing discipline of Audio User Experience (Audio UX), demonstrating how carefully designed sounds help products communicate naturally, build trust and express personality. Drawing upon projects for companies including Google, Tesla and Postmates, he argued that successful product sound design extends far beyond creating attractive audio. It begins with understanding the people for whom they are designed. Throughout the session, one principle emerged repeatedly. Every sound should communicate with purpose.

    Moore introduced his work by describing the remarkable breadth of modern Audio UX. Working from California’s Bay Area, he collaborates with companies developing products across robotics, automotive technology, connected devices, consumer electronics and digital services. Although these industries appear very different, they all share a common challenge. Products increasingly communicate with people through sound, requiring designers to think carefully about what those sounds communicate and how they contribute to the wider identity of a brand. Rather than approaching each project as an isolated collection of sound effects, Moore described building coherent sonic systems that extend across products, marketing, physical environments and user interactions. Individual sounds matter, though they become most effective when they form part of a larger and recognisable design language.

    This broader perspective also explains why strategy sits at the beginning of every project rather than at the end. Before designing a single sound, his team seeks to understand the objectives of the product, the identity of the organisation and the experience that users should ultimately have. Brand workshops, creative discussions and detailed reviews of existing sounds all contribute towards this early stage of development. Competitor analysis also plays an important role. Understanding how other companies sound allows designers to identify opportunities for meaningful differentiation rather than unintentionally reproducing familiar ideas. The objective is not simply to sound different. It is to create a sonic identity that genuinely reflects the values and personality of the organisation. Sound therefore becomes a strategic design material rather than a decorative addition introduced once the product has already been completed.

    One of the most thought-provoking ideas introduced during the presentation concerned what Moore described as connected audio ecosystems. Many organisations continue to commission isolated sounds for individual products or services, yet users increasingly encounter the same company across multiple devices and environments. A person may hear a notification on a smartphone, interact with a smart speaker at home, use an in-car navigation system and later encounter advertising or public installations produced by the same organisation. Rather than allowing each experience to develop independently, Moore argued that they should all share recognisable sonic characteristics. Consistent instrumentation, similar timbral qualities and carefully related musical ideas allow users to recognise a brand without needing to see a logo or screen. Sound therefore becomes another component of brand identity, working alongside visual design to create familiarity and trust.

    Google’s product ecosystem provided one of the clearest illustrations of this philosophy. Moore described how his work began during the development of Google Glass, a product that sought to make an unfamiliar technology feel approachable. Rather than emphasising futuristic electronic sounds, the design drew upon simple acoustic instruments such as piano, chimes and mallet percussion. These familiar timbres helped ground an otherwise unfamiliar experience, making the product feel more human and intuitive. As Google’s product portfolio expanded through devices such as Pixel phones, Google Pay and automotive systems, this underlying sonic character evolved while remaining recognisably connected. Different products naturally demanded different technical solutions and frequency ranges, though the overall identity remained remarkably consistent. Moore argued that brands evolve in much the same way as people do. Their sonic identities should therefore develop over time while retaining a recognisable sense of continuity.

    Perhaps the most unexpected principle discussed during the session concerned silence. Designers often assume that every interaction requires another notification, another confirmation or another layer of feedback. Moore challenged this instinct directly. As products increasingly incorporate sound into everyday life, designers also acquire a responsibility not to make the world unnecessarily louder. Like negative space in graphic design, silence performs an important communicative function. It creates contrast, draws attention to genuinely important events and prevents users from becoming overwhelmed by constant auditory stimulation. Successful Audio UX therefore depends not only upon knowing which sounds should exist, but equally upon recognising which moments deserve silence instead.

    He illustrated this philosophy through the development of Sense, a sleep monitoring device designed to help users understand the quality of their sleep. Conventional alarm clocks often rely upon abrupt, attention-grabbing sounds that force people awake almost instantly. Moore saw an opportunity to rethink that experience entirely. Instead of beginning loudly, the alarms gradually evolved over time, introducing increasing musical complexity, richer timbres and subtle changes in tempo. Lighter sleepers could wake during the earliest stages, while heavier sleepers would gradually encounter a more energetic composition. Even error tones and voice interactions were designed using soft, restrained timbres that preserved the calm atmosphere of the bedroom rather than disrupting it. The project demonstrated that product sounds need not simply communicate efficiently. They can also influence the emotional quality of everyday experiences.

    Moore then introduced one of his central design philosophies: communicative and expressive design. Throughout the presentation, he repeatedly distinguished between creating sounds that merely reinforce a brand and creating sounds that genuinely help people understand what is happening. Branding undoubtedly matters, though communication always takes priority. Every sound should first convey meaning. Only then should it contribute towards a wider sonic identity. This perspective encourages designers to think carefully about urgency, expectation and human perception rather than treating every notification as another opportunity for creative expression. Product sounds exist to guide behaviour as much as they exist to establish identity.

    Tesla provided a particularly revealing case study. Moore described developing different categories of sounds according to the urgency of the information they needed to convey. Low-priority events, such as incoming calls, were designed to emerge gradually using softer timbres and lower levels of perceptual urgency. Medium-priority notifications, including seatbelt reminders, employed greater repetition and brighter timbres, encouraging users to respond without becoming unnecessarily stressful. High-priority warnings, including forward collision alerts, demanded a very different approach. Higher frequency content, more percussive attacks and rapid repetition ensured that these sounds immediately captured attention during situations where rapid action could prevent an accident. Rather than relying upon arbitrary aesthetic decisions, Moore demonstrated how pitch, repetition, harmonic content and timbre can all be manipulated systematically to communicate different levels of urgency. Sound becomes a carefully designed language through which products communicate urgency, intention and behaviour.

    By this stage, a clear philosophy had emerged. Audio UX is not simply concerned with creating pleasant sounds or memorable sonic logos. It asks how products should communicate with the people who use them every day. Strategy, branding, silence, musical structure and perceptual psychology all contribute towards that objective, though none of them represents the ultimate goal. Every design decision serves the relationship between people and technology. Once sound is understood as a form of communication rather than decoration, the challenge shifts from asking what a product should sound like to asking what it should say. That question became even more significant when Moore turned to the rapidly developing world of robotics.

    The second half of Moore’s presentation shifted from broad design principles towards a detailed case study that demonstrated how those ideas are applied in practice. The project centred on Serve, the autonomous delivery robot developed by Postmates. Rather than treating the robot simply as another product requiring notification sounds, Moore used it to explore a far broader question. How should an intelligent machine communicate with people as it moves through shared public spaces? The answer, he suggested, depends upon much more than selecting attractive sounds. It requires understanding personality, context, expectation and human behaviour long before the first sound is ever designed.

    Like every project discussed earlier in the presentation, the design process began with strategy rather than sound. Before any recording or composition took place, the team explored what the robot represented, how people would encounter it and the personality it should express. Several distinct sonic directions were developed around different interpretations of the brand before being refined through successive design reviews and evaluated within the robot itself. Moore emphasised that successful Audio UX develops through continual iteration rather than moments of inspiration. Sounds that appear convincing inside a studio may behave very differently once reproduced by a moving robot navigating busy streets, restaurants and crowded pavements. Testing therefore becomes an integral part of the creative process rather than simply a means of checking technical performance.

    One of the most revealing aspects of the project concerned personality. Popular culture has encouraged audiences to expect robots to communicate through futuristic electronic sounds or highly expressive synthetic voices. Moore deliberately avoided both extremes. The ambition was to create a robot that felt warm, approachable and reassuring without pretending to possess human intelligence or emotional awareness. At the same time, the team resisted the temptation to rely upon recorded speech, recognising that a natural voice would create expectations that the technology could not consistently fulfil. Instead, they searched for a middle ground in which sound suggested character without imitating humanity. This balance between familiarity and honesty reflected one of the most thoughtful ideas running throughout the presentation. Good Audio UX should communicate clearly without misleading users about a product’s capabilities.

    Developing that personality required exploration rather than immediate certainty. Moore described creating several contrasting sonic directions, each expressing a different interpretation of the robot’s identity. Some embraced more mechanical qualities that acknowledged the machine’s physical presence. Others explored vocal-like synthesis capable of suggesting expression without becoming literal speech. A third direction employed simple sine-wave tones that created a calmer, softer and more abstract character. Rather than choosing a favourite instinctively, these alternatives became prototypes through which designers could observe how people responded emotionally to different sonic identities. The final design combined warmth, clarity and subtle expressiveness, producing a robot that felt approachable without becoming theatrical or sentimental. The process illustrated an important principle: successful sound design rarely emerges fully formed. It develops through comparison, evaluation and refinement.

    Attention then shifted from the robot’s overall personality to the design of individual interactions. Different situations demanded different styles of communication. Interactions inside restaurants, where staff loaded deliveries into the robot, prioritised efficiency through short, direct auditory cues that confirmed actions without interrupting the workflow. Encounters with members of the public required a gentler approach. Longer note durations, more relaxed phrasing and softer musical gestures created the impression of patience rather than urgency. In effect, the robot adapted its acoustic behaviour according to the social environment in which it operated, much as people instinctively alter their own behaviour between professional and public settings.

    A particularly memorable example centred on one of the simplest interactions imaginable: saying, “Excuse me.” Rather than relying upon recorded speech, Moore developed a brief auditory gesture that politely attracted attention before allowing the robot to continue its journey. The intention was not to surprise pedestrians or demand an immediate response. Instead, the sound functioned more like a courteous acknowledgement of another person’s presence. This small interaction captured a principle that extended throughout the presentation. Effective communication often depends upon restraint rather than intensity. Products should seek attention only when attention genuinely needs to be given.

    Safety presented a different set of priorities. Warning sounds must communicate immediately and unambiguously, leaving little room for ambiguity or interpretation. Here, Moore returned to ideas introduced earlier in the presentation. Instead of inventing unfamiliar sonic languages, the design frequently drew upon acoustic references that people already understood from everyday experience. Turn indicators, movement cues and other operational sounds retained familiar characteristics while remaining consistent with the robot’s wider sonic identity. This reduced the need for users to learn entirely new sonic conventions. Familiar sounds could be interpreted almost instinctively, allowing people to respond appropriately without consciously analysing what they had heard.

    Perhaps the most technically demanding challenge involved the robot’s continuous movement through public space. Moore explored several possible solutions, including humming, whistling and slowly evolving tonal textures that he described as “glowing.” Each communicated the robot’s presence in a slightly different way. Some attracted attention more effectively, while others blended more comfortably into the surrounding soundscape. Extensive user testing, including sessions involving blind participants, revealed that restrained harmonic complexity and carefully controlled modulation proved more effective than more elaborate alternatives. Yet Moore resisted the temptation to increase the amount of sound simply to improve awareness. His longer-term ambition was quite the opposite. Intelligent products should become quieter rather than louder. If a robot recognises that nobody is nearby, there may be no need for it to produce sound at all.

    This idea provides a fitting conclusion to Moore’s broader philosophy of Audio UX. The discipline is not concerned with filling products with attractive sounds or memorable sonic logos. It asks how technology can communicate clearly, respectfully and appropriately with the people who use it. Whether designing for autonomous robots, electric vehicles, smartphones or medical devices, the same principles continue to apply. Strategy comes before implementation. Communication matters more than novelty. Personality should remain authentic. Silence deserves to be designed as carefully as sound itself. When those ideas come together successfully, sound ceases to be decoration and becomes an essential part of the conversation between people and technology.

  • How Does Sound Affect Us? Julian Treasure on Listening, Wellbeing, and Designing with Our Ears

    Julian Treasure

    How does sound affect us?

    Most people think about sound only when it becomes a problem. We notice the neighbour’s loud music, the traffic outside a bedroom window, the distracting conversation in an open-plan office or the shrill alarm that interrupts an otherwise quiet day. Far less attention is paid to the countless sounds that quietly shape our emotions, influence our behaviour and affect our health from one moment to the next. During his online guest lecture for Edinburgh Napier University, Julian Treasure argued that this oversight represents one of the greatest shortcomings of modern design. Buildings, products and public spaces are often designed primarily for the eye, while the ear receives remarkably little attention. Yet sound continually influences the way people think, work, communicate and feel. Throughout the presentation, one message emerged repeatedly. If we wish to design better experiences, we must learn to design with our ears.

    Treasure began by asking what sound actually affects. The answer, he suggested, is surprisingly simple. Sound influences our happiness, our effectiveness and our wellbeing, along with those of everybody who shares the environments we create. This observation immediately shifts the discussion away from traditional concerns about noise control or acoustic specifications. Sound becomes a human issue rather than merely a technical one. The quality of an acoustic environment influences far more than whether a room sounds pleasant. It affects how effectively people communicate, how comfortably they work, how safely they respond to hazards and how they experience the spaces in which they spend their lives. Sound therefore deserves to be regarded as one of the fundamental materials of design rather than an afterthought considered once construction has already been completed.

    To explain why sound exerts such profound influence, Treasure described four principal ways in which it affects human beings. The first is physiological. Unlike vision, which depends upon the direction in which we happen to be looking, hearing continuously monitors the environment around us. Human beings cannot close their ears in the way they close their eyes. Throughout evolution, this has made hearing our primary warning system, continually searching for signs of danger beyond the limits of our vision. As a consequence, sound reaches deeply into the nervous system with remarkable immediacy. Sudden or unpleasant sounds trigger hormonal responses associated with stress and vigilance, while calmer acoustic environments encourage relaxation. Treasure illustrated this contrast using familiar examples. An unexpected loud noise immediately increases physiological arousal, whereas gentle natural sounds such as breaking waves often slow breathing and encourage a sense of calm. These responses are not matters of personal preference alone. Sound influences heart rate, hormone secretion, breathing patterns and even patterns of brain activity, quietly shaping the body’s internal rhythms throughout the day.

    The discussion then moved beyond physiology towards psychology. Music provides perhaps the most familiar illustration of this relationship. People instinctively choose particular music to celebrate, to concentrate, to relax or to reflect, recognising that different sounds evoke different emotional states. Treasure argued that natural sounds often produce similarly powerful responses. Birdsong, for example, tends to create feelings of safety and reassurance. Rather than being arbitrary preferences, these reactions may reflect deep evolutionary associations developed over thousands of generations. Birds sing when environmental conditions are relatively safe, allowing those sounds to become unconsciously associated with security. Although listeners rarely analyse these processes consciously, they nevertheless influence emotional experience in subtle yet persistent ways. For sound designers, this observation carries important implications. Designing an acoustic environment involves much more than controlling sound levels. It also requires understanding the emotional associations that different sounds naturally evoke.

    Treasure’s argument became even more relevant to contemporary workplaces when he turned to the cognitive effects of sound. Human attention is a limited resource. People often imagine that they can listen to several conversations simultaneously, though the reality proves rather different. Treasure observed that the human brain possesses only a limited capacity for processing speech, making it extremely difficult to concentrate when nearby conversations compete for attention. Open-plan offices provide a familiar example. Designers frequently value openness, flexibility and visual communication, yet the resulting soundscape often undermines the very productivity these environments seek to encourage. Relevant speech continually draws attention away from the task at hand, interrupting concentration and increasing mental effort. Research cited during the presentation suggests that productivity can fall dramatically under these conditions, illustrating that acoustic design contributes directly to cognitive performance rather than simply influencing comfort. Decisions about the sonic character of workplaces therefore become decisions about how effectively people can think.

    By this stage, a broader pattern had already become clear. Sound is not simply something that accompanies our activities. It shapes them. Physiological responses, emotional reactions and cognitive performance all depend, to varying degrees, upon the acoustic environments within which people live and work. This perspective challenges a long-standing tendency to regard sound as secondary to visual design. Treasure instead presented listening as a central consideration for architects, designers, engineers and sound professionals alike. Before deciding how a space should look, he suggested, we should also ask how it will sound, and how those sounds will influence the people who experience them every day. That question would remain at the heart of the remainder of the discussion.

    Having established that sound influences our physiology, psychology and cognition, Treasure turned to its effect upon behaviour. This influence often operates below the level of conscious awareness, making it particularly easy to overlook. Most people assume they make decisions independently of their acoustic surroundings, yet evidence suggests otherwise. Unpleasant environments encourage people to leave sooner, while attractive soundscapes invite them to remain longer. Treasure illustrated this with a striking study of consumer behaviour. In a supermarket displaying French and German wines with identical visual presentation, researchers changed nothing except the background music. On days when French music was played, French wine substantially outsold German wine. When German music replaced it, purchasing patterns reversed. Customers generally remained unaware that the music had influenced their choices, demonstrating that sound can shape behaviour without requiring conscious attention. For designers, retailers and architects alike, this example reinforced an important point. The acoustic environment is never simply a backdrop. It actively participates in shaping human decisions.

    The implications extend far beyond retail spaces. Treasure argued that every designed environment communicates through sound, whether intentionally or otherwise. A restaurant may create an atmosphere that encourages relaxed conversation, while another overwhelms diners with reverberation and competing voices. A hospital waiting room may reduce anxiety through carefully considered acoustics, or increase it through intrusive alarms and mechanical noise. An office may support concentration, or continually undermine it through poorly managed speech privacy. In each case, the acoustic environment becomes part of the overall design, influencing how people behave within the space. Designers therefore make decisions about human experience whenever they make decisions about sound, even if those decisions consist of ignoring it altogether.

    Treasure observed that this neglect reflects a broader imbalance within contemporary design practice. Buildings are routinely judged according to their appearance, products are evaluated through their visual form, and digital technologies devote enormous attention to graphical interfaces. Comparatively little thought is often given to how these same environments sound. This imbalance is surprising when one considers that hearing operates continuously. We can choose where to look, though we cannot simply decide to stop hearing the world around us. Sound therefore accompanies every activity, continually influencing perception in ways that visual design alone cannot achieve. Rather than treating acoustics as a specialist concern addressed late in a project, Treasure encouraged students to recognise listening as a fundamental design consideration from the very beginning.

    This perspective resonates strongly with professional sound design. Whether creating a film soundtrack, designing interface sounds, producing a virtual instrument or developing an interactive game, practitioners rarely add sound simply to occupy silence. Every sound communicates information, guides attention or influences emotional response. Treasure’s presentation broadened this principle beyond media production into everyday life. The same questions that sound designers ask while constructing a soundtrack also apply to architecture, product design and urban planning. What should the listener notice? Which sounds deserve emphasis? Which should remain unobtrusive? How can sound support rather than distract from the intended experience? The boundaries between sound design and environmental design begin to blur once listening itself becomes the central concern.

    Perhaps the most compelling aspect of Treasure’s argument lay in its optimism. If sound can undermine wellbeing, productivity and behaviour, it can equally improve them. Pleasant acoustic environments encourage relaxation, reduce physiological stress and support clearer thinking. Appropriate sound can strengthen communication, promote social interaction and make public spaces more welcoming. Rather than presenting acoustics as a matter of reducing unwanted noise, Treasure reframed the discussion in positive terms. The objective is not simply to remove bad sound, but to create environments in which good sound actively contributes to human wellbeing. This shift in perspective encourages designers to think creatively about the role sound can play rather than treating it solely as a problem to be controlled.

    These ideas naturally led towards a broader discussion of listening itself. If sound exerts such profound influence over human experience, then the ability to listen carefully becomes an essential professional skill rather than an incidental personal habit. Treasure suggested that hearing and listening are not the same activity. Hearing occurs automatically, while listening demands conscious attention, intention and practice. In an increasingly noisy world filled with competing sources of information, the ability to listen thoughtfully may be becoming more valuable rather than less. This distinction between passive hearing and active listening would ultimately form the foundation of his concluding message, not only for sound designers but for anyone responsible for creating environments in which other people live, work and communicate.

    Having demonstrated that sound influences physiology, emotion, cognition and behaviour, Treasure turned towards a more practical question. If sound has such profound effects upon human experience, what should designers actually do differently? His answer was strikingly optimistic. Rather than treating acoustics as a problem to be solved, he encouraged students to think of sound as a resource that can be shaped deliberately to improve people’s lives. Well-designed soundscapes do more than reduce unwanted noise. They encourage particular patterns of behaviour, support communication and create environments in which people feel healthier, calmer and more engaged. Designing with sound therefore becomes an act of positive intervention rather than damage limitation.

    Treasure illustrated this philosophy through a series of real-world projects. Airports, shopping centres and public spaces all benefited from carefully designed soundscapes that considered not only what people heard, but how those sounds influenced the way they behaved. Introducing natural sounds and thoughtfully composed musical environments increased customer satisfaction, encouraged visitors to remain longer and, in several cases, improved commercial performance. In one public space, the introduction of a biophilic soundscape was even associated with a measurable reduction in crime. These examples reinforced a central point running throughout the presentation. Sound does not merely accompany human activity. It shapes it. Decisions about the acoustic environment therefore become decisions about wellbeing, behaviour and social experience rather than simply matters of technical acoustics.

    Although these examples came from architecture and environmental design, their relevance extends directly to professional sound design. Every soundtrack contains foreground and background elements competing for the listener’s attention. Treasure encouraged students to think carefully about the role each sound should play within that wider acoustic picture. Not every sound deserves prominence, and not every moment benefits from additional music or greater complexity. Like a visual composition, an effective soundscape depends upon balance, hierarchy and clarity. He also encouraged designers to draw inspiration from natural environments, particularly through the thoughtful use of biophilic sound and adaptive or generative soundscapes that evolve over time rather than repeating mechanically. Different spaces support different activities, and their sonic character should reflect those differing purposes. Designing for concentration requires different acoustic decisions from designing for relaxation, learning or social interaction.

    The discussion naturally led back to listening itself. Treasure argued that hearing should never be confused with listening. Hearing is automatic. Listening is intentional. It requires attention, effort and continual practice. In an age characterised by constant distraction and increasingly complex acoustic environments, the ability to listen carefully becomes one of the most valuable professional skills a sound designer can develop. Technical expertise undoubtedly remains important, though it cannot substitute for careful listening. The most sophisticated recording equipment or software offers little value if the designer fails to recognise what listeners actually experience. Listening therefore becomes both a creative skill and an ethical responsibility. Before changing the sound of the world, designers must first learn to hear it properly.

    Treasure concluded by describing what he called the four foundations of effective listening: being conscious, committed, compassionate and curious. Conscious listening requires recognising that listening is an active process rather than a passive consequence of hearing. Commitment acknowledges that good listening demands time, attention and intention. Compassion encourages genuine understanding of other people through careful listening, particularly when viewpoints differ from our own. Curiosity reminds us that every sound and every conversation offers an opportunity to learn something new. Although these principles were presented in the context of listening, they also describe many of the qualities that distinguish thoughtful sound designers. Successful practitioners remain attentive, purposeful, empathetic and continually curious about how people experience the acoustic world around them.

    Treasure’s final appeal brought together everything that had preceded it. He encouraged students to become champions of listening and, above all, to “design with your ears.” This simple phrase encapsulated the wider philosophy running throughout the presentation. Sound should never be regarded as an afterthought added once visual design has been completed. It is one of the primary ways in which people experience the world. Every building, product, public space and interactive system possesses an acoustic identity that influences those who encounter it. Whether designing a film soundtrack, a hospital, a mobile application or a railway station, the same principle applies. The sounds we create shape the lives of the people who hear them.

    Taken together, Treasure’s presentation offered a compelling vision of contemporary sound design. It challenged the traditional tendency to regard sound as secondary to vision and instead positioned listening at the centre of human experience. Physiological responses, emotional wellbeing, cognitive performance, behaviour and communication all depend, to varying degrees, upon the acoustic environments we inhabit. For sound designers, this represents both an opportunity and a responsibility. Every decision about sound has consequences extending beyond aesthetics alone. Designing well therefore means more than creating compelling audio. It means understanding how people listen, recognising how profoundly sound affects everyday life and applying that knowledge to create environments in which individuals and communities can genuinely flourish.

  • How Do You Design a Virtual Instrument? Alejandro Cabrera on Sampling, Sound Design, and Building Kontakt Libraries

    Alejandro Cabrera

    How do you design a virtual instrument?

    Every virtual instrument begins long before the first note is recorded. Musicians often experience sample libraries as polished products that load instantly inside a digital audio workstation, responding naturally to every performance. Hidden behind that apparent simplicity lies an extraordinary amount of planning, recording, editing and technical development. During an online guest lecture for Edinburgh Napier University, Sound Design alumnus Alejandro Cabrera drew upon his professional experience developing sample libraries at 8Dio to reveal how professional virtual instruments are created. Although he used Kontakt to illustrate many of the techniques, his wider message extended well beyond any individual software platform. Successful sound design depends as much upon preparation, organisation and critical listening as it does upon recording itself.

    Cabrera began by challenging a common misconception. Building a virtual instrument is not simply a matter of recording every note and loading the resulting files into a sampler. Instead, it is a carefully structured process comprising pre-production, recording, editing and software development, with each stage influencing everything that follows. Recording sessions may occupy only a small proportion of the overall project, yet their success depends almost entirely upon the decisions made beforehand. Choosing the instrument, selecting an appropriate recording space, determining microphone configurations, deciding which articulations should be captured and calculating the number of samples required all take place before the recording engineer presses record. By the time the first note is performed, many of the most significant creative decisions have already been made.

    Planning emerged as one of the defining themes of Cabrera’s presentation. Recording studios are expensive environments in which every unnecessary decision consumes valuable time. Arriving without a detailed recording plan risks producing inconsistent material, overlooking essential articulations or capturing far more audio than the finished instrument will ever require. To avoid these problems, Cabrera demonstrated the production sheets used to calculate precisely how many samples each instrument will need. The combination of notes, microphone positions, dynamic layers, articulations and recorded variations quickly expands into thousands of individual files. Even a comparatively modest instrument can generate an unexpectedly large collection of audio once every variation has been considered. Careful preparation therefore becomes far more than administrative organisation. It provides the framework upon which the entire virtual instrument will later be constructed.

    This emphasis upon preparation reflects a broader principle that extends well beyond sample library development. Whether recording Foley, ambience, dialogue or musical instruments, professional sound designers rarely begin by placing microphones in front of a source and hoping for the best. They begin by asking what the finished project needs to achieve. Every technical decision should support that objective. Microphone placement depends upon the character of the instrument, the intended listening experience and the amount of flexibility required during production. Recording an intimate acoustic instrument demands different decisions from sampling a full drum kit with multiple microphone positions, while noisy environments require different strategies from carefully controlled studio spaces. Cabrera encouraged students to think of recording not as an isolated technical exercise, but as one stage within a much larger design process in which every decision influences those that follow.

    One particularly revealing discussion centred upon how rapidly complexity increases once realism becomes the goal. Professional sample libraries rarely rely upon a single recording of each note. Different playing dynamics, alternative articulations, multiple microphone positions and repeated performances all contribute towards creating an instrument that responds naturally to the performer. Cabrera introduced concepts such as velocity layers and round robins, not simply as software features, but as perceptual design decisions. Human listeners detect repeated sounds remarkably quickly. Replaying exactly the same recording whenever a note is triggered produces an artificial, mechanical quality that immediately reveals the illusion. Recording carefully controlled variations allows the instrument to remain convincing even during repeated passages, illustrating that realism often depends less upon producing more sound than upon introducing meaningful variation. The objective is not to simulate every possible performance. It is to create enough believable variation that musicians stop thinking about the technology and simply play.

    By this point, a recurring theme had become unmistakable. Building a convincing virtual instrument is not primarily a software problem. It is a sound design problem. The quality of the finished library depends upon understanding the instrument, anticipating how musicians will perform with it and making thoughtful decisions long before the first recording session begins. Technology undoubtedly provides the tools, though preparation, organisation and critical listening determine how successfully those tools can ultimately be used.

    Once the recordings have been completed, the project enters what is often the longest and least visible stage of development. Thousands of individual recordings must be reviewed, edited and organised before they can become a playable instrument. Cabrera emphasised that this work extends far beyond removing unwanted noise or trimming the beginnings and endings of files. Every sample must behave consistently alongside every other sample, allowing the finished instrument to respond naturally regardless of how it is played. Editing therefore becomes a continuation of the design process rather than a separate technical activity. Decisions made at this stage shape the responsiveness of the instrument every bit as much as the recordings themselves.

    Organisation proved equally important. A professional sample library may contain many thousands of individual audio files representing different notes, articulations, dynamic levels, microphone positions and performance variations. Without a rigorous naming convention and carefully structured file management, even relatively modest projects quickly become difficult to maintain. Cabrera demonstrated how systematic organisation supports every subsequent stage of development. Samples can be located immediately, revisions become easier to implement and future updates remain manageable long after the original recording sessions have finished. Good organisation rarely attracts attention, yet it underpins almost every successful production workflow.

    The discussion then turned to Kontakt, the software platform used to assemble these recordings into fully playable virtual instruments. Rather than presenting Kontakt as a collection of technical features, Cabrera used it to demonstrate a broader principle. Software should serve the behaviour of the instrument rather than dictate it. Every mapping decision, performance control and scripting choice exists to make the instrument respond in ways that feel intuitive to the musician. The objective is not simply to trigger recordings accurately, but to create the impression that a real instrument is responding naturally to performance. Technology becomes valuable only when it disappears behind the experience of playing.

    This philosophy also shaped Cabrera’s discussion of scripting. Many musicians never see the programming that sits beneath the graphical interface, yet these invisible systems determine how the instrument behaves. Scripts decide which recordings should be triggered, how different articulations are selected, how repeated notes vary over time and how controls respond to the performer. Much of the intelligence within a modern virtual instrument therefore lies not in the recordings themselves, but in the logic that governs their behaviour. Sound design, software engineering and user experience become closely interconnected, each contributing towards the illusion that the performer is interacting with a coherent musical instrument rather than a collection of audio files.

    Throughout the discussion, Cabrera consistently resisted the temptation to equate realism with complexity. Recording more samples, adding more controls or increasing the number of available options does not automatically produce a better instrument. Every additional recording increases editing time, complicates organisation and places greater demands upon storage, processing power and the musician using the library. The more important question concerns value rather than volume. Which additional recordings genuinely improve the playing experience, and which merely increase complexity without offering meaningful benefit? Successful virtual instruments emerge through thoughtful selection rather than unlimited accumulation.

    These decisions reflect a much broader principle within sound design. Whether recording dialogue, creating Foley, designing interactive game audio or developing sample libraries, practitioners continually shape the listener’s experience by deciding which details deserve attention and which can remain implicit. Technology undoubtedly expands the range of available possibilities, though it rarely removes the need for editorial judgement. Every successful project depends upon identifying the information that listeners or performers genuinely need, then presenting it clearly without unnecessary complication. The objective is not technical excess, but meaningful communication.

    The discussion also highlighted the collaborative nature of professional practice. Developing a virtual instrument combines disciplines that are often treated separately within education and industry. Recording engineers, musicians, software developers, editors, interface designers and producers each contribute different forms of expertise, yet the finished instrument succeeds only when those contributions work together coherently. Cabrera’s examples demonstrated that professional sound design rarely develops in isolation. The most effective solutions emerge when technical and creative perspectives continually inform one another throughout the production process rather than being treated as independent stages.

    Taken together, these discussions revealed that virtual instruments represent far more than collections of recorded sounds. They are carefully designed systems that combine acoustics, performance, recording, editing and software into a single expressive tool. Every decision, from the earliest planning documents to the final user interface, contributes towards the illusion that a performer is interacting with a living instrument rather than triggering digital recordings. For sound designers, perhaps that is the most enduring lesson. The success of a design is rarely determined by the sophistication of its technology alone. It depends upon how completely the technology disappears, allowing creativity, expression and musical performance to take centre stage.

  • How Does a Crowd Find Its Voice? David Monteath on Crowd ADR, Performance, and Creating Believable Worlds

    David Monteath

    How does a crowd find its voice?

    When audiences watch a film or television programme, their attention naturally settles upon the principal actors. Far less notice is taken of the countless background voices that transform a collection of images into a believable social world. Conversations drifting through a restaurant, murmured discussions in an office, distant arguments in a crowded street or the indistinct atmosphere of a busy marketplace all contribute to the impression that life continues beyond the central characters. Remove those voices, and even the most carefully photographed scene can feel strangely artificial. During his online guest lecture for Edinburgh Napier University, David Monteath returned to the University as a Sound Design alumnus to explore the specialised craft of crowd ADR. Drawing upon more than three decades working as an actor and voice artist, he demonstrated that believable crowd performances depend upon observation, improvisation and an understanding of dramatic context rather than simply recording large numbers of voices. One principle underpinned the discussion. Context is king.

    Rather than replacing the dialogue of principal actors, crowd ADR creates the sense that an entire world exists beyond them. A small group of performers may become the customers in a restaurant, the spectators at a football match, the passengers waiting on a railway platform or the crowd gathered in a courtroom. Individual conversations overlap, reactions ripple through the group and emotional responses emerge at precisely the right moments, creating the impression that every person visible on screen possesses a life extending beyond the immediate story. Audiences rarely notice these performances consciously, yet they immediately recognise when they are missing. Scenes that lack convincing crowd performances often feel unexpectedly empty, regardless of how carefully they have been photographed or edited.

    Monteath repeatedly challenged the assumption that this work consists simply of creating background noise. Crowd ADR is first and foremost a form of acting. Every performance responds to the circumstances of the scene, the relationships between characters and the emotional atmosphere established by the director. People waiting quietly in a hospital corridor behave differently from supporters leaving a football stadium. Conversations in an expensive restaurant differ from those heard in a busy café, while voices surrounding a royal procession carry a very different energy from those accompanying a political protest. Every reaction, interruption and fragment of conversation exists to support the dramatic reality of the scene rather than to attract attention in its own right. Authenticity emerges from understanding how people genuinely behave in different situations, not from making scenes louder or busier.

    This emphasis upon dramatic context shaped every practical discussion throughout the lecture. Monteath encouraged students to think beyond individual words and instead consider the circumstances in which those words are spoken. Before deciding how loudly to speak, how quickly to react or even what might be said, performers first need to understand where they are, who surrounds them and what is happening within the story. The same phrase may require entirely different delivery depending upon whether it takes place in a library, an airport, a football ground or the middle of a battlefield. Successful crowd performers therefore begin by observing people. Everyday behaviour, casual conversations, shared laughter, hesitation, disagreement and excitement all provide material that can later be adapted naturally within the recording studio. The objective is not to invent behaviour, but to recognise and recreate it convincingly.

    Perhaps the most revealing insight from this opening part of the lecture concerned the relationship between realism and audibility. Many beginning sound designers instinctively assume that important sounds should always be heard clearly. Monteath argued for almost the opposite approach. Successful crowd ADR often succeeds precisely when audiences remain largely unaware of it. Background voices should usually be felt rather than heard, contributing movement, texture and emotional energy without competing with the principal dialogue. Monteath returned repeatedly to the idea that audiences should sense the presence of a living world long before they consciously identify individual voices. Crowd ADR achieves its greatest success not when listeners admire the performance, but when they accept the world on screen without ever questioning how it came to life.

    One of the most valuable themes running through the lecture concerned the difference between sounding natural and sounding believable. These ideas are not always identical. Performers working in crowd ADR rarely speak at the same level they would use in everyday conversation, yet exaggeration can become equally unconvincing. Monteath described the continual process of judging how voices should sit within the perspective of the scene. A performer passing close to the camera requires a different vocal presence from someone crossing the background several metres away, while conversations taking place outdoors demand a different energy from those occurring in confined interior spaces. Every decision depends upon dramatic perspective rather than fixed performance rules. Context, once again, determines everything. For sound designers, these distinctions become equally important during editing and mixing. A crowd recording that sounds entirely convincing in isolation may feel unexpectedly prominent once placed alongside production dialogue, Foley and ambience. Perspective therefore emerges through the relationship between every element of the soundtrack rather than through any individual recording considered on its own.

    This attention to perspective extends beyond volume alone. Monteath discussed the subtle adjustments people make instinctively when speaking in different environments. Outside, voices naturally rise in level before settling into an appropriate projection as people unconsciously judge the surrounding space. He compared this process to a form of echolocation. Speakers continually test their surroundings, modifying projection almost instantly until their voice feels appropriate for the environment. Recording inside a studio removes many of the environmental cues that normally guide these unconscious adjustments, requiring performers to recreate them deliberately. The challenge is not simply to speak more loudly for an exterior scene, but to reproduce the natural behaviour that accompanies speaking outdoors. Audiences rarely analyse these details consciously, though they recognise immediately when they feel unconvincing. Successful crowd ADR therefore depends upon recreating patterns of human behaviour rather than merely increasing vocal intensity.

    The physical demands of crowd ADR also proved far greater than many students had expected. Scenes involving panic, conflict or large-scale action often require sustained shouting over many hours, placing considerable strain on performers’ voices. Monteath reflected upon sessions in which actors had pushed themselves to the point of temporary vocal exhaustion, particularly when recording intense battle scenes. Curiously, he observed that shouting repeatedly inside a recording studio often proves more tiring than raising the voice naturally outdoors. In everyday life, people instinctively project according to their surroundings. Within the artificial environment of a studio, performers can find themselves holding unnecessary tension in the throat in ways that feel surprisingly unnatural. Maintaining vocal health therefore becomes an important professional skill alongside acting itself. It also reflects another aspect of professional sound design that audiences rarely consider. Recordings capable of conveying fear, excitement or urgency often depend upon performers sustaining physically demanding work throughout lengthy recording sessions while preserving consistency from one take to the next.

    The discussion of large battle sequences illustrated another revealing aspect of the profession. Crowd performers may spend an entire day creating layers of screams, reactions and movement for scenes involving hundreds or even thousands of people, fully aware that much of their work will eventually disappear beneath music, sound effects and the principal action. Monteath recalled recording material for a major battle sequence in Game of Thrones, where hours of physically demanding vocal performances ultimately became almost imperceptible within the finished soundtrack. Rather than expressing disappointment, he presented this as an inevitable consequence of professional sound design. The objective was never for individual performances to stand out. Their purpose was to contribute energy, scale and credibility to the scene, even if audiences remained almost entirely unaware of their presence. The irony is that some of the hardest work in post-production often becomes the least conspicuous in the finished mix.

    Monteath’s recurring phrase, “Context is king,” captures this philosophy particularly well. Every vocal decision derives from the dramatic situation rather than from the performer. Voices rise or fall according to the surrounding environment, emotional reactions emerge in response to the unfolding action and every fragment of conversation exists to reinforce the illusion that life extends beyond the principal characters. Successful crowd ADR is therefore measured not by how clearly individual voices are heard, but by how convincingly they allow audiences to believe in the world unfolding around them. Like many aspects of professional sound design, its greatest achievement lies in remaining almost invisible while making the fictional world feel entirely real.

    The lecture concluded with a discussion that moved beyond recording techniques and towards the broader decisions that shape professional sound design. One student described the challenge of creating the atmosphere for a bank robbery scene. Adding more and more voices had seemed the obvious solution, yet the result quickly became cluttered and distracted from the drama. Monteath’s response illustrated once again why crowd ADR depends upon judgement rather than quantity. Real crowds rarely behave as a single, unified group. Even in moments of fear, surprise or excitement, different people react at different times and in different ways. Some remain silent, others whisper, a few call out, while many simply watch events unfold. Attempting to represent every visible person with an equally prominent vocal performance often produces a soundtrack that feels less realistic rather than more so. Believability emerges through carefully judged variation, allowing individual reactions to appear and disappear naturally instead of competing continuously for the listener’s attention.

    This observation extends well beyond crowd ADR. Throughout post-production, sound designers continually decide what deserves the audience’s attention and what should remain part of the wider acoustic environment. A convincing soundtrack is not created through the accumulation of detail, but through the careful organisation of that detail into a coherent dramatic experience. Crowd performances occupy a role similar to ambience, Foley and environmental sound. They establish context, scale and emotional texture without constantly demanding attention. Their purpose is not to demonstrate how much work has been carried out, but to convince audiences that the world extending beyond the principal characters already exists. Like every other element of a soundtrack, their success depends upon supporting the story rather than competing with it.

    Towards the end of the lecture, discussion turned to the growing influence of artificial intelligence within the voice industry. Monteath acknowledged that AI is already beginning to affect areas such as commercial voice-over, where some clients have started experimenting with synthetic voices. He regarded crowd ADR rather differently. While aspects of the work may eventually become automated, authentic crowd performance depends upon subtle variations that emerge naturally whenever people work together. Voices change over the course of a recording session as performers become tired. Emotional intensity shifts between takes. Individual personalities influence rhythm, timing and vocal colour in ways that are difficult to predict or reproduce consistently. These variations might appear inconvenient from a purely technical perspective, yet they contribute directly to the richness, unpredictability and authenticity that audiences instinctively recognise as human. Technology will continue to evolve, though observation, collaboration and performance remain at the heart of believable sound design.

    For sound design students, perhaps the most valuable lesson lay in the way Monteath described his profession. Crowd ADR may appear to occupy the margins of post-production, hidden beneath dialogue, music and sound effects, yet it influences how audiences perceive almost every scene they watch. Every murmur in the background of a restaurant, every distant conversation in a station concourse and every carefully judged reaction during a moment of crisis contributes to the illusion that life continues beyond the frame. These performances do not simply fill silence. They create social spaces that feel inhabited, allowing viewers to concentrate on the story without questioning the reality of the world surrounding it.

    Throughout the lecture, Monteath returned repeatedly to one deceptively simple principle: “Context is king.” Crowd ADR succeeds not through memorable performances or individually recognisable voices, but through creating the impression that every environment extends beyond the limits of the frame. Every carefully judged laugh, argument, whispered conversation and fleeting reaction reinforces a believable social world without distracting from the principal narrative. For sound designers, this represents a broader lesson that reaches far beyond dialogue replacement. Successful audio is rarely measured by how noticeable it becomes. More often, it is measured by how completely it allows audiences to believe in the world they are experiencing. Crowd ADR exemplifies that philosophy. It remains one of the least visible aspects of professional sound design, yet it is also one of the crafts that most quietly transforms moving images into convincing places inhabited by believable people.