Author: iainmcgregor

  • How Do You Preserve the Sound of a Place? Damian Murphy on OpenAir, Acoustic Heritage, and Reconstructing Lost Spaces

    Damian Murphy

    How do you preserve the sound of a place?

    A building can survive through photographs, architectural drawings, maps and written descriptions. Its dimensions can be measured, materials catalogued and appearance reconstructed long after the original structure has disappeared. Yet places are not experienced through vision alone. A cathedral changes the sound of a choir, a tiled chamber changes the sound of a voice, snow alters the acoustic behaviour of a forest, and the hard surfaces of a mausoleum can allow sound to continue long after its source has stopped. Architecture surrounds every action with reflections, resonances and reverberation, but this part of a place can disappear without leaving anything visible behind.

    During his online guest lecture for Edinburgh Napier University, Professor Damian Murphy of the University of York explored almost two decades of work investigating how the acoustics of places can be measured, preserved, reconstructed and experienced. Much of this work centres upon OpenAir, the Open Acoustic Impulse Response Library, which contains acoustic measurements gathered from buildings, landscapes and other environments. Its contents range from cathedrals, churches and theatres to industrial buildings, caves, forests and vehicles, connecting acoustic science with music production, spatial audio, games, archaeology, heritage and historical research. Murphy moved between these different places and applications through a question that became larger as the examples accumulated. What can the sound of a place tell us that its image cannot?

    Murphy began with a ruin. The Temple of Decision stands on a hill overlooking the Falkland Estate in Fife, where artists David Chapman and Louise K. Wilson had been commissioned to explore the landscape through sound. Archive material could offer clues about the temple’s former appearance, while historical research could provide fragments of information about its use, but the surviving structure could no longer reveal how the intact building had sounded. What would it have been like to speak inside the room? How might voices have behaved around a table? How would conversation, movement or a fire have interacted with its surfaces? The artists posed a question that provided a starting point for Murphy’s lecture: in the absence of clear echoes, how can we know a place?

    Answering that question first requires an understanding of what an acoustic environment contributes to anything heard within it. Imagine a short, sharp sound produced inside a room. A listener initially receives the direct sound travelling along the shortest path from source to receiver. Early reflections arrive shortly afterwards from nearby surfaces, followed by increasingly complex patterns as energy travels through the room, interacting repeatedly with walls, floor, ceiling and objects before gradually decaying. Together, these components form a room impulse response for a particular relationship between a source and receiver. Direct sound carries information about the source and its distance, while early reflections contribute to perceptions of geometry and position. Later reverberation communicates qualities associated with volume, materials and enclosure. Change the architecture and the response changes. Move the source or listener and it changes again. An impulse response is therefore not a complete acoustic identity for a building, but a record of how sound travelled between particular positions under particular conditions.

    Once captured, that relationship can be used for more than numerical analysis. Through convolution, a recording made without significant room acoustics can be combined with an impulse response measured elsewhere. Murphy demonstrated the process using a four-part vocal ensemble recorded in an anechoic chamber. The singers had never performed in York Minster, yet convolution with a measured response allowed their dry voices to acquire characteristics of the cathedral. York Minster has a reverberation time of approximately eight seconds through part of the mid-frequency range, compared with around half a second for a typical living room. Voices behave very differently in each environment. Notes overlap, transitions blur and the building continues sounding after the performers have stopped producing sound. The acoustic is not simply decoration placed around a performance. It changes the temporal relationships through which that performance is heard.

    Auralisation, however, introduces a distinction between recreating acoustic conditions and recreating experience. Singers performing in an anechoic chamber do not behave as they would inside a highly reverberant cathedral. Performers hear themselves and adapt. Tempo, articulation, phrasing, dynamics and pauses can change in response to sound returning from the room, while musicians continually adjust to one another through the acoustic environment they share. Convolution can reproduce the effect of a measured response upon a recording, but it cannot retrospectively create the performance that might have developed inside that space. Murphy acknowledged this limitation directly when discussing the York Minster example. The anechoic performance was not the performance the singers would have given in the cathedral, and even the spacing of phrases in the demonstration had been altered to allow the reverberation to emerge more clearly.

    The distinction matters beyond the simulation of reverberation. A room impulse response can describe how energy travels between defined positions, but people are not passive sound sources or microphones. They move, listen selectively, change their behaviour and respond to what they hear. Preserving a response gives researchers evidence about the acoustic conditions of a place. What people did in response to those conditions remains a different question.

    OpenAir developed from an ambition to preserve such evidence and make it available for others to explore. Murphy traced one important influence to Angelo Farina’s work on recording concert halls for posterity. Improvements in measurement techniques and the emergence of practical convolution reverberation created an opportunity to document significant spaces not merely through reverberation times and other summary values, but through impulse responses that could be analysed, reproduced and used creatively. OpenAir extended this principle into a growing archive, with a measurement system designed to collect spatially rich data that could remain useful beyond the immediate research question.

    Early measurement sessions used a Genelec S30D loudspeaker to excite the space while microphones captured its response. A computer-controlled turntable allowed measurements at regular angular intervals, and an ambisonic Soundfield microphone was combined with a Neumann KM100 cardioid microphone to provide spatial information alongside material suitable for different forms of analysis and production. Measurements could be repeated across several source and receiver positions, preserving a set of relationships rather than presenting each building through one supposedly definitive response. Ambisonics was particularly valuable for an archive whose future applications could not be predicted. A first-order ambisonic recording represents a three-dimensional sound field through an omnidirectional component and three directional components, separating the captured information from one fixed loudspeaker arrangement. Material can later be decoded for different reproduction systems or manipulated in ways that may not have been anticipated when the recording was made.

    Flexibility matters when access to a significant site may last only a few hours. Researchers need to gather material rich enough to support questions that have not yet been formulated and technologies that may change long after a measurement session has ended. During the discussion after the lecture, Murphy described more recent work at St Paul’s Cathedral, undertaken with a composer who wanted impulse responses from the building. Three researchers had only three hours to move equipment through the enormous space and capture responses from locations including the nave, a stairwell and the Whispering Gallery. Practical decisions about where to measure become part of preservation itself. As Murphy observed, there is no single sound of a large building. Different positions offer different acoustic experiences, and any archive necessarily records selected relationships within a much larger field of possibilities.

    As OpenAir expanded, its growing range of places made the idea of acoustic preservation less straightforward. York Minster was an obvious candidate, since its long reverberation is closely connected with experiences of worship, tourism and musical performance. Other sites raised different questions about what deserves to be preserved and why. When the former Terry’s chocolate factory in York closed, Murphy and his colleagues gained access before redevelopment and measured spaces including a warehouse and the former typists’ room, a striking interior dominated by glass and wood. The activities for which these spaces had been designed had already disappeared. An empty typists’ room could still be photographed, but its appearance prompted another question: what might it have sounded like when filled with the overlapping mechanical activity of typewriters?

    A subterranean reactor hall beneath the Royal Institute of Technology in Stockholm preserved another relationship between architecture and former activity, while measurements in historic churches allowed acoustic theories to be tested rather than merely repeated. At St Andrew’s Church in Lyddington, the team examined jars embedded within the walls, architectural features sometimes interpreted through theories of resonant vessels extending back to the Roman architect Vitruvius. Measurement provided little evidence that the jars were making a substantial contribution to the acoustic character of the church. Their form did not correspond closely with the behaviour expected of effective Helmholtz resonators. Acoustic research could challenge explanations attached to historic architecture as well as document spaces admired for their sound.

    York Theatre Royal shifted attention from buildings as fixed objects towards places in changing states. Murphy’s team measured the auditorium before refurbishment and returned afterwards to document its altered acoustic. They also captured measurements with an audience present during a pantomime, recognising that an occupied theatre does not behave acoustically like the same room when empty. Seats, clothing and bodies absorb and scatter sound, making occupancy part of the acoustic system rather than simply a group of listeners placed within it. Materials age, spaces are repurposed and environmental conditions change. Preserving an acoustic environment may therefore involve documenting several states of the same place rather than searching for one definitive response.

    Outdoor and semi-outdoor locations created different practical problems, and the original measurement system could not simply be carried everywhere. On the Falkland Estate, the Bottle Dungeon could be reached through a trapdoor, but access was sufficiently awkward that the usual equipment was impractical. A balloon was attached to one stick, a pin to another and a microphone lowered into the space. Bursting the balloon remotely provided the excitation needed to capture a response. The improvised arrangement was far removed from the computer-controlled turntable used elsewhere, yet it addressed the same underlying need: introduce a suitable sound into an inaccessible environment and record how the space transforms it.

    Landscapes demanded further adaptation. Equipment had to become portable and independent of mains electricity as researchers travelled into the Yorkshire Dales to measure caves and gorges. Work in Finland examined the same forest under different seasonal conditions. Researchers used GPS alongside ribbons tied around trees to return to the same positions after deep snow had transformed the landscape. Geographically, it remained the same forest. Acoustically, it had changed. Snow altered the interaction between sound, ground and surrounding environment, demonstrating that acoustic character can change while location remains constant.

    A collaboration with Codemasters carried that thinking into interactive media. Looking at an archive rich in churches and historic interiors, the developer asked a practical question: what about the environments needed for games? The collaboration encouraged further work on landscapes and led to experiments for GRID Autosport involving vehicle interiors, where the acoustic problem was unusually complex. Codemasters wanted to represent a car gradually falling apart during a race, so the team needed measurements capable of describing changing states, including doors opening or disappearing and the boot being open. The experience of being inside a racing car also comes from more than airborne sound. Engine vibration travels mechanically through the structure and contributes to what an occupant hears and feels.

    Conventional room measurement could not fully represent that relationship, so Murphy and his colleagues experimented with using the engine itself as part of the measurement process. Revving the car caused the structure to vibrate as it would during use, after which signal-processing methods were applied to separate the excitation from the resulting cabin response and derive an approximation of the impulse response. The approach sought to preserve something more specific than the reverberation of a small enclosure. It attempted to capture the interaction between a vibrating machine, its structure, the enclosed air and the listener inside it.

    Games also demonstrated how preservation can become creative infrastructure. An impulse response gathered for research may later help construct a virtual environment, become part of a music-production tool or support a question that had not existed when the measurement was made. Murphy described OpenAir material finding its way into software used by musicians and audio practitioners, allowing measurements collected years earlier to acquire new purposes. Distribution under a Creative Commons licence reflected this wider ambition. Researchers can analyse the acoustic behaviour of a building, composers can use the same response creatively, sound designers can place fictional events inside measured environments and developers can incorporate selected material into new tools.

    OpenAir originally allowed members of the wider community to upload their own measurements. As contributions accumulated, variations in recording quality became difficult to ignore. The team eventually reviewed the existing material, retained the strongest contributions and moved towards a more curated model in which new contributors contact the team directly. Open access expands what an archive can become, but reuse also depends upon confidence in how the material was produced.

    These measurements deal with places that still exist, even when they are changing. The Temple of Decision presents a different problem. Its original interior has already gone. Once a room has disappeared, there is nothing left to measure. Its former acoustic behaviour has to be approached indirectly through surviving evidence and modelling.

    Using information about a lost structure, researchers can construct a three-dimensional geometric representation and simulate the propagation of sound within it. Virtual rays travel through the model, reflecting between surfaces to generate impulse responses for selected source and receiver positions. Those responses can then be analysed or used to process voices and other recordings, allowing listeners to hear an interpretation of how sound might have behaved inside architecture that no longer survives. Such a result differs fundamentally from measuring an existing building. Geometry may be uncertain, material properties need to be estimated and every modelling method introduces limitations. Historical auralisation produces an evidence-based acoustic proposition rather than a recording recovered from the past.

    Work on St Mary’s Abbey in York made both the possibilities and limitations of this approach audible. The former church survives as a ruin, while archaeological and architectural evidence allowed Murphy’s team to construct a three-dimensional representation suitable for acoustic modelling. Simulated impulse responses could then be compared with measurements from York Minster, a surviving building with some comparable characteristics. Estimated reverberation times occupied a similar range, but the reconstructed St Mary’s sounded noticeably brighter. The difference did not necessarily reveal a historical distinction between the buildings. Murphy explained that the ray-tracing model used for St Mary’s was less effective at reproducing low-frequency behaviour than the physical measurement system used in York Minster. Part of what listeners heard therefore belonged to the method of reconstruction itself.

    A model may sound convincing while still containing audible consequences of the technique used to create it. Plausibility has to emerge from evidence, comparison and methodological transparency rather than from the apparent realism of the result alone. Once those boundaries are understood, a model can make relationships perceptible in ways that drawings and numerical data cannot. During a public performance among the ruins of St Mary’s Abbey, a live choir was captured and processed through impulse responses derived from the reconstructed church, then projected to an audience gathered at the site. Present-day voices sounded among the physical remains while the model returned an interpretation of the missing acoustic architecture. Research data became part of an experience connecting a surviving place with a vanished interior.

    Reconstructing an abbey or temple can help audiences imagine how a lost building might have shaped music and speech. Murphy’s more recent work moved towards a question with wider historical consequences. If architecture changes what people can hear, could reconstruction help investigate who had access to speech in the past?

    Working with historians and art historians, Murphy’s team investigated spaces associated with the historic House of Commons within the former Palace of Westminster. Among the most revealing parts of that history was the experience of women who listened to parliamentary debate from a roof space above the chamber. Known as the Ventilator, this space allowed women excluded from formal political participation to gather above the House of Commons and listen through the architectural structure separating them from the debate below. Historical evidence could establish that they were there. Acoustic reconstruction allowed another question to be asked: what could they actually understand?

    Answering it required more than recreating the debating chamber as an isolated room. Researchers had to consider the chamber, roof void, Ventilator and routes through which speech travelled. The women listening above did not have a direct line of sight to the speaker, making their experience a problem of acoustic transmission through a complex architectural arrangement rather than ordinary listening within a single enclosure. Murphy connected this project with broader questions of directionality, speech transmission and listening position, while acknowledging that objective measures cannot reproduce every aspect of human attention or historical experience.

    No surviving recording can reveal exactly what those listeners heard. The team therefore combined historical reconstruction with comparative measurement, examining surviving spaces connected with parliamentary history or comparable in geometry, scale, period or use. Measurements from rooms at the University of Oxford, York Guildhall and the present House of Commons chamber provided contexts against which aspects of the model could be considered. None could prove how the lost chamber sounded, but together they allowed reverberation and speech intelligibility to be examined across different positions and occupancy conditions. A question about architectural acoustics had become inseparable from a question about political access.

    Hearing is a form of access. Architecture, distance, reverberation, occupancy and barriers influence whether speech remains intelligible, while attention, familiarity and expectations affect what can be understood from imperfect information. A person may be physically close to political debate while remaining acoustically separated from it. Reconstructing the conditions of listening can therefore contribute to historical questions about participation and exclusion that visual records alone cannot answer.

    The Temple of Decision now appears less like an isolated case and more like the beginning of a much larger enquiry. In the absence of clear echoes, how can we know a place? We can measure what survives, compare different conditions, preserve spaces before they change and build models from evidence when the original architecture has disappeared. We can listen to those models while remaining clear about what they can and cannot establish. Most of all, we can ask questions that become difficult to formulate when architecture is treated only as something seen.

    A photograph can preserve the appearance of a parliamentary chamber from one position. An architectural plan can show where walls, doors, galleries and roof spaces were located. Written testimony can tell us that people gathered somewhere to listen. Acoustic research adds another layer by asking how sound travelled between those positions, how reverberation affected speech and whether architecture enabled or obstructed understanding. The same thinking can ask how a ruined abbey shaped musical performance, how a theatre changed during refurbishment or how winter transformed the acoustic behaviour of a forest.

    Preserving the sound of a place is not simply an attempt to save an attractive reverberation before it disappears. Places participate in human activity. They change how people speak, perform and listen. Their surfaces and geometry influence whether voices remain intimate or become collective, whether musical phrases overlap or remain distinct, and whether somebody standing beyond a barrier can understand words spoken elsewhere. Acoustic conditions can shape behaviour, access and participation without leaving a visible trace.

    The Temple of Decision can still be visited. St Mary’s Abbey remains visible as a ruin. The old House of Commons chamber has disappeared, while the former typists’ room at Terry’s chocolate factory no longer performs its original function. A forest changes when snow arrives and changes again when it melts. Even a building that survives intact contains many acoustic relationships, only some of which can be measured during the limited hours when researchers have access.

    Sound is especially vulnerable to disappearance. Once a room changes, an audience leaves or a building is lost, its former acoustic behaviour cannot simply be photographed. Murphy’s lecture showed that it is not entirely beyond preservation. What remains may be a measurement, a model, a comparison or a carefully documented uncertainty, each offering a different way of understanding a place as more than a visual container for history.

    In the absence of clear echoes, we may never know a place completely. We can, however, preserve evidence of how sound moved through it, reconstruct plausible relationships when the original has disappeared and ask what those acoustics meant for the people who performed, spoke and listened there. An impulse response may last only a few seconds, yet within it can remain an acoustic trace of a cathedral, a factory, a theatre, a cave, a forest under snow or a room about to change forever. When the room itself has already disappeared, a model can return a possibility rather than a certainty: a voice reflecting from lost walls, a choir inhabiting a ruined abbey or political speech travelling towards listeners hidden above a chamber from which they were excluded.

    Preserving the sound of a place means preserving another way of understanding what happened there. Walls determine more than what people can see. They shape what can be heard, how clearly it can be understood and who is able to listen.

  • How Do You Design Sound for an Experience You Cannot Control? Wylie Stateman on Storytelling, Simplicity, and the Future of the Soundtrack

    Wylie Stateman

    How do you design sound for an experience you cannot control?

    A filmmaker can frame an image. The edges of the screen define what the audience sees, while composition, focus, lighting and editing direct attention within it. Sound is less obedient. It extends beyond the frame, surrounds the audience, enters rooms with different acoustics and reaches listeners through systems ranging from enormous cinema installations to headphones, televisions, laptops and tiny mobile-phone speakers. A soundtrack may be created with extraordinary precision, yet nobody involved in its production can completely control how, where or at what level it will finally be heard.

    During his online guest lecture for Edinburgh Napier University, supervising sound editor and sound designer Wylie Stateman explored the creative and professional consequences of working with such an elusive medium. Drawing upon a career spanning more than four decades and collaborations with filmmakers including Quentin Tarantino, Oliver Stone and John Hughes, he described sound as an art form positioned between science and subjectivity. Its physical behaviour can be measured, yet its meaning depends upon perception, attention, expectation and context. His work has extended from feature films and television to advertising, audiobooks and theme-park attractions, but a consistent question connects these different forms: how can sound professionals create an experience that remains clear, emotionally purposeful and dramatically coherent when both production and listening contain so much uncertainty?

    For Stateman, answering that question requires sound practitioners to think beyond individual sounds. A designer may create remarkable material, but the audience experiences relationships: dialogue against music, effects within an environment, silence before an impact and the entire soundtrack through a particular playback system at a particular level. Creating those relationships at scale requires collaboration, while maintaining them requires somebody to understand the complete experience. Throughout the lecture, Stateman moved repeatedly between these levels, from the organisation of large creative teams to the placement of a single sound and from the controlled environment of the mixing stage to the unpredictability of a listener pressing play somewhere else. Connecting them was a consistent philosophy. Complexity is unavoidable behind the scenes, but it should produce clarity for the audience.

    Stateman began with the unusual nature of sound itself. A visual composition can be stopped and examined. People can point towards a particular area of an image, discuss its composition and compare alternatives while the material remains stationary. Sound exists through time. A mixer listens, makes an adjustment, returns to an earlier point and listens again. Creative refinement becomes a repeated movement backwards and forwards through the material, making the number of meaningful passes completed within a working day directly relevant to the sophistication of the result.

    Such a process makes both technology and collaboration important, but neither can replace judgement. Stateman described audio as a discipline in which scientific knowledge and subjective interpretation continually meet. Vibration can be measured objectively, while a listener’s response to it cannot be reduced so easily. Designers therefore work simultaneously with physical systems and human perception. Loudspeakers, auditoria, codecs and playback levels matter, but so do memory, expectation, emotion and attention. Even the most technically controlled production process eventually encounters a listener whose response cannot be engineered with the same certainty as the system delivering the sound.

    Perhaps this uncertainty helps explain why Stateman spoke so strongly about collaboration. Rather than presenting professional success as the achievement of a solitary creative individual, he described a career built through relationships with people whose abilities complemented his own. Soundelux, the company he established with fellow sound designer Lon Bender, grew from a small operation into an organisation employing hundreds of people across several cities. Its development depended upon far more than creative talent. Business management, accounting, engineering, technology, sales and production all had to support the work of designers, editors and mixers.

    The lesson for students was not that everyone should attempt to build a large company. Stateman’s broader point concerned complementary ability. Nobody needs to become equally skilled at every aspect of creative and professional life. Someone with little interest in finance needs a trustworthy person who understands it. A creative specialist working on complex productions benefits from engineers and technologists capable of turning ideas into reliable systems. Partnerships become valuable when they extend what a group can imagine and accomplish rather than merely reproducing the same abilities several times.

    Professional value also emerges through reliability. Stateman described a valuable colleague in strikingly simple terms: somebody who can understand a problem, take responsibility for it and allow everyone else to stop worrying about it. Creative ability matters enormously, but large productions depend upon trust. A person who solves a problem without creating several new ones becomes increasingly valuable to the people around them. Careers are built not only through the quality of isolated work, but through the confidence that others can place in somebody when a difficult problem arrives.

    Filmmaking itself operates through the same interdependence. Every specialist inevitably perceives the project through a particular discipline. Sound designers think about sound, composers about music, cinematographers about light and composition, costume designers about clothing and production designers about the physical world. Stateman compared this to a collection of unreliable narrators, each understanding the film from a particular perspective. The director’s responsibility is to bring those partial perspectives together into a coherent experience.

    Sound professionals therefore need both commitment to their own discipline and awareness of the larger work. Stateman’s long relationships with individual filmmakers allowed such understanding to develop across several projects. By the time of Once Upon a Time in Hollywood, he had worked with Quentin Tarantino on seven films. Repeated collaboration created trust and a shared shorthand, allowing Stateman to take substantial creative ownership of the soundtrack while remaining clear that every decision ultimately served the director’s film.

    That balance between ownership and service runs through much of professional sound design. Creative contribution requires conviction. A designer who merely waits for instructions cannot provide the full value of specialist expertise. Yet conviction must remain connected to the filmmaker’s intention rather than becoming an opportunity to demonstrate technique. Stateman’s account suggested a form of authorship that is confident without becoming possessive. Sound teams need enough ownership to make strong decisions, while recognising that those decisions belong within a larger work whose purpose they must understand.

    Understanding that purpose becomes increasingly important as the available technology grows more powerful. More capability does not automatically justify more activity, and one of Stateman’s clearest demonstrations came from work beyond conventional cinema. Soundelux developed audio and show-control systems for large theme-park attractions, including a Terminator 2 experience designed to move hundreds of visitors through repeated performances every day. Audiences passed through a pre-show before entering a theatre containing three 3D IMAX screens and a motion base. Reliability was essential, but technical reliability alone could not create an engaging experience.

    Despite having access to the motion platform throughout the attraction, the production reserved its major movement for a single moment. Audiences were allowed to become comfortable before the platform suddenly dropped. Their physical reaction was powerful precisely through its rarity. Continual movement would have made the mechanism familiar and reduced its dramatic value. Behind one carefully timed surprise sat engineering, amplification, show control, multiple screens, motion systems and the work of a large team. None of that complexity needed to become the audience’s concern. They experienced the result.

    The same principle shaped Stateman’s approach to film sound. A system may offer hundreds of channels, extensive spatial control and enormous dynamic range, but the designer still needs to decide which possibilities deserve to be used. Creative sophistication can appear through restraint, and technical capability becomes valuable when it helps direct attention rather than continually demanding it.

    When asked how he decides what an audience should attend to, Stateman returned to simplicity. Film combines visual and auditory information, but listeners cannot process every available element with equal attention. A dense soundtrack may contain extraordinary detail while communicating very little. The designer’s task is therefore not to make everything audible at once, but to create a clear path through the scene.

    Dialogue often provides the starting point. Human listeners extract extraordinary amounts of information from voices. Words communicate explicit meaning, while rhythm, pitch, timbre, hesitation and vocal effort reveal character and emotion. If audiences struggle to understand what somebody is saying, an essential layer of narrative and performance has been weakened. Music offers another route through the experience, leading or following emotional movement, preparing audiences for change or allowing feeling to emerge after an event. Environments and sound effects establish space, physicality, scale, tension and perspective. None of these categories possesses permanent priority. What matters is understanding what the audience should receive from a particular moment and organising the soundtrack accordingly.

    Stateman’s preferred method was additive rather than deductive. Instead of filling a soundtrack with every plausible sound and gradually removing whatever causes problems, begin with what is essential. Establish the dialogue. Introduce music when the scene requires it. Add environment to create space and context. Bring in effects deliberately, allowing each contribution to justify its presence. This approach makes purpose part of the design process before complexity accumulates.

    Spatial sound presents the same challenge on another scale. Contemporary systems allow designers to place and move sounds through increasingly elaborate speaker arrangements, but movement acquires meaning only through its relationship with the story. Surrounding listeners with constant activity can make spatial information less expressive rather than more. If every sound moves, movement itself loses significance. Space becomes useful when its behaviour supports attention, perspective or dramatic intention.

    For Once Upon a Time in Hollywood, the absence of a conventional composed score created a distinctive set of possibilities. Music arrived through records, radio and material connected closely with the period represented by the film. Sound design could occupy areas that might otherwise have belonged to score, contributing low-frequency energy, changes in texture and transitions that influenced mood without announcing themselves as musical cues. The team did not begin by asking how every capability of contemporary soundtrack production could be demonstrated. They considered what would make the film feel connected to 1969.

    That thinking also informed Stateman’s use of Dolby Atmos. He regarded the format as an impressive creative environment, but not as a reason to send objects continually moving through the auditorium. Early experimentation with expanded spatial systems could become cluttered when additional speakers were treated as spaces waiting to be filled. Stateman instead described a stable foundation with carefully selected individual elements used when spatial movement genuinely contributed to the experience.

    For a film drawing heavily upon period recordings, radio and two-channel music sources, aggressive object movement could have conflicted with the aesthetic world being created. A technically advanced format can sometimes serve a film most effectively by concealing its sophistication. As with the single movement of the Terminator 2 platform, possibility acquires value through selection.

    Working with very different directors reinforced Stateman’s resistance to universal solutions. The sonic worlds of John Hughes and Oliver Stone required radically different forms of expression. Moving between comedy and films such as JFK, Born on the Fourth of July and Natural Born Killers prevented one successful method from becoming a formula applied repeatedly. An early mentor had encouraged Stateman to approach every project as a new problem and continue trying different methods rather than relying upon established answers.

    Experience, from this perspective, should expand the vocabulary available to a practitioner rather than narrow the range of possible responses. A Quentin Tarantino film does not require the same sonic logic as an Oliver Stone film, and neither can be approached as a variation of a John Hughes comedy. Even repeated work with the same director changes as the project changes. Trust allows a shorthand to develop, but familiarity should not turn into repetition.

    Comedy offered a useful example of this contextual thinking. Familiarity can establish a pattern, while an unexpected interruption creates surprise. Yet the same broad relationship between expectation and disruption can also support drama, suspense and horror. A technique has no fixed emotional meaning outside its context. What matters is the audience’s developing expectation and the moment at which the soundtrack confirms, delays or overturns it.

    Such decisions require somebody to understand more than individual sounds, which led towards the most distinctive professional idea in Stateman’s lecture. He argued that contemporary practitioners should think of themselves not only as sound designers, but as sound directors, sound producers and sound designers. These are not three disconnected jobs. They represent different perspectives on a shared responsibility for the complete sonic experience.

    The sound director understands intention. This person can discuss the desired experience with the filmmaker, interpret creative needs and establish an overall direction for the soundtrack. The sound producer understands resources. Schedules, budgets, staffing and priorities determine what can be achieved and where effort should be concentrated. The sound designer turns those intentions and resources into creative work. Stateman regarded the combination of these abilities as a professional ideal.

    Directors themselves often move through a similar expansion of responsibility. A writer may become a director to shape the interpretation of the material, then become a producer to gain greater influence over resources and priorities. Stateman argued that sound professionals can develop in a comparable way. Technical expertise remains essential, but understanding intention, organising people and resources and making creative decisions gives the practitioner a more substantial role in shaping the work.

    His career offers several examples of this expanded responsibility. Building large creative teams required production thinking. Designing complex theme-park attractions demanded an understanding of playback systems and engineering. Long-term relationships with directors depended upon interpreting intention rather than waiting for technical instructions. Global distribution and localisation required the sound team to consider what happens after the supposedly finished soundtrack leaves the mixing stage.

    At that point, the problem of control becomes larger. Stateman may work in a highly controlled Atmos environment containing an extensive loudspeaker system, while an audience member eventually hears the result through earbuds on a train. Another listener watches through a television with speakers facing towards a wall. Somebody else lies in bed listening quietly while another person sleeps nearby. A theatrical soundtrack mixed at a high reference level may later be heard at a dramatically lower domestic level.

    No single master can guarantee an identical perceptual experience across all of those conditions. Low-level detail that remains clear in a cinema may disappear when the entire soundtrack is turned down. Wide dynamics that feel exciting in a theatre can become impractical late at night. Dialogue that is intelligible through one system may become difficult to follow through another. Translation therefore concerns more than whether a format can technically fold down from one speaker configuration to another. It concerns what survives perceptually when the listener, environment and playback level change.

    Even cinemas introduce uncertainty. Stateman described discussions with theatre owners about films being played below their intended reference levels. Their explanation was practical. A film begins at the specified level, somebody complains, and the level is reduced. Further complaints lead to further reductions until complaints stop. From the exhibitor’s perspective, this is a rational response to the audience in the room.

    Production can create pressure in the opposite direction. Filmmakers listening in a controlled mixing environment may repeatedly ask for greater excitement, prompting levels to rise until the result satisfies the room. One part of the system pushes upwards while another turns the finished film down. The problem cannot be solved by allowing dialogue, music and effects departments to maximise their own material independently. Somebody needs to maintain responsibility for the dynamic shape of the complete experience.

    Dynamics, in this sense, organise attention across time. Genuine intensity requires quieter material around it. Low-level information needs to remain meaningful under realistic playback conditions, while moments of scale need enough contrast to feel exceptional. The same principle that made one movement of the Terminator 2 platform effective applies to the soundtrack more broadly. Constant intensity reduces the expressive power of intensity itself.

    Global distribution introduces another kind of variation. A soundtrack may need to function across many languages, each with different rhythms, durations and vocal characteristics. Preserving quality across those versions is not merely an administrative task completed after the creative work. It is an audio-design problem involving performance, mixing, workflow and technology.

    Stateman’s approach was to break an apparently overwhelming challenge into smaller problems without losing sight of the whole. A modern feature soundtrack can require enormous quantities of editorial and mixing time, making meaningful control by one person impossible. Large teams become necessary, yet specialisation creates the risk that individuals see only their own component. The sound director-producer-designer model provides a means of maintaining overall intention while complex work is distributed among specialists.

    Localisation makes the limitations of a single fixed soundtrack particularly visible. Different languages alter timing and vocal behaviour, while different audiences and distribution systems introduce further variables. Streaming services now possess forms of information and control unavailable to earlier theatrical systems. Different languages and formats can already be delivered to different users. Stateman’s discussion suggested a logical extension of that capability: audio systems that respond more directly to individual listeners and the circumstances in which they are listening.

    Rather than treating a soundtrack as one fixed master expected to survive every possible situation, elements could remain available for intelligent recombination. Dialogue might receive greater prominence where needed. Dynamic behaviour could change for quiet listening. Different relationships between elements might be created according to the user’s environment, hearing ability or preferences.

    Such possibilities do not necessarily weaken creative intention. They raise a more fundamental question about what preserving intention actually means. A listener who cannot understand the dialogue is not receiving the intended experience merely through being sent the same electrical signal as everyone else. Someone listening quietly in bed occupies a different perceptual situation from an audience inside a calibrated cinema. Identical delivery does not guarantee equivalent experience.

    Adaptive audio therefore extends rather than abandons Stateman’s earlier principles. If the purpose of sound design is to guide attention, communicate story and create emotional relationships, then changes in listening circumstances matter. The challenge is to determine which aspects of an experience must remain stable and which can change to preserve them. A theme-park attraction, theatrical soundtrack, localised release and adaptive audio system appear very different, yet each requires designers to think about the complete journey between creative intention and audience experience.

    Across Stateman’s lecture, simplicity emerged not as an absence of complexity, but as its successful organisation. Hundreds of people may contribute to a project. Thousands of hours may be spent editing and mixing. Sophisticated engineering may sit behind the delivery system. The audience does not need to experience any of that as complication. They need to understand a voice, feel a change in atmosphere, anticipate an event or be surprised by a sudden movement.

    A film soundtrack can contain vast numbers of tracks, edits, recordings and processing decisions while resolving into a clear moment of attention. Achieving that clarity requires more than technical skill. Designers need to understand what matters within the scene. Producers need to organise the resources that make the work possible. Sound directors need to maintain the relationship between individual decisions and the complete experience. Teams need complementary expertise and enough trust for difficult problems to be handed to people capable of solving them.

    Stateman’s lecture ultimately presented sound design as the organisation of attention through time. The designer decides what matters now, what can wait, what should disappear and what should arrive unexpectedly. The producer makes those decisions achievable. The director maintains an understanding of why they matter. Technology expands the available possibilities, but cannot decide which ones belong in the experience.

    That distinction becomes more important as audio technology develops. Spatial formats offer increasingly detailed control inside theatres and listening rooms. Streaming services distribute content globally and create new demands for localisation. Object-based systems can preserve elements beyond a fixed master, while adaptive delivery may eventually allow soundtracks to respond more intelligently to listeners and environments. Each development increases possibility, but also increases the need for judgement.

    Somewhere beyond all of those systems, a listener is trying to follow a story. They may be sitting inside an elaborate cinema, travelling on a train, watching a laptop or lying quietly in bed. The sound professional cannot control every condition surrounding that experience, but can understand the relationships that matter. Dialogue can remain clear. Attention can be guided. Complexity can be organised. Dynamics can create contrast rather than exhaustion. Technology can serve intention rather than advertise itself.

    The response to uncertainty is not complete control. It is to design intelligently for variation while remaining clear about intention. Build teams capable of solving problems that no individual could solve alone. Understand the whole experience while taking responsibility for one part of it. Use complexity behind the scenes to create clarity for the listener. Know the story well enough to recognise when a convention should be followed, when it should be overturned and when the most powerful use of a technology is to leave it silent.

    A soundtrack may be shaped inside a carefully controlled room, but it does not remain there. It travels through cinemas, languages, formats, devices, environments and listeners. Every stage changes the circumstances in which the work will be experienced. Stateman’s lecture suggested that the future of sound design lies not in pretending those differences can be eliminated, but in understanding them well enough to preserve what matters. The mix may be finished when it leaves the studio. The experience begins when somebody, somewhere, presses play.

  • How Do You Design the Sound of Reality? Sefi Carmel on Documentary Sound, Perception, and the Ethics of Construction

    Sefi Carmel

    How do you design the sound of reality?

    A whale dives beneath the surface of the Atlantic Ocean. Filmed from a distant boat, its tail disappears into the water and a splash is heard. The moment appears entirely natural, yet the camera crew may have been hundreds of metres away, surrounded by engine noise and incapable of recording anything resembling the sound presented in the finished film. Perhaps the splash came from a sound-effects library. Perhaps somebody dropped an object into water. Perhaps a Foley artist moved a scuba flipper through a bathtub. Does adding that sound make the documentary less truthful, or does it help audiences experience an event that genuinely occurred but could not be captured adequately during filming?

    During his online guest lecture for Edinburgh Napier University, London-based sound designer, composer and dubbing mixer Sefi Carmel explored the creative and ethical questions surrounding soundtrack creation for documentaries. Drawing upon experience of mixing more than one hundred documentaries, he challenged the assumption that factual filmmaking requires a fundamentally different sonic vocabulary from drama. Dialogue, music, atmospheres, spot effects, Foley and abstract sound design can all contribute towards documentary storytelling. The crucial issue is not whether a sound was recorded at the moment shown on screen, but what its addition asks the audience to believe.

    Throughout the lecture, Carmel returned to a broad understanding of the soundtrack. Everything emerging from the speakers belongs to a single composition created in relationship with the image. Dialogue, music and effects may be separated technically, but audiences experience their combined movement through time and space. Like music, a soundtrack arranges rhythm, dynamics and timbre to communicate emotions and ideas. Documentary sound therefore involves much more than cleaning interviews and placing music underneath them. It is the construction of an audiovisual experience whose materials may come from reality while their organisation remains an act of filmmaking.

    Carmel began by questioning a belief he had held as a young sound designer. News reporting and documentary filmmaking had once appeared to occupy similar positions on a spectrum between reality and fiction. At one extreme, news aspires towards an account of events with minimal manipulation. At the other, drama openly asks audiences to accept scripted performances, constructed edits, designed sound effects and music intended to influence emotion. Documentary can initially appear closer to the first model, yet Carmel’s experience of the form led him towards a different conclusion. A feature-length documentary still has to hold an audience for sixty, ninety or more minutes. It communicates ideas, develops relationships, establishes places, controls pace and creates emotional movement. Those demands make it filmmaking rather than an extended news report.

    Understanding documentary in those terms considerably widens its creative possibilities. Music can shape emotional interpretation, sound effects can strengthen actions and atmospheres can establish locations that production recordings fail to communicate. Archive footage can be reconstructed into a convincing audiovisual world, Foley can restore physical detail and abstract sound design can emphasise an important transition or idea. Documentary makers have access to almost the complete filmmaking toolbox. With that freedom comes an ethical problem. If documentary claims a relationship with reality, how far can its soundtrack depart from literal recording before enhancement becomes deception?

    The whale provides a useful test. The animal really entered the water and its tail really created a splash, even though the filmmakers could not capture that sound from their position. Adding a plausible splash does not invent the event. It reconstructs an acoustic consequence already implied by the image, allowing audiences to experience the animal’s movement and scale more immediately without asking them to believe that something happened when it did not. Matters become more difficult when sound influences interpretation rather than restoring an unheard event. Carmel drew his clearest ethical line around speech. Reconstructing the likely sound of an action differs substantially from editing somebody’s words in a manner that changes their meaning or context. The latter can alter the evidence from which audiences understand an event. For Carmel, creative sound design remains legitimate when it serves the film without becoming a gross lie.

    A much less serious encounter with perceptual truth came during his work on a documentary about the Winter Olympics. Faced with a distant shot of a skier slaloming down a snowy slope, the director wanted the movement to have greater sonic presence. Library searches failed to produce a suitable recording, so Carmel created one himself. The director liked the result and repeatedly asked what had produced it. Carmel initially refused to answer, but eventually revealed the source: a knife scraping across toast. The sound had worked perfectly while the director perceived it as skiing. Once its origin became known, however, the illusion collapsed. The director could hear only toast and eventually asked for it to be removed. Nothing in the waveform had changed. Image, expectation and context had previously allowed one event to become another, while knowledge of the source created a new and apparently irreversible interpretation.

    Listeners do not identify sounds solely through their acoustic properties. Visual information, expectation, context and prior knowledge shape perception. A recording does not acquire credibility merely through sharing a physical origin with the object shown on screen, and a constructed sound does not automatically become deceptive through coming from somewhere else. Credibility emerges from the relationship between sound, image and meaning. Documentary sound design occupies a space between physical truth and perceptual credibility, requiring the designer to consider where a sound came from alongside what audiences understand it to represent.

    Reality presents another difficulty long before questions of creative enhancement arise. Documentary crews work in environments that cannot be controlled in the manner of a drama production. Interviews happen near roads, beneath aircraft routes, inside noisy buildings and beside machinery. Important moments may occur only once, leaving post-production dependent upon whatever the location recordist managed to capture. Documentary also has limited access to ADR. Replacing a contributor’s voice with a later studio performance can compromise spontaneity, authenticity and practicality. An imperfect recording of an essential contribution therefore has to be made intelligible and aesthetically acceptable even when the original conditions were hostile.

    Some problems are comparatively manageable. Low-frequency rumble can be reduced, while constant noises such as air conditioning, electrical hum or an aircraft cabin may respond well to noise reduction. Variable interference presents a much harder problem. An accelerating motorcycle can move through the same frequency regions as speech while continually changing its spectral character, making removal difficult without damaging the voice. Aggressive processing introduces further compromises. Equalisation can isolate the most intelligible part of a voice while leaving it thin and unnatural. Noise reduction can remove interference at the cost of audible artefacts. A line may become technically clearer but aesthetically less convincing. Restoration is not a contest to remove the greatest possible quantity of unwanted sound. Every intervention changes the material audiences hear.

    Carmel connected this problem to the idea of aesthetic disturbance. Drawing upon Ludwig Wittgenstein, he described aesthetics partly through the recognition of something being wrong: a picture hanging at an angle, for example, produces a disturbance that disappears when the relationship is corrected. Documentary sound can create similar disturbances. A close image accompanied by an unexpectedly distant, reverberant voice may produce audiovisual dissonance. Distortion attracts attention, while harshness, brittleness or excessive boxiness can make listening uncomfortable even when every word remains understandable. Intelligibility is necessary but insufficient. Dialogue also needs to feel congruent with the image and surrounding soundtrack, and a technically rescued voice can still undermine a scene if its perspective, spectrum or acoustic character appears disconnected from what viewers see.

    Modern restoration tools have expanded what can be recovered. Carmel discussed noise reduction, de-clicking, de-crackling, de-clipping, dereverberation and spectral repair as processes capable of rescuing recordings that might once have been considered unusable. Constant noise can sometimes be reduced substantially, while sophisticated interpolation may reconstruct clipped or distorted speech with surprising effectiveness. Greater capability does not remove the need for judgement. Restoration can erase information that belongs to the story. Crackle on an old recording may be technically undesirable while simultaneously communicating historical distance. The noise of an old shellac disc or archive recording can help audiences understand that they are hearing material from another period. Removing every imperfection may weaken the narrative. Technical possibility matters less than understanding what the existing sound already communicates.

    This tension between repair and preservation led to one of Carmel’s central principles: ideally, every layer of the soundtrack should be capable of telling the story. Dialogue carries narrative through words. An atmosphere can communicate location, weather, time of day and activity without spoken explanation. Music can reveal emotional direction. A forest atmosphere containing birds, wind and running water immediately places listeners within a particular kind of environment. Each layer contributes different information, and the complete soundtrack emerges from their interaction rather than from one dominant element surrounded by decoration.

    Documentary production rarely provides ideal materials, so those layers often have to support one another. Aggressively cleaned dialogue recorded on a windy location may no longer contain enough environmental information to make its setting believable. Carefully constructed atmospheres can return that context. Sound effects can reinforce actions the original recording failed to capture, while music can support emotional movement that damaged or fragmented production sound cannot carry alone. None of these layers needs to become conspicuous. A weak recording may become convincing once placed inside an appropriately designed environment, since audiences hear relationships between elements rather than evaluating every track in isolation.

    Voiceover introduces another distinctive element within those relationships. Casting is often a directorial decision, though Carmel argued that sound professionals can contribute valuable thinking when invited into the process. The appropriate voice depends upon the subject, intended audience and emotional character of the film. A documentary about the rise of a young pop group requires a different vocal identity from one exploring humpback whales in the North Atlantic. Performance matters as much as casting. Tone, energy, pacing and character shape the audience’s relationship with the film, while the chosen voice influences the space available for music, effects and atmosphere. Voiceover is part of the soundtrack’s composition, not information simply placed above it.

    Music performs an equally integrated role. Carmel distinguished between music existing within the world shown on screen and music functioning as score. Source music might come from a visible performer, radio, jukebox or other plausible location within the scene, while score operates outside that visible world and shapes emotion from another level. Documentary makers can also blur the distinction deliberately. An old rock-and-roll song treated with band-limiting and room reverberation might appear to come from a jukebox in a diner, allowing music to reinforce the setting, contribute historical or cultural information and perhaps help mask weaknesses in the location sound.

    Selecting or composing music requires an understanding of everything else occupying the scene. Dialogue-heavy sequences need music capable of supporting speech without continually competing for attention. Dense lead instruments or vocals can occupy perceptual and spectral territory similar to the human voice, forcing the music much lower in the mix. More restrained arrangements can create emotional colour while leaving space for narration and interviews. Other sequences allow music to carry more of the narrative. An expansive aerial view with little dialogue may support a large thematic statement that would overwhelm an intimate interview. Musical effectiveness in isolation matters less than the role a piece needs to perform at a particular moment and the relationships it forms with the rest of the soundtrack.

    Location dialogue complicates those relationships further. A controlled voiceover recording tends to maintain comparatively stable level and performance, while spontaneous speech can vary considerably. Contributors may begin sentences with energy and trail away as thoughts conclude. Music automation must respond to those changing patterns. A static reduction may leave quieter words obscured or make stronger phrases feel unnecessarily exposed. Mixing becomes a continual negotiation between intelligibility and musical continuity, with the soundtrack moving around the natural behaviour of voices that were never performed for the convenience of the mixer.

    Atmospheres perform several roles simultaneously. Most obviously, they tell audiences where they are. Traffic, birds, wind, room tone, distant machinery or human activity can define an environment before viewers consciously analyse the image. They also smooth editorial transitions. Documentary scenes are frequently assembled from material recorded at different moments, positions or even days, and a continuous environmental bed can help separate pieces of location sound feel as though they belong to a coherent space. Carmel identified another, less obvious function: atmospheres can contribute spectral balance. If a scene feels sonically empty within a particular frequency region, an appropriate environmental layer can help create a more aesthetically satisfying whole. The choice still needs to make narrative sense, but storytelling and sonic composition overlap here. Atmosphere can provide information, continuity and texture at the same time.

    Spot effects operate on a more local scale. A car door closes, a telephone is placed down or a gun fires. Brief synchronised events can reinforce visible actions and restore details absent from production recordings. Archive footage provides particularly rich opportunities for this kind of reconstruction, especially when historical images arrive without usable synchronous sound. Old footage may be silent or accompanied by narration and music unsuitable for the contemporary documentary. A battlefield sequence showing tanks, artillery and soldiers therefore presents the sound designer with an empty world that needs to be rebuilt.

    For Carmel, archive reconstruction can be approached with the same dramatic ambition used in fiction. Tanks can advance through the frame, gunfire can occupy different distances and artillery can establish scale, while wind across an exposed landscape gives the environment continuity. The designer might process the soundtrack to suggest historical recording technology or create a vivid modern sound world that places the audience imaginatively inside the event. A deliberately aged soundtrack reminds viewers that they are encountering archive material, while a contemporary reconstruction can reduce historical distance and make an event feel immediate. Neither choice is neutral. Both interpret history, leaving the designer to consider the relationship the documentary seeks to create between the audience and the past.

    Foley can contribute in much the same way, although documentary schedules and budgets rarely permit complete coverage. Selective use can still transform significant moments. A historical reconstruction showing chainmail being worn, armour handled or a sword drawn may deserve detailed physical sound even when none was captured during filming. If the moment carries narrative importance, there is no reason to reject Foley merely through an assumption that documentary sound must remain limited to location recordings. Abstract sound design extends the same principle further. Drones, impacts and heavily processed transitions can strengthen important ideas or structural moments, creating unease, giving a cut greater dramatic force or helping an audience experience a transition emotionally as well as intellectually.

    For Carmel, the legitimacy of these devices depends upon purpose. Dramatic sound should strengthen the film rather than substitute manipulation for argument. Documentary editing, cinematography and music already influence how audiences understand material, and sound design participates in the same process. Ethical responsibility lies in recognising what each construction communicates rather than pretending that construction does not occur. Documentaries are built from real people, events, evidence and places, yet films do not emerge automatically from those materials. Someone selects shots, orders sequences, chooses where music begins, decides when silence matters and determines which details audiences hear. Soundtrack creation is part of that authorship.

    Creative decisions are only part of the work. The documentary must also survive the circumstances in which it will be heard. Carmel emphasised that mixing begins with the destination. A theatrical documentary, television broadcast, online film and festival screening present different playback conditions and technical expectations. Format, dynamics, equalisation and level decisions need to reflect those contexts. A large theatrical environment can support substantial low-frequency extension and wider dynamics, while television playback may occur through much smaller speakers and in less controlled surroundings. Processing that creates clarity in one context can become harsh or excessive in another.

    Platform awareness connects technical delivery directly with audience experience. Spectral energy inaudible on small television speakers can still consume headroom. Extreme dynamics may work beautifully in a cinema while causing viewers at home to continually adjust volume. Loudness standards formalise part of that relationship, particularly for broadcast delivery, but Carmel’s broader point concerned the distribution of intensity across the film. Loudness can be understood as a budget. If every moment is treated as maximally intense, little room remains for genuine peaks. Dynamic planning becomes another form of storytelling, allowing contrast to carry dramatic meaning rather than treating level merely as a compliance problem.

    Compression and limiting require similar contextual judgement. A theatrical mix may use comparatively subtle master processing, preserving headroom and contrast, while television and online material may tolerate or require greater control. One version is not inherently superior to another. Each mix needs to function within the medium for which it is intended, preserving the film’s intentions under different listening conditions.

    Deliverables extend that responsibility beyond the primary audience. Documentary films may travel between territories and require new narration or dubbed dialogue. Music and effects tracks need to support localisation rather than reproduce automation created around the timing of the original language. Carmel explained the value of providing undipped music for this purpose. In an English version, music may be reduced beneath a particular phrase and raised again when the speaker stops. A translated version may take longer or shorter to communicate the same idea. If the music stem already contains automation tied to the English timing, the foreign-language mixer inherits a structure that no longer fits. Providing material without those dialogue-specific reductions allows the new mix to respond properly to the translated performance.

    A soundtrack therefore exists as more than a finished mix. It may need to survive new languages, platforms and contexts, and good delivery anticipates the work of people who will encounter the material later. Carmel ended with an even simpler responsibility: check the work. Quality control may appear less intellectually exciting than documentary ethics or perceptual sound design, yet it protects every creative decision made before delivery. A mixer can spend days constructing a sophisticated soundtrack and still send an unusable file through a routing mistake or export error. Recording or exporting something does not prove that the expected material exists in the resulting file. Listen to it. Watch it. Check it.

    That practical instruction sits neatly alongside the lecture’s larger argument. Documentary soundtrack creation moves constantly between interpretation and responsibility. Designers can construct sounds that were never recorded, rebuild silent archives, use Foley, shape emotion through music and introduce dramatic sonic devices. Those freedoms demand judgement. Does a sound restore an experience, clarify it, interpret it or falsify it? What does it ask the audience to believe? Does it support the film’s argument without altering the meaning of its evidence?

    Carmel’s lecture presented documentary sound as an art of constructing experience from incomplete reality. Location recordings arrive damaged. Cameras capture actions from distances at which their sounds cannot be heard. Archive images survive after their original acoustic worlds have disappeared. Interviews need to coexist with music, while fragmented scenes require atmospheres capable of making them feel continuous. The documentary soundtrack is built through responses to these absences. Sometimes the response is technological: remove a constant noise, repair distortion or restore intelligibility. Sometimes it is editorial: create a continuous atmosphere around fragmented material. Sometimes it is performative: add Foley to a significant physical action. Sometimes it is musical, shaping the emotional direction of a sequence. At other moments, the solution may be a completely unrelated object whose acoustic behaviour happens to make an image believable.

    A knife scraping toast can become a skier moving across snow, at least until somebody learns the secret. Reality does not arrive in post-production as a complete audiovisual object waiting to be preserved. It arrives as recordings, images, testimony, fragments and absences. Filmmakers decide how those materials should be organised into an experience that audiences can follow and understand. Sound designers participate in that process by reconstructing relationships between actions and consequences, voices and spaces, images and expectations.

    Construction and dishonesty are not the same thing. A documentary soundtrack can be richly designed while remaining faithful to the people and events it represents. Literal accuracy may sometimes be essential. At other moments, perceptual credibility communicates an experience more effectively than an unusable or absent location recording ever could. A splash can give weight to a whale entering the ocean. An atmosphere can return a damaged interview to its environment. Designed sound can give silent archive footage physical immediacy. Music can reveal emotional relationships without changing the evidence shown on screen.

    Documentary sound occupies the space between what happened, what could be recorded and what audiences need in order to understand and feel the film. Carmel’s lecture showed that this space is not a technical inconvenience to be hidden. Much of the creative work begins there. The documentary sound designer cannot preserve every sound of reality, since many were never captured in the first place. The responsibility is to decide what should be repaired, what can be reconstructed, what needs to remain imperfect and what must never be changed.

  • How Do You Make a Game Feel Dangerous? Will Morton on Emotion, Attention, and Designing Sound for the Player

    Will Morton

    How do you make a game feel dangerous?

    A gun can sound enormous and still fail to make a gunfight frightening. Every weapon may have a powerful attack, convincing mechanical detail and an impressive environmental tail, yet the player can remain strangely detached from the danger. Solving that problem may have little to do with redesigning the weapon itself. Bullets pass close to the head. Impacts strike nearby walls with exaggerated force. Brickwork breaks apart, fragments scatter and the environment appears to react violently to the threat. During his online guest lecture for Edinburgh Napier University, game audio designer Will Morton explored how sound can shape emotion, focus attention and guide players through complex interactive experiences. Drawing upon twelve years at Rockstar North, where his work included the Grand Theft Auto series, Red Dead Redemption and L.A. Noire, followed by the establishment of Solid Audio Works with fellow former Rockstar audio specialist Craig Connor, Morton presented game sound design as a discipline of selection. Thousands of sounds may exist within a game, but their value depends upon knowing which ones matter at any particular moment. Throughout the lecture, one principle repeatedly emerged. A designer must ask not only what a game world should sound like, but what the player needs to hear in order to feel what the game intends them to feel.

    Morton began by placing creative decisions within the realities of AAA game production. Large games are expensive, technically constrained and continually changing. Development rarely follows a fixed design from beginning to end. Features evolve, producers reconsider decisions and new requests arrive late in production. Platform restrictions impose further limits through storage, memory, streaming performance and processing capability. Scale introduces organisational complexity as well. Larger games require larger teams, while experienced creative specialists can find increasing amounts of their time absorbed by scheduling, administration and coordination. Sound design develops inside a moving system of technical, financial and organisational constraints. Success requires more than imagining an ideal soundtrack. Designers must create one capable of surviving years of changing requirements.

    Historical changes in technology have altered the scale of those constraints without eliminating them. Morton recalled Commodore 64 composer Martin Galway fitting numerous sound effects into approximately one kilobyte of memory. Restrictions of that magnitude appear almost comic from the perspective of contemporary production, yet modern open-world games can still leave audio teams fighting for storage and memory. Vastly greater resources are now available, while games simultaneously attempt to represent entire cities, landscapes and fictional worlds. Technological abundance creates new possibilities, but ambition expands alongside it. Designers still have to decide where limited resources will make the greatest contribution.

    Money introduces another set of choices. Sound effects require designers, recording equipment, locations, editing time, libraries, software and continually changing computer systems. Dialogue adds writers, actors, directors, studios, recording staff and extensive editing. Music may involve composition, licensing, performers, orchestras, specialist recording facilities and interactive implementation. Morton’s overview exposed the consequences hidden behind apparently simple creative ambitions. Another recording variation, character voice or interactive music feature consumes time, money, memory and attention that cannot be spent elsewhere.

    Dialogue provided one of the clearest examples of complexity hiding behind familiar production tasks. Morton spent much of his Rockstar career dividing his time between sound design and dialogue before the scale of Grand Theft Auto V led him to work entirely as dialogue supervisor. For story-heavy games, dialogue cannot be treated as a sequence of lines requested by designers and recorded by actors. Repetition needs consideration. Lines must make sense across changing gameplay situations. Story information may unfold over many hours or days of play, while players can interrupt, delay or alter the circumstances in which dialogue occurs. A dialogue designer therefore needs to understand the interactive structure of the game as deeply as the individual performances being recorded.

    Direction presents a related challenge. Experience in film and television can produce excellent performances, yet game dialogue introduces problems absent from linear media. Cutscenes may operate much like conventional scenes, while in-game dialogue can occur within changing gameplay circumstances and across story structures experienced differently by individual players. Directors need more than an ability to elicit compelling performances. Detailed knowledge of the script, game and eventual context of each line becomes essential. A performance recorded in isolation must later remain convincing within circumstances that may not even be visible inside the recording studio.

    Morton also challenged assumptions about professional recording. Technically excellent dialogue can be captured with comparatively modest equipment in a carefully treated space. A suitable microphone, simple accessories and an acoustically controlled room may produce results approaching those of a far more expensive studio. Recording quality, however, forms only one part of a professional session. High-profile performers need confidence that they are participating in a serious production. Environment, organisation and treatment of the actor can influence trust in the project and relationships with agents and management. Professionalism encompasses the experience surrounding a recording alongside the technical quality of the resulting file.

    From these production realities, Morton moved towards the central creative argument of the lecture. Designers can easily assume that everything visible in a game should automatically produce a sound. Faced with an unsounded world, teams begin filling every action with detail. Cars receive engines and collisions. Characters acquire footsteps. Objects gain interactions. Environments fill with ambiences. A technically comprehensive and entirely plausible soundtrack gradually emerges, yet plausibility alone cannot guarantee clarity, excitement or emotional effect.

    Practical concerns provide one reason for restraint. Every additional sound requires creation, editing, implementation, memory and testing. Artistic considerations are even more important. If everything demands attention simultaneously, nothing receives focus. Morton encouraged designers to identify what deserves to be heard and what can remain absent. Important sounds require room to breathe. Dynamics rely upon contrast, while emotional emphasis depends upon moving attention between elements. Silence and omission become active design decisions.

    Player experience consequently takes priority over literal acoustic reconstruction. Real events often sound less dramatic than audiences expect. A gunshot captured from a particular position may seem surprisingly small. A suppressed weapon does not necessarily produce the familiar cinematic whisper audiences have learned to associate with it. A real minigun may collapse into an almost continuous mechanical roar instead of revealing every stage of its operation. Decades of film, television and games have established sonic conventions that now influence how audiences expect objects and events to behave. Designers work within that accumulated perceptual history.

    Morton demonstrated the point by comparing cinematic and real recordings of suppressed firearms. A familiar designed version sounded short, controlled and immediately recognisable, carrying characteristics audiences strongly associate with a silenced weapon. Real recordings behaved quite differently. Neither approach offered a universal answer. Documentary representation might favour acoustic accuracy, while an action game may need immediate recognition, excitement and dramatic satisfaction. Before designing a sound, Morton considers the role accuracy should play, whether an event needs to feel larger than reality and how strongly audience expectations should influence the result.

    Miniguns provided an even clearer illustration. Morton discussed the famous weapon sequence in Predator, where mechanical movement, spinning and other details reinforce the spectacle shown on screen. Real recordings present a very different impression, dominated by the extraordinary density of rapid gunfire. Grand Theft Auto: Vice City and Grand Theft Auto V offered further interpretations, each constructing the weapon differently, while Terminator 2 adopted another cinematic approach that remained closer to aspects of the real sound. Comparing them did not reveal a correct minigun. Instead, their differences showed how each design serves the experience of a particular production.

    Reality and relevance are therefore separate considerations. A game designed as escapism may gain little from reproducing everyday acoustic experience with complete fidelity. Players do not always need to hear events from the perspective of ordinary observers. They need a version that communicates the event’s importance within the game. Sound design selects, enlarges, simplifies and reshapes reality according to dramatic purpose.

    Attention can be directed just as effectively through subtraction. Morton used a remotely detonated explosive from Grand Theft Auto IV: The Ballad of Gay Tony to demonstrate how removing surrounding sound can make a single event dominate awareness. Immediately before the explosion, the wider soundtrack recedes and the bomb’s warning becomes the focus. He compared the moment with the seismic charge sequence from Star Wars: Episode II, where a brief interruption in the expected sound field creates anticipation and gives the subsequent event greater impact. Both sequences derive part of their power from absence.

    Focus involves more than increasing the level of an important sound. Every event competes with its surroundings. Removing distractions can achieve more than additional layers, greater loudness or further spectral exaggeration. Presence acquires meaning through contrast with absence. A fraction of a second of reduced activity can prepare an event more effectively than a continuously dense soundtrack.

    Morton’s most revealing example concerned a producer asking for guns to sound more dangerous. Existing weapon sounds already seemed successful to the audio team. They contained convincing mechanical detail, a strong initial attack, substantial body and satisfying environmental tails. Reworking those qualities did not solve the problem. Although the request appeared to concern the sound of the guns, the underlying dissatisfaction was emotional.

    An unexpected experience away from the studio suggested another approach. During a paintball game, Morton found himself sheltering behind structures made from metal oil drums. The paintball markers themselves produced relatively insignificant sounds. Fear came from sudden, violent impacts striking the metal around him. His immediate environment appeared to contract as attention focused upon nearby impacts, resonances and the sense that projectiles were arriving from directions he could not fully control. The weapon itself was not frightening. Being under fire was.

    Recognising that distinction transformed the design problem. Bullet passes became more prominent, with their levels responding more strongly to proximity. Impacts against brick and concrete gained force. Low-frequency energy added physical weight, while debris and crumbling material made the environment appear to react to nearby gunfire. Details that might be acoustically subordinate to a real gunshot were deliberately brought forward. Stronger weapon sounds had never been the answer. Gunfights needed to communicate vulnerability and danger.

    Here Morton identified a broader professional skill. Directors and producers rarely describe every sound problem in acoustic terms. They may ask for something to be louder, bigger, darker, faster or more dangerous while expressing dissatisfaction with an emotional result. Following the literal wording can send a designer towards the wrong solution. Morton argued for interpreting the intention behind the request. Someone asking for a more dangerous gun may actually be asking to feel vulnerable. Once the desired experience becomes clear, the designer can decide which part of the sound world needs to change.

    Creative collaboration therefore requires translation between intention and acoustic action. Sound professionals develop specialised vocabularies for frequency, dynamics, spatial behaviour, envelope and processing. Producers may describe experiences through emotion, imagery or metaphor. Useful information can exist within both forms of language. Professional expertise includes moving between them and identifying the experience concealed inside an apparently vague request.

    Gunfire also demonstrates how game audio operates through systems of cause and consequence. A weapon consists of more than the sound emitted when a trigger is pulled. Projectiles move through space. Near misses pass the player. Bullets strike walls. Debris falls afterwards. Environments respond differently according to material and distance. Danger emerges from relationships between these events. Concentrating exclusively upon the source can leave the wider experience emotionally incomplete.

    Similar systems govern the game mix. Morton described mixing a large game as a potentially overwhelming process involving thousands of assets and potentially hundreds of simultaneous channels, all changing according to gameplay. Exact combinations cannot be predicted in advance as they can within a linear soundtrack. Dialogue, music, ambience, vehicles, weapons, footsteps and environmental interactions combine differently according to player behaviour. Even excellent individual assets can produce an exhausting result when their relationships are poorly controlled.

    Morton recalled playing games whose soundtracks felt like continuous acoustic assault. Constant density can be as tiring as excessive level. When every category remains active and prominent, listeners receive no opportunity for recovery and little indication of where attention belongs. More detail can therefore produce less communication.

    Working with Craig Connor, Morton developed a useful analogy for managing this complexity: approach the game as a music producer approaches a song. Every element needs an appropriate place and enough space around it. The comparison does not imply imposing a static musical mix upon an interactive system. Its value lies in encouraging relational thinking. Sounds acquire meaning through their positions alongside other elements rather than through isolated perfection.

    Arrangement offers another useful parallel. Music producers rarely expect every instrument to occupy the foreground continuously. Parts enter and leave, density changes and contrast creates structure. Individual elements occupy different spectral, spatial and dynamic roles. Interactive audio can apply similar principles while responding continuously to player action, with priority changing according to what the player is doing and what the game needs to communicate.

    Waiting until asset production is complete before attempting a final mix creates serious problems. By the time thousands of sounds have accumulated, assumptions about level, density and priority may already be embedded throughout the project. Assets designed without a meaningful reference mix can prove difficult to reconcile, particularly when each has been created to sound impressive in isolation.

    Morton recommended an iterative alternative. Early in development, designers can choose a small but representative collection of sounds and make them work together convincingly. A useful reference might include something loud, such as a gunshot or explosion, alongside quieter material such as footsteps. Approximate ambience establishes the environmental bed, while dialogue examples can cover a range from quiet speech through ordinary conversation to shouting. Mixing these elements early creates a working scale for everything that follows.

    New assets can then be designed in relation to an existing sonic framework. Designers know roughly where a sound needs to sit and how much space surrounds it. Original reference assets may eventually be replaced, but their early role remains valuable. They establish relationships before growing complexity makes fundamental decisions harder to change.

    Mixing, from this perspective, becomes part of sound design rather than a finishing process. Working without context encourages every gun, vehicle, impact and interaction to become enormous, detailed and impressive. Once combined, they compete. An evolving reference mix permits more varied decisions. Some sounds can remain small. Others can be narrow, distant or restrained. Detail can be reserved for moments when players have enough space to perceive it.

    Focus also connects sound directly to gameplay. Players continually decide where to look, where to move and what to do next. Audio can support those decisions by drawing attention towards useful information or allowing distractions to recede. An approaching threat, important character, changing environment or imminent event may receive temporary priority. Elements contributing little to the current experience can move into the background.

    Informational and emotional focus can coexist. A soundtrack can communicate danger without identifying an enemy’s exact position, or create suspense without explaining precisely what will happen next. Morton’s examples showed how game audio shapes a player’s state of mind while remaining part of the fictional world. Excitement, suspense, drama and humour are designed responses. A weapon feels powerful partly through its sound. An approaching explosion gains anticipation from the quiet preceding it. A gunfight becomes dangerous when incoming fire appears to tear apart the world immediately around the player.

    Realism consequently remains flexible. Effective sounds may preserve recognisable aspects of reality while exaggerating qualities useful to the experience. A weapon can retain enough mechanical identity to remain believable while gaining additional weight. An impact can exceed its real counterpart without feeling inappropriate to the image. A suppressed firearm can satisfy an established cultural expectation even when that expectation differs from literal acoustic reality.

    Morton resisted turning these observations into universal rules. Predator and Terminator 2 can present radically different miniguns while both succeeding on their own terms. A game pursuing realism may demand a different balance from an exaggerated action title. Intimate narrative experiences may use restraint where large-scale spectacle needs greater sonic scale. Designers need to understand what their particular project is asking players to experience.

    Selectivity also shapes resource allocation. No production can pursue every possible recording session, dialogue variation or interactive feature. Large games generate an almost unlimited number of potential tasks while budgets and schedules remain finite. Designers select which events deserve sound, which sounds deserve prominence, which systems justify development time and which details will genuinely improve the experience.

    Morton’s discussion of sound libraries introduced a longer view of professional practice. Building a useful collection of original recordings is expensive, but opportunities to capture interesting material should be taken when possible. A sound recorded today may remain unused for years before finding its purpose. Game audio professionals develop habits of listening beyond individual projects, continually noticing potential material in the world around them.

    His paintball experience represents an even deeper form of professional listening. Morton did not return merely with a useful recording. He had experienced a relationship between threat, proximity and environmental impact that changed how he understood a design problem. Everyday experiences can reveal how attention and emotion respond to sound. Listening professionally involves recognising such relationships as well as collecting interesting timbres.

    Technical constraints and psychological effects remain closely connected throughout Morton’s practice. Memory budgets, recording costs, dialogue systems, asset creation and mixing strategies eventually converge upon one question: what experience is the player having? Sophisticated technology has limited value when it does not support that experience. Conversely, a simple decision such as briefly removing surrounding sound can transform a moment when it directs attention effectively.

    By the end of the lecture, Morton had presented AAA game audio as a discipline balanced between enormous complexity and deliberate restraint. Contemporary designers have access to more memory, processing, channels and real-time synthesis than earlier generations could have imagined. Additional technical capacity, however, does not remove the need to choose. More available sounds do not make more audible sounds desirable. Processing power cannot determine where attention belongs. Larger worlds make focus increasingly important.

    His account also challenged the familiar suggestion that audiences notice game audio only when it fails. Players respond to powerful sound design even when they cannot describe every mechanism behind it. They feel the danger of a gunfight, anticipate an explosion during a sudden moment of quiet and recognise when a world has enough space to breathe. Good audio can elevate games whose code, visuals and other systems already represent enormous investment. Its contribution reaches far beyond correcting problems. Sound helps determine the emotional character of the experience.

    Morton’s lecture ultimately revealed game sound design as the management of attention through an interactive world. Designers decide when reality should be preserved and when expectation should take priority. Silence can become more powerful than another layer. A producer’s request needs to be interpreted through the emotion it seeks to achieve. A gun may already sound excellent while the gunfight surrounding it remains ineffective. Thousands of assets can exist within a game, yet the success of the soundtrack depends upon knowing which ones deserve attention at a particular moment.

    Perhaps the most important question is simply what the player needs to hear now. Sometimes the answer is a spectacular weapon. At another moment, it is the violent impact of a projectile against the wall beside them. Elsewhere, a quiet footstep, distant ambience or line of dialogue needs enough space to be understood. Occasionally, almost everything should disappear. Game audio becomes powerful when sound shapes experience rather than catalogues events. A virtual world may contain thousands of possible sounds. The art lies in choosing the ones that make the player feel something.

  • How Do You Perform a World Through Sound? Jason Swanscott on Foley, Improvisation, and Making Movement Believable

    Jason Swanscott

    How do you perform a world through sound?

    A character crosses a room, adjusts a coat, places a glass on a table and sits down. Nothing about the sequence appears extraordinary, yet its soundtrack may contain dozens of carefully performed details. Every footstep must carry the correct weight. Fabric must move with the body rather than merely rustle somewhere in the background. The glass must sound appropriate for its size and surface, while the chair should respond convincingly to the character’s movement. None of these sounds is likely to attract conscious attention, yet without them the scene can feel strangely empty. During his online guest lecture for Edinburgh Napier University, veteran Foley artist Jason Swanscott explored the extraordinary craft behind these apparently ordinary sounds. Drawing upon almost three decades of work across film, television and games, he revealed Foley as something far richer than synchronising footsteps and props to pictures. It is a form of performance in which movement, character, materials, objects and physical spaces are interpreted through sound. Throughout the session, one idea repeatedly emerged. Foley does not simply reproduce what can be seen on screen. It performs the physical behaviour of an entire world.

    Swanscott began by establishing the three principal elements of Foley: movement, footsteps and spot effects. This distinction provides a useful practical framework, though each category immediately reveals the interpretive nature of the work. Movement concerns the subtle sounds produced as characters shift their bodies and clothing. Footsteps recreate their interaction with the ground. Spot effects encompass the countless objects they touch, lift, open, close, carry and manipulate. Together, these tracks restore a physical presence that may be missing from production recordings or require greater definition within the finished soundtrack. Their purpose is not to make the soundtrack busier. They make bodies, materials and objects feel as though they occupy the world shown on screen.

    Movement often begins with fabric, but choosing a material is only the first decision. A suit jacket does not behave like a tweed coat, while cotton, silk and heavier period fabrics each possess different textures and patterns of movement. Swanscott watches how the character moves and performs those movements through the selected material. A slight adjustment in a chair requires a different gesture from somebody running, fighting or struggling into a coat. Sound follows action rather than existing as a continuous layer of generic clothing noise. The artist watches the body, interprets its movement and recreates its acoustic consequences.

    Footsteps make the relationship between sound and performance even clearer. Correct shoes and surfaces matter enormously. Foley stages therefore contain collections of footwear and recording surfaces capable of representing different periods, occupations, characters and environments. Leather-soled shoes suggest a different person from combat boots or high heels, while wood, gravel, tile and cobblestone each change the relationship between the performer and the ground. Owning the appropriate shoe and standing on the appropriate surface, however, does not guarantee a convincing footstep. The Foley artist must perform the character. Weight, pace, hesitation, confidence, exhaustion and emotional state all influence the rhythm and physical force of movement. A footstep can communicate who the character is and how they are moving through the scene.

    Spot effects extend this performance across the material world. A hand lifts a glass, a door opens, a weapon is drawn or an object falls onto a surface. Some actions can be recreated using objects similar to those shown in the image. Others require considerable invention. A sword may acquire additional metallic resonance to give its movement greater presence. A body impact may emerge from materials whose physical origins bear little resemblance to the action on screen. Swanscott’s task is to produce the sound that makes the visible action believable, regardless of whether the studio prop resembles its apparent source.

    One of the most interesting themes running through the lecture concerned the difference between physical truth and perceptual truth. Real events do not always produce the sounds audiences expect from them. A literal recording may appear weak, ambiguous or dramatically inappropriate once placed against an image. Foley therefore occupies a curious territory between reality and expectation. The artist frequently uses an object that is physically wrong to create a sound that feels perceptually right.

    Familiar examples are entertaining precisely through the unlikely relationship between source and result. Celery and cabbage can provide organic fractures and breaks. Bananas and watermelons can contribute the resistance and wetness required for sounds involving cutting flesh. Water, washing-up liquid and wallpaper paste can create viscous materials for blood and other bodily effects. Wet newspaper and other organic materials may provide textures required for medical and forensic scenes. Real offal can also be used, yet Swanscott’s examples demonstrate that literal materials hold no automatic claim to greater credibility. Screen reality is constructed through judgement rather than fidelity to the original source.

    This becomes especially apparent in genres where audiences possess strong expectations about events that few people have experienced directly. A cinematic punch must communicate weight and violence immediately. Different materials and performances can be combined to create that impression of force. Leather, boxing gloves, weighted impacts and other layers contribute qualities that the image appears to demand. Listeners perceive a body being struck even though the sound may have been constructed from objects with entirely different identities. Foley succeeds when attention remains on the event rather than the materials used to create it.

    Different productions also demand distinct styles of performance. A period drama such as Emma requires attention to the weight and behaviour of long dresses, delicate gloves, heeled footwear, polished floors and carefully handled objects. Its physical world is controlled and refined, so the Foley performance must support that quality. An action film such as Kick-Ass demands something entirely different. Fights require aggressive bodily performance, impacts, falls, furniture movement and layers of activity that communicate speed and force. Horror presents another challenge, where small sounds can become disproportionately important. A creaking surface, distant movement or restrained scrape may contribute more tension than an obviously dramatic effect. Foley changes with the storytelling language of the production.

    Across these genres, Swanscott must interpret images remarkably quickly. Foley artists frequently begin work without having watched the production in advance. They may receive limited notes and occasional guidance about particular details such as footwear or costume, while much of the decision-making happens as the material appears before them. A contemporary drama, period production, action sequence or fantasy world immediately suggests different requirements, yet the artist must continually decide what deserves performance, which materials might work and how much detail the scene can support.

    Such a working method reveals a form of expertise that is difficult to reduce to written instructions. Swanscott’s practice depends upon an accumulated relationship between sight, movement and sound. Watching a character walk prompts immediate decisions about shoe, surface, rhythm and weight. While performing those footsteps, he may already be noticing the objects that will require attention during the later spot-effects pass. Several forms of attention overlap. He is performing the current sound, watching synchronisation, interpreting character and preparing mentally for sounds that have yet to be recorded.

    Pressure makes this embodied expertise particularly visible. Swanscott contrasted the schedules available to different productions. A forty-five-minute ITV drama episode might receive only two days of Foley work, while a longer BBC production such as Silent Witness may allow around five. Feature films can provide considerably longer schedules, with major productions sometimes allowing several weeks. More time allows experimentation, comparison and repeated refinement. Fast television requires strong decisions almost immediately.

    On a two-day television schedule, movement and footsteps may occupy the first day, with spot effects following on the second. Every character still needs to move through the programme. Shoes still need to match, surfaces still need to change and objects still need to acquire physical presence. The schedule compresses the time available without removing the detail. Speed in this context means more than moving quickly. It depends upon recognising patterns, anticipating requirements and drawing upon a sufficiently developed physical and sonic vocabulary that experimentation can happen almost instantaneously.

    Years of this practice turn Foley stages into unusual archives of material culture. Shoes, fabrics, doors, glasses, tools, weapons, furniture and fragments of apparently unremarkable objects accumulate over time. Their value bears little relationship to monetary worth or original purpose. An object becomes valuable when the artist recognises useful sonic behaviour within it. Swanscott recalled finding a discarded sink outside a studio and bringing it inside. Its metallic resonance later proved ideal for a sound effect. A small paving slab could suggest the movement of a vast stone door. Neither object is limited by its physical scale or intended function once the artist begins to imagine what its sound might become.

    Foley therefore requires a peculiar form of auditory imagination. Swanscott has developed knowledge of how materials behave, then uses that knowledge to recognise relationships between sounds and images that may initially appear unrelated. He hears what an object can become when performed differently, recorded from another perspective or combined with another layer.

    One of the clearest examples came from The Legend of Tarzan. Swanscott recreated the heavy, padded movement of gorillas using boxing gloves. The relationship becomes logical when considered through physical qualities rather than visual resemblance: both can produce broad, soft impacts carrying considerable apparent weight. Discovering that relationship requires an artist to think in terms of behaviour. The boxing glove becomes useful through what it can perform.

    Other problems require materials that can be made to behave in particular ways. A snake moving across a surface can be suggested through materials such as cornflour-filled pillowcases, where the shifting internal texture creates the required movement. Magical or supernatural objects may combine recognisable physical impacts with resonant elements that suggest something beyond ordinary reality. In The Witcher, Swanscott described combining physical sounds with the resonance of a singing wine glass to give an enchanted object both material weight and a supernatural quality. Listeners need to believe that the object occupies physical space while also sensing that it belongs to a world governed by unfamiliar forces. Foley can communicate both qualities at once.

    Science fiction extends this challenge further. Productions such as Intergalactic require sounds for technologies without direct real-world equivalents. Spaceship doors, airlocks, damaged suits and magnetic footwear still need to communicate material, force and function. Heavy work boots combined with a rubbery surface can suggest magnetic contact, while compressed air and vocal techniques can contribute to the impression of air escaping from a damaged suit. Even within imaginary worlds, invention remains grounded in physical performance.

    Visually synthetic productions make this physical grounding especially valuable. Computer-generated creatures, magical objects and futuristic environments remove any possibility of simply recording what happened during filming. Audiences nevertheless expect those worlds to possess internal physical coherence. A creature needs weight. A machine needs resistance. Clothing must move with bodies, and objects must appear to interact with surfaces. Familiar materials give impossible worlds a physical logic that listeners recognise instinctively.

    Sometimes the object alone cannot create that credibility. Acoustic space can become part of the performance. Swanscott discussed work on Captain Phillips, where scenes inside the enclosed lifeboat required a claustrophobic acoustic quality. Performing the Foley conventionally within an open recording space would not reproduce the confined reflections demanded by the scene. The team instead worked inside an overturned water tank, using the enclosure itself to create the required resonance.

    That example broadens the definition of the Foley instrument. The performer and prop exist within an acoustic relationship shaped by surfaces, boundaries, reflections and microphone perspective. For the lifeboat, claustrophobia emerged through the physical conditions of recording. The space itself became part of the performance.

    Large scenes demand another kind of construction. A soldier running through a battlefield is represented by far more than footsteps on dirt. Clothing moves, equipment shifts, buckles rattle and weapons interact with the body. Each component may be relatively simple in isolation, yet together they create the physical complexity of a person carrying weight through an environment. Foley gives the body layers. Listeners may never consciously identify each buckle or movement, but their combined behaviour makes the character feel physically present.

    Close collaboration with sound editors allows this layered physicality to remain flexible. The artist performs movement, footsteps and spot effects, while editors organise, refine and prepare those recordings for the wider soundtrack. Individual elements can be adjusted independently, allowing movement, footsteps and object interactions to find appropriate relationships within the final mix. Additional sound effects or processing may supplement the performance, while the Foley recording provides an organic foundation tied directly to the rhythm of the image.

    That connection to performance becomes especially important when production dialogue has been replaced. ADR can separate dialogue from the incidental sounds originally captured on set. Footsteps, clothing and object interaction may then need to be reconstructed around the replacement dialogue so that the scene regains its physical continuity. Foley restores relationships between bodies and environments that post-production processes may have separated.

    Swanscott’s examples also challenged any assumption that every visible action requires equal sonic emphasis. Foley serves storytelling, and different scenes demand different levels of detail and exaggeration. Restrained handling of porcelain in a period drama communicates something very different from the aggressive materiality of an action sequence. Horror may depend upon an isolated movement sound emerging from otherwise restrained surroundings. Fantasy may require a familiar impact combined with an unfamiliar resonance. Every scene asks the artist to decide how an event should sound and how strongly the audience should experience it.

    Improvisation is central to these decisions, though professional improvisation is far from random. Years spent learning the behaviour of materials, surfaces and objects allow relationships to be tested quickly when an unexpected problem appears. A discarded sink becomes useful through recognition of its resonant potential. Boxing gloves become gorilla feet through their combination of softness and apparent weight. What looks spontaneous from outside the studio rests upon a deeply developed vocabulary of physical sound.

    Swanscott’s own body is equally important. Fight sequences may require vigorous physical movement. Footsteps must be performed with the rhythm and weight of somebody whose body may differ considerably from that of the artist. A delicate character, exhausted soldier and threatening pursuer cannot all be represented through the same neutral walking pattern. The Foley artist inhabits movement without being visible. His body becomes an interpretive instrument.

    Precision and expressiveness therefore exist together. Synchronisation matters. A footstep landing visibly out of time can break the connection between sound and image. Accurate timing, however, cannot rescue a performance with the wrong weight, rhythm or intention. Like other forms of performance, Foley uses technical control to enable expression. The artist must arrive at the correct moment while making that moment belong to the character.

    Interactive media changes the structure within which these performances operate. Swanscott’s work extends into video games, where sounds cannot always be performed against a fixed sequence experienced identically by every player. Games such as Alien: Isolation require movement and environmental sounds that respond dynamically to player behaviour. Footsteps and interactions need variation, while sounds must remain convincing when triggered in different sequences and contexts. The direct relationship between performer and fixed picture becomes a system of potential relationships between recorded material and player action.

    Embodied performance remains central even within that nonlinear structure. Weight, texture and physical credibility still matter. Repetition must be controlled so that repeated actions do not expose the limited number of recordings behind them. Direction and distance gain importance as sounds operate within spatial environments. The player may determine when a footstep occurs, while the Foley artist still determines what that step communicates.

    Changes in production systems place pressure on this kind of embodied practice. Swanscott described compressed schedules alongside continuing expectations for highly detailed soundtracks. Productions that once allowed longer periods for Foley may now expect comparable quality in fewer days. Outsourcing and remote working can reduce the close feedback once shared between Foley stages and other post-production departments, while sound libraries provide increasingly convenient alternatives for some categories of material.

    A library effect and a Foley performance differ most fundamentally in their relationship to a particular moment. Foley is created for a specific body, action and dramatic context. Pace, force, rhythm and texture can change in direct response to the image. Libraries can contain exceptional recordings, but the relationship between a pre-existing sound and a new movement must be constructed afterwards. Foley creates that relationship through performance.

    Technology continues to change how this embodied knowledge is captured and used. Digital workflows have transformed recording and editing. New microphones can reveal forms of vibration and underwater sound that conventional recording methods may miss. Synthetic imagery creates opportunities for invention, while games require nonlinear systems rather than fixed sequences. Swanscott’s practice has adapted continually to these changes. At its centre remains a physical act: watching movement, understanding its dramatic purpose and finding a way to perform its sound.

    Preserving the profession therefore involves more than documenting a catalogue of techniques. A list can identify suitable shoes, useful surfaces and familiar prop substitutions, but it cannot easily communicate the embodied timing required to perform another person’s movement or the instinct that allows an artist to hear a discarded object and recognise its potential. Swanscott emphasised the importance of mentorship and training within a craft learned through watching, listening, trying, failing and gradually developing a relationship with materials.

    Perhaps this is the most revealing way to understand Foley. Spectacular examples naturally attract attention: vegetables become broken bones, boxing gloves become gorillas and an overturned water tank becomes a lifeboat. Yet the deeper skill lies in the decisions connecting those materials to the image. A watermelon has no inherent cinematic meaning. Its value emerges through how it is performed, where the microphone is placed, which qualities of its sound are useful and whether the resulting texture belongs within the world of the scene.

    Swanscott’s lecture consequently revealed a discipline of transformation. Fabric becomes bodily movement. Shoes become character. Small objects become enormous mechanisms. Familiar materials give physical credibility to creatures and technologies that have never existed. A recording space becomes part of a fictional environment, while a performer standing on a Foley stage inhabits the movement of somebody in another place, another period or an entirely imaginary world.

    Finished Foley conceals this constant process of interpretation. Audiences see a character walk and hear footsteps. They see clothing move and hear fabric. They watch an object fall and accept its weight. The connection appears inevitable, even though every sound may have required decisions about material, performance, surface, timing, perspective and dramatic emphasis. As the performance becomes more convincing, the labour required to create that relationship becomes less visible.

    By the end of the lecture, Swanscott had transformed the Foley stage from a room full of shoes, fabrics, surfaces and strange objects into something closer to a physical imagination of the screen. Nothing in the room has only one identity. A boxing glove can become a gorilla. A paving slab can become monumental architecture. A discarded sink can acquire a new purpose through its resonance. The artist’s task is to watch movement, understand character and hear possibilities hidden inside ordinary materials. Foley gives bodies weight, objects substance and imaginary places a physical life. When the work succeeds, audiences do not hear somebody performing a world in a studio. They simply believe that the world was always there.

  • How Do You Make ADR Sound Like It Was Never Replaced? Paul Carden and Chris Navarro on Performance, Technology, and the Art of Dialogue Replacement

    Paul Carden and Chris Navarro

    How do you make ADR sound like it was never replaced?

    A line of dialogue may last only a few seconds, yet replacing it convincingly can require an extraordinary combination of preparation, performance, technology and judgement. The original production recording must first be identified as unusable, the replacement carefully documented and prepared, the actor returned to the emotional and physical circumstances of a performance recorded months earlier, and the new dialogue captured so that it matches the timing, vocal quality, microphone perspective and acoustic character of the original scene. If the process succeeds, the audience should never know that any of this work happened. During their joint online guest lecture for Edinburgh Napier University, ADR supervisor Paul Carden and ADR mixer Chris Navarro demonstrated this complete process from beginning to end. Rather than discussing Automated Dialogue Replacement only in theory, they created a deliberately compromised line of production dialogue, prepared it for replacement, recorded it on an ADR stage and evaluated the result. Their demonstration revealed a process in which meticulous preparation and technical fluency serve a deceptively simple objective: allowing everyone involved to concentrate upon the performance.

    Carden began not in a recording studio, but outside beside a car. His objective was to demonstrate how an apparently simple piece of dialogue could become unusable during production. Wearing a lavalier microphone, he performed the five-word line “I’m late for work” while getting into the vehicle. The exercise immediately revealed how vulnerable production dialogue can be. Clothing could obscure the microphone, a zipped jacket could change the sound, movement could complicate recording and the closing car door could mask the line itself. The example was deliberately simple. Carden was standing almost still, knew exactly what he was going to say and had constructed the situation specifically for the demonstration. On a real production, actors may be walking, running, interacting with objects and performing emotionally demanding scenes while the production sound team manages multiple radio microphones and one or more booms. From this perspective, the surprise is often not that some lines require replacement, but that so much production dialogue survives at all.

    The demonstration also established one of the central tensions within ADR. Production dialogue is not simply speech recorded on location. It contains a performance created at a particular moment, within a particular physical environment, as part of an interaction with other performers. Months later, an actor may arrive on an ADR stage while already immersed in an entirely different project and be asked to recreate a few seconds from that earlier performance. The recording environment offers little of the original context. The actor stands in a comparatively neutral room, watches an image on a screen, listens for cues and attempts to reproduce not merely the words, but the emotional and physical conditions under which those words were originally spoken. Carden emphasised that this is one reason actors can find ADR difficult. Matching a line involves returning to a performance that may no longer feel immediate or familiar.

    Before any actor reaches the stage, however, the line must become an ADR cue. Carden demonstrated this preparation process by locating the damaged dialogue within the picture, defining the cue and documenting the reason for replacement. The cue was assigned a unique identifier, linked to its scene information and accompanied by notes explaining that the car door had closed over the line. He also demonstrated the value of professional flexibility. Where a production line was imperfect but potentially usable, he might identify the replacement as optional rather than forcing an unnecessary argument over whether it must be replaced. The cue sheet then carried the information required by the recording stage, including project details, version information, character and actor names, cue number, timecode and recording information. A five-word performance therefore arrived at the stage supported by an extensive information system designed to ensure that everybody was working on the correct material.

    Version control was particularly important. Carden described production teams as living and dying by version dates, reflecting the practical reality that picture changes can quickly make carefully prepared cues inaccurate. Cue numbering performs a similarly essential role. Each take is voice-slated so that the recording itself retains its identity even if paperwork becomes separated from the audio. These details may appear administrative when viewed from outside professional production, yet they protect the continuity of the entire process. An ADR session may contain hundreds of replacement lines recorded across multiple days, actors and facilities. The actor should not have to think about whether the cue is correctly identified or whether the stage is working from the right picture version. Preparation creates the stability within which performance can happen.

    Carden also offered one deceptively simple instruction for aspiring ADR mixers: keep recording until somebody clearly indicates that the take has ended. An actor may continue the performance, a director may decide to record another version immediately or the session may move into wild recording without a formal interruption. Stopping too early can lose useful material for no meaningful benefit. Storage is cheap; an unrecoverable performance is not. The advice reflected a wider principle that would continue through the lecture. Good ADR practice depends upon remaining attentive to what is happening in the room rather than allowing the machinery of recording to dictate the session.

    The second stage of the demonstration moved into Navarro’s ADR facility. Even though the project consisted of a single line created for the lecture, he established the session as though it were a conventional production. Folder structures, documents, Pro Tools sessions and media were organised according to the same principles he would use for a project involving one session or hundreds. His ADR template was already configured for the ordinary demands of recording, while remaining adaptable when unusual situations arose, such as multiple actors performing together with several microphones each. A good template does not prescribe every session in advance. It removes predictable technical work so that attention remains available for situations that cannot be predicted.

    Even the imported picture introduced a practical lesson. Navarro noticed that the system was responding sluggishly and identified the compressed H.264 picture as a likely cause, since the computer had to decode it continuously during playback. The moment was minor, but revealing. Professional technical fluency often appears not through dramatic troubleshooting, but through the ability to recognise quickly why a system is behaving unexpectedly and continue without allowing the problem to dominate the room. Earlier, Carden had offered similarly pragmatic advice about computer failure: save the work, restart the system and continue rather than allowing panic to consume valuable time. Both speakers treated technical expertise as calm familiarity rather than technological display.

    Microphone selection then revealed how ADR matching differs from conventional voice recording. The objective is not to capture the most beautiful possible version of the actor’s voice. It is to create a recording that can inhabit the existing production soundtrack without attracting attention. Navarro showed that the stage had previously been configured for voice-over work, with microphones positioned closely and directly to produce the clear, present sound required for narration. ADR demanded a different approach. Carden had recorded the original line using a lavalier, so the replacement needed to reproduce the qualities of that production perspective rather than simply offering a technically superior recording.

    Carden explained that ADR sessions commonly record both boom and lavalier microphones, ideally using the same models employed during production. Even when the original line appears to come predominantly from one microphone, alternatives can prove unexpectedly useful. Navarro demonstrated this by positioning a shotgun microphone close to Carden but substantially off-axis. Pointing it directly at the performer would have produced a cleaner and more extended sound, but that was not necessarily desirable. The off-axis position rejected particular frequencies and created a tonal character that could potentially sit closer to the lavalier recording. The result was counterintuitive: the boom microphone could, under those conditions, sound more like the production lavalier than the replacement lavalier itself. Matching therefore began not with assumptions about microphone categories, but with listening.

    This distinction between technical quality and contextual suitability runs throughout professional ADR. A beautifully recorded line can fail if it sounds too clean, too close, too rich or too controlled for the image and surrounding production dialogue. Conversely, a microphone position that would appear unconventional in another recording context may provide exactly the spectral character required for a convincing match. The mixer must understand microphones well enough to use their imperfections deliberately. The question is never simply which microphone sounds best. It is which recording can become part of the scene without revealing the process that created it.

    Once Carden stepped in front of the microphone, the demonstration moved from equipment towards performance. The familiar system of three audible beeps established the timing, with the actor beginning the line where an imaginary fourth beep would occur. In principle, the task appeared straightforward. In practice, Carden repeatedly entered early, became self-conscious about the timing and discovered how quickly an apparently trivial five-word line could become difficult once performance, synchronisation and technical awareness competed for attention. His reaction gave the students an unusually useful demonstration of the psychological demands placed upon actors during ADR. Knowing the line is not enough. Understanding the timing is not enough. Even achieving synchronisation is not enough. The replacement still has to sound as though it belongs to the original performance.

    At one point, Carden produced a take that was synchronised correctly but immediately recognised that the performance itself was wrong. Navarro’s response was subtle. He did not ask for shouting or a dramatically louder delivery. He suggested only slightly more projection. The difference was not simply level. A small change in the physical production of the voice altered its tone and gave the line the additional presence heard in the original recording. When the replacement was compared directly with production dialogue, the improvement became obvious. The exercise demonstrated that ADR matching cannot be reduced to waveform alignment, pitch correction or equalisation. Performance changes the spectrum of the voice before any microphone or processor becomes involved.

    Navarro later developed this point in greater depth. Two performances can have similar apparent loudness and pitch while differing substantially in vocal quality. An actor performing naturally on set may have a relaxed throat and a particular physical relationship with the surrounding scene. On the ADR stage, tension, self-consciousness or the effort to satisfy technical instructions can change the voice. A performer may become tighter, brighter or less natural even while reproducing the words and timing accurately. Matching therefore requires attention to qualities that are difficult to describe numerically. Projection, pitch, volume and rhythm matter, but so do muscular relaxation, breath and the physical origin of the voice.

    This presents the ADR mixer with a delicate problem. The mixer may hear precisely what is preventing a line from matching, yet communicating every technical observation to the actor may make the performance worse. Navarro warned that performers can absorb only so many notes before they begin thinking about the mechanics of speech rather than the character. A useful intervention must therefore translate technical listening into language that supports performance. Sometimes the correct decision is to offer a suggestion. Sometimes it is to communicate through the director. Sometimes it is to recognise that an imperfection is not important enough to justify disturbing the creative balance of the room. Hearing a problem and knowing whether to mention it are separate professional skills.

    The lecture also demonstrated an older approach to ADR recording that remains remarkably useful. Navarro sampled the original production line and played it repeatedly into Carden’s headphones. Instead of concentrating simultaneously upon picture, beeps and performance, Carden could hear the original line and immediately reproduce it, repeating the process several times. Navarro then edited the resulting takes into position and compared them with the production recording. The method offered a direct reference for timing, intonation, volume and vocal character, allowing the performer to respond to the sound of the original performance rather than attempting to reconstruct every detail intellectually.

    Navarro’s explanation of this technique revealed an important insight into perception. Picture can sometimes become a distraction. An actor watching for a mouth movement may wait until it becomes visually apparent, by which time the correct moment to begin has already passed. If the original audio is correctly synchronised to picture, matching its rhythm and timing can naturally reproduce picture sync. The performer can therefore concentrate upon hearing and responding rather than continually monitoring several streams of information at once. Despite the age of the sampling technique, Navarro regarded it as one of the most useful tools available on the stage. Its continued value comes not from technological sophistication, but from the way it simplifies the performer’s task.

    The same objective shaped Navarro’s control room. His system contained extensive routing, multiple microphone inputs, separate monitoring paths and a heavily customised control surface, yet the purpose of this complexity was to make the session itself feel simple. He might need to manage separate mixes for the control room, recording stage, actor headphones, supervisor headphones and remote participants, with each requiring different material during rehearsal, recording and playback. Attempting to reconfigure every route manually between passes would slow the session and create repeated opportunities for error. Navarro therefore designed systems that allowed complex changes to happen immediately.

    His customised keypad provided a particularly revealing example. Individual buttons could trigger sequences of macros that armed tracks, initiated recording and performed repetitive editing operations once a take had finished. Navarro had spent months developing the core system and continued refining it whenever he noticed himself repeating unnecessary keyboard operations. He had begun his career as an ADR recordist and understood the division of labour on a two-person stage, where one person could manage recordings and files while the mixer concentrated upon the performers. Working alone, he used automation to reproduce some of that support, delegating repetitive technical operations to macros so that his attention could remain directed towards the stage.

    This led to one of Navarro’s most important ideas: the ADR mixer must work at the speed of creativity. An actor or director may suddenly discover a new approach to a line and want to record it immediately. If the mixer responds by asking everyone to wait while tracks are configured, routes changed or files prepared, the idea may lose its immediacy. A delay of only a few seconds can alter the atmosphere of a performance. Technical speed therefore has a human purpose. The mixer learns the system thoroughly enough that machinery does not interrupt thought.

    Navarro connected this principle to advice he had received about respected ADR mixer Tommy O’Connell. Asked what distinguished O’Connell’s work, a sound editor gave a simple answer: he anticipates. Useful actions are completed before anyone needs to request them. Navarro interpreted this not as a mysterious talent, but as the result of attention. If the mixer is watching the stage, listening to conversations and understanding the direction in which a session is moving, preparation for the next action can begin before a formal instruction arrives. Anticipation therefore depends upon both technical readiness and social awareness. A mixer whose attention is buried in the workstation may complete every requested task correctly while still remaining one step behind the session.

    Carden’s preparation of the cue and Navarro’s management of the stage reveal two sides of the same professional process. Carden reduces uncertainty before the session begins through accurate cueing, documentation, version control and communication. Navarro reduces friction during the session through templates, routing, automation and anticipation. Neither form of preparation is intended to make the process more rigid. Both preserve the possibility that an actor, director or mixer can respond immediately when something unexpected and valuable occurs.

    Microphone signal flow offered another example of this balance. Navarro described a deliberately clean recording path from microphone through preamplifier into Pro Tools, avoiding unnecessary outboard processing. Within the workstation, compression and equalisation could be used sparingly, but he warned against making irreversible decisions without good reason. Recording completely flat preserves maximum flexibility for the re-recording mixer, while careful decisions made during the session can still produce useful, committed tracks. Heavy processing may be difficult to undo later. A technically impressive recording decision is not valuable if it reduces the options available to the production.

    Navarro’s session was configured to make microphone comparison immediate. Several microphone inputs could remain available, feeding dedicated record tracks at appropriate levels. A boom and lavalier could be recorded simultaneously as discrete channels, preserving both perspectives for later evaluation. The workflow reflected the uncertainty inherent in matching. The microphone expected to work best may not produce the most convincing result once the line is placed against production. Recording alternatives gives later editors and mixers material from which to construct the most believable transition.

    His discussion of recording format was equally pragmatic. For conventional ADR, he worked at the established production standard of 48 kHz, 24-bit audio. Higher sample rates could be valuable for sound effects intended for extensive manipulation, where recordings might later be slowed or processed heavily. ADR serves a different purpose. Recording at unnecessarily high rates would increase storage requirements before the material was eventually converted to the format used by the production. The appropriate technical choice depends upon what will happen to the sound. More data is not automatically more useful.

    As the lecture progressed, it became increasingly clear that the most difficult parts of ADR were not contained within any equipment specification. Navarro estimated that learning the technical fundamentals and becoming comfortable with Pro Tools took years, yet he distinguished that competence from the broader ability to run a session. Recording dialogue in sync with picture is ultimately a technical process that can be learned. Managing actors, directors, supervisors, performance, uncertainty and the emotional energy of the room represents a different level of expertise.

    Actors may arrive nervous. Directors may be highly collaborative, completely self-sufficient or resistant to suggestions. A performer may struggle with a line while becoming increasingly aware of the difficulty. The mixer must understand how much intervention the situation can support. Navarro described the importance of creating a calm and comfortable environment, particularly when clients are unfamiliar with the process. Confidence can be communicated without dominance. If someone is uncertain, the mixer can explain what will happen, answer questions and demonstrate that the session is under control. Technical authority becomes useful when it reduces anxiety rather than displaying superiority.

    Carden and Navarro also acknowledged the professional judgement involved in deciding when not to intervene. A mixer may recognise an aspect of a performance that could be improved, yet nobody has asked for technical input and the issue may not be important enough to disrupt the session. Navarro framed this as a balancing act. Would the intervention materially improve the line? Is the director open to suggestions? Will another note distract the actor from a performance that is already working? Expertise includes recognising problems, while professional maturity requires distinguishing consequential problems from imperfections that do not matter.

    The ADR stage therefore brings technical precision and emotional sensitivity into the same process. The mixer must hear minute changes in vocal quality, understand microphone behaviour, maintain synchronisation, manage several monitoring environments and operate the recording system almost instinctively. At the same time, attention must remain on body language, conversation, confidence, frustration and creative momentum. Mastery of the technology is essential precisely so that the technology no longer consumes the attention needed elsewhere.

    Carden’s demonstration also placed the scale of professional ADR into perspective. Their five-word line required production recording, cue preparation, documentation, stage setup, microphone selection, rehearsal, multiple takes, alternative recording methods, editing and comparison. A skilled actor might complete around ten or twelve conventional cues in an hour, while a feature film may contain 150 or 200 ADR lines. Complicated performances naturally take longer. What appears to an audience as a few moments of seamless dialogue can therefore represent days of concentrated work.

    Yet the success of that work is measured largely through its invisibility. When Navarro compared Carden’s replacement line with the production recording, the evaluation concerned relationships: volume, tone, vocal quality, perspective and the way the replacement interacted with the surrounding scene. The door slam that had originally damaged the line was restored around the new performance, returning the replacement to the event from which it had temporarily been separated. Heard alone, the ADR recording was merely a voice on a stage. Placed back into the scene, it became part of an action.

    This transformation captures the philosophy shared by both speakers. ADR is often described as the replacement of unusable dialogue, but their demonstration showed that replacement is only the beginning of the problem. The objective is reconstruction. The actor reconstructs a performance. The mixer reconstructs a microphone perspective. Editors reconstruct timing and continuity. The final soundtrack reconstructs the relationship between voice, action and environment so successfully that the audience experiences a single uninterrupted moment.

    The joint nature of the lecture made this especially clear. Carden approached ADR through the complete production process, showing how a line travels from location problem to documented cue and finally to the recording stage. Navarro approached the same process from inside the control room, revealing the technical systems, listening skills and interpersonal judgement required to capture a convincing replacement. Their perspectives met at the point where professional preparation serves human performance.

    By the end of the session, the deliberately damaged line had become something much larger than a technical demonstration. It revealed why ADR demands more than synchronisation, why the cleanest microphone is not always the correct microphone, why a performer can match timing while missing the voice, why an old sampler can remain useful in a modern digital workflow and why a highly automated control room can make a session feel more human rather than less. Above all, Carden and Navarro showed that successful ADR depends upon where attention is directed. The actor should be thinking about performance rather than machinery. The director should be thinking about the scene rather than routing. The mixer should be watching and listening to the room rather than fighting the workstation. The audience, finally, should be thinking about none of these things. When every part of the process works together, they simply hear a character speak.

  • How Do You Mix Television Sound Under Pressure? Frank Morrone on Dialogue, Workflow, and the Art of Re-Recording

    Frank Morrone

    How do you mix television sound under pressure?

    A television soundtrack may contain hundreds of dialogue recordings, sound effects, Foley performances, backgrounds, ADR takes and music stems, all competing for space within a mix that must remain clear, emotionally convincing and technically suitable for broadcast. The audience should never become aware of that complexity. They should simply understand every line, believe every environment and remain absorbed in the story. During his online guest lecture for Edinburgh Napier University, re-recording mixer Frank Morrone explored the craft behind achieving that apparent simplicity. Drawing upon a career in film and television that began in 1979, and projects including Lost, The Strain, Sleepy Hollow and Criminal Minds: Beyond Borders, he revealed a discipline shaped equally by technology, organisation, collaboration and judgement. Throughout the session, one principle emerged repeatedly. The most effective mixing workflows allow enormous technical complexity to disappear behind the story.

    Morrone began by tracing a career that developed across several different areas of professional audio. His earliest work took place in music studios, recording jazz and orchestral film scores before following those recordings into the dubbing theatre and becoming increasingly interested in the way complete soundtracks were assembled. A move into post-production allowed him to work across dialogue editing, music editing and Foley recording before concentrating upon re-recording mixing. That breadth of experience shaped the collaborative philosophy running throughout the lecture. Unlike a music recording session, where one engineer may remain closely involved from recording through to the final mix, film and television sound brings together work created by many different specialists. The dub stage is where those contributions finally meet. Successful mixing therefore depends upon understanding not only the material itself, but the people, processes and decisions that produced it.

    Television makes this collaboration particularly demanding. Morrone described an industry in which track counts have continued to increase while schedules and budgets have become progressively tighter. Sophisticated surround mixes must be created rapidly, and emerging formats add further complexity without removing the need to support conventional playback systems. His response is not simply to work faster. It is to design workflows that remove unnecessary decisions from the mixing stage. As soon as he joins a project, he communicates with the supervising sound editor about track requirements and provides a starting template so that incoming material already fits an established structure. Organisation begins before the mixer enters the room. Under severe time pressure, the ability to find, control and compare material immediately becomes part of the creative process itself.

    This reveals something deeper about Morrone’s understanding of expertise. His templates, early conversations with sound editors, knowledge of production microphones, preparation of alternative takes and habit of printing completed passes all anticipate problems before they are allowed to interrupt the mix. The same thinking extends beyond the dubbing theatre. He considers how broadcast processing will react to dynamics, how a mix will translate to domestic systems and how material created for one format will behave when heard through another. Professional experience, in this sense, is not simply the ability to solve problems quickly. It is the ability to recognise where problems are likely to emerge and construct a workflow in which many of them have already been addressed before they become urgent.

    The scale of Lost provided a striking illustration. Morrone showed the students sessions containing extraordinary numbers of elements, including dozens of tracks dedicated to the Smoke Monster alone. Its identity emerged from a deliberately ambiguous combination of animal voices, pneumatic machinery, roller-coaster wheels and other contrasting sources, creating something that resisted being understood as either entirely organic or entirely mechanical. Those effects existed alongside hard effects, backgrounds, Foley, production dialogue, ADR, group recordings and substantial music deliveries. The technology available at the time imposed strict limits upon voices and processing, requiring careful decisions about resource allocation as well as creative balance. Complexity could not simply be solved by adding more processing. The session itself had to be organised so that the mixers could navigate it instinctively.

    Custom fader layouts and VCA groups became essential to that process. Morrone described arranging controls so that principal dialogue could immediately be balanced against ADR, group recordings and music, while different categories of material remained independently accessible. Music sources could be separated from score, while dialogue in different languages could be isolated for international deliverables. Every layer of organisation reduced the time between hearing a problem and solving it. This became especially important during pilot season, when mixers might receive unusually elaborate material without first having time to develop a workflow around the programme. A sufficiently flexible template must already be capable of accommodating whatever arrives. Preparation, in this context, creates the conditions in which creative decisions can still be made under pressure.

    The production of Lost also demonstrated the tension between creative ambition and delivery requirements. Morrone recalled the exceptional resources devoted to the programme, including a large pilot budget and Michael Giacchino’s insistence upon recording a live orchestra for each episode. Yet the soundtrack still had to survive the restrictions of television broadcast. The team therefore created a more dynamic version for DVD before producing a contained broadcast mix designed to survive transmission processing. A similar approach was later adopted on The Strain. The distinction mattered. A mix can remain technically within specification and still behave poorly when subsequent broadcast processing responds to excessive dynamics. Re-recording therefore requires mixers to think beyond the dubbing theatre. They are mixing for every system through which the programme will eventually reach its audience.

    Dialogue occupied the centre of Morrone’s approach. His first objective is always to preserve the production performance wherever possible. When ADR has been recorded, he wants to know why. A line replaced for performance reasons presents a different problem from one replaced to solve a technical fault. If the director wanted a different performance, Morrone respects that decision while keeping the original available as an alternative. If the problem was technical, he first explores whether the production recording can be repaired. Modern restoration tools have dramatically expanded what can be rescued, reducing the need to replace performances that may possess subtleties difficult to recreate months later in an ADR studio.

    When ADR is necessary, matching involves far more than applying equalisation and reverb. Morrone keeps production dialogue available alongside the replacement so that he can compare transitions directly, matching tone, acoustic environment, pacing and performance. He prefers to receive several strong takes and recordings from both boom and lavalier microphones, recognising that the microphone apparently closest to the production perspective is not always the easiest to integrate. Understanding which microphones and wireless systems were used during production can also provide valuable clues, particularly where transmission systems have imparted their own sonic characteristics. ADR matching consequently becomes a process of reconstructing relationships rather than searching for a single corrective setting.

    Technology has transformed that work, though Morrone repeatedly warned against allowing powerful restoration tools to encourage excessive processing. Noise reduction, spectral repair, ambience matching, EQ matching and dereverberation can rescue material that would previously have required replacement. Yet he deliberately removes less noise than might appear necessary when dialogue is heard in isolation. Once backgrounds, effects and music return, much of the remaining noise may be perceptually masked. Processing that sounds impressively clean in solo can leave dialogue lifeless and constricted in the finished mix. Morrone therefore keeps copies of original material and sometimes returns to less processed versions during the final mix. The objective is not the cleanest possible dialogue track. It is dialogue that remains natural and convincing within the complete soundtrack.

    His attitude towards restoration reveals a broader philosophy of technology. Morrone began working with magnetic tape, a Cat 43 noise reduction unit and a notch filter, and he clearly values the extraordinary capabilities available to contemporary mixers. Yet greater technical power has not removed the need for judgement. In many respects, it has increased it. The ability to remove more noise does not mean that more noise should be removed, just as the ability to place sound almost anywhere within an immersive field does not mean that every available position should be used. New tools expand the range of possible decisions. They do not determine which decisions are appropriate. Throughout Morrone’s lecture, technical capability remained subordinate to perception.

    This distinction between isolation and context became one of the lecture’s most important ideas. Morrone described television mixing as beginning with dialogue and music, establishing the foundation against which the effects mixer can develop backgrounds and action. Once those elements come together, the mixers decide what should drive the scene. Some moments belong to music, others to effects, while still others require a more subjective perspective or a deliberate reduction of material. Experienced mixing partners develop an almost instinctive understanding of one another’s decisions. If an effect obscures a line during an early pass, Morrone knows that a trusted colleague will create space for it during refinement. Collaboration is therefore more than the division of tracks between two people. It is a shared process of deciding where the audience should listen.

    The idea that sounds must be judged in context reaches far beyond balancing dialogue against effects. Dialogue that appears too noisy when soloed may become entirely convincing once the world of the scene surrounds it. ADR that attracts attention after twenty repeated comparisons may pass unnoticed when the audience encounters it once within a continuous performance. Group recordings that sound absurd under isolated scrutiny may perform their role perfectly when placed at the correct distance behind principal dialogue. Morrone’s examples repeatedly challenged the assumption that individual elements should be perfected independently. The meaningful unit of evaluation is ultimately the audience’s experience of the scene. A soundtrack succeeds through relationships between sounds rather than the isolated perfection of its components.

    Normally, the priority within those relationships remains dialogue. Morrone described it as both the foundation of the mix and the principal carrier of storytelling. His concern with intelligibility, however, extends beyond maintaining a technical hierarchy between dialogue, music and effects. It is fundamentally about preserving attention. His most revealing test is not a meter reading but a listener asking what somebody has just said. At that moment, the audience member has been pulled out of the story and required to think about the failure of the soundtrack. Clear dialogue therefore supports immersion precisely by avoiding attention to itself. The better the mix communicates, the less the audience needs to think about the process of communication.

    Yet the rule is not absolute. For one sequence in The Family, a child was being used as bait to attract a kidnapper in a crowded shopping centre. The environment needed to feel genuinely busy, yet the available collection of separate crowd recordings and reverberant elements never created a convincing whole. Morrone took a six-channel recorder into a real shopping centre, captured the food court from different perspectives and brought the recordings back into the mix. The result allowed the environment to crowd the dialogue slightly, which was precisely what the scene required. Clarity remains fundamental, but realism sometimes depends upon controlled difficulty.

    The shopping-centre sequence illustrates an important tension within Morrone’s approach. His practice is built upon strong principles, but those principles do not become inflexible rules. Dialogue normally takes priority, except when allowing the environment to interfere with it makes the dramatic situation more believable. Restoration should preserve intelligibility, except when excessive cleaning destroys naturalness. Acoustic simulation is valuable, except when recording the real physical relationship between people and space produces a more convincing result. Expertise therefore involves knowing the rules well enough to understand what they protect, then recognising the moments when the needs of the scene justify bending them.

    This willingness to leave the dubbing theatre and record real spaces appeared repeatedly throughout the session. Morrone discussed impulse responses, convolution reverbs and carefully developed presets for rooms, vehicles and other environments, all of which provide useful starting points for worldizing sound. Yet he remained pragmatic about their limitations. A real location does not necessarily sound like the space suggested by the finished image, and even a carefully captured impulse response may require additional reflection, delay or reverberation before it feels convincing. Cars present an especially difficult problem, combining strong early reflections from glass with highly absorptive surfaces elsewhere. Experience gradually produces a library of useful starting points, though listening still determines the final result.

    Sometimes no simulation is as convincing as returning to the physical situation itself. Morrone described a scene in which a group of foster children were supposed to be creating chaos upstairs while a conversation took place below. Studio-recorded group voices did not reproduce the peculiar combination of footfalls, structural transmission and reflections travelling down a staircase into another room. His solution was direct. He gathered children on the upper floor of a house, recorded from downstairs and captured the complete acoustic event as it occurred. No increasingly elaborate chain of processing was required. The physical relationship between performers, building and microphone provided what the scene needed.

    The pressures of television production make such judgement particularly important. Morrone compared television mixing to boot camp. Schedules leave little room for hesitation, and mixers must develop workflows capable of producing strong results quickly. Once a scene has been successfully mixed, he prefers to print it rather than trusting that automation will remain untouched throughout later work. Accidental writes, changed sends and technical errors can occur even in experienced hands. Printing completed work provides security and allows later changes to be punched into established stems. Efficiency does not mean rushing blindly. It means reducing the opportunities for avoidable problems to consume the limited time available.

    Long sessions also introduce a more human limitation: hearing fatigue. Morrone described working on the highly dynamic soundtrack of The Strain, where sustained exposure to loud material forced him to think deliberately about auditory recovery. His solution was simple but important. He left the room for short periods, walked and allowed his ears to recover while his mixing partner continued working. The two mixers could alternate demanding passes, giving each other opportunities for rest without stopping the session. After decades of mixing, one of the useful discoveries was not another plug-in or processor, but the value of leaving the chair for ten minutes. Professional listening depends upon recognising the limits of the listener.

    Morrone’s discussion of client relationships revealed another dimension of the re-recording mix that students may rarely encounter in technical demonstrations. Mixers are working with directors, producers and other clients who may have strong preferences that differ from their own. Morrone described situations in which clients wanted music loud enough to compete with dialogue. His responsibility was to explain the likely consequences, demonstrate how the mix translated at lower levels and on smaller monitors, and search for a compromise that preserved the client’s intention while protecting intelligibility as far as possible. The mixer offers expertise, but does not own the programme. Knowing which decisions are worth challenging and which require accommodation is part of the craft.

    His description of deciding whether an issue represents a “hill to die on” reveals a sophisticated understanding of professional authority. Expertise does not give the mixer unlimited control over the work, nor does collaboration require the abandonment of professional judgement. The mixer must advocate for the audience, explain likely consequences and make alternatives audible, while recognising that the final creative intention belongs to the client. Professional judgement therefore includes negotiation. Sometimes expertise means defending a decision. Sometimes it means finding a compromise that neither side initially imagined. Occasionally, it means implementing a choice that remains contrary to personal taste while ensuring that it works as successfully as possible.

    Technical fluency plays an important role in maintaining those relationships. When a client requests a change, Morrone wants to make it immediately, play it once and continue. Searching through tracks or repeatedly troubleshooting a familiar process changes the atmosphere of the room and interrupts attention to the programme itself. A well-designed session keeps the conversation focused upon storytelling and communication. The deeper the mixer’s command of the tools, routing and session layout, the less those systems intrude upon the creative discussion.

    This is another form of transparency. Morrone repeatedly returned to the idea that technology should disappear from the client’s experience, yet transparency does not mean that technology has become unimportant. The opposite is closer to the truth. Considerable technical knowledge is required to make complex systems feel immediate. Templates, custom layouts, routing, monitoring, printed stems and intimate familiarity with the workstation create an environment in which a creative request can become an audible result without breaking the flow of the session. Mastery becomes visible through the absence of friction.

    The growth of immersive sound has expanded this challenge further. Morrone described mixers as working between extremes, from Dolby Atmos and sophisticated home theatres to stereo playback and mobile devices with earbuds. His philosophy was to begin with the best mix possible in the most capable format, then ensure that it translates successfully into simpler ones. Immersive technology may offer extraordinary spatial possibilities, but Morrone remained cautious about using novelty without considering perception. In particular, he defended dialogue remaining anchored to the centre. Moving voices between speakers can alter timbre and create distracting changes that audiences may notice without understanding their source. New formats create possibilities, though they do not invalidate principles developed through decades of listening.

    His argument about dialogue placement is particularly revealing. Audiences do not need to identify the technical source of a problem in order to experience discomfort. Morrone described viewers sensing that something was wrong when dialogue moved around an immersive field, even when they could not explain precisely what disturbed them. This places an unusual responsibility upon the mixer. Professional listening must sometimes diagnose experiences that ordinary listeners can feel but cannot name. The purpose of expertise is not to dismiss those responses as technically uninformed, but to understand the perceptual conditions that produced them.

    The same concern for translation shaped his approach to low frequencies. Subwoofers vary enormously between listening environments, and domestic listeners frequently adjust them far beyond calibrated levels. Morrone therefore warned against depending entirely upon the LFE channel for the weight of a soundtrack. Low-frequency energy can also be carried through the main channels, creating a result that remains powerful across a wider range of playback systems. His experience of hearing a domestic subwoofer struggle with low-frequency material from Sleepy Hollow reinforced the point. A soundtrack must survive real listening environments, not merely sound impressive on a perfectly calibrated dubbing stage.

    As the discussion widened, Morrone considered the future possibility of mixes that adapt more intelligently to different devices and contexts. Streaming, immersive audio, virtual reality and personalised playback were already creating pressure for soundtracks to function across radically different systems. Yet his underlying philosophy remained remarkably consistent. Whatever the format, begin with the strongest possible mix, preserve the storytelling hierarchy and understand how human perception responds to the result. Technology changes quickly. The responsibility to guide attention and communicate narrative does not.

    Questions from the students returned the discussion to practical preparation. Morrone strongly supported editors delivering material that is already sensibly balanced before it reaches the stage. Dialogue editors who use clip gain to create consistent levels save valuable mixing time, while backgrounds and effects arriving close to useful operating levels allow the mixer to begin creatively rather than first correcting avoidable problems. He described requesting particular editors for demanding productions precisely for this reason. Good preparation is noticed. In a professional environment where a pilot may need to be mixed in only a few days, the person who consistently delivers well-organised, intelligently balanced material becomes someone mixers actively want on the next project.

    His discussion of ADR offered another deceptively simple lesson about perception. Morrone sometimes works on difficult replacement lines privately through headphones while the effects mixer is making a pass. The client then hears the finished line only once in context rather than listening to it repeated dozens of times during adjustment. Repetition directs attention towards the repair and teaches the listener exactly where to expect it. The same awareness shaped his humorous rule about never soloing loop group in front of a client. Background conversations that work perfectly as part of a scene may sound absurd when isolated and scrutinised. What matters is not whether every element survives examination on its own, but whether it performs its intended role in the scene.

    These examples reveal that Morrone is not simply mixing sound. He is managing attention, expectation and knowledge. Once listeners have been taught where an edit exists, they may hear it differently. Once an element has been isolated, they may judge it according to criteria that have little relevance to its actual purpose. Perception is shaped not only by acoustic information, but by what listeners have been encouraged to notice. Part of the mixer’s craft therefore lies in protecting the audience’s experience from unnecessary awareness of the mechanisms used to construct it.

    Yet group recording could also become a powerful storytelling tool. On Criminal Minds: Beyond Borders, episodes moved between international locations while much of the production remained based in Los Angeles. Carefully performed local-language group recordings, combined with music and other environmental elements, became essential to establishing each location convincingly. Here, group material could be brought forward rather than hidden. There was no universal rule governing how loudly an element should be mixed. Its appropriate level depended upon what the scene needed to communicate.

    By the end of the session, Morrone’s account of re-recording mixing had moved far beyond faders, plug-ins and delivery specifications. The technology matters enormously, as do templates, routing, restoration tools, monitoring and control surfaces, but those things serve a larger process. A mixer must understand performance, storytelling, perception, collaboration, translation and the subtle politics of working with clients under pressure. The session may contain hundreds of tracks, yet the audience should hear a coherent world rather than the complexity required to construct it. Morrone’s lecture revealed a craft built upon anticipation, contextual judgement and the careful management of attention. Preparation preserves the possibility of creativity under pressure. Technical knowledge allows technology to disappear from the conversation. Rules provide essential foundations, while experience reveals when the needs of a scene require them to bend. Great television sound is not created by making every element impressive or every recording perfect in isolation. It emerges from understanding what the audience needs to hear, recognising what they should never need to notice, and making hundreds of individual decisions feel like one continuous experience.

  • How Do You Build a Sonic Brand? Sean Thornton on Audio Branding, User Experience, and Designing with Sound

    Sean Thornton

    How do you build a sonic brand?

    Most people can recognise a familiar logo within a fraction of a second. Distinctive colours, typography and visual symbols allow organisations to establish an identity that remains remarkably consistent across products, advertising and digital services. Sound, however, has often been treated very differently. Many organisations commission an audio logo, perhaps a short mnemonic played at the end of an advertisement, and regard the task as complete. During his online guest lecture for Edinburgh Napier University, Sean Thornton challenged this assumption directly. Drawing upon more than a decade working in audio branding and as co-founder of Audio UX, he argued that effective sonic branding is not created through isolated assets but through coherent systems. Throughout the lecture, one idea repeatedly emerged. Sound should be treated as a complete design language rather than a collection of individual sounds.

    Thornton began by reflecting upon his own professional journey. Although he initially studied music with ambitions of becoming a film composer, early experiences within game audio and commercial music gradually led him towards a field that scarcely existed as a recognised discipline when he was a student. That progression illustrated one of the lecture’s broader themes. Creative careers rarely follow predictable paths. Formal education provides an important foundation, though many professional specialisms emerge only through experimentation, curiosity and the willingness to explore opportunities beyond conventional career routes. Audio branding, he suggested, represents precisely this kind of evolving discipline, combining composition, sound design, psychology, branding and user experience into a field that continues to redefine itself.

    Before discussing design methods, Thornton carefully distinguished between several closely related concepts. Audio branding, in its simplest form, involves the strategic use of sound to establish or reinforce identity. Yet he argued that this definition captures only part of the picture. Organisations increasingly communicate across websites, mobile applications, physical products, voice assistants, advertising, podcasts and public spaces. Users encounter brands through countless interactions rather than through a single advertisement or logo. If every one of those interactions produces unrelated sounds, opportunities to build familiarity and trust are quickly lost. Holistic audio branding therefore asks a different question. Instead of considering how an individual sound performs in isolation, it asks how every auditory experience contributes towards a coherent and recognisable identity. The objective is not consistency through repetition, but consistency through design.

    This broader perspective naturally led to Thornton’s discussion of Audio User Experience, or Audio UX, which underpins his company’s approach. The distinction is subtle yet important. Audio branding concerns the sounds themselves. Audio UX concerns how people experience those sounds. A perfectly crafted sonic identity still fails if it frustrates, distracts or overwhelms the people who encounter it. Throughout the lecture, Thornton repeatedly encouraged students to approach sound through empathy before aesthetics. Designers should not begin by asking whether a sound expresses the personality of a brand. They should first ask whether it genuinely improves the experience of the person hearing it. Successful sonic branding therefore begins with people rather than brands. When the user experience has been designed well, brand recognition becomes a natural consequence rather than the primary objective.

    This philosophy becomes much clearer when viewed alongside visual identity design. Few organisations would expect a designer to create only a logo while ignoring colours, typography, photography and layout. Successful visual identities are built from systems of related elements that can be applied consistently across many different contexts. Thornton argued that sound should be approached in exactly the same way. Instead of delivering a handful of isolated assets, audio designers should create flexible frameworks capable of supporting products, advertising, digital interfaces, physical environments and future developments that may not yet exist. Individual sounds remain important, though their real value lies in how they work together. Sonic branding therefore becomes less about composing memorable cues and more about designing a language that other designers can continue to use.

    One of the lecture’s most distinctive ideas concerned the difference between creating assets and creating design systems. Traditional projects often revolve around a short audio logo, a notification sound or a piece of advertising music delivered as a finished product. Thornton argued that this approach solves only today’s problem. Instead, he proposed building reusable sonic building blocks from which many future experiences could emerge. Characteristic instrumental colours, recurring melodic gestures, harmonic language, rhythmic behaviour and carefully selected textures become components within a wider design framework rather than fixed compositions. The designer is no longer delivering a finished collection of sounds. They are creating the vocabulary from which an organisation’s sonic identity can continue to grow.

    Thornton illustrated this philosophy through the development of a comprehensive sonic identity for USA Today. Rather than beginning with musical preferences or fashionable production styles, the project started by asking a more fundamental question. What should one of America’s largest news organisations sound like? The resulting concept centred upon the idea of a sonic mosaic that reflected the diversity of voices, cultures and experiences represented within the publication. Recordings of instruments associated with different regions of the United States were combined with contributions from a wide variety of people, gradually creating a distinctive palette of sonic materials. These recordings were never intended simply to become finished pieces of music. They formed the foundation of a much larger design system from which future compositions, interface sounds, broadcast material and branded experiences could all develop while remaining recognisably connected. The case study demonstrated that successful audio branding begins long before composition. It begins by defining the ideas, relationships and design principles from which every subsequent sound will emerge.

    Perhaps the most forward-looking aspect of Thornton’s presentation concerned what happens after a sonic identity has been created. Visual brands are supported by style guides, governance and clear documentation that help maintain consistency across years of future development. Thornton argued that sonic identities require the same level of stewardship. Designers should not simply hand over a collection of audio files. They should provide the principles, frameworks and guidance that allow other teams to implement those sounds consistently across new products, services and technologies. Designing a sonic brand therefore involves far more than composing memorable sounds. It means creating a living design language capable of evolving without losing its identity. That broader challenge would become the focus of the remainder of the lecture, as Thornton explored how these systems are implemented, managed and refined over time.

    Having established the principles behind holistic audio branding, Thornton turned to the practical question that ultimately determines whether those ideas succeed. Designing a sonic identity is only the beginning. The greater challenge lies in ensuring that it remains useful, recognisable and adaptable as organisations evolve. Throughout the second half of the lecture, he argued that the true value of an audio brand emerges not from individual assets, but from the systems that allow those assets to be deployed intelligently across countless future experiences.

    The USA Today project provided a detailed illustration of this philosophy in practice. Rather than delivering a fixed collection of completed sounds, Thornton’s team developed a series of reusable audio components that could be assembled in different ways depending upon the context. This modular approach reflected the same design thinking used throughout modern visual branding. Designers rarely create a new colour palette or typeface for every campaign. Instead, they draw from a shared system that allows individual pieces of communication to remain distinctive while still feeling unmistakably connected. Thornton argued that sound should operate according to exactly the same principle. By working with carefully designed components instead of inflexible assets, organisations gain the freedom to create new experiences without continually reinventing their sonic identity. The result is a brand that remains recognisable while continuing to evolve alongside new products, technologies and audiences.

    One of the most revealing aspects of this system involved thinking about sound through the language of design rather than music. Thornton described key signatures as global design patterns capable of giving different pieces of audio a shared tonal identity. Characteristic textures functioned in much the same way as colours within a visual brand, while processed recordings became recognisable building blocks that could be reused in multiple contexts. Even small musical gestures could perform the same role as graphic motifs or visual icons, quietly reinforcing identity without demanding attention. This perspective shifted the discussion away from individual compositions towards relationships between sonic elements. Listeners may never consciously recognise these patterns, yet together they help establish familiarity across an increasingly diverse collection of products and services.

    The USA Today identity demonstrated how this philosophy could be implemented across very different media. A podcast introduction, a smart speaker news briefing and a broadcast sequence all required different durations, different pacing and different levels of branding. Treating each as an independent project would inevitably weaken consistency. Instead, every experience drew from the same underlying design language while adapting its presentation to suit the situation. Thornton described, for example, how a brief branded audio fingerprint proved more appropriate for voice assistants than a longer musical introduction. Listeners asking for the latest headlines wanted information quickly. A concise sonic reminder of the brand communicated identity without delaying access to the content. User experience and branding were therefore not competing priorities. Each strengthened the other.

    This naturally led to one of Thornton’s strongest practical messages. Designing an effective sonic identity is only half of the task. The identity must also be implemented consistently by the many people who will use it in the future. A beautifully designed audio logo achieves very little if nobody understands when it should be used, when silence would be more appropriate or how new material should relate to the existing system. Implementation therefore becomes a creative activity rather than an administrative one. Clear principles, documentation and governance allow organisations to grow their sonic identities confidently without gradually losing the coherence that made them distinctive in the first place.

    Voice itself formed another important component of this broader system. Thornton argued that organisations should think beyond selecting an appropriate voice actor. Voice also encompasses conversational style, language, pacing and the structure of spoken interactions. These decisions become particularly significant as brands increasingly communicate through voice assistants, automated services and accessibility technologies. Designing effective conversations therefore requires the same careful planning applied to music and sound design. Every spoken interaction contributes towards the wider experience of the brand, making conversation design an integral part of holistic audio branding rather than a separate discipline.

    Accessibility emerged as a recurring theme throughout this discussion. Thornton repeatedly returned to the idea that good sonic design should improve experiences for everyone rather than simply strengthening recognition. Voice interfaces, carefully considered auditory feedback and appropriately designed sonic cues can all help create more inclusive interactions for people with different sensory abilities. Importantly, these improvements should not be regarded as specialist features added for a small minority of users. Designing with accessibility in mind frequently produces better experiences for everybody. Rather than seeing accessibility as a constraint, Thornton encouraged students to view it as an opportunity to create richer, more intuitive and more human-centred experiences. Once again, the lecture returned to its central principle. Successful sonic branding begins with people rather than brands.

    Towards the end of the lecture, Thornton reflected upon the future of the discipline with considerable optimism. Advances in spatial audio, wearable technology and increasingly precise location-aware devices present exciting opportunities to rethink how organisations communicate through sound. Yet these possibilities also introduce important ethical questions. Thornton imagined a future in which highly personalised spatial advertisements might follow listeners through everyday environments, responding continuously to their location and behaviour. Although technically possible, he questioned whether such applications would genuinely improve people’s lives. New technologies, he argued, should first be viewed as opportunities to create more meaningful experiences rather than simply more opportunities for marketing. The question is never merely what audio branding can do, but what it should do.

    This concern led naturally to a broader critique of current practice. Thornton observed that many organisations continue to treat sonic branding as a superficial exercise, commissioning generic three-note logos or brief musical signatures without considering the wider user experience. As more companies embrace sound, he believes the greatest opportunity lies not in creating more branded audio, but in creating better branded audio. Holistic systems, careful implementation and genuine concern for the listener allow organisations to stand out far more effectively than louder or more intrusive branding ever could. Thoughtful design therefore becomes both a competitive advantage and an ethical responsibility.

    The lecture concluded by returning to one of its earliest themes. Brands are not static objects. They evolve continually alongside the people who use them. Thornton argued that sonic identities should therefore be managed in much the same way as living organisms. They require ongoing stewardship, periodic refinement and careful adaptation as technologies, expectations and cultural contexts change. An audio brand is never truly finished. It develops through continual listening, evaluation and improvement. Ultimately, Thornton challenged students to stop thinking about audio branding as the design of sounds and instead see it as the design of relationships. Every carefully considered interaction strengthens the connection between people and the organisations they encounter. When approached in this way, sonic branding becomes less about recognition and more about creating experiences that people genuinely value.

  • How Do You Design the Sound of a Blockbuster Game? Michael Caisley on Creativity, Recording, and Crafting the Sound of Call of Duty

    Michael Caisley

    How do you design the sound of a blockbuster game?

    Modern video games are built from extraordinarily complex systems. Artificial intelligence, physics, animation, graphics and networking all operate simultaneously to create worlds that respond continuously to the player’s decisions. Sound design must function within that same complexity. Unlike film, where every frame is predetermined, game audio unfolds differently every time someone plays. Thousands of individual sounds interact dynamically, responding to changing environments, player behaviour and gameplay events without losing clarity or dramatic impact. During his online guest lecture for Edinburgh Napier University, Michael Caisley drew upon his experience as Senior Sound Designer on Call of Duty: Advanced Warfare to explore how one of the industry’s largest productions approached this challenge. Throughout the session, one principle emerged repeatedly. Great game audio is designed as a complete system rather than a collection of individual sound effects.

    This philosophy shaped every stage of the project’s development. Rather than asking how individual weapons, footsteps or explosions should sound, the audio team began with a broader question. How should the player experience the world? Every recording, editing decision and implementation technique ultimately served that objective. Sound design therefore became an exercise in shaping perception rather than simply producing assets. Individual recordings remained important, though their true value emerged only through the relationships they formed with every other element of the soundtrack. The player never experiences sounds in isolation. They experience an acoustic world.

    Caisley explained that this perspective influenced one of the team’s earliest decisions. Although Call of Duty already possessed an established sonic identity developed across multiple successful titles, the audio team resisted the temptation simply to inherit those conventions. Instead, they treated Advanced Warfare as an opportunity to rethink the game’s entire sound philosophy from first principles. Existing assets, familiar production techniques and long-standing implementation methods were all reconsidered. Their ambition was not to reject the past, but to ensure that every creative decision continued to serve the experience they wanted players to have. Innovation therefore emerged through careful questioning rather than change for its own sake.

    That philosophy also transformed the relationship between sound design and implementation. In many production pipelines, sound designers create assets that are later integrated into the game by other specialists. Caisley described a markedly different approach. Sound designers remained responsible for implementation inside the game itself, allowing them to shape how recordings behaved once they became part of the interactive experience. The timing of a sound, the circumstances under which it played, the way it interacted with other events and its contribution to the overall mix all became part of the design process. Creating an excellent recording represented only the beginning. The player’s experience ultimately depended upon how successfully that recording functioned within the wider system. Implementation was therefore not separate from sound design. It was an essential part of it.

    The same systems-oriented thinking naturally extended to recording. Rather than relying primarily upon commercial sound libraries, the team invested heavily in producing original recordings specifically for the game. Specialist libraries remained valuable resources, particularly carefully curated collections produced by experienced field recordists, though Caisley consistently argued that original recording provides opportunities to discover sounds that nobody else possesses. More importantly, recording becomes a creative process rather than simply a method of gathering raw material. Unexpected textures, unusual perspectives and subtle acoustic details often emerge only when designers capture sounds for themselves. Distinctive game audio begins long before editing or implementation. It begins with listening carefully to the world.

    One particularly revealing example involved footsteps. Traditional Foley often records isolated footsteps on carefully prepared surfaces inside controlled studio environments. Caisley questioned whether this approach remained appropriate for a first-person game in which movement is experienced continuously through the player rather than observed from an external viewpoint. Instead, the team carried lightweight portable recorders into forests, hillsides and outdoor locations, capturing complete performances that naturally progressed from walking to running and sprinting. Rather than constructing movement artificially from disconnected recordings, they captured the changing rhythm, effort and momentum that emerge naturally when people move through real environments. The resulting recordings felt noticeably more convincing, illustrating that authenticity sometimes depends less upon technical precision than upon preserving the natural behaviour of the performer.

    The recording equipment itself reflected the same practical philosophy. Caisley encouraged students not to become preoccupied with expensive technology at the expense of creative opportunity. Much of the team’s field recording relied upon compact portable recorders that could be deployed quickly whenever an interesting sound presented itself. Mounted directly onto lightweight boom poles, these systems reduced handling noise while allowing recording sessions to remain flexible and spontaneous. The lesson extended far beyond the specific equipment being used. Interesting sounds rarely arrive when it is convenient to record them. Designers therefore benefit from tools that allow them to respond immediately rather than waiting for ideal conditions or elaborate recording setups. Creativity, he suggested, often rewards preparedness more than perfection.

    The same willingness to question established practice shaped the recording of weapons. Rather than organising one large recording session intended to capture every firearm in a single location, the team divided the work across numerous smaller sessions. This approach simplified logistics, though its greatest benefit proved creative rather than organisational. Each session could be reviewed afterwards, allowing the team to identify opportunities for improvement before returning to record additional material. Different environments also introduced naturally varying acoustic characteristics, providing a richer collection of perspectives than a single location could have offered. Recording therefore became an iterative process in which every session informed the next. The objective was not simply to accumulate material, but to refine the sonic identity of the game through continual experimentation.

    Perhaps the most important lesson from this stage of the lecture concerned the relationship between individual sounds and the finished player experience. Caisley observed that players rarely remember isolated recordings. They remember moments. The impact of those moments depends upon countless design decisions working together, from recording and editing through implementation, mixing and gameplay design. The audio team’s objective was therefore never to create the loudest explosion or the most detailed weapon recording. It was to build a soundtrack in which every element supported the player’s understanding of the world. Call of Duty: Advanced Warfare consequently adopted a more dynamic approach to mixing, allowing important sounds to occupy the foreground while leaving space for the rest of the soundtrack to breathe. Restraint became every bit as valuable as spectacle. The most memorable moments did not emerge from individual sound effects alone. They emerged from a coherent acoustic world in which every element strengthened the player’s belief that the environment around them was responsive, believable and alive.

    Having established the technical foundations of the project, Caisley turned towards the creative decisions that ultimately give a game its identity. Recording and implementation provide the raw materials, though they do not determine how a player experiences a moment. That depends upon judgement. Throughout the remainder of the session, he returned repeatedly to an idea that sounds deceptively simple but lies at the heart of professional sound design. Every sound reflects a design decision. The role of the sound designer is not merely to create convincing audio, but to decide what deserves to be heard, when it should be heard and, just as importantly, what should remain absent.

    This philosophy shaped the way Caisley approached almost every design problem. Instead of searching immediately for the perfect recording, he preferred to build what he described as palettes of possibilities. Families of related sounds sharing particular textures, movements and tonal characteristics were assembled through recording, processing and experimentation. Organic recordings of motors, impacts, machinery and environmental sounds were manipulated repeatedly, gradually forming a collection of materials from which the final design could emerge. Creativity therefore developed through exploration instead of beginning with a predetermined solution. Designers rarely know exactly what they are searching for at the start of a project. They discover it by experimenting until unexpected relationships begin to reveal themselves.

    His workflow reflected the same exploratory mindset. Projects often began in apparent disorder, with sounds accumulating rapidly as multiple ideas were investigated simultaneously. Immediate organisation was deliberately given lower priority than experimentation. Once a broad range of possibilities had been created, the process shifted towards careful refinement. Caisley compared this approach to sculpting. A sculptor begins with a block of material and gradually removes everything that does not belong until the final form becomes visible. Sound design, he suggested, often develops in exactly the same way. Instead of continually asking what should be added, designers should also ask what can be removed.

    This idea challenges one of the most common assumptions made by new sound designers. Richer sound does not necessarily result from adding more layers. As recordings accumulate, frequency masking increases, textures become crowded and important details begin to disappear. Caisley described repeatedly muting, removing and simplifying elements until only those making a genuine contribution remained. Equalisation, dynamics processing, timing adjustments and careful layering all supported this process, though none represented the objective in itself. Their purpose was to improve clarity, strengthen communication and ensure that every remaining sound justified its place within the mix. Professional sound design therefore depends less upon the quantity of material than upon the quality of the decisions shaping it.

    A particularly memorable example came from a sequence in which the player escapes across a glass roof before an ally destroys the structure beneath pursuing enemies. The obvious solution might appear to involve recording increasingly dramatic glass impacts before combining them into one spectacular crash. Caisley approached the problem very differently. The event was divided into a sequence of distinct dramatic stages. Initial bullet impacts, subtle structural weakening, growing instability and the final collapse each received their own carefully judged sonic treatment. Texture, pacing and silence changed gradually as the scene unfolded, allowing players to follow the progression of the collapse as a connected series of events rather than experiencing a single overwhelming burst of noise. The sequence derived its dramatic impact from the way the sound evolved over time, allowing the narrative of the scene to unfold naturally through listening as well as through the visuals.

    The same attention to dramatic pacing shaped Caisley’s approach to synchronisation. Students often assume that every visible action should be matched precisely by an accompanying sound. Professional practice, he suggested, is considerably more nuanced. Delaying one sound slightly, allowing another to emerge first or simplifying an otherwise crowded moment can produce a stronger dramatic effect than strict synchronisation alone. Rhythm, pacing, expectation and contrast all become compositional tools that guide the player’s attention. Instead of following every visual event mechanically, sound design helps determine what players notice, what they anticipate and how they interpret the unfolding action. Games therefore rely upon many of the same principles of dramatic storytelling found in music and cinema, while remaining responsive to player interaction.

    Equally revealing was Caisley’s discussion of realism. Throughout the lecture, he challenged the assumption that authentic sound must originate from authentic sources. Recording larger explosions does not necessarily produce better explosions, nor does striking more metal automatically create more convincing mechanical impacts. Professional sound designers routinely combine recordings whose original sources bear little resemblance to the finished result. Environmental ambiences, machinery, organic textures and countless unexpected recordings may all contribute qualities that literal recording alone cannot provide. What ultimately matters is not the origin of the sound, but whether it supports the player’s perception of the world. Believability depends upon the finished experience rather than literal accuracy.

    Technical processing formed part of this broader creative process rather than existing as an end in itself. Equalisation, compression, distortion and other processing tools undoubtedly shape the final soundtrack, though Caisley resisted presenting them as universal recipes. Every adjustment served a specific purpose within the wider composition. Heavy compression might transform an otherwise unremarkable recording into the perfect supporting layer. Subtle timing adjustments could reveal details previously hidden within the mix. Equalisation often preserved recordings that might otherwise have been discarded. Considered individually, many processed sounds appeared incomplete or even unattractive. Their value emerged only through their relationship with every other element. As throughout the lecture, the emphasis remained firmly upon systems rather than isolated sounds.

    Towards the end of the session, Caisley reflected upon the qualities that distinguish successful sound designers from merely competent technicians. Technical expertise undoubtedly matters, though he argued that curiosity, collaboration and the willingness to accept constructive criticism exert a far greater influence over long-term professional development. Working alongside experienced colleagues continually challenges assumptions and exposes designers to alternative ways of thinking. Equally valuable is the habit of listening analytically to other people’s work. Rather than deciding whether an entire game succeeds or fails, Caisley encouraged students to identify individual moments that demonstrate particularly thoughtful creative decisions. Examining one successful interaction in depth often teaches far more than making broad judgements about an entire soundtrack. Developing as a sound designer therefore depends as much upon careful listening as upon creating new sounds.

    Taken together, Caisley’s presentation revealed that blockbuster game audio is built as much through judgement as through technology. Recording, editing, implementation and mixing undoubtedly provide the necessary tools, though those tools acquire meaning only through the decisions that shape them. Every sound exists in relation to every other sound, every moment contributes to a larger dramatic experience and every creative choice influences how players understand the world around them. Sound design is not the art of creating more sound, but of making better decisions. Technology provides the tools. Careful listening, thoughtful judgement and an understanding of human perception transform those tools into interactive experiences that players instinctively accept as real.

  • How Do Robots Communicate Through Sound? Connor Moore on Audio UX, Robotics, and Designing Meaningful Interactions

    Connor Moore

    How do robots communicate through sound?

    People increasingly interact with technology through sound. Smartphones acknowledge completed payments, electric vehicles alert drivers to potential hazards, wearable devices provide subtle notifications and intelligent products communicate through a growing vocabulary of tones, chimes and alerts. Yet these sounds rarely receive the same attention as visual design. During his online guest lecture for Edinburgh Napier University, Connor Moore explored the growing discipline of Audio User Experience (Audio UX), demonstrating how carefully designed sounds help products communicate naturally, build trust and express personality. Drawing upon projects for companies including Google, Tesla and Postmates, he argued that successful product sound design extends far beyond creating attractive audio. It begins with understanding the people for whom they are designed. Throughout the session, one principle emerged repeatedly. Every sound should communicate with purpose.

    Moore introduced his work by describing the remarkable breadth of modern Audio UX. Working from California’s Bay Area, he collaborates with companies developing products across robotics, automotive technology, connected devices, consumer electronics and digital services. Although these industries appear very different, they all share a common challenge. Products increasingly communicate with people through sound, requiring designers to think carefully about what those sounds communicate and how they contribute to the wider identity of a brand. Rather than approaching each project as an isolated collection of sound effects, Moore described building coherent sonic systems that extend across products, marketing, physical environments and user interactions. Individual sounds matter, though they become most effective when they form part of a larger and recognisable design language.

    This broader perspective also explains why strategy sits at the beginning of every project rather than at the end. Before designing a single sound, his team seeks to understand the objectives of the product, the identity of the organisation and the experience that users should ultimately have. Brand workshops, creative discussions and detailed reviews of existing sounds all contribute towards this early stage of development. Competitor analysis also plays an important role. Understanding how other companies sound allows designers to identify opportunities for meaningful differentiation rather than unintentionally reproducing familiar ideas. The objective is not simply to sound different. It is to create a sonic identity that genuinely reflects the values and personality of the organisation. Sound therefore becomes a strategic design material rather than a decorative addition introduced once the product has already been completed.

    One of the most thought-provoking ideas introduced during the presentation concerned what Moore described as connected audio ecosystems. Many organisations continue to commission isolated sounds for individual products or services, yet users increasingly encounter the same company across multiple devices and environments. A person may hear a notification on a smartphone, interact with a smart speaker at home, use an in-car navigation system and later encounter advertising or public installations produced by the same organisation. Rather than allowing each experience to develop independently, Moore argued that they should all share recognisable sonic characteristics. Consistent instrumentation, similar timbral qualities and carefully related musical ideas allow users to recognise a brand without needing to see a logo or screen. Sound therefore becomes another component of brand identity, working alongside visual design to create familiarity and trust.

    Google’s product ecosystem provided one of the clearest illustrations of this philosophy. Moore described how his work began during the development of Google Glass, a product that sought to make an unfamiliar technology feel approachable. Rather than emphasising futuristic electronic sounds, the design drew upon simple acoustic instruments such as piano, chimes and mallet percussion. These familiar timbres helped ground an otherwise unfamiliar experience, making the product feel more human and intuitive. As Google’s product portfolio expanded through devices such as Pixel phones, Google Pay and automotive systems, this underlying sonic character evolved while remaining recognisably connected. Different products naturally demanded different technical solutions and frequency ranges, though the overall identity remained remarkably consistent. Moore argued that brands evolve in much the same way as people do. Their sonic identities should therefore develop over time while retaining a recognisable sense of continuity.

    Perhaps the most unexpected principle discussed during the session concerned silence. Designers often assume that every interaction requires another notification, another confirmation or another layer of feedback. Moore challenged this instinct directly. As products increasingly incorporate sound into everyday life, designers also acquire a responsibility not to make the world unnecessarily louder. Like negative space in graphic design, silence performs an important communicative function. It creates contrast, draws attention to genuinely important events and prevents users from becoming overwhelmed by constant auditory stimulation. Successful Audio UX therefore depends not only upon knowing which sounds should exist, but equally upon recognising which moments deserve silence instead.

    He illustrated this philosophy through the development of Sense, a sleep monitoring device designed to help users understand the quality of their sleep. Conventional alarm clocks often rely upon abrupt, attention-grabbing sounds that force people awake almost instantly. Moore saw an opportunity to rethink that experience entirely. Instead of beginning loudly, the alarms gradually evolved over time, introducing increasing musical complexity, richer timbres and subtle changes in tempo. Lighter sleepers could wake during the earliest stages, while heavier sleepers would gradually encounter a more energetic composition. Even error tones and voice interactions were designed using soft, restrained timbres that preserved the calm atmosphere of the bedroom rather than disrupting it. The project demonstrated that product sounds need not simply communicate efficiently. They can also influence the emotional quality of everyday experiences.

    Moore then introduced one of his central design philosophies: communicative and expressive design. Throughout the presentation, he repeatedly distinguished between creating sounds that merely reinforce a brand and creating sounds that genuinely help people understand what is happening. Branding undoubtedly matters, though communication always takes priority. Every sound should first convey meaning. Only then should it contribute towards a wider sonic identity. This perspective encourages designers to think carefully about urgency, expectation and human perception rather than treating every notification as another opportunity for creative expression. Product sounds exist to guide behaviour as much as they exist to establish identity.

    Tesla provided a particularly revealing case study. Moore described developing different categories of sounds according to the urgency of the information they needed to convey. Low-priority events, such as incoming calls, were designed to emerge gradually using softer timbres and lower levels of perceptual urgency. Medium-priority notifications, including seatbelt reminders, employed greater repetition and brighter timbres, encouraging users to respond without becoming unnecessarily stressful. High-priority warnings, including forward collision alerts, demanded a very different approach. Higher frequency content, more percussive attacks and rapid repetition ensured that these sounds immediately captured attention during situations where rapid action could prevent an accident. Rather than relying upon arbitrary aesthetic decisions, Moore demonstrated how pitch, repetition, harmonic content and timbre can all be manipulated systematically to communicate different levels of urgency. Sound becomes a carefully designed language through which products communicate urgency, intention and behaviour.

    By this stage, a clear philosophy had emerged. Audio UX is not simply concerned with creating pleasant sounds or memorable sonic logos. It asks how products should communicate with the people who use them every day. Strategy, branding, silence, musical structure and perceptual psychology all contribute towards that objective, though none of them represents the ultimate goal. Every design decision serves the relationship between people and technology. Once sound is understood as a form of communication rather than decoration, the challenge shifts from asking what a product should sound like to asking what it should say. That question became even more significant when Moore turned to the rapidly developing world of robotics.

    The second half of Moore’s presentation shifted from broad design principles towards a detailed case study that demonstrated how those ideas are applied in practice. The project centred on Serve, the autonomous delivery robot developed by Postmates. Rather than treating the robot simply as another product requiring notification sounds, Moore used it to explore a far broader question. How should an intelligent machine communicate with people as it moves through shared public spaces? The answer, he suggested, depends upon much more than selecting attractive sounds. It requires understanding personality, context, expectation and human behaviour long before the first sound is ever designed.

    Like every project discussed earlier in the presentation, the design process began with strategy rather than sound. Before any recording or composition took place, the team explored what the robot represented, how people would encounter it and the personality it should express. Several distinct sonic directions were developed around different interpretations of the brand before being refined through successive design reviews and evaluated within the robot itself. Moore emphasised that successful Audio UX develops through continual iteration rather than moments of inspiration. Sounds that appear convincing inside a studio may behave very differently once reproduced by a moving robot navigating busy streets, restaurants and crowded pavements. Testing therefore becomes an integral part of the creative process rather than simply a means of checking technical performance.

    One of the most revealing aspects of the project concerned personality. Popular culture has encouraged audiences to expect robots to communicate through futuristic electronic sounds or highly expressive synthetic voices. Moore deliberately avoided both extremes. The ambition was to create a robot that felt warm, approachable and reassuring without pretending to possess human intelligence or emotional awareness. At the same time, the team resisted the temptation to rely upon recorded speech, recognising that a natural voice would create expectations that the technology could not consistently fulfil. Instead, they searched for a middle ground in which sound suggested character without imitating humanity. This balance between familiarity and honesty reflected one of the most thoughtful ideas running throughout the presentation. Good Audio UX should communicate clearly without misleading users about a product’s capabilities.

    Developing that personality required exploration rather than immediate certainty. Moore described creating several contrasting sonic directions, each expressing a different interpretation of the robot’s identity. Some embraced more mechanical qualities that acknowledged the machine’s physical presence. Others explored vocal-like synthesis capable of suggesting expression without becoming literal speech. A third direction employed simple sine-wave tones that created a calmer, softer and more abstract character. Rather than choosing a favourite instinctively, these alternatives became prototypes through which designers could observe how people responded emotionally to different sonic identities. The final design combined warmth, clarity and subtle expressiveness, producing a robot that felt approachable without becoming theatrical or sentimental. The process illustrated an important principle: successful sound design rarely emerges fully formed. It develops through comparison, evaluation and refinement.

    Attention then shifted from the robot’s overall personality to the design of individual interactions. Different situations demanded different styles of communication. Interactions inside restaurants, where staff loaded deliveries into the robot, prioritised efficiency through short, direct auditory cues that confirmed actions without interrupting the workflow. Encounters with members of the public required a gentler approach. Longer note durations, more relaxed phrasing and softer musical gestures created the impression of patience rather than urgency. In effect, the robot adapted its acoustic behaviour according to the social environment in which it operated, much as people instinctively alter their own behaviour between professional and public settings.

    A particularly memorable example centred on one of the simplest interactions imaginable: saying, “Excuse me.” Rather than relying upon recorded speech, Moore developed a brief auditory gesture that politely attracted attention before allowing the robot to continue its journey. The intention was not to surprise pedestrians or demand an immediate response. Instead, the sound functioned more like a courteous acknowledgement of another person’s presence. This small interaction captured a principle that extended throughout the presentation. Effective communication often depends upon restraint rather than intensity. Products should seek attention only when attention genuinely needs to be given.

    Safety presented a different set of priorities. Warning sounds must communicate immediately and unambiguously, leaving little room for ambiguity or interpretation. Here, Moore returned to ideas introduced earlier in the presentation. Instead of inventing unfamiliar sonic languages, the design frequently drew upon acoustic references that people already understood from everyday experience. Turn indicators, movement cues and other operational sounds retained familiar characteristics while remaining consistent with the robot’s wider sonic identity. This reduced the need for users to learn entirely new sonic conventions. Familiar sounds could be interpreted almost instinctively, allowing people to respond appropriately without consciously analysing what they had heard.

    Perhaps the most technically demanding challenge involved the robot’s continuous movement through public space. Moore explored several possible solutions, including humming, whistling and slowly evolving tonal textures that he described as “glowing.” Each communicated the robot’s presence in a slightly different way. Some attracted attention more effectively, while others blended more comfortably into the surrounding soundscape. Extensive user testing, including sessions involving blind participants, revealed that restrained harmonic complexity and carefully controlled modulation proved more effective than more elaborate alternatives. Yet Moore resisted the temptation to increase the amount of sound simply to improve awareness. His longer-term ambition was quite the opposite. Intelligent products should become quieter rather than louder. If a robot recognises that nobody is nearby, there may be no need for it to produce sound at all.

    This idea provides a fitting conclusion to Moore’s broader philosophy of Audio UX. The discipline is not concerned with filling products with attractive sounds or memorable sonic logos. It asks how technology can communicate clearly, respectfully and appropriately with the people who use it. Whether designing for autonomous robots, electric vehicles, smartphones or medical devices, the same principles continue to apply. Strategy comes before implementation. Communication matters more than novelty. Personality should remain authentic. Silence deserves to be designed as carefully as sound itself. When those ideas come together successfully, sound ceases to be decoration and becomes an essential part of the conversation between people and technology.