Category: Audio Perception

  • How Do You Preserve the Sound of a Place? Damian Murphy on OpenAir, Acoustic Heritage, and Reconstructing Lost Spaces

    Damian Murphy

    How do you preserve the sound of a place?

    A building can survive through photographs, architectural drawings, maps and written descriptions. Its dimensions can be measured, materials catalogued and appearance reconstructed long after the original structure has disappeared. Yet places are not experienced through vision alone. A cathedral changes the sound of a choir, a tiled chamber changes the sound of a voice, snow alters the acoustic behaviour of a forest, and the hard surfaces of a mausoleum can allow sound to continue long after its source has stopped. Architecture surrounds every action with reflections, resonances and reverberation, but this part of a place can disappear without leaving anything visible behind.

    During his online guest lecture for Edinburgh Napier University, Professor Damian Murphy of the University of York explored almost two decades of work investigating how the acoustics of places can be measured, preserved, reconstructed and experienced. Much of this work centres upon OpenAir, the Open Acoustic Impulse Response Library, which contains acoustic measurements gathered from buildings, landscapes and other environments. Its contents range from cathedrals, churches and theatres to industrial buildings, caves, forests and vehicles, connecting acoustic science with music production, spatial audio, games, archaeology, heritage and historical research. Murphy moved between these different places and applications through a question that became larger as the examples accumulated. What can the sound of a place tell us that its image cannot?

    Murphy began with a ruin. The Temple of Decision stands on a hill overlooking the Falkland Estate in Fife, where artists David Chapman and Louise K. Wilson had been commissioned to explore the landscape through sound. Archive material could offer clues about the temple’s former appearance, while historical research could provide fragments of information about its use, but the surviving structure could no longer reveal how the intact building had sounded. What would it have been like to speak inside the room? How might voices have behaved around a table? How would conversation, movement or a fire have interacted with its surfaces? The artists posed a question that provided a starting point for Murphy’s lecture: in the absence of clear echoes, how can we know a place?

    Answering that question first requires an understanding of what an acoustic environment contributes to anything heard within it. Imagine a short, sharp sound produced inside a room. A listener initially receives the direct sound travelling along the shortest path from source to receiver. Early reflections arrive shortly afterwards from nearby surfaces, followed by increasingly complex patterns as energy travels through the room, interacting repeatedly with walls, floor, ceiling and objects before gradually decaying. Together, these components form a room impulse response for a particular relationship between a source and receiver. Direct sound carries information about the source and its distance, while early reflections contribute to perceptions of geometry and position. Later reverberation communicates qualities associated with volume, materials and enclosure. Change the architecture and the response changes. Move the source or listener and it changes again. An impulse response is therefore not a complete acoustic identity for a building, but a record of how sound travelled between particular positions under particular conditions.

    Once captured, that relationship can be used for more than numerical analysis. Through convolution, a recording made without significant room acoustics can be combined with an impulse response measured elsewhere. Murphy demonstrated the process using a four-part vocal ensemble recorded in an anechoic chamber. The singers had never performed in York Minster, yet convolution with a measured response allowed their dry voices to acquire characteristics of the cathedral. York Minster has a reverberation time of approximately eight seconds through part of the mid-frequency range, compared with around half a second for a typical living room. Voices behave very differently in each environment. Notes overlap, transitions blur and the building continues sounding after the performers have stopped producing sound. The acoustic is not simply decoration placed around a performance. It changes the temporal relationships through which that performance is heard.

    Auralisation, however, introduces a distinction between recreating acoustic conditions and recreating experience. Singers performing in an anechoic chamber do not behave as they would inside a highly reverberant cathedral. Performers hear themselves and adapt. Tempo, articulation, phrasing, dynamics and pauses can change in response to sound returning from the room, while musicians continually adjust to one another through the acoustic environment they share. Convolution can reproduce the effect of a measured response upon a recording, but it cannot retrospectively create the performance that might have developed inside that space. Murphy acknowledged this limitation directly when discussing the York Minster example. The anechoic performance was not the performance the singers would have given in the cathedral, and even the spacing of phrases in the demonstration had been altered to allow the reverberation to emerge more clearly.

    The distinction matters beyond the simulation of reverberation. A room impulse response can describe how energy travels between defined positions, but people are not passive sound sources or microphones. They move, listen selectively, change their behaviour and respond to what they hear. Preserving a response gives researchers evidence about the acoustic conditions of a place. What people did in response to those conditions remains a different question.

    OpenAir developed from an ambition to preserve such evidence and make it available for others to explore. Murphy traced one important influence to Angelo Farina’s work on recording concert halls for posterity. Improvements in measurement techniques and the emergence of practical convolution reverberation created an opportunity to document significant spaces not merely through reverberation times and other summary values, but through impulse responses that could be analysed, reproduced and used creatively. OpenAir extended this principle into a growing archive, with a measurement system designed to collect spatially rich data that could remain useful beyond the immediate research question.

    Early measurement sessions used a Genelec S30D loudspeaker to excite the space while microphones captured its response. A computer-controlled turntable allowed measurements at regular angular intervals, and an ambisonic Soundfield microphone was combined with a Neumann KM100 cardioid microphone to provide spatial information alongside material suitable for different forms of analysis and production. Measurements could be repeated across several source and receiver positions, preserving a set of relationships rather than presenting each building through one supposedly definitive response. Ambisonics was particularly valuable for an archive whose future applications could not be predicted. A first-order ambisonic recording represents a three-dimensional sound field through an omnidirectional component and three directional components, separating the captured information from one fixed loudspeaker arrangement. Material can later be decoded for different reproduction systems or manipulated in ways that may not have been anticipated when the recording was made.

    Flexibility matters when access to a significant site may last only a few hours. Researchers need to gather material rich enough to support questions that have not yet been formulated and technologies that may change long after a measurement session has ended. During the discussion after the lecture, Murphy described more recent work at St Paul’s Cathedral, undertaken with a composer who wanted impulse responses from the building. Three researchers had only three hours to move equipment through the enormous space and capture responses from locations including the nave, a stairwell and the Whispering Gallery. Practical decisions about where to measure become part of preservation itself. As Murphy observed, there is no single sound of a large building. Different positions offer different acoustic experiences, and any archive necessarily records selected relationships within a much larger field of possibilities.

    As OpenAir expanded, its growing range of places made the idea of acoustic preservation less straightforward. York Minster was an obvious candidate, since its long reverberation is closely connected with experiences of worship, tourism and musical performance. Other sites raised different questions about what deserves to be preserved and why. When the former Terry’s chocolate factory in York closed, Murphy and his colleagues gained access before redevelopment and measured spaces including a warehouse and the former typists’ room, a striking interior dominated by glass and wood. The activities for which these spaces had been designed had already disappeared. An empty typists’ room could still be photographed, but its appearance prompted another question: what might it have sounded like when filled with the overlapping mechanical activity of typewriters?

    A subterranean reactor hall beneath the Royal Institute of Technology in Stockholm preserved another relationship between architecture and former activity, while measurements in historic churches allowed acoustic theories to be tested rather than merely repeated. At St Andrew’s Church in Lyddington, the team examined jars embedded within the walls, architectural features sometimes interpreted through theories of resonant vessels extending back to the Roman architect Vitruvius. Measurement provided little evidence that the jars were making a substantial contribution to the acoustic character of the church. Their form did not correspond closely with the behaviour expected of effective Helmholtz resonators. Acoustic research could challenge explanations attached to historic architecture as well as document spaces admired for their sound.

    York Theatre Royal shifted attention from buildings as fixed objects towards places in changing states. Murphy’s team measured the auditorium before refurbishment and returned afterwards to document its altered acoustic. They also captured measurements with an audience present during a pantomime, recognising that an occupied theatre does not behave acoustically like the same room when empty. Seats, clothing and bodies absorb and scatter sound, making occupancy part of the acoustic system rather than simply a group of listeners placed within it. Materials age, spaces are repurposed and environmental conditions change. Preserving an acoustic environment may therefore involve documenting several states of the same place rather than searching for one definitive response.

    Outdoor and semi-outdoor locations created different practical problems, and the original measurement system could not simply be carried everywhere. On the Falkland Estate, the Bottle Dungeon could be reached through a trapdoor, but access was sufficiently awkward that the usual equipment was impractical. A balloon was attached to one stick, a pin to another and a microphone lowered into the space. Bursting the balloon remotely provided the excitation needed to capture a response. The improvised arrangement was far removed from the computer-controlled turntable used elsewhere, yet it addressed the same underlying need: introduce a suitable sound into an inaccessible environment and record how the space transforms it.

    Landscapes demanded further adaptation. Equipment had to become portable and independent of mains electricity as researchers travelled into the Yorkshire Dales to measure caves and gorges. Work in Finland examined the same forest under different seasonal conditions. Researchers used GPS alongside ribbons tied around trees to return to the same positions after deep snow had transformed the landscape. Geographically, it remained the same forest. Acoustically, it had changed. Snow altered the interaction between sound, ground and surrounding environment, demonstrating that acoustic character can change while location remains constant.

    A collaboration with Codemasters carried that thinking into interactive media. Looking at an archive rich in churches and historic interiors, the developer asked a practical question: what about the environments needed for games? The collaboration encouraged further work on landscapes and led to experiments for GRID Autosport involving vehicle interiors, where the acoustic problem was unusually complex. Codemasters wanted to represent a car gradually falling apart during a race, so the team needed measurements capable of describing changing states, including doors opening or disappearing and the boot being open. The experience of being inside a racing car also comes from more than airborne sound. Engine vibration travels mechanically through the structure and contributes to what an occupant hears and feels.

    Conventional room measurement could not fully represent that relationship, so Murphy and his colleagues experimented with using the engine itself as part of the measurement process. Revving the car caused the structure to vibrate as it would during use, after which signal-processing methods were applied to separate the excitation from the resulting cabin response and derive an approximation of the impulse response. The approach sought to preserve something more specific than the reverberation of a small enclosure. It attempted to capture the interaction between a vibrating machine, its structure, the enclosed air and the listener inside it.

    Games also demonstrated how preservation can become creative infrastructure. An impulse response gathered for research may later help construct a virtual environment, become part of a music-production tool or support a question that had not existed when the measurement was made. Murphy described OpenAir material finding its way into software used by musicians and audio practitioners, allowing measurements collected years earlier to acquire new purposes. Distribution under a Creative Commons licence reflected this wider ambition. Researchers can analyse the acoustic behaviour of a building, composers can use the same response creatively, sound designers can place fictional events inside measured environments and developers can incorporate selected material into new tools.

    OpenAir originally allowed members of the wider community to upload their own measurements. As contributions accumulated, variations in recording quality became difficult to ignore. The team eventually reviewed the existing material, retained the strongest contributions and moved towards a more curated model in which new contributors contact the team directly. Open access expands what an archive can become, but reuse also depends upon confidence in how the material was produced.

    These measurements deal with places that still exist, even when they are changing. The Temple of Decision presents a different problem. Its original interior has already gone. Once a room has disappeared, there is nothing left to measure. Its former acoustic behaviour has to be approached indirectly through surviving evidence and modelling.

    Using information about a lost structure, researchers can construct a three-dimensional geometric representation and simulate the propagation of sound within it. Virtual rays travel through the model, reflecting between surfaces to generate impulse responses for selected source and receiver positions. Those responses can then be analysed or used to process voices and other recordings, allowing listeners to hear an interpretation of how sound might have behaved inside architecture that no longer survives. Such a result differs fundamentally from measuring an existing building. Geometry may be uncertain, material properties need to be estimated and every modelling method introduces limitations. Historical auralisation produces an evidence-based acoustic proposition rather than a recording recovered from the past.

    Work on St Mary’s Abbey in York made both the possibilities and limitations of this approach audible. The former church survives as a ruin, while archaeological and architectural evidence allowed Murphy’s team to construct a three-dimensional representation suitable for acoustic modelling. Simulated impulse responses could then be compared with measurements from York Minster, a surviving building with some comparable characteristics. Estimated reverberation times occupied a similar range, but the reconstructed St Mary’s sounded noticeably brighter. The difference did not necessarily reveal a historical distinction between the buildings. Murphy explained that the ray-tracing model used for St Mary’s was less effective at reproducing low-frequency behaviour than the physical measurement system used in York Minster. Part of what listeners heard therefore belonged to the method of reconstruction itself.

    A model may sound convincing while still containing audible consequences of the technique used to create it. Plausibility has to emerge from evidence, comparison and methodological transparency rather than from the apparent realism of the result alone. Once those boundaries are understood, a model can make relationships perceptible in ways that drawings and numerical data cannot. During a public performance among the ruins of St Mary’s Abbey, a live choir was captured and processed through impulse responses derived from the reconstructed church, then projected to an audience gathered at the site. Present-day voices sounded among the physical remains while the model returned an interpretation of the missing acoustic architecture. Research data became part of an experience connecting a surviving place with a vanished interior.

    Reconstructing an abbey or temple can help audiences imagine how a lost building might have shaped music and speech. Murphy’s more recent work moved towards a question with wider historical consequences. If architecture changes what people can hear, could reconstruction help investigate who had access to speech in the past?

    Working with historians and art historians, Murphy’s team investigated spaces associated with the historic House of Commons within the former Palace of Westminster. Among the most revealing parts of that history was the experience of women who listened to parliamentary debate from a roof space above the chamber. Known as the Ventilator, this space allowed women excluded from formal political participation to gather above the House of Commons and listen through the architectural structure separating them from the debate below. Historical evidence could establish that they were there. Acoustic reconstruction allowed another question to be asked: what could they actually understand?

    Answering it required more than recreating the debating chamber as an isolated room. Researchers had to consider the chamber, roof void, Ventilator and routes through which speech travelled. The women listening above did not have a direct line of sight to the speaker, making their experience a problem of acoustic transmission through a complex architectural arrangement rather than ordinary listening within a single enclosure. Murphy connected this project with broader questions of directionality, speech transmission and listening position, while acknowledging that objective measures cannot reproduce every aspect of human attention or historical experience.

    No surviving recording can reveal exactly what those listeners heard. The team therefore combined historical reconstruction with comparative measurement, examining surviving spaces connected with parliamentary history or comparable in geometry, scale, period or use. Measurements from rooms at the University of Oxford, York Guildhall and the present House of Commons chamber provided contexts against which aspects of the model could be considered. None could prove how the lost chamber sounded, but together they allowed reverberation and speech intelligibility to be examined across different positions and occupancy conditions. A question about architectural acoustics had become inseparable from a question about political access.

    Hearing is a form of access. Architecture, distance, reverberation, occupancy and barriers influence whether speech remains intelligible, while attention, familiarity and expectations affect what can be understood from imperfect information. A person may be physically close to political debate while remaining acoustically separated from it. Reconstructing the conditions of listening can therefore contribute to historical questions about participation and exclusion that visual records alone cannot answer.

    The Temple of Decision now appears less like an isolated case and more like the beginning of a much larger enquiry. In the absence of clear echoes, how can we know a place? We can measure what survives, compare different conditions, preserve spaces before they change and build models from evidence when the original architecture has disappeared. We can listen to those models while remaining clear about what they can and cannot establish. Most of all, we can ask questions that become difficult to formulate when architecture is treated only as something seen.

    A photograph can preserve the appearance of a parliamentary chamber from one position. An architectural plan can show where walls, doors, galleries and roof spaces were located. Written testimony can tell us that people gathered somewhere to listen. Acoustic research adds another layer by asking how sound travelled between those positions, how reverberation affected speech and whether architecture enabled or obstructed understanding. The same thinking can ask how a ruined abbey shaped musical performance, how a theatre changed during refurbishment or how winter transformed the acoustic behaviour of a forest.

    Preserving the sound of a place is not simply an attempt to save an attractive reverberation before it disappears. Places participate in human activity. They change how people speak, perform and listen. Their surfaces and geometry influence whether voices remain intimate or become collective, whether musical phrases overlap or remain distinct, and whether somebody standing beyond a barrier can understand words spoken elsewhere. Acoustic conditions can shape behaviour, access and participation without leaving a visible trace.

    The Temple of Decision can still be visited. St Mary’s Abbey remains visible as a ruin. The old House of Commons chamber has disappeared, while the former typists’ room at Terry’s chocolate factory no longer performs its original function. A forest changes when snow arrives and changes again when it melts. Even a building that survives intact contains many acoustic relationships, only some of which can be measured during the limited hours when researchers have access.

    Sound is especially vulnerable to disappearance. Once a room changes, an audience leaves or a building is lost, its former acoustic behaviour cannot simply be photographed. Murphy’s lecture showed that it is not entirely beyond preservation. What remains may be a measurement, a model, a comparison or a carefully documented uncertainty, each offering a different way of understanding a place as more than a visual container for history.

    In the absence of clear echoes, we may never know a place completely. We can, however, preserve evidence of how sound moved through it, reconstruct plausible relationships when the original has disappeared and ask what those acoustics meant for the people who performed, spoke and listened there. An impulse response may last only a few seconds, yet within it can remain an acoustic trace of a cathedral, a factory, a theatre, a cave, a forest under snow or a room about to change forever. When the room itself has already disappeared, a model can return a possibility rather than a certainty: a voice reflecting from lost walls, a choir inhabiting a ruined abbey or political speech travelling towards listeners hidden above a chamber from which they were excluded.

    Preserving the sound of a place means preserving another way of understanding what happened there. Walls determine more than what people can see. They shape what can be heard, how clearly it can be understood and who is able to listen.

  • How Do You Design Sound for an Experience You Cannot Control? Wylie Stateman on Storytelling, Simplicity, and the Future of the Soundtrack

    Wylie Stateman

    How do you design sound for an experience you cannot control?

    A filmmaker can frame an image. The edges of the screen define what the audience sees, while composition, focus, lighting and editing direct attention within it. Sound is less obedient. It extends beyond the frame, surrounds the audience, enters rooms with different acoustics and reaches listeners through systems ranging from enormous cinema installations to headphones, televisions, laptops and tiny mobile-phone speakers. A soundtrack may be created with extraordinary precision, yet nobody involved in its production can completely control how, where or at what level it will finally be heard.

    During his online guest lecture for Edinburgh Napier University, supervising sound editor and sound designer Wylie Stateman explored the creative and professional consequences of working with such an elusive medium. Drawing upon a career spanning more than four decades and collaborations with filmmakers including Quentin Tarantino, Oliver Stone and John Hughes, he described sound as an art form positioned between science and subjectivity. Its physical behaviour can be measured, yet its meaning depends upon perception, attention, expectation and context. His work has extended from feature films and television to advertising, audiobooks and theme-park attractions, but a consistent question connects these different forms: how can sound professionals create an experience that remains clear, emotionally purposeful and dramatically coherent when both production and listening contain so much uncertainty?

    For Stateman, answering that question requires sound practitioners to think beyond individual sounds. A designer may create remarkable material, but the audience experiences relationships: dialogue against music, effects within an environment, silence before an impact and the entire soundtrack through a particular playback system at a particular level. Creating those relationships at scale requires collaboration, while maintaining them requires somebody to understand the complete experience. Throughout the lecture, Stateman moved repeatedly between these levels, from the organisation of large creative teams to the placement of a single sound and from the controlled environment of the mixing stage to the unpredictability of a listener pressing play somewhere else. Connecting them was a consistent philosophy. Complexity is unavoidable behind the scenes, but it should produce clarity for the audience.

    Stateman began with the unusual nature of sound itself. A visual composition can be stopped and examined. People can point towards a particular area of an image, discuss its composition and compare alternatives while the material remains stationary. Sound exists through time. A mixer listens, makes an adjustment, returns to an earlier point and listens again. Creative refinement becomes a repeated movement backwards and forwards through the material, making the number of meaningful passes completed within a working day directly relevant to the sophistication of the result.

    Such a process makes both technology and collaboration important, but neither can replace judgement. Stateman described audio as a discipline in which scientific knowledge and subjective interpretation continually meet. Vibration can be measured objectively, while a listener’s response to it cannot be reduced so easily. Designers therefore work simultaneously with physical systems and human perception. Loudspeakers, auditoria, codecs and playback levels matter, but so do memory, expectation, emotion and attention. Even the most technically controlled production process eventually encounters a listener whose response cannot be engineered with the same certainty as the system delivering the sound.

    Perhaps this uncertainty helps explain why Stateman spoke so strongly about collaboration. Rather than presenting professional success as the achievement of a solitary creative individual, he described a career built through relationships with people whose abilities complemented his own. Soundelux, the company he established with fellow sound designer Lon Bender, grew from a small operation into an organisation employing hundreds of people across several cities. Its development depended upon far more than creative talent. Business management, accounting, engineering, technology, sales and production all had to support the work of designers, editors and mixers.

    The lesson for students was not that everyone should attempt to build a large company. Stateman’s broader point concerned complementary ability. Nobody needs to become equally skilled at every aspect of creative and professional life. Someone with little interest in finance needs a trustworthy person who understands it. A creative specialist working on complex productions benefits from engineers and technologists capable of turning ideas into reliable systems. Partnerships become valuable when they extend what a group can imagine and accomplish rather than merely reproducing the same abilities several times.

    Professional value also emerges through reliability. Stateman described a valuable colleague in strikingly simple terms: somebody who can understand a problem, take responsibility for it and allow everyone else to stop worrying about it. Creative ability matters enormously, but large productions depend upon trust. A person who solves a problem without creating several new ones becomes increasingly valuable to the people around them. Careers are built not only through the quality of isolated work, but through the confidence that others can place in somebody when a difficult problem arrives.

    Filmmaking itself operates through the same interdependence. Every specialist inevitably perceives the project through a particular discipline. Sound designers think about sound, composers about music, cinematographers about light and composition, costume designers about clothing and production designers about the physical world. Stateman compared this to a collection of unreliable narrators, each understanding the film from a particular perspective. The director’s responsibility is to bring those partial perspectives together into a coherent experience.

    Sound professionals therefore need both commitment to their own discipline and awareness of the larger work. Stateman’s long relationships with individual filmmakers allowed such understanding to develop across several projects. By the time of Once Upon a Time in Hollywood, he had worked with Quentin Tarantino on seven films. Repeated collaboration created trust and a shared shorthand, allowing Stateman to take substantial creative ownership of the soundtrack while remaining clear that every decision ultimately served the director’s film.

    That balance between ownership and service runs through much of professional sound design. Creative contribution requires conviction. A designer who merely waits for instructions cannot provide the full value of specialist expertise. Yet conviction must remain connected to the filmmaker’s intention rather than becoming an opportunity to demonstrate technique. Stateman’s account suggested a form of authorship that is confident without becoming possessive. Sound teams need enough ownership to make strong decisions, while recognising that those decisions belong within a larger work whose purpose they must understand.

    Understanding that purpose becomes increasingly important as the available technology grows more powerful. More capability does not automatically justify more activity, and one of Stateman’s clearest demonstrations came from work beyond conventional cinema. Soundelux developed audio and show-control systems for large theme-park attractions, including a Terminator 2 experience designed to move hundreds of visitors through repeated performances every day. Audiences passed through a pre-show before entering a theatre containing three 3D IMAX screens and a motion base. Reliability was essential, but technical reliability alone could not create an engaging experience.

    Despite having access to the motion platform throughout the attraction, the production reserved its major movement for a single moment. Audiences were allowed to become comfortable before the platform suddenly dropped. Their physical reaction was powerful precisely through its rarity. Continual movement would have made the mechanism familiar and reduced its dramatic value. Behind one carefully timed surprise sat engineering, amplification, show control, multiple screens, motion systems and the work of a large team. None of that complexity needed to become the audience’s concern. They experienced the result.

    The same principle shaped Stateman’s approach to film sound. A system may offer hundreds of channels, extensive spatial control and enormous dynamic range, but the designer still needs to decide which possibilities deserve to be used. Creative sophistication can appear through restraint, and technical capability becomes valuable when it helps direct attention rather than continually demanding it.

    When asked how he decides what an audience should attend to, Stateman returned to simplicity. Film combines visual and auditory information, but listeners cannot process every available element with equal attention. A dense soundtrack may contain extraordinary detail while communicating very little. The designer’s task is therefore not to make everything audible at once, but to create a clear path through the scene.

    Dialogue often provides the starting point. Human listeners extract extraordinary amounts of information from voices. Words communicate explicit meaning, while rhythm, pitch, timbre, hesitation and vocal effort reveal character and emotion. If audiences struggle to understand what somebody is saying, an essential layer of narrative and performance has been weakened. Music offers another route through the experience, leading or following emotional movement, preparing audiences for change or allowing feeling to emerge after an event. Environments and sound effects establish space, physicality, scale, tension and perspective. None of these categories possesses permanent priority. What matters is understanding what the audience should receive from a particular moment and organising the soundtrack accordingly.

    Stateman’s preferred method was additive rather than deductive. Instead of filling a soundtrack with every plausible sound and gradually removing whatever causes problems, begin with what is essential. Establish the dialogue. Introduce music when the scene requires it. Add environment to create space and context. Bring in effects deliberately, allowing each contribution to justify its presence. This approach makes purpose part of the design process before complexity accumulates.

    Spatial sound presents the same challenge on another scale. Contemporary systems allow designers to place and move sounds through increasingly elaborate speaker arrangements, but movement acquires meaning only through its relationship with the story. Surrounding listeners with constant activity can make spatial information less expressive rather than more. If every sound moves, movement itself loses significance. Space becomes useful when its behaviour supports attention, perspective or dramatic intention.

    For Once Upon a Time in Hollywood, the absence of a conventional composed score created a distinctive set of possibilities. Music arrived through records, radio and material connected closely with the period represented by the film. Sound design could occupy areas that might otherwise have belonged to score, contributing low-frequency energy, changes in texture and transitions that influenced mood without announcing themselves as musical cues. The team did not begin by asking how every capability of contemporary soundtrack production could be demonstrated. They considered what would make the film feel connected to 1969.

    That thinking also informed Stateman’s use of Dolby Atmos. He regarded the format as an impressive creative environment, but not as a reason to send objects continually moving through the auditorium. Early experimentation with expanded spatial systems could become cluttered when additional speakers were treated as spaces waiting to be filled. Stateman instead described a stable foundation with carefully selected individual elements used when spatial movement genuinely contributed to the experience.

    For a film drawing heavily upon period recordings, radio and two-channel music sources, aggressive object movement could have conflicted with the aesthetic world being created. A technically advanced format can sometimes serve a film most effectively by concealing its sophistication. As with the single movement of the Terminator 2 platform, possibility acquires value through selection.

    Working with very different directors reinforced Stateman’s resistance to universal solutions. The sonic worlds of John Hughes and Oliver Stone required radically different forms of expression. Moving between comedy and films such as JFK, Born on the Fourth of July and Natural Born Killers prevented one successful method from becoming a formula applied repeatedly. An early mentor had encouraged Stateman to approach every project as a new problem and continue trying different methods rather than relying upon established answers.

    Experience, from this perspective, should expand the vocabulary available to a practitioner rather than narrow the range of possible responses. A Quentin Tarantino film does not require the same sonic logic as an Oliver Stone film, and neither can be approached as a variation of a John Hughes comedy. Even repeated work with the same director changes as the project changes. Trust allows a shorthand to develop, but familiarity should not turn into repetition.

    Comedy offered a useful example of this contextual thinking. Familiarity can establish a pattern, while an unexpected interruption creates surprise. Yet the same broad relationship between expectation and disruption can also support drama, suspense and horror. A technique has no fixed emotional meaning outside its context. What matters is the audience’s developing expectation and the moment at which the soundtrack confirms, delays or overturns it.

    Such decisions require somebody to understand more than individual sounds, which led towards the most distinctive professional idea in Stateman’s lecture. He argued that contemporary practitioners should think of themselves not only as sound designers, but as sound directors, sound producers and sound designers. These are not three disconnected jobs. They represent different perspectives on a shared responsibility for the complete sonic experience.

    The sound director understands intention. This person can discuss the desired experience with the filmmaker, interpret creative needs and establish an overall direction for the soundtrack. The sound producer understands resources. Schedules, budgets, staffing and priorities determine what can be achieved and where effort should be concentrated. The sound designer turns those intentions and resources into creative work. Stateman regarded the combination of these abilities as a professional ideal.

    Directors themselves often move through a similar expansion of responsibility. A writer may become a director to shape the interpretation of the material, then become a producer to gain greater influence over resources and priorities. Stateman argued that sound professionals can develop in a comparable way. Technical expertise remains essential, but understanding intention, organising people and resources and making creative decisions gives the practitioner a more substantial role in shaping the work.

    His career offers several examples of this expanded responsibility. Building large creative teams required production thinking. Designing complex theme-park attractions demanded an understanding of playback systems and engineering. Long-term relationships with directors depended upon interpreting intention rather than waiting for technical instructions. Global distribution and localisation required the sound team to consider what happens after the supposedly finished soundtrack leaves the mixing stage.

    At that point, the problem of control becomes larger. Stateman may work in a highly controlled Atmos environment containing an extensive loudspeaker system, while an audience member eventually hears the result through earbuds on a train. Another listener watches through a television with speakers facing towards a wall. Somebody else lies in bed listening quietly while another person sleeps nearby. A theatrical soundtrack mixed at a high reference level may later be heard at a dramatically lower domestic level.

    No single master can guarantee an identical perceptual experience across all of those conditions. Low-level detail that remains clear in a cinema may disappear when the entire soundtrack is turned down. Wide dynamics that feel exciting in a theatre can become impractical late at night. Dialogue that is intelligible through one system may become difficult to follow through another. Translation therefore concerns more than whether a format can technically fold down from one speaker configuration to another. It concerns what survives perceptually when the listener, environment and playback level change.

    Even cinemas introduce uncertainty. Stateman described discussions with theatre owners about films being played below their intended reference levels. Their explanation was practical. A film begins at the specified level, somebody complains, and the level is reduced. Further complaints lead to further reductions until complaints stop. From the exhibitor’s perspective, this is a rational response to the audience in the room.

    Production can create pressure in the opposite direction. Filmmakers listening in a controlled mixing environment may repeatedly ask for greater excitement, prompting levels to rise until the result satisfies the room. One part of the system pushes upwards while another turns the finished film down. The problem cannot be solved by allowing dialogue, music and effects departments to maximise their own material independently. Somebody needs to maintain responsibility for the dynamic shape of the complete experience.

    Dynamics, in this sense, organise attention across time. Genuine intensity requires quieter material around it. Low-level information needs to remain meaningful under realistic playback conditions, while moments of scale need enough contrast to feel exceptional. The same principle that made one movement of the Terminator 2 platform effective applies to the soundtrack more broadly. Constant intensity reduces the expressive power of intensity itself.

    Global distribution introduces another kind of variation. A soundtrack may need to function across many languages, each with different rhythms, durations and vocal characteristics. Preserving quality across those versions is not merely an administrative task completed after the creative work. It is an audio-design problem involving performance, mixing, workflow and technology.

    Stateman’s approach was to break an apparently overwhelming challenge into smaller problems without losing sight of the whole. A modern feature soundtrack can require enormous quantities of editorial and mixing time, making meaningful control by one person impossible. Large teams become necessary, yet specialisation creates the risk that individuals see only their own component. The sound director-producer-designer model provides a means of maintaining overall intention while complex work is distributed among specialists.

    Localisation makes the limitations of a single fixed soundtrack particularly visible. Different languages alter timing and vocal behaviour, while different audiences and distribution systems introduce further variables. Streaming services now possess forms of information and control unavailable to earlier theatrical systems. Different languages and formats can already be delivered to different users. Stateman’s discussion suggested a logical extension of that capability: audio systems that respond more directly to individual listeners and the circumstances in which they are listening.

    Rather than treating a soundtrack as one fixed master expected to survive every possible situation, elements could remain available for intelligent recombination. Dialogue might receive greater prominence where needed. Dynamic behaviour could change for quiet listening. Different relationships between elements might be created according to the user’s environment, hearing ability or preferences.

    Such possibilities do not necessarily weaken creative intention. They raise a more fundamental question about what preserving intention actually means. A listener who cannot understand the dialogue is not receiving the intended experience merely through being sent the same electrical signal as everyone else. Someone listening quietly in bed occupies a different perceptual situation from an audience inside a calibrated cinema. Identical delivery does not guarantee equivalent experience.

    Adaptive audio therefore extends rather than abandons Stateman’s earlier principles. If the purpose of sound design is to guide attention, communicate story and create emotional relationships, then changes in listening circumstances matter. The challenge is to determine which aspects of an experience must remain stable and which can change to preserve them. A theme-park attraction, theatrical soundtrack, localised release and adaptive audio system appear very different, yet each requires designers to think about the complete journey between creative intention and audience experience.

    Across Stateman’s lecture, simplicity emerged not as an absence of complexity, but as its successful organisation. Hundreds of people may contribute to a project. Thousands of hours may be spent editing and mixing. Sophisticated engineering may sit behind the delivery system. The audience does not need to experience any of that as complication. They need to understand a voice, feel a change in atmosphere, anticipate an event or be surprised by a sudden movement.

    A film soundtrack can contain vast numbers of tracks, edits, recordings and processing decisions while resolving into a clear moment of attention. Achieving that clarity requires more than technical skill. Designers need to understand what matters within the scene. Producers need to organise the resources that make the work possible. Sound directors need to maintain the relationship between individual decisions and the complete experience. Teams need complementary expertise and enough trust for difficult problems to be handed to people capable of solving them.

    Stateman’s lecture ultimately presented sound design as the organisation of attention through time. The designer decides what matters now, what can wait, what should disappear and what should arrive unexpectedly. The producer makes those decisions achievable. The director maintains an understanding of why they matter. Technology expands the available possibilities, but cannot decide which ones belong in the experience.

    That distinction becomes more important as audio technology develops. Spatial formats offer increasingly detailed control inside theatres and listening rooms. Streaming services distribute content globally and create new demands for localisation. Object-based systems can preserve elements beyond a fixed master, while adaptive delivery may eventually allow soundtracks to respond more intelligently to listeners and environments. Each development increases possibility, but also increases the need for judgement.

    Somewhere beyond all of those systems, a listener is trying to follow a story. They may be sitting inside an elaborate cinema, travelling on a train, watching a laptop or lying quietly in bed. The sound professional cannot control every condition surrounding that experience, but can understand the relationships that matter. Dialogue can remain clear. Attention can be guided. Complexity can be organised. Dynamics can create contrast rather than exhaustion. Technology can serve intention rather than advertise itself.

    The response to uncertainty is not complete control. It is to design intelligently for variation while remaining clear about intention. Build teams capable of solving problems that no individual could solve alone. Understand the whole experience while taking responsibility for one part of it. Use complexity behind the scenes to create clarity for the listener. Know the story well enough to recognise when a convention should be followed, when it should be overturned and when the most powerful use of a technology is to leave it silent.

    A soundtrack may be shaped inside a carefully controlled room, but it does not remain there. It travels through cinemas, languages, formats, devices, environments and listeners. Every stage changes the circumstances in which the work will be experienced. Stateman’s lecture suggested that the future of sound design lies not in pretending those differences can be eliminated, but in understanding them well enough to preserve what matters. The mix may be finished when it leaves the studio. The experience begins when somebody, somewhere, presses play.

  • How Do You Design the Sound of Reality? Sefi Carmel on Documentary Sound, Perception, and the Ethics of Construction

    Sefi Carmel

    How do you design the sound of reality?

    A whale dives beneath the surface of the Atlantic Ocean. Filmed from a distant boat, its tail disappears into the water and a splash is heard. The moment appears entirely natural, yet the camera crew may have been hundreds of metres away, surrounded by engine noise and incapable of recording anything resembling the sound presented in the finished film. Perhaps the splash came from a sound-effects library. Perhaps somebody dropped an object into water. Perhaps a Foley artist moved a scuba flipper through a bathtub. Does adding that sound make the documentary less truthful, or does it help audiences experience an event that genuinely occurred but could not be captured adequately during filming?

    During his online guest lecture for Edinburgh Napier University, London-based sound designer, composer and dubbing mixer Sefi Carmel explored the creative and ethical questions surrounding soundtrack creation for documentaries. Drawing upon experience of mixing more than one hundred documentaries, he challenged the assumption that factual filmmaking requires a fundamentally different sonic vocabulary from drama. Dialogue, music, atmospheres, spot effects, Foley and abstract sound design can all contribute towards documentary storytelling. The crucial issue is not whether a sound was recorded at the moment shown on screen, but what its addition asks the audience to believe.

    Throughout the lecture, Carmel returned to a broad understanding of the soundtrack. Everything emerging from the speakers belongs to a single composition created in relationship with the image. Dialogue, music and effects may be separated technically, but audiences experience their combined movement through time and space. Like music, a soundtrack arranges rhythm, dynamics and timbre to communicate emotions and ideas. Documentary sound therefore involves much more than cleaning interviews and placing music underneath them. It is the construction of an audiovisual experience whose materials may come from reality while their organisation remains an act of filmmaking.

    Carmel began by questioning a belief he had held as a young sound designer. News reporting and documentary filmmaking had once appeared to occupy similar positions on a spectrum between reality and fiction. At one extreme, news aspires towards an account of events with minimal manipulation. At the other, drama openly asks audiences to accept scripted performances, constructed edits, designed sound effects and music intended to influence emotion. Documentary can initially appear closer to the first model, yet Carmel’s experience of the form led him towards a different conclusion. A feature-length documentary still has to hold an audience for sixty, ninety or more minutes. It communicates ideas, develops relationships, establishes places, controls pace and creates emotional movement. Those demands make it filmmaking rather than an extended news report.

    Understanding documentary in those terms considerably widens its creative possibilities. Music can shape emotional interpretation, sound effects can strengthen actions and atmospheres can establish locations that production recordings fail to communicate. Archive footage can be reconstructed into a convincing audiovisual world, Foley can restore physical detail and abstract sound design can emphasise an important transition or idea. Documentary makers have access to almost the complete filmmaking toolbox. With that freedom comes an ethical problem. If documentary claims a relationship with reality, how far can its soundtrack depart from literal recording before enhancement becomes deception?

    The whale provides a useful test. The animal really entered the water and its tail really created a splash, even though the filmmakers could not capture that sound from their position. Adding a plausible splash does not invent the event. It reconstructs an acoustic consequence already implied by the image, allowing audiences to experience the animal’s movement and scale more immediately without asking them to believe that something happened when it did not. Matters become more difficult when sound influences interpretation rather than restoring an unheard event. Carmel drew his clearest ethical line around speech. Reconstructing the likely sound of an action differs substantially from editing somebody’s words in a manner that changes their meaning or context. The latter can alter the evidence from which audiences understand an event. For Carmel, creative sound design remains legitimate when it serves the film without becoming a gross lie.

    A much less serious encounter with perceptual truth came during his work on a documentary about the Winter Olympics. Faced with a distant shot of a skier slaloming down a snowy slope, the director wanted the movement to have greater sonic presence. Library searches failed to produce a suitable recording, so Carmel created one himself. The director liked the result and repeatedly asked what had produced it. Carmel initially refused to answer, but eventually revealed the source: a knife scraping across toast. The sound had worked perfectly while the director perceived it as skiing. Once its origin became known, however, the illusion collapsed. The director could hear only toast and eventually asked for it to be removed. Nothing in the waveform had changed. Image, expectation and context had previously allowed one event to become another, while knowledge of the source created a new and apparently irreversible interpretation.

    Listeners do not identify sounds solely through their acoustic properties. Visual information, expectation, context and prior knowledge shape perception. A recording does not acquire credibility merely through sharing a physical origin with the object shown on screen, and a constructed sound does not automatically become deceptive through coming from somewhere else. Credibility emerges from the relationship between sound, image and meaning. Documentary sound design occupies a space between physical truth and perceptual credibility, requiring the designer to consider where a sound came from alongside what audiences understand it to represent.

    Reality presents another difficulty long before questions of creative enhancement arise. Documentary crews work in environments that cannot be controlled in the manner of a drama production. Interviews happen near roads, beneath aircraft routes, inside noisy buildings and beside machinery. Important moments may occur only once, leaving post-production dependent upon whatever the location recordist managed to capture. Documentary also has limited access to ADR. Replacing a contributor’s voice with a later studio performance can compromise spontaneity, authenticity and practicality. An imperfect recording of an essential contribution therefore has to be made intelligible and aesthetically acceptable even when the original conditions were hostile.

    Some problems are comparatively manageable. Low-frequency rumble can be reduced, while constant noises such as air conditioning, electrical hum or an aircraft cabin may respond well to noise reduction. Variable interference presents a much harder problem. An accelerating motorcycle can move through the same frequency regions as speech while continually changing its spectral character, making removal difficult without damaging the voice. Aggressive processing introduces further compromises. Equalisation can isolate the most intelligible part of a voice while leaving it thin and unnatural. Noise reduction can remove interference at the cost of audible artefacts. A line may become technically clearer but aesthetically less convincing. Restoration is not a contest to remove the greatest possible quantity of unwanted sound. Every intervention changes the material audiences hear.

    Carmel connected this problem to the idea of aesthetic disturbance. Drawing upon Ludwig Wittgenstein, he described aesthetics partly through the recognition of something being wrong: a picture hanging at an angle, for example, produces a disturbance that disappears when the relationship is corrected. Documentary sound can create similar disturbances. A close image accompanied by an unexpectedly distant, reverberant voice may produce audiovisual dissonance. Distortion attracts attention, while harshness, brittleness or excessive boxiness can make listening uncomfortable even when every word remains understandable. Intelligibility is necessary but insufficient. Dialogue also needs to feel congruent with the image and surrounding soundtrack, and a technically rescued voice can still undermine a scene if its perspective, spectrum or acoustic character appears disconnected from what viewers see.

    Modern restoration tools have expanded what can be recovered. Carmel discussed noise reduction, de-clicking, de-crackling, de-clipping, dereverberation and spectral repair as processes capable of rescuing recordings that might once have been considered unusable. Constant noise can sometimes be reduced substantially, while sophisticated interpolation may reconstruct clipped or distorted speech with surprising effectiveness. Greater capability does not remove the need for judgement. Restoration can erase information that belongs to the story. Crackle on an old recording may be technically undesirable while simultaneously communicating historical distance. The noise of an old shellac disc or archive recording can help audiences understand that they are hearing material from another period. Removing every imperfection may weaken the narrative. Technical possibility matters less than understanding what the existing sound already communicates.

    This tension between repair and preservation led to one of Carmel’s central principles: ideally, every layer of the soundtrack should be capable of telling the story. Dialogue carries narrative through words. An atmosphere can communicate location, weather, time of day and activity without spoken explanation. Music can reveal emotional direction. A forest atmosphere containing birds, wind and running water immediately places listeners within a particular kind of environment. Each layer contributes different information, and the complete soundtrack emerges from their interaction rather than from one dominant element surrounded by decoration.

    Documentary production rarely provides ideal materials, so those layers often have to support one another. Aggressively cleaned dialogue recorded on a windy location may no longer contain enough environmental information to make its setting believable. Carefully constructed atmospheres can return that context. Sound effects can reinforce actions the original recording failed to capture, while music can support emotional movement that damaged or fragmented production sound cannot carry alone. None of these layers needs to become conspicuous. A weak recording may become convincing once placed inside an appropriately designed environment, since audiences hear relationships between elements rather than evaluating every track in isolation.

    Voiceover introduces another distinctive element within those relationships. Casting is often a directorial decision, though Carmel argued that sound professionals can contribute valuable thinking when invited into the process. The appropriate voice depends upon the subject, intended audience and emotional character of the film. A documentary about the rise of a young pop group requires a different vocal identity from one exploring humpback whales in the North Atlantic. Performance matters as much as casting. Tone, energy, pacing and character shape the audience’s relationship with the film, while the chosen voice influences the space available for music, effects and atmosphere. Voiceover is part of the soundtrack’s composition, not information simply placed above it.

    Music performs an equally integrated role. Carmel distinguished between music existing within the world shown on screen and music functioning as score. Source music might come from a visible performer, radio, jukebox or other plausible location within the scene, while score operates outside that visible world and shapes emotion from another level. Documentary makers can also blur the distinction deliberately. An old rock-and-roll song treated with band-limiting and room reverberation might appear to come from a jukebox in a diner, allowing music to reinforce the setting, contribute historical or cultural information and perhaps help mask weaknesses in the location sound.

    Selecting or composing music requires an understanding of everything else occupying the scene. Dialogue-heavy sequences need music capable of supporting speech without continually competing for attention. Dense lead instruments or vocals can occupy perceptual and spectral territory similar to the human voice, forcing the music much lower in the mix. More restrained arrangements can create emotional colour while leaving space for narration and interviews. Other sequences allow music to carry more of the narrative. An expansive aerial view with little dialogue may support a large thematic statement that would overwhelm an intimate interview. Musical effectiveness in isolation matters less than the role a piece needs to perform at a particular moment and the relationships it forms with the rest of the soundtrack.

    Location dialogue complicates those relationships further. A controlled voiceover recording tends to maintain comparatively stable level and performance, while spontaneous speech can vary considerably. Contributors may begin sentences with energy and trail away as thoughts conclude. Music automation must respond to those changing patterns. A static reduction may leave quieter words obscured or make stronger phrases feel unnecessarily exposed. Mixing becomes a continual negotiation between intelligibility and musical continuity, with the soundtrack moving around the natural behaviour of voices that were never performed for the convenience of the mixer.

    Atmospheres perform several roles simultaneously. Most obviously, they tell audiences where they are. Traffic, birds, wind, room tone, distant machinery or human activity can define an environment before viewers consciously analyse the image. They also smooth editorial transitions. Documentary scenes are frequently assembled from material recorded at different moments, positions or even days, and a continuous environmental bed can help separate pieces of location sound feel as though they belong to a coherent space. Carmel identified another, less obvious function: atmospheres can contribute spectral balance. If a scene feels sonically empty within a particular frequency region, an appropriate environmental layer can help create a more aesthetically satisfying whole. The choice still needs to make narrative sense, but storytelling and sonic composition overlap here. Atmosphere can provide information, continuity and texture at the same time.

    Spot effects operate on a more local scale. A car door closes, a telephone is placed down or a gun fires. Brief synchronised events can reinforce visible actions and restore details absent from production recordings. Archive footage provides particularly rich opportunities for this kind of reconstruction, especially when historical images arrive without usable synchronous sound. Old footage may be silent or accompanied by narration and music unsuitable for the contemporary documentary. A battlefield sequence showing tanks, artillery and soldiers therefore presents the sound designer with an empty world that needs to be rebuilt.

    For Carmel, archive reconstruction can be approached with the same dramatic ambition used in fiction. Tanks can advance through the frame, gunfire can occupy different distances and artillery can establish scale, while wind across an exposed landscape gives the environment continuity. The designer might process the soundtrack to suggest historical recording technology or create a vivid modern sound world that places the audience imaginatively inside the event. A deliberately aged soundtrack reminds viewers that they are encountering archive material, while a contemporary reconstruction can reduce historical distance and make an event feel immediate. Neither choice is neutral. Both interpret history, leaving the designer to consider the relationship the documentary seeks to create between the audience and the past.

    Foley can contribute in much the same way, although documentary schedules and budgets rarely permit complete coverage. Selective use can still transform significant moments. A historical reconstruction showing chainmail being worn, armour handled or a sword drawn may deserve detailed physical sound even when none was captured during filming. If the moment carries narrative importance, there is no reason to reject Foley merely through an assumption that documentary sound must remain limited to location recordings. Abstract sound design extends the same principle further. Drones, impacts and heavily processed transitions can strengthen important ideas or structural moments, creating unease, giving a cut greater dramatic force or helping an audience experience a transition emotionally as well as intellectually.

    For Carmel, the legitimacy of these devices depends upon purpose. Dramatic sound should strengthen the film rather than substitute manipulation for argument. Documentary editing, cinematography and music already influence how audiences understand material, and sound design participates in the same process. Ethical responsibility lies in recognising what each construction communicates rather than pretending that construction does not occur. Documentaries are built from real people, events, evidence and places, yet films do not emerge automatically from those materials. Someone selects shots, orders sequences, chooses where music begins, decides when silence matters and determines which details audiences hear. Soundtrack creation is part of that authorship.

    Creative decisions are only part of the work. The documentary must also survive the circumstances in which it will be heard. Carmel emphasised that mixing begins with the destination. A theatrical documentary, television broadcast, online film and festival screening present different playback conditions and technical expectations. Format, dynamics, equalisation and level decisions need to reflect those contexts. A large theatrical environment can support substantial low-frequency extension and wider dynamics, while television playback may occur through much smaller speakers and in less controlled surroundings. Processing that creates clarity in one context can become harsh or excessive in another.

    Platform awareness connects technical delivery directly with audience experience. Spectral energy inaudible on small television speakers can still consume headroom. Extreme dynamics may work beautifully in a cinema while causing viewers at home to continually adjust volume. Loudness standards formalise part of that relationship, particularly for broadcast delivery, but Carmel’s broader point concerned the distribution of intensity across the film. Loudness can be understood as a budget. If every moment is treated as maximally intense, little room remains for genuine peaks. Dynamic planning becomes another form of storytelling, allowing contrast to carry dramatic meaning rather than treating level merely as a compliance problem.

    Compression and limiting require similar contextual judgement. A theatrical mix may use comparatively subtle master processing, preserving headroom and contrast, while television and online material may tolerate or require greater control. One version is not inherently superior to another. Each mix needs to function within the medium for which it is intended, preserving the film’s intentions under different listening conditions.

    Deliverables extend that responsibility beyond the primary audience. Documentary films may travel between territories and require new narration or dubbed dialogue. Music and effects tracks need to support localisation rather than reproduce automation created around the timing of the original language. Carmel explained the value of providing undipped music for this purpose. In an English version, music may be reduced beneath a particular phrase and raised again when the speaker stops. A translated version may take longer or shorter to communicate the same idea. If the music stem already contains automation tied to the English timing, the foreign-language mixer inherits a structure that no longer fits. Providing material without those dialogue-specific reductions allows the new mix to respond properly to the translated performance.

    A soundtrack therefore exists as more than a finished mix. It may need to survive new languages, platforms and contexts, and good delivery anticipates the work of people who will encounter the material later. Carmel ended with an even simpler responsibility: check the work. Quality control may appear less intellectually exciting than documentary ethics or perceptual sound design, yet it protects every creative decision made before delivery. A mixer can spend days constructing a sophisticated soundtrack and still send an unusable file through a routing mistake or export error. Recording or exporting something does not prove that the expected material exists in the resulting file. Listen to it. Watch it. Check it.

    That practical instruction sits neatly alongside the lecture’s larger argument. Documentary soundtrack creation moves constantly between interpretation and responsibility. Designers can construct sounds that were never recorded, rebuild silent archives, use Foley, shape emotion through music and introduce dramatic sonic devices. Those freedoms demand judgement. Does a sound restore an experience, clarify it, interpret it or falsify it? What does it ask the audience to believe? Does it support the film’s argument without altering the meaning of its evidence?

    Carmel’s lecture presented documentary sound as an art of constructing experience from incomplete reality. Location recordings arrive damaged. Cameras capture actions from distances at which their sounds cannot be heard. Archive images survive after their original acoustic worlds have disappeared. Interviews need to coexist with music, while fragmented scenes require atmospheres capable of making them feel continuous. The documentary soundtrack is built through responses to these absences. Sometimes the response is technological: remove a constant noise, repair distortion or restore intelligibility. Sometimes it is editorial: create a continuous atmosphere around fragmented material. Sometimes it is performative: add Foley to a significant physical action. Sometimes it is musical, shaping the emotional direction of a sequence. At other moments, the solution may be a completely unrelated object whose acoustic behaviour happens to make an image believable.

    A knife scraping toast can become a skier moving across snow, at least until somebody learns the secret. Reality does not arrive in post-production as a complete audiovisual object waiting to be preserved. It arrives as recordings, images, testimony, fragments and absences. Filmmakers decide how those materials should be organised into an experience that audiences can follow and understand. Sound designers participate in that process by reconstructing relationships between actions and consequences, voices and spaces, images and expectations.

    Construction and dishonesty are not the same thing. A documentary soundtrack can be richly designed while remaining faithful to the people and events it represents. Literal accuracy may sometimes be essential. At other moments, perceptual credibility communicates an experience more effectively than an unusable or absent location recording ever could. A splash can give weight to a whale entering the ocean. An atmosphere can return a damaged interview to its environment. Designed sound can give silent archive footage physical immediacy. Music can reveal emotional relationships without changing the evidence shown on screen.

    Documentary sound occupies the space between what happened, what could be recorded and what audiences need in order to understand and feel the film. Carmel’s lecture showed that this space is not a technical inconvenience to be hidden. Much of the creative work begins there. The documentary sound designer cannot preserve every sound of reality, since many were never captured in the first place. The responsibility is to decide what should be repaired, what can be reconstructed, what needs to remain imperfect and what must never be changed.

  • How Do You Make a Game Feel Dangerous? Will Morton on Emotion, Attention, and Designing Sound for the Player

    Will Morton

    How do you make a game feel dangerous?

    A gun can sound enormous and still fail to make a gunfight frightening. Every weapon may have a powerful attack, convincing mechanical detail and an impressive environmental tail, yet the player can remain strangely detached from the danger. Solving that problem may have little to do with redesigning the weapon itself. Bullets pass close to the head. Impacts strike nearby walls with exaggerated force. Brickwork breaks apart, fragments scatter and the environment appears to react violently to the threat. During his online guest lecture for Edinburgh Napier University, game audio designer Will Morton explored how sound can shape emotion, focus attention and guide players through complex interactive experiences. Drawing upon twelve years at Rockstar North, where his work included the Grand Theft Auto series, Red Dead Redemption and L.A. Noire, followed by the establishment of Solid Audio Works with fellow former Rockstar audio specialist Craig Connor, Morton presented game sound design as a discipline of selection. Thousands of sounds may exist within a game, but their value depends upon knowing which ones matter at any particular moment. Throughout the lecture, one principle repeatedly emerged. A designer must ask not only what a game world should sound like, but what the player needs to hear in order to feel what the game intends them to feel.

    Morton began by placing creative decisions within the realities of AAA game production. Large games are expensive, technically constrained and continually changing. Development rarely follows a fixed design from beginning to end. Features evolve, producers reconsider decisions and new requests arrive late in production. Platform restrictions impose further limits through storage, memory, streaming performance and processing capability. Scale introduces organisational complexity as well. Larger games require larger teams, while experienced creative specialists can find increasing amounts of their time absorbed by scheduling, administration and coordination. Sound design develops inside a moving system of technical, financial and organisational constraints. Success requires more than imagining an ideal soundtrack. Designers must create one capable of surviving years of changing requirements.

    Historical changes in technology have altered the scale of those constraints without eliminating them. Morton recalled Commodore 64 composer Martin Galway fitting numerous sound effects into approximately one kilobyte of memory. Restrictions of that magnitude appear almost comic from the perspective of contemporary production, yet modern open-world games can still leave audio teams fighting for storage and memory. Vastly greater resources are now available, while games simultaneously attempt to represent entire cities, landscapes and fictional worlds. Technological abundance creates new possibilities, but ambition expands alongside it. Designers still have to decide where limited resources will make the greatest contribution.

    Money introduces another set of choices. Sound effects require designers, recording equipment, locations, editing time, libraries, software and continually changing computer systems. Dialogue adds writers, actors, directors, studios, recording staff and extensive editing. Music may involve composition, licensing, performers, orchestras, specialist recording facilities and interactive implementation. Morton’s overview exposed the consequences hidden behind apparently simple creative ambitions. Another recording variation, character voice or interactive music feature consumes time, money, memory and attention that cannot be spent elsewhere.

    Dialogue provided one of the clearest examples of complexity hiding behind familiar production tasks. Morton spent much of his Rockstar career dividing his time between sound design and dialogue before the scale of Grand Theft Auto V led him to work entirely as dialogue supervisor. For story-heavy games, dialogue cannot be treated as a sequence of lines requested by designers and recorded by actors. Repetition needs consideration. Lines must make sense across changing gameplay situations. Story information may unfold over many hours or days of play, while players can interrupt, delay or alter the circumstances in which dialogue occurs. A dialogue designer therefore needs to understand the interactive structure of the game as deeply as the individual performances being recorded.

    Direction presents a related challenge. Experience in film and television can produce excellent performances, yet game dialogue introduces problems absent from linear media. Cutscenes may operate much like conventional scenes, while in-game dialogue can occur within changing gameplay circumstances and across story structures experienced differently by individual players. Directors need more than an ability to elicit compelling performances. Detailed knowledge of the script, game and eventual context of each line becomes essential. A performance recorded in isolation must later remain convincing within circumstances that may not even be visible inside the recording studio.

    Morton also challenged assumptions about professional recording. Technically excellent dialogue can be captured with comparatively modest equipment in a carefully treated space. A suitable microphone, simple accessories and an acoustically controlled room may produce results approaching those of a far more expensive studio. Recording quality, however, forms only one part of a professional session. High-profile performers need confidence that they are participating in a serious production. Environment, organisation and treatment of the actor can influence trust in the project and relationships with agents and management. Professionalism encompasses the experience surrounding a recording alongside the technical quality of the resulting file.

    From these production realities, Morton moved towards the central creative argument of the lecture. Designers can easily assume that everything visible in a game should automatically produce a sound. Faced with an unsounded world, teams begin filling every action with detail. Cars receive engines and collisions. Characters acquire footsteps. Objects gain interactions. Environments fill with ambiences. A technically comprehensive and entirely plausible soundtrack gradually emerges, yet plausibility alone cannot guarantee clarity, excitement or emotional effect.

    Practical concerns provide one reason for restraint. Every additional sound requires creation, editing, implementation, memory and testing. Artistic considerations are even more important. If everything demands attention simultaneously, nothing receives focus. Morton encouraged designers to identify what deserves to be heard and what can remain absent. Important sounds require room to breathe. Dynamics rely upon contrast, while emotional emphasis depends upon moving attention between elements. Silence and omission become active design decisions.

    Player experience consequently takes priority over literal acoustic reconstruction. Real events often sound less dramatic than audiences expect. A gunshot captured from a particular position may seem surprisingly small. A suppressed weapon does not necessarily produce the familiar cinematic whisper audiences have learned to associate with it. A real minigun may collapse into an almost continuous mechanical roar instead of revealing every stage of its operation. Decades of film, television and games have established sonic conventions that now influence how audiences expect objects and events to behave. Designers work within that accumulated perceptual history.

    Morton demonstrated the point by comparing cinematic and real recordings of suppressed firearms. A familiar designed version sounded short, controlled and immediately recognisable, carrying characteristics audiences strongly associate with a silenced weapon. Real recordings behaved quite differently. Neither approach offered a universal answer. Documentary representation might favour acoustic accuracy, while an action game may need immediate recognition, excitement and dramatic satisfaction. Before designing a sound, Morton considers the role accuracy should play, whether an event needs to feel larger than reality and how strongly audience expectations should influence the result.

    Miniguns provided an even clearer illustration. Morton discussed the famous weapon sequence in Predator, where mechanical movement, spinning and other details reinforce the spectacle shown on screen. Real recordings present a very different impression, dominated by the extraordinary density of rapid gunfire. Grand Theft Auto: Vice City and Grand Theft Auto V offered further interpretations, each constructing the weapon differently, while Terminator 2 adopted another cinematic approach that remained closer to aspects of the real sound. Comparing them did not reveal a correct minigun. Instead, their differences showed how each design serves the experience of a particular production.

    Reality and relevance are therefore separate considerations. A game designed as escapism may gain little from reproducing everyday acoustic experience with complete fidelity. Players do not always need to hear events from the perspective of ordinary observers. They need a version that communicates the event’s importance within the game. Sound design selects, enlarges, simplifies and reshapes reality according to dramatic purpose.

    Attention can be directed just as effectively through subtraction. Morton used a remotely detonated explosive from Grand Theft Auto IV: The Ballad of Gay Tony to demonstrate how removing surrounding sound can make a single event dominate awareness. Immediately before the explosion, the wider soundtrack recedes and the bomb’s warning becomes the focus. He compared the moment with the seismic charge sequence from Star Wars: Episode II, where a brief interruption in the expected sound field creates anticipation and gives the subsequent event greater impact. Both sequences derive part of their power from absence.

    Focus involves more than increasing the level of an important sound. Every event competes with its surroundings. Removing distractions can achieve more than additional layers, greater loudness or further spectral exaggeration. Presence acquires meaning through contrast with absence. A fraction of a second of reduced activity can prepare an event more effectively than a continuously dense soundtrack.

    Morton’s most revealing example concerned a producer asking for guns to sound more dangerous. Existing weapon sounds already seemed successful to the audio team. They contained convincing mechanical detail, a strong initial attack, substantial body and satisfying environmental tails. Reworking those qualities did not solve the problem. Although the request appeared to concern the sound of the guns, the underlying dissatisfaction was emotional.

    An unexpected experience away from the studio suggested another approach. During a paintball game, Morton found himself sheltering behind structures made from metal oil drums. The paintball markers themselves produced relatively insignificant sounds. Fear came from sudden, violent impacts striking the metal around him. His immediate environment appeared to contract as attention focused upon nearby impacts, resonances and the sense that projectiles were arriving from directions he could not fully control. The weapon itself was not frightening. Being under fire was.

    Recognising that distinction transformed the design problem. Bullet passes became more prominent, with their levels responding more strongly to proximity. Impacts against brick and concrete gained force. Low-frequency energy added physical weight, while debris and crumbling material made the environment appear to react to nearby gunfire. Details that might be acoustically subordinate to a real gunshot were deliberately brought forward. Stronger weapon sounds had never been the answer. Gunfights needed to communicate vulnerability and danger.

    Here Morton identified a broader professional skill. Directors and producers rarely describe every sound problem in acoustic terms. They may ask for something to be louder, bigger, darker, faster or more dangerous while expressing dissatisfaction with an emotional result. Following the literal wording can send a designer towards the wrong solution. Morton argued for interpreting the intention behind the request. Someone asking for a more dangerous gun may actually be asking to feel vulnerable. Once the desired experience becomes clear, the designer can decide which part of the sound world needs to change.

    Creative collaboration therefore requires translation between intention and acoustic action. Sound professionals develop specialised vocabularies for frequency, dynamics, spatial behaviour, envelope and processing. Producers may describe experiences through emotion, imagery or metaphor. Useful information can exist within both forms of language. Professional expertise includes moving between them and identifying the experience concealed inside an apparently vague request.

    Gunfire also demonstrates how game audio operates through systems of cause and consequence. A weapon consists of more than the sound emitted when a trigger is pulled. Projectiles move through space. Near misses pass the player. Bullets strike walls. Debris falls afterwards. Environments respond differently according to material and distance. Danger emerges from relationships between these events. Concentrating exclusively upon the source can leave the wider experience emotionally incomplete.

    Similar systems govern the game mix. Morton described mixing a large game as a potentially overwhelming process involving thousands of assets and potentially hundreds of simultaneous channels, all changing according to gameplay. Exact combinations cannot be predicted in advance as they can within a linear soundtrack. Dialogue, music, ambience, vehicles, weapons, footsteps and environmental interactions combine differently according to player behaviour. Even excellent individual assets can produce an exhausting result when their relationships are poorly controlled.

    Morton recalled playing games whose soundtracks felt like continuous acoustic assault. Constant density can be as tiring as excessive level. When every category remains active and prominent, listeners receive no opportunity for recovery and little indication of where attention belongs. More detail can therefore produce less communication.

    Working with Craig Connor, Morton developed a useful analogy for managing this complexity: approach the game as a music producer approaches a song. Every element needs an appropriate place and enough space around it. The comparison does not imply imposing a static musical mix upon an interactive system. Its value lies in encouraging relational thinking. Sounds acquire meaning through their positions alongside other elements rather than through isolated perfection.

    Arrangement offers another useful parallel. Music producers rarely expect every instrument to occupy the foreground continuously. Parts enter and leave, density changes and contrast creates structure. Individual elements occupy different spectral, spatial and dynamic roles. Interactive audio can apply similar principles while responding continuously to player action, with priority changing according to what the player is doing and what the game needs to communicate.

    Waiting until asset production is complete before attempting a final mix creates serious problems. By the time thousands of sounds have accumulated, assumptions about level, density and priority may already be embedded throughout the project. Assets designed without a meaningful reference mix can prove difficult to reconcile, particularly when each has been created to sound impressive in isolation.

    Morton recommended an iterative alternative. Early in development, designers can choose a small but representative collection of sounds and make them work together convincingly. A useful reference might include something loud, such as a gunshot or explosion, alongside quieter material such as footsteps. Approximate ambience establishes the environmental bed, while dialogue examples can cover a range from quiet speech through ordinary conversation to shouting. Mixing these elements early creates a working scale for everything that follows.

    New assets can then be designed in relation to an existing sonic framework. Designers know roughly where a sound needs to sit and how much space surrounds it. Original reference assets may eventually be replaced, but their early role remains valuable. They establish relationships before growing complexity makes fundamental decisions harder to change.

    Mixing, from this perspective, becomes part of sound design rather than a finishing process. Working without context encourages every gun, vehicle, impact and interaction to become enormous, detailed and impressive. Once combined, they compete. An evolving reference mix permits more varied decisions. Some sounds can remain small. Others can be narrow, distant or restrained. Detail can be reserved for moments when players have enough space to perceive it.

    Focus also connects sound directly to gameplay. Players continually decide where to look, where to move and what to do next. Audio can support those decisions by drawing attention towards useful information or allowing distractions to recede. An approaching threat, important character, changing environment or imminent event may receive temporary priority. Elements contributing little to the current experience can move into the background.

    Informational and emotional focus can coexist. A soundtrack can communicate danger without identifying an enemy’s exact position, or create suspense without explaining precisely what will happen next. Morton’s examples showed how game audio shapes a player’s state of mind while remaining part of the fictional world. Excitement, suspense, drama and humour are designed responses. A weapon feels powerful partly through its sound. An approaching explosion gains anticipation from the quiet preceding it. A gunfight becomes dangerous when incoming fire appears to tear apart the world immediately around the player.

    Realism consequently remains flexible. Effective sounds may preserve recognisable aspects of reality while exaggerating qualities useful to the experience. A weapon can retain enough mechanical identity to remain believable while gaining additional weight. An impact can exceed its real counterpart without feeling inappropriate to the image. A suppressed firearm can satisfy an established cultural expectation even when that expectation differs from literal acoustic reality.

    Morton resisted turning these observations into universal rules. Predator and Terminator 2 can present radically different miniguns while both succeeding on their own terms. A game pursuing realism may demand a different balance from an exaggerated action title. Intimate narrative experiences may use restraint where large-scale spectacle needs greater sonic scale. Designers need to understand what their particular project is asking players to experience.

    Selectivity also shapes resource allocation. No production can pursue every possible recording session, dialogue variation or interactive feature. Large games generate an almost unlimited number of potential tasks while budgets and schedules remain finite. Designers select which events deserve sound, which sounds deserve prominence, which systems justify development time and which details will genuinely improve the experience.

    Morton’s discussion of sound libraries introduced a longer view of professional practice. Building a useful collection of original recordings is expensive, but opportunities to capture interesting material should be taken when possible. A sound recorded today may remain unused for years before finding its purpose. Game audio professionals develop habits of listening beyond individual projects, continually noticing potential material in the world around them.

    His paintball experience represents an even deeper form of professional listening. Morton did not return merely with a useful recording. He had experienced a relationship between threat, proximity and environmental impact that changed how he understood a design problem. Everyday experiences can reveal how attention and emotion respond to sound. Listening professionally involves recognising such relationships as well as collecting interesting timbres.

    Technical constraints and psychological effects remain closely connected throughout Morton’s practice. Memory budgets, recording costs, dialogue systems, asset creation and mixing strategies eventually converge upon one question: what experience is the player having? Sophisticated technology has limited value when it does not support that experience. Conversely, a simple decision such as briefly removing surrounding sound can transform a moment when it directs attention effectively.

    By the end of the lecture, Morton had presented AAA game audio as a discipline balanced between enormous complexity and deliberate restraint. Contemporary designers have access to more memory, processing, channels and real-time synthesis than earlier generations could have imagined. Additional technical capacity, however, does not remove the need to choose. More available sounds do not make more audible sounds desirable. Processing power cannot determine where attention belongs. Larger worlds make focus increasingly important.

    His account also challenged the familiar suggestion that audiences notice game audio only when it fails. Players respond to powerful sound design even when they cannot describe every mechanism behind it. They feel the danger of a gunfight, anticipate an explosion during a sudden moment of quiet and recognise when a world has enough space to breathe. Good audio can elevate games whose code, visuals and other systems already represent enormous investment. Its contribution reaches far beyond correcting problems. Sound helps determine the emotional character of the experience.

    Morton’s lecture ultimately revealed game sound design as the management of attention through an interactive world. Designers decide when reality should be preserved and when expectation should take priority. Silence can become more powerful than another layer. A producer’s request needs to be interpreted through the emotion it seeks to achieve. A gun may already sound excellent while the gunfight surrounding it remains ineffective. Thousands of assets can exist within a game, yet the success of the soundtrack depends upon knowing which ones deserve attention at a particular moment.

    Perhaps the most important question is simply what the player needs to hear now. Sometimes the answer is a spectacular weapon. At another moment, it is the violent impact of a projectile against the wall beside them. Elsewhere, a quiet footstep, distant ambience or line of dialogue needs enough space to be understood. Occasionally, almost everything should disappear. Game audio becomes powerful when sound shapes experience rather than catalogues events. A virtual world may contain thousands of possible sounds. The art lies in choosing the ones that make the player feel something.

  • How Do You Perform a World Through Sound? Jason Swanscott on Foley, Improvisation, and Making Movement Believable

    Jason Swanscott

    How do you perform a world through sound?

    A character crosses a room, adjusts a coat, places a glass on a table and sits down. Nothing about the sequence appears extraordinary, yet its soundtrack may contain dozens of carefully performed details. Every footstep must carry the correct weight. Fabric must move with the body rather than merely rustle somewhere in the background. The glass must sound appropriate for its size and surface, while the chair should respond convincingly to the character’s movement. None of these sounds is likely to attract conscious attention, yet without them the scene can feel strangely empty. During his online guest lecture for Edinburgh Napier University, veteran Foley artist Jason Swanscott explored the extraordinary craft behind these apparently ordinary sounds. Drawing upon almost three decades of work across film, television and games, he revealed Foley as something far richer than synchronising footsteps and props to pictures. It is a form of performance in which movement, character, materials, objects and physical spaces are interpreted through sound. Throughout the session, one idea repeatedly emerged. Foley does not simply reproduce what can be seen on screen. It performs the physical behaviour of an entire world.

    Swanscott began by establishing the three principal elements of Foley: movement, footsteps and spot effects. This distinction provides a useful practical framework, though each category immediately reveals the interpretive nature of the work. Movement concerns the subtle sounds produced as characters shift their bodies and clothing. Footsteps recreate their interaction with the ground. Spot effects encompass the countless objects they touch, lift, open, close, carry and manipulate. Together, these tracks restore a physical presence that may be missing from production recordings or require greater definition within the finished soundtrack. Their purpose is not to make the soundtrack busier. They make bodies, materials and objects feel as though they occupy the world shown on screen.

    Movement often begins with fabric, but choosing a material is only the first decision. A suit jacket does not behave like a tweed coat, while cotton, silk and heavier period fabrics each possess different textures and patterns of movement. Swanscott watches how the character moves and performs those movements through the selected material. A slight adjustment in a chair requires a different gesture from somebody running, fighting or struggling into a coat. Sound follows action rather than existing as a continuous layer of generic clothing noise. The artist watches the body, interprets its movement and recreates its acoustic consequences.

    Footsteps make the relationship between sound and performance even clearer. Correct shoes and surfaces matter enormously. Foley stages therefore contain collections of footwear and recording surfaces capable of representing different periods, occupations, characters and environments. Leather-soled shoes suggest a different person from combat boots or high heels, while wood, gravel, tile and cobblestone each change the relationship between the performer and the ground. Owning the appropriate shoe and standing on the appropriate surface, however, does not guarantee a convincing footstep. The Foley artist must perform the character. Weight, pace, hesitation, confidence, exhaustion and emotional state all influence the rhythm and physical force of movement. A footstep can communicate who the character is and how they are moving through the scene.

    Spot effects extend this performance across the material world. A hand lifts a glass, a door opens, a weapon is drawn or an object falls onto a surface. Some actions can be recreated using objects similar to those shown in the image. Others require considerable invention. A sword may acquire additional metallic resonance to give its movement greater presence. A body impact may emerge from materials whose physical origins bear little resemblance to the action on screen. Swanscott’s task is to produce the sound that makes the visible action believable, regardless of whether the studio prop resembles its apparent source.

    One of the most interesting themes running through the lecture concerned the difference between physical truth and perceptual truth. Real events do not always produce the sounds audiences expect from them. A literal recording may appear weak, ambiguous or dramatically inappropriate once placed against an image. Foley therefore occupies a curious territory between reality and expectation. The artist frequently uses an object that is physically wrong to create a sound that feels perceptually right.

    Familiar examples are entertaining precisely through the unlikely relationship between source and result. Celery and cabbage can provide organic fractures and breaks. Bananas and watermelons can contribute the resistance and wetness required for sounds involving cutting flesh. Water, washing-up liquid and wallpaper paste can create viscous materials for blood and other bodily effects. Wet newspaper and other organic materials may provide textures required for medical and forensic scenes. Real offal can also be used, yet Swanscott’s examples demonstrate that literal materials hold no automatic claim to greater credibility. Screen reality is constructed through judgement rather than fidelity to the original source.

    This becomes especially apparent in genres where audiences possess strong expectations about events that few people have experienced directly. A cinematic punch must communicate weight and violence immediately. Different materials and performances can be combined to create that impression of force. Leather, boxing gloves, weighted impacts and other layers contribute qualities that the image appears to demand. Listeners perceive a body being struck even though the sound may have been constructed from objects with entirely different identities. Foley succeeds when attention remains on the event rather than the materials used to create it.

    Different productions also demand distinct styles of performance. A period drama such as Emma requires attention to the weight and behaviour of long dresses, delicate gloves, heeled footwear, polished floors and carefully handled objects. Its physical world is controlled and refined, so the Foley performance must support that quality. An action film such as Kick-Ass demands something entirely different. Fights require aggressive bodily performance, impacts, falls, furniture movement and layers of activity that communicate speed and force. Horror presents another challenge, where small sounds can become disproportionately important. A creaking surface, distant movement or restrained scrape may contribute more tension than an obviously dramatic effect. Foley changes with the storytelling language of the production.

    Across these genres, Swanscott must interpret images remarkably quickly. Foley artists frequently begin work without having watched the production in advance. They may receive limited notes and occasional guidance about particular details such as footwear or costume, while much of the decision-making happens as the material appears before them. A contemporary drama, period production, action sequence or fantasy world immediately suggests different requirements, yet the artist must continually decide what deserves performance, which materials might work and how much detail the scene can support.

    Such a working method reveals a form of expertise that is difficult to reduce to written instructions. Swanscott’s practice depends upon an accumulated relationship between sight, movement and sound. Watching a character walk prompts immediate decisions about shoe, surface, rhythm and weight. While performing those footsteps, he may already be noticing the objects that will require attention during the later spot-effects pass. Several forms of attention overlap. He is performing the current sound, watching synchronisation, interpreting character and preparing mentally for sounds that have yet to be recorded.

    Pressure makes this embodied expertise particularly visible. Swanscott contrasted the schedules available to different productions. A forty-five-minute ITV drama episode might receive only two days of Foley work, while a longer BBC production such as Silent Witness may allow around five. Feature films can provide considerably longer schedules, with major productions sometimes allowing several weeks. More time allows experimentation, comparison and repeated refinement. Fast television requires strong decisions almost immediately.

    On a two-day television schedule, movement and footsteps may occupy the first day, with spot effects following on the second. Every character still needs to move through the programme. Shoes still need to match, surfaces still need to change and objects still need to acquire physical presence. The schedule compresses the time available without removing the detail. Speed in this context means more than moving quickly. It depends upon recognising patterns, anticipating requirements and drawing upon a sufficiently developed physical and sonic vocabulary that experimentation can happen almost instantaneously.

    Years of this practice turn Foley stages into unusual archives of material culture. Shoes, fabrics, doors, glasses, tools, weapons, furniture and fragments of apparently unremarkable objects accumulate over time. Their value bears little relationship to monetary worth or original purpose. An object becomes valuable when the artist recognises useful sonic behaviour within it. Swanscott recalled finding a discarded sink outside a studio and bringing it inside. Its metallic resonance later proved ideal for a sound effect. A small paving slab could suggest the movement of a vast stone door. Neither object is limited by its physical scale or intended function once the artist begins to imagine what its sound might become.

    Foley therefore requires a peculiar form of auditory imagination. Swanscott has developed knowledge of how materials behave, then uses that knowledge to recognise relationships between sounds and images that may initially appear unrelated. He hears what an object can become when performed differently, recorded from another perspective or combined with another layer.

    One of the clearest examples came from The Legend of Tarzan. Swanscott recreated the heavy, padded movement of gorillas using boxing gloves. The relationship becomes logical when considered through physical qualities rather than visual resemblance: both can produce broad, soft impacts carrying considerable apparent weight. Discovering that relationship requires an artist to think in terms of behaviour. The boxing glove becomes useful through what it can perform.

    Other problems require materials that can be made to behave in particular ways. A snake moving across a surface can be suggested through materials such as cornflour-filled pillowcases, where the shifting internal texture creates the required movement. Magical or supernatural objects may combine recognisable physical impacts with resonant elements that suggest something beyond ordinary reality. In The Witcher, Swanscott described combining physical sounds with the resonance of a singing wine glass to give an enchanted object both material weight and a supernatural quality. Listeners need to believe that the object occupies physical space while also sensing that it belongs to a world governed by unfamiliar forces. Foley can communicate both qualities at once.

    Science fiction extends this challenge further. Productions such as Intergalactic require sounds for technologies without direct real-world equivalents. Spaceship doors, airlocks, damaged suits and magnetic footwear still need to communicate material, force and function. Heavy work boots combined with a rubbery surface can suggest magnetic contact, while compressed air and vocal techniques can contribute to the impression of air escaping from a damaged suit. Even within imaginary worlds, invention remains grounded in physical performance.

    Visually synthetic productions make this physical grounding especially valuable. Computer-generated creatures, magical objects and futuristic environments remove any possibility of simply recording what happened during filming. Audiences nevertheless expect those worlds to possess internal physical coherence. A creature needs weight. A machine needs resistance. Clothing must move with bodies, and objects must appear to interact with surfaces. Familiar materials give impossible worlds a physical logic that listeners recognise instinctively.

    Sometimes the object alone cannot create that credibility. Acoustic space can become part of the performance. Swanscott discussed work on Captain Phillips, where scenes inside the enclosed lifeboat required a claustrophobic acoustic quality. Performing the Foley conventionally within an open recording space would not reproduce the confined reflections demanded by the scene. The team instead worked inside an overturned water tank, using the enclosure itself to create the required resonance.

    That example broadens the definition of the Foley instrument. The performer and prop exist within an acoustic relationship shaped by surfaces, boundaries, reflections and microphone perspective. For the lifeboat, claustrophobia emerged through the physical conditions of recording. The space itself became part of the performance.

    Large scenes demand another kind of construction. A soldier running through a battlefield is represented by far more than footsteps on dirt. Clothing moves, equipment shifts, buckles rattle and weapons interact with the body. Each component may be relatively simple in isolation, yet together they create the physical complexity of a person carrying weight through an environment. Foley gives the body layers. Listeners may never consciously identify each buckle or movement, but their combined behaviour makes the character feel physically present.

    Close collaboration with sound editors allows this layered physicality to remain flexible. The artist performs movement, footsteps and spot effects, while editors organise, refine and prepare those recordings for the wider soundtrack. Individual elements can be adjusted independently, allowing movement, footsteps and object interactions to find appropriate relationships within the final mix. Additional sound effects or processing may supplement the performance, while the Foley recording provides an organic foundation tied directly to the rhythm of the image.

    That connection to performance becomes especially important when production dialogue has been replaced. ADR can separate dialogue from the incidental sounds originally captured on set. Footsteps, clothing and object interaction may then need to be reconstructed around the replacement dialogue so that the scene regains its physical continuity. Foley restores relationships between bodies and environments that post-production processes may have separated.

    Swanscott’s examples also challenged any assumption that every visible action requires equal sonic emphasis. Foley serves storytelling, and different scenes demand different levels of detail and exaggeration. Restrained handling of porcelain in a period drama communicates something very different from the aggressive materiality of an action sequence. Horror may depend upon an isolated movement sound emerging from otherwise restrained surroundings. Fantasy may require a familiar impact combined with an unfamiliar resonance. Every scene asks the artist to decide how an event should sound and how strongly the audience should experience it.

    Improvisation is central to these decisions, though professional improvisation is far from random. Years spent learning the behaviour of materials, surfaces and objects allow relationships to be tested quickly when an unexpected problem appears. A discarded sink becomes useful through recognition of its resonant potential. Boxing gloves become gorilla feet through their combination of softness and apparent weight. What looks spontaneous from outside the studio rests upon a deeply developed vocabulary of physical sound.

    Swanscott’s own body is equally important. Fight sequences may require vigorous physical movement. Footsteps must be performed with the rhythm and weight of somebody whose body may differ considerably from that of the artist. A delicate character, exhausted soldier and threatening pursuer cannot all be represented through the same neutral walking pattern. The Foley artist inhabits movement without being visible. His body becomes an interpretive instrument.

    Precision and expressiveness therefore exist together. Synchronisation matters. A footstep landing visibly out of time can break the connection between sound and image. Accurate timing, however, cannot rescue a performance with the wrong weight, rhythm or intention. Like other forms of performance, Foley uses technical control to enable expression. The artist must arrive at the correct moment while making that moment belong to the character.

    Interactive media changes the structure within which these performances operate. Swanscott’s work extends into video games, where sounds cannot always be performed against a fixed sequence experienced identically by every player. Games such as Alien: Isolation require movement and environmental sounds that respond dynamically to player behaviour. Footsteps and interactions need variation, while sounds must remain convincing when triggered in different sequences and contexts. The direct relationship between performer and fixed picture becomes a system of potential relationships between recorded material and player action.

    Embodied performance remains central even within that nonlinear structure. Weight, texture and physical credibility still matter. Repetition must be controlled so that repeated actions do not expose the limited number of recordings behind them. Direction and distance gain importance as sounds operate within spatial environments. The player may determine when a footstep occurs, while the Foley artist still determines what that step communicates.

    Changes in production systems place pressure on this kind of embodied practice. Swanscott described compressed schedules alongside continuing expectations for highly detailed soundtracks. Productions that once allowed longer periods for Foley may now expect comparable quality in fewer days. Outsourcing and remote working can reduce the close feedback once shared between Foley stages and other post-production departments, while sound libraries provide increasingly convenient alternatives for some categories of material.

    A library effect and a Foley performance differ most fundamentally in their relationship to a particular moment. Foley is created for a specific body, action and dramatic context. Pace, force, rhythm and texture can change in direct response to the image. Libraries can contain exceptional recordings, but the relationship between a pre-existing sound and a new movement must be constructed afterwards. Foley creates that relationship through performance.

    Technology continues to change how this embodied knowledge is captured and used. Digital workflows have transformed recording and editing. New microphones can reveal forms of vibration and underwater sound that conventional recording methods may miss. Synthetic imagery creates opportunities for invention, while games require nonlinear systems rather than fixed sequences. Swanscott’s practice has adapted continually to these changes. At its centre remains a physical act: watching movement, understanding its dramatic purpose and finding a way to perform its sound.

    Preserving the profession therefore involves more than documenting a catalogue of techniques. A list can identify suitable shoes, useful surfaces and familiar prop substitutions, but it cannot easily communicate the embodied timing required to perform another person’s movement or the instinct that allows an artist to hear a discarded object and recognise its potential. Swanscott emphasised the importance of mentorship and training within a craft learned through watching, listening, trying, failing and gradually developing a relationship with materials.

    Perhaps this is the most revealing way to understand Foley. Spectacular examples naturally attract attention: vegetables become broken bones, boxing gloves become gorillas and an overturned water tank becomes a lifeboat. Yet the deeper skill lies in the decisions connecting those materials to the image. A watermelon has no inherent cinematic meaning. Its value emerges through how it is performed, where the microphone is placed, which qualities of its sound are useful and whether the resulting texture belongs within the world of the scene.

    Swanscott’s lecture consequently revealed a discipline of transformation. Fabric becomes bodily movement. Shoes become character. Small objects become enormous mechanisms. Familiar materials give physical credibility to creatures and technologies that have never existed. A recording space becomes part of a fictional environment, while a performer standing on a Foley stage inhabits the movement of somebody in another place, another period or an entirely imaginary world.

    Finished Foley conceals this constant process of interpretation. Audiences see a character walk and hear footsteps. They see clothing move and hear fabric. They watch an object fall and accept its weight. The connection appears inevitable, even though every sound may have required decisions about material, performance, surface, timing, perspective and dramatic emphasis. As the performance becomes more convincing, the labour required to create that relationship becomes less visible.

    By the end of the lecture, Swanscott had transformed the Foley stage from a room full of shoes, fabrics, surfaces and strange objects into something closer to a physical imagination of the screen. Nothing in the room has only one identity. A boxing glove can become a gorilla. A paving slab can become monumental architecture. A discarded sink can acquire a new purpose through its resonance. The artist’s task is to watch movement, understand character and hear possibilities hidden inside ordinary materials. Foley gives bodies weight, objects substance and imaginary places a physical life. When the work succeeds, audiences do not hear somebody performing a world in a studio. They simply believe that the world was always there.

  • How Do You Make ADR Sound Like It Was Never Replaced? Paul Carden and Chris Navarro on Performance, Technology, and the Art of Dialogue Replacement

    Paul Carden and Chris Navarro

    How do you make ADR sound like it was never replaced?

    A line of dialogue may last only a few seconds, yet replacing it convincingly can require an extraordinary combination of preparation, performance, technology and judgement. The original production recording must first be identified as unusable, the replacement carefully documented and prepared, the actor returned to the emotional and physical circumstances of a performance recorded months earlier, and the new dialogue captured so that it matches the timing, vocal quality, microphone perspective and acoustic character of the original scene. If the process succeeds, the audience should never know that any of this work happened. During their joint online guest lecture for Edinburgh Napier University, ADR supervisor Paul Carden and ADR mixer Chris Navarro demonstrated this complete process from beginning to end. Rather than discussing Automated Dialogue Replacement only in theory, they created a deliberately compromised line of production dialogue, prepared it for replacement, recorded it on an ADR stage and evaluated the result. Their demonstration revealed a process in which meticulous preparation and technical fluency serve a deceptively simple objective: allowing everyone involved to concentrate upon the performance.

    Carden began not in a recording studio, but outside beside a car. His objective was to demonstrate how an apparently simple piece of dialogue could become unusable during production. Wearing a lavalier microphone, he performed the five-word line “I’m late for work” while getting into the vehicle. The exercise immediately revealed how vulnerable production dialogue can be. Clothing could obscure the microphone, a zipped jacket could change the sound, movement could complicate recording and the closing car door could mask the line itself. The example was deliberately simple. Carden was standing almost still, knew exactly what he was going to say and had constructed the situation specifically for the demonstration. On a real production, actors may be walking, running, interacting with objects and performing emotionally demanding scenes while the production sound team manages multiple radio microphones and one or more booms. From this perspective, the surprise is often not that some lines require replacement, but that so much production dialogue survives at all.

    The demonstration also established one of the central tensions within ADR. Production dialogue is not simply speech recorded on location. It contains a performance created at a particular moment, within a particular physical environment, as part of an interaction with other performers. Months later, an actor may arrive on an ADR stage while already immersed in an entirely different project and be asked to recreate a few seconds from that earlier performance. The recording environment offers little of the original context. The actor stands in a comparatively neutral room, watches an image on a screen, listens for cues and attempts to reproduce not merely the words, but the emotional and physical conditions under which those words were originally spoken. Carden emphasised that this is one reason actors can find ADR difficult. Matching a line involves returning to a performance that may no longer feel immediate or familiar.

    Before any actor reaches the stage, however, the line must become an ADR cue. Carden demonstrated this preparation process by locating the damaged dialogue within the picture, defining the cue and documenting the reason for replacement. The cue was assigned a unique identifier, linked to its scene information and accompanied by notes explaining that the car door had closed over the line. He also demonstrated the value of professional flexibility. Where a production line was imperfect but potentially usable, he might identify the replacement as optional rather than forcing an unnecessary argument over whether it must be replaced. The cue sheet then carried the information required by the recording stage, including project details, version information, character and actor names, cue number, timecode and recording information. A five-word performance therefore arrived at the stage supported by an extensive information system designed to ensure that everybody was working on the correct material.

    Version control was particularly important. Carden described production teams as living and dying by version dates, reflecting the practical reality that picture changes can quickly make carefully prepared cues inaccurate. Cue numbering performs a similarly essential role. Each take is voice-slated so that the recording itself retains its identity even if paperwork becomes separated from the audio. These details may appear administrative when viewed from outside professional production, yet they protect the continuity of the entire process. An ADR session may contain hundreds of replacement lines recorded across multiple days, actors and facilities. The actor should not have to think about whether the cue is correctly identified or whether the stage is working from the right picture version. Preparation creates the stability within which performance can happen.

    Carden also offered one deceptively simple instruction for aspiring ADR mixers: keep recording until somebody clearly indicates that the take has ended. An actor may continue the performance, a director may decide to record another version immediately or the session may move into wild recording without a formal interruption. Stopping too early can lose useful material for no meaningful benefit. Storage is cheap; an unrecoverable performance is not. The advice reflected a wider principle that would continue through the lecture. Good ADR practice depends upon remaining attentive to what is happening in the room rather than allowing the machinery of recording to dictate the session.

    The second stage of the demonstration moved into Navarro’s ADR facility. Even though the project consisted of a single line created for the lecture, he established the session as though it were a conventional production. Folder structures, documents, Pro Tools sessions and media were organised according to the same principles he would use for a project involving one session or hundreds. His ADR template was already configured for the ordinary demands of recording, while remaining adaptable when unusual situations arose, such as multiple actors performing together with several microphones each. A good template does not prescribe every session in advance. It removes predictable technical work so that attention remains available for situations that cannot be predicted.

    Even the imported picture introduced a practical lesson. Navarro noticed that the system was responding sluggishly and identified the compressed H.264 picture as a likely cause, since the computer had to decode it continuously during playback. The moment was minor, but revealing. Professional technical fluency often appears not through dramatic troubleshooting, but through the ability to recognise quickly why a system is behaving unexpectedly and continue without allowing the problem to dominate the room. Earlier, Carden had offered similarly pragmatic advice about computer failure: save the work, restart the system and continue rather than allowing panic to consume valuable time. Both speakers treated technical expertise as calm familiarity rather than technological display.

    Microphone selection then revealed how ADR matching differs from conventional voice recording. The objective is not to capture the most beautiful possible version of the actor’s voice. It is to create a recording that can inhabit the existing production soundtrack without attracting attention. Navarro showed that the stage had previously been configured for voice-over work, with microphones positioned closely and directly to produce the clear, present sound required for narration. ADR demanded a different approach. Carden had recorded the original line using a lavalier, so the replacement needed to reproduce the qualities of that production perspective rather than simply offering a technically superior recording.

    Carden explained that ADR sessions commonly record both boom and lavalier microphones, ideally using the same models employed during production. Even when the original line appears to come predominantly from one microphone, alternatives can prove unexpectedly useful. Navarro demonstrated this by positioning a shotgun microphone close to Carden but substantially off-axis. Pointing it directly at the performer would have produced a cleaner and more extended sound, but that was not necessarily desirable. The off-axis position rejected particular frequencies and created a tonal character that could potentially sit closer to the lavalier recording. The result was counterintuitive: the boom microphone could, under those conditions, sound more like the production lavalier than the replacement lavalier itself. Matching therefore began not with assumptions about microphone categories, but with listening.

    This distinction between technical quality and contextual suitability runs throughout professional ADR. A beautifully recorded line can fail if it sounds too clean, too close, too rich or too controlled for the image and surrounding production dialogue. Conversely, a microphone position that would appear unconventional in another recording context may provide exactly the spectral character required for a convincing match. The mixer must understand microphones well enough to use their imperfections deliberately. The question is never simply which microphone sounds best. It is which recording can become part of the scene without revealing the process that created it.

    Once Carden stepped in front of the microphone, the demonstration moved from equipment towards performance. The familiar system of three audible beeps established the timing, with the actor beginning the line where an imaginary fourth beep would occur. In principle, the task appeared straightforward. In practice, Carden repeatedly entered early, became self-conscious about the timing and discovered how quickly an apparently trivial five-word line could become difficult once performance, synchronisation and technical awareness competed for attention. His reaction gave the students an unusually useful demonstration of the psychological demands placed upon actors during ADR. Knowing the line is not enough. Understanding the timing is not enough. Even achieving synchronisation is not enough. The replacement still has to sound as though it belongs to the original performance.

    At one point, Carden produced a take that was synchronised correctly but immediately recognised that the performance itself was wrong. Navarro’s response was subtle. He did not ask for shouting or a dramatically louder delivery. He suggested only slightly more projection. The difference was not simply level. A small change in the physical production of the voice altered its tone and gave the line the additional presence heard in the original recording. When the replacement was compared directly with production dialogue, the improvement became obvious. The exercise demonstrated that ADR matching cannot be reduced to waveform alignment, pitch correction or equalisation. Performance changes the spectrum of the voice before any microphone or processor becomes involved.

    Navarro later developed this point in greater depth. Two performances can have similar apparent loudness and pitch while differing substantially in vocal quality. An actor performing naturally on set may have a relaxed throat and a particular physical relationship with the surrounding scene. On the ADR stage, tension, self-consciousness or the effort to satisfy technical instructions can change the voice. A performer may become tighter, brighter or less natural even while reproducing the words and timing accurately. Matching therefore requires attention to qualities that are difficult to describe numerically. Projection, pitch, volume and rhythm matter, but so do muscular relaxation, breath and the physical origin of the voice.

    This presents the ADR mixer with a delicate problem. The mixer may hear precisely what is preventing a line from matching, yet communicating every technical observation to the actor may make the performance worse. Navarro warned that performers can absorb only so many notes before they begin thinking about the mechanics of speech rather than the character. A useful intervention must therefore translate technical listening into language that supports performance. Sometimes the correct decision is to offer a suggestion. Sometimes it is to communicate through the director. Sometimes it is to recognise that an imperfection is not important enough to justify disturbing the creative balance of the room. Hearing a problem and knowing whether to mention it are separate professional skills.

    The lecture also demonstrated an older approach to ADR recording that remains remarkably useful. Navarro sampled the original production line and played it repeatedly into Carden’s headphones. Instead of concentrating simultaneously upon picture, beeps and performance, Carden could hear the original line and immediately reproduce it, repeating the process several times. Navarro then edited the resulting takes into position and compared them with the production recording. The method offered a direct reference for timing, intonation, volume and vocal character, allowing the performer to respond to the sound of the original performance rather than attempting to reconstruct every detail intellectually.

    Navarro’s explanation of this technique revealed an important insight into perception. Picture can sometimes become a distraction. An actor watching for a mouth movement may wait until it becomes visually apparent, by which time the correct moment to begin has already passed. If the original audio is correctly synchronised to picture, matching its rhythm and timing can naturally reproduce picture sync. The performer can therefore concentrate upon hearing and responding rather than continually monitoring several streams of information at once. Despite the age of the sampling technique, Navarro regarded it as one of the most useful tools available on the stage. Its continued value comes not from technological sophistication, but from the way it simplifies the performer’s task.

    The same objective shaped Navarro’s control room. His system contained extensive routing, multiple microphone inputs, separate monitoring paths and a heavily customised control surface, yet the purpose of this complexity was to make the session itself feel simple. He might need to manage separate mixes for the control room, recording stage, actor headphones, supervisor headphones and remote participants, with each requiring different material during rehearsal, recording and playback. Attempting to reconfigure every route manually between passes would slow the session and create repeated opportunities for error. Navarro therefore designed systems that allowed complex changes to happen immediately.

    His customised keypad provided a particularly revealing example. Individual buttons could trigger sequences of macros that armed tracks, initiated recording and performed repetitive editing operations once a take had finished. Navarro had spent months developing the core system and continued refining it whenever he noticed himself repeating unnecessary keyboard operations. He had begun his career as an ADR recordist and understood the division of labour on a two-person stage, where one person could manage recordings and files while the mixer concentrated upon the performers. Working alone, he used automation to reproduce some of that support, delegating repetitive technical operations to macros so that his attention could remain directed towards the stage.

    This led to one of Navarro’s most important ideas: the ADR mixer must work at the speed of creativity. An actor or director may suddenly discover a new approach to a line and want to record it immediately. If the mixer responds by asking everyone to wait while tracks are configured, routes changed or files prepared, the idea may lose its immediacy. A delay of only a few seconds can alter the atmosphere of a performance. Technical speed therefore has a human purpose. The mixer learns the system thoroughly enough that machinery does not interrupt thought.

    Navarro connected this principle to advice he had received about respected ADR mixer Tommy O’Connell. Asked what distinguished O’Connell’s work, a sound editor gave a simple answer: he anticipates. Useful actions are completed before anyone needs to request them. Navarro interpreted this not as a mysterious talent, but as the result of attention. If the mixer is watching the stage, listening to conversations and understanding the direction in which a session is moving, preparation for the next action can begin before a formal instruction arrives. Anticipation therefore depends upon both technical readiness and social awareness. A mixer whose attention is buried in the workstation may complete every requested task correctly while still remaining one step behind the session.

    Carden’s preparation of the cue and Navarro’s management of the stage reveal two sides of the same professional process. Carden reduces uncertainty before the session begins through accurate cueing, documentation, version control and communication. Navarro reduces friction during the session through templates, routing, automation and anticipation. Neither form of preparation is intended to make the process more rigid. Both preserve the possibility that an actor, director or mixer can respond immediately when something unexpected and valuable occurs.

    Microphone signal flow offered another example of this balance. Navarro described a deliberately clean recording path from microphone through preamplifier into Pro Tools, avoiding unnecessary outboard processing. Within the workstation, compression and equalisation could be used sparingly, but he warned against making irreversible decisions without good reason. Recording completely flat preserves maximum flexibility for the re-recording mixer, while careful decisions made during the session can still produce useful, committed tracks. Heavy processing may be difficult to undo later. A technically impressive recording decision is not valuable if it reduces the options available to the production.

    Navarro’s session was configured to make microphone comparison immediate. Several microphone inputs could remain available, feeding dedicated record tracks at appropriate levels. A boom and lavalier could be recorded simultaneously as discrete channels, preserving both perspectives for later evaluation. The workflow reflected the uncertainty inherent in matching. The microphone expected to work best may not produce the most convincing result once the line is placed against production. Recording alternatives gives later editors and mixers material from which to construct the most believable transition.

    His discussion of recording format was equally pragmatic. For conventional ADR, he worked at the established production standard of 48 kHz, 24-bit audio. Higher sample rates could be valuable for sound effects intended for extensive manipulation, where recordings might later be slowed or processed heavily. ADR serves a different purpose. Recording at unnecessarily high rates would increase storage requirements before the material was eventually converted to the format used by the production. The appropriate technical choice depends upon what will happen to the sound. More data is not automatically more useful.

    As the lecture progressed, it became increasingly clear that the most difficult parts of ADR were not contained within any equipment specification. Navarro estimated that learning the technical fundamentals and becoming comfortable with Pro Tools took years, yet he distinguished that competence from the broader ability to run a session. Recording dialogue in sync with picture is ultimately a technical process that can be learned. Managing actors, directors, supervisors, performance, uncertainty and the emotional energy of the room represents a different level of expertise.

    Actors may arrive nervous. Directors may be highly collaborative, completely self-sufficient or resistant to suggestions. A performer may struggle with a line while becoming increasingly aware of the difficulty. The mixer must understand how much intervention the situation can support. Navarro described the importance of creating a calm and comfortable environment, particularly when clients are unfamiliar with the process. Confidence can be communicated without dominance. If someone is uncertain, the mixer can explain what will happen, answer questions and demonstrate that the session is under control. Technical authority becomes useful when it reduces anxiety rather than displaying superiority.

    Carden and Navarro also acknowledged the professional judgement involved in deciding when not to intervene. A mixer may recognise an aspect of a performance that could be improved, yet nobody has asked for technical input and the issue may not be important enough to disrupt the session. Navarro framed this as a balancing act. Would the intervention materially improve the line? Is the director open to suggestions? Will another note distract the actor from a performance that is already working? Expertise includes recognising problems, while professional maturity requires distinguishing consequential problems from imperfections that do not matter.

    The ADR stage therefore brings technical precision and emotional sensitivity into the same process. The mixer must hear minute changes in vocal quality, understand microphone behaviour, maintain synchronisation, manage several monitoring environments and operate the recording system almost instinctively. At the same time, attention must remain on body language, conversation, confidence, frustration and creative momentum. Mastery of the technology is essential precisely so that the technology no longer consumes the attention needed elsewhere.

    Carden’s demonstration also placed the scale of professional ADR into perspective. Their five-word line required production recording, cue preparation, documentation, stage setup, microphone selection, rehearsal, multiple takes, alternative recording methods, editing and comparison. A skilled actor might complete around ten or twelve conventional cues in an hour, while a feature film may contain 150 or 200 ADR lines. Complicated performances naturally take longer. What appears to an audience as a few moments of seamless dialogue can therefore represent days of concentrated work.

    Yet the success of that work is measured largely through its invisibility. When Navarro compared Carden’s replacement line with the production recording, the evaluation concerned relationships: volume, tone, vocal quality, perspective and the way the replacement interacted with the surrounding scene. The door slam that had originally damaged the line was restored around the new performance, returning the replacement to the event from which it had temporarily been separated. Heard alone, the ADR recording was merely a voice on a stage. Placed back into the scene, it became part of an action.

    This transformation captures the philosophy shared by both speakers. ADR is often described as the replacement of unusable dialogue, but their demonstration showed that replacement is only the beginning of the problem. The objective is reconstruction. The actor reconstructs a performance. The mixer reconstructs a microphone perspective. Editors reconstruct timing and continuity. The final soundtrack reconstructs the relationship between voice, action and environment so successfully that the audience experiences a single uninterrupted moment.

    The joint nature of the lecture made this especially clear. Carden approached ADR through the complete production process, showing how a line travels from location problem to documented cue and finally to the recording stage. Navarro approached the same process from inside the control room, revealing the technical systems, listening skills and interpersonal judgement required to capture a convincing replacement. Their perspectives met at the point where professional preparation serves human performance.

    By the end of the session, the deliberately damaged line had become something much larger than a technical demonstration. It revealed why ADR demands more than synchronisation, why the cleanest microphone is not always the correct microphone, why a performer can match timing while missing the voice, why an old sampler can remain useful in a modern digital workflow and why a highly automated control room can make a session feel more human rather than less. Above all, Carden and Navarro showed that successful ADR depends upon where attention is directed. The actor should be thinking about performance rather than machinery. The director should be thinking about the scene rather than routing. The mixer should be watching and listening to the room rather than fighting the workstation. The audience, finally, should be thinking about none of these things. When every part of the process works together, they simply hear a character speak.

  • How Do You Mix Television Sound Under Pressure? Frank Morrone on Dialogue, Workflow, and the Art of Re-Recording

    Frank Morrone

    How do you mix television sound under pressure?

    A television soundtrack may contain hundreds of dialogue recordings, sound effects, Foley performances, backgrounds, ADR takes and music stems, all competing for space within a mix that must remain clear, emotionally convincing and technically suitable for broadcast. The audience should never become aware of that complexity. They should simply understand every line, believe every environment and remain absorbed in the story. During his online guest lecture for Edinburgh Napier University, re-recording mixer Frank Morrone explored the craft behind achieving that apparent simplicity. Drawing upon a career in film and television that began in 1979, and projects including Lost, The Strain, Sleepy Hollow and Criminal Minds: Beyond Borders, he revealed a discipline shaped equally by technology, organisation, collaboration and judgement. Throughout the session, one principle emerged repeatedly. The most effective mixing workflows allow enormous technical complexity to disappear behind the story.

    Morrone began by tracing a career that developed across several different areas of professional audio. His earliest work took place in music studios, recording jazz and orchestral film scores before following those recordings into the dubbing theatre and becoming increasingly interested in the way complete soundtracks were assembled. A move into post-production allowed him to work across dialogue editing, music editing and Foley recording before concentrating upon re-recording mixing. That breadth of experience shaped the collaborative philosophy running throughout the lecture. Unlike a music recording session, where one engineer may remain closely involved from recording through to the final mix, film and television sound brings together work created by many different specialists. The dub stage is where those contributions finally meet. Successful mixing therefore depends upon understanding not only the material itself, but the people, processes and decisions that produced it.

    Television makes this collaboration particularly demanding. Morrone described an industry in which track counts have continued to increase while schedules and budgets have become progressively tighter. Sophisticated surround mixes must be created rapidly, and emerging formats add further complexity without removing the need to support conventional playback systems. His response is not simply to work faster. It is to design workflows that remove unnecessary decisions from the mixing stage. As soon as he joins a project, he communicates with the supervising sound editor about track requirements and provides a starting template so that incoming material already fits an established structure. Organisation begins before the mixer enters the room. Under severe time pressure, the ability to find, control and compare material immediately becomes part of the creative process itself.

    This reveals something deeper about Morrone’s understanding of expertise. His templates, early conversations with sound editors, knowledge of production microphones, preparation of alternative takes and habit of printing completed passes all anticipate problems before they are allowed to interrupt the mix. The same thinking extends beyond the dubbing theatre. He considers how broadcast processing will react to dynamics, how a mix will translate to domestic systems and how material created for one format will behave when heard through another. Professional experience, in this sense, is not simply the ability to solve problems quickly. It is the ability to recognise where problems are likely to emerge and construct a workflow in which many of them have already been addressed before they become urgent.

    The scale of Lost provided a striking illustration. Morrone showed the students sessions containing extraordinary numbers of elements, including dozens of tracks dedicated to the Smoke Monster alone. Its identity emerged from a deliberately ambiguous combination of animal voices, pneumatic machinery, roller-coaster wheels and other contrasting sources, creating something that resisted being understood as either entirely organic or entirely mechanical. Those effects existed alongside hard effects, backgrounds, Foley, production dialogue, ADR, group recordings and substantial music deliveries. The technology available at the time imposed strict limits upon voices and processing, requiring careful decisions about resource allocation as well as creative balance. Complexity could not simply be solved by adding more processing. The session itself had to be organised so that the mixers could navigate it instinctively.

    Custom fader layouts and VCA groups became essential to that process. Morrone described arranging controls so that principal dialogue could immediately be balanced against ADR, group recordings and music, while different categories of material remained independently accessible. Music sources could be separated from score, while dialogue in different languages could be isolated for international deliverables. Every layer of organisation reduced the time between hearing a problem and solving it. This became especially important during pilot season, when mixers might receive unusually elaborate material without first having time to develop a workflow around the programme. A sufficiently flexible template must already be capable of accommodating whatever arrives. Preparation, in this context, creates the conditions in which creative decisions can still be made under pressure.

    The production of Lost also demonstrated the tension between creative ambition and delivery requirements. Morrone recalled the exceptional resources devoted to the programme, including a large pilot budget and Michael Giacchino’s insistence upon recording a live orchestra for each episode. Yet the soundtrack still had to survive the restrictions of television broadcast. The team therefore created a more dynamic version for DVD before producing a contained broadcast mix designed to survive transmission processing. A similar approach was later adopted on The Strain. The distinction mattered. A mix can remain technically within specification and still behave poorly when subsequent broadcast processing responds to excessive dynamics. Re-recording therefore requires mixers to think beyond the dubbing theatre. They are mixing for every system through which the programme will eventually reach its audience.

    Dialogue occupied the centre of Morrone’s approach. His first objective is always to preserve the production performance wherever possible. When ADR has been recorded, he wants to know why. A line replaced for performance reasons presents a different problem from one replaced to solve a technical fault. If the director wanted a different performance, Morrone respects that decision while keeping the original available as an alternative. If the problem was technical, he first explores whether the production recording can be repaired. Modern restoration tools have dramatically expanded what can be rescued, reducing the need to replace performances that may possess subtleties difficult to recreate months later in an ADR studio.

    When ADR is necessary, matching involves far more than applying equalisation and reverb. Morrone keeps production dialogue available alongside the replacement so that he can compare transitions directly, matching tone, acoustic environment, pacing and performance. He prefers to receive several strong takes and recordings from both boom and lavalier microphones, recognising that the microphone apparently closest to the production perspective is not always the easiest to integrate. Understanding which microphones and wireless systems were used during production can also provide valuable clues, particularly where transmission systems have imparted their own sonic characteristics. ADR matching consequently becomes a process of reconstructing relationships rather than searching for a single corrective setting.

    Technology has transformed that work, though Morrone repeatedly warned against allowing powerful restoration tools to encourage excessive processing. Noise reduction, spectral repair, ambience matching, EQ matching and dereverberation can rescue material that would previously have required replacement. Yet he deliberately removes less noise than might appear necessary when dialogue is heard in isolation. Once backgrounds, effects and music return, much of the remaining noise may be perceptually masked. Processing that sounds impressively clean in solo can leave dialogue lifeless and constricted in the finished mix. Morrone therefore keeps copies of original material and sometimes returns to less processed versions during the final mix. The objective is not the cleanest possible dialogue track. It is dialogue that remains natural and convincing within the complete soundtrack.

    His attitude towards restoration reveals a broader philosophy of technology. Morrone began working with magnetic tape, a Cat 43 noise reduction unit and a notch filter, and he clearly values the extraordinary capabilities available to contemporary mixers. Yet greater technical power has not removed the need for judgement. In many respects, it has increased it. The ability to remove more noise does not mean that more noise should be removed, just as the ability to place sound almost anywhere within an immersive field does not mean that every available position should be used. New tools expand the range of possible decisions. They do not determine which decisions are appropriate. Throughout Morrone’s lecture, technical capability remained subordinate to perception.

    This distinction between isolation and context became one of the lecture’s most important ideas. Morrone described television mixing as beginning with dialogue and music, establishing the foundation against which the effects mixer can develop backgrounds and action. Once those elements come together, the mixers decide what should drive the scene. Some moments belong to music, others to effects, while still others require a more subjective perspective or a deliberate reduction of material. Experienced mixing partners develop an almost instinctive understanding of one another’s decisions. If an effect obscures a line during an early pass, Morrone knows that a trusted colleague will create space for it during refinement. Collaboration is therefore more than the division of tracks between two people. It is a shared process of deciding where the audience should listen.

    The idea that sounds must be judged in context reaches far beyond balancing dialogue against effects. Dialogue that appears too noisy when soloed may become entirely convincing once the world of the scene surrounds it. ADR that attracts attention after twenty repeated comparisons may pass unnoticed when the audience encounters it once within a continuous performance. Group recordings that sound absurd under isolated scrutiny may perform their role perfectly when placed at the correct distance behind principal dialogue. Morrone’s examples repeatedly challenged the assumption that individual elements should be perfected independently. The meaningful unit of evaluation is ultimately the audience’s experience of the scene. A soundtrack succeeds through relationships between sounds rather than the isolated perfection of its components.

    Normally, the priority within those relationships remains dialogue. Morrone described it as both the foundation of the mix and the principal carrier of storytelling. His concern with intelligibility, however, extends beyond maintaining a technical hierarchy between dialogue, music and effects. It is fundamentally about preserving attention. His most revealing test is not a meter reading but a listener asking what somebody has just said. At that moment, the audience member has been pulled out of the story and required to think about the failure of the soundtrack. Clear dialogue therefore supports immersion precisely by avoiding attention to itself. The better the mix communicates, the less the audience needs to think about the process of communication.

    Yet the rule is not absolute. For one sequence in The Family, a child was being used as bait to attract a kidnapper in a crowded shopping centre. The environment needed to feel genuinely busy, yet the available collection of separate crowd recordings and reverberant elements never created a convincing whole. Morrone took a six-channel recorder into a real shopping centre, captured the food court from different perspectives and brought the recordings back into the mix. The result allowed the environment to crowd the dialogue slightly, which was precisely what the scene required. Clarity remains fundamental, but realism sometimes depends upon controlled difficulty.

    The shopping-centre sequence illustrates an important tension within Morrone’s approach. His practice is built upon strong principles, but those principles do not become inflexible rules. Dialogue normally takes priority, except when allowing the environment to interfere with it makes the dramatic situation more believable. Restoration should preserve intelligibility, except when excessive cleaning destroys naturalness. Acoustic simulation is valuable, except when recording the real physical relationship between people and space produces a more convincing result. Expertise therefore involves knowing the rules well enough to understand what they protect, then recognising the moments when the needs of the scene justify bending them.

    This willingness to leave the dubbing theatre and record real spaces appeared repeatedly throughout the session. Morrone discussed impulse responses, convolution reverbs and carefully developed presets for rooms, vehicles and other environments, all of which provide useful starting points for worldizing sound. Yet he remained pragmatic about their limitations. A real location does not necessarily sound like the space suggested by the finished image, and even a carefully captured impulse response may require additional reflection, delay or reverberation before it feels convincing. Cars present an especially difficult problem, combining strong early reflections from glass with highly absorptive surfaces elsewhere. Experience gradually produces a library of useful starting points, though listening still determines the final result.

    Sometimes no simulation is as convincing as returning to the physical situation itself. Morrone described a scene in which a group of foster children were supposed to be creating chaos upstairs while a conversation took place below. Studio-recorded group voices did not reproduce the peculiar combination of footfalls, structural transmission and reflections travelling down a staircase into another room. His solution was direct. He gathered children on the upper floor of a house, recorded from downstairs and captured the complete acoustic event as it occurred. No increasingly elaborate chain of processing was required. The physical relationship between performers, building and microphone provided what the scene needed.

    The pressures of television production make such judgement particularly important. Morrone compared television mixing to boot camp. Schedules leave little room for hesitation, and mixers must develop workflows capable of producing strong results quickly. Once a scene has been successfully mixed, he prefers to print it rather than trusting that automation will remain untouched throughout later work. Accidental writes, changed sends and technical errors can occur even in experienced hands. Printing completed work provides security and allows later changes to be punched into established stems. Efficiency does not mean rushing blindly. It means reducing the opportunities for avoidable problems to consume the limited time available.

    Long sessions also introduce a more human limitation: hearing fatigue. Morrone described working on the highly dynamic soundtrack of The Strain, where sustained exposure to loud material forced him to think deliberately about auditory recovery. His solution was simple but important. He left the room for short periods, walked and allowed his ears to recover while his mixing partner continued working. The two mixers could alternate demanding passes, giving each other opportunities for rest without stopping the session. After decades of mixing, one of the useful discoveries was not another plug-in or processor, but the value of leaving the chair for ten minutes. Professional listening depends upon recognising the limits of the listener.

    Morrone’s discussion of client relationships revealed another dimension of the re-recording mix that students may rarely encounter in technical demonstrations. Mixers are working with directors, producers and other clients who may have strong preferences that differ from their own. Morrone described situations in which clients wanted music loud enough to compete with dialogue. His responsibility was to explain the likely consequences, demonstrate how the mix translated at lower levels and on smaller monitors, and search for a compromise that preserved the client’s intention while protecting intelligibility as far as possible. The mixer offers expertise, but does not own the programme. Knowing which decisions are worth challenging and which require accommodation is part of the craft.

    His description of deciding whether an issue represents a “hill to die on” reveals a sophisticated understanding of professional authority. Expertise does not give the mixer unlimited control over the work, nor does collaboration require the abandonment of professional judgement. The mixer must advocate for the audience, explain likely consequences and make alternatives audible, while recognising that the final creative intention belongs to the client. Professional judgement therefore includes negotiation. Sometimes expertise means defending a decision. Sometimes it means finding a compromise that neither side initially imagined. Occasionally, it means implementing a choice that remains contrary to personal taste while ensuring that it works as successfully as possible.

    Technical fluency plays an important role in maintaining those relationships. When a client requests a change, Morrone wants to make it immediately, play it once and continue. Searching through tracks or repeatedly troubleshooting a familiar process changes the atmosphere of the room and interrupts attention to the programme itself. A well-designed session keeps the conversation focused upon storytelling and communication. The deeper the mixer’s command of the tools, routing and session layout, the less those systems intrude upon the creative discussion.

    This is another form of transparency. Morrone repeatedly returned to the idea that technology should disappear from the client’s experience, yet transparency does not mean that technology has become unimportant. The opposite is closer to the truth. Considerable technical knowledge is required to make complex systems feel immediate. Templates, custom layouts, routing, monitoring, printed stems and intimate familiarity with the workstation create an environment in which a creative request can become an audible result without breaking the flow of the session. Mastery becomes visible through the absence of friction.

    The growth of immersive sound has expanded this challenge further. Morrone described mixers as working between extremes, from Dolby Atmos and sophisticated home theatres to stereo playback and mobile devices with earbuds. His philosophy was to begin with the best mix possible in the most capable format, then ensure that it translates successfully into simpler ones. Immersive technology may offer extraordinary spatial possibilities, but Morrone remained cautious about using novelty without considering perception. In particular, he defended dialogue remaining anchored to the centre. Moving voices between speakers can alter timbre and create distracting changes that audiences may notice without understanding their source. New formats create possibilities, though they do not invalidate principles developed through decades of listening.

    His argument about dialogue placement is particularly revealing. Audiences do not need to identify the technical source of a problem in order to experience discomfort. Morrone described viewers sensing that something was wrong when dialogue moved around an immersive field, even when they could not explain precisely what disturbed them. This places an unusual responsibility upon the mixer. Professional listening must sometimes diagnose experiences that ordinary listeners can feel but cannot name. The purpose of expertise is not to dismiss those responses as technically uninformed, but to understand the perceptual conditions that produced them.

    The same concern for translation shaped his approach to low frequencies. Subwoofers vary enormously between listening environments, and domestic listeners frequently adjust them far beyond calibrated levels. Morrone therefore warned against depending entirely upon the LFE channel for the weight of a soundtrack. Low-frequency energy can also be carried through the main channels, creating a result that remains powerful across a wider range of playback systems. His experience of hearing a domestic subwoofer struggle with low-frequency material from Sleepy Hollow reinforced the point. A soundtrack must survive real listening environments, not merely sound impressive on a perfectly calibrated dubbing stage.

    As the discussion widened, Morrone considered the future possibility of mixes that adapt more intelligently to different devices and contexts. Streaming, immersive audio, virtual reality and personalised playback were already creating pressure for soundtracks to function across radically different systems. Yet his underlying philosophy remained remarkably consistent. Whatever the format, begin with the strongest possible mix, preserve the storytelling hierarchy and understand how human perception responds to the result. Technology changes quickly. The responsibility to guide attention and communicate narrative does not.

    Questions from the students returned the discussion to practical preparation. Morrone strongly supported editors delivering material that is already sensibly balanced before it reaches the stage. Dialogue editors who use clip gain to create consistent levels save valuable mixing time, while backgrounds and effects arriving close to useful operating levels allow the mixer to begin creatively rather than first correcting avoidable problems. He described requesting particular editors for demanding productions precisely for this reason. Good preparation is noticed. In a professional environment where a pilot may need to be mixed in only a few days, the person who consistently delivers well-organised, intelligently balanced material becomes someone mixers actively want on the next project.

    His discussion of ADR offered another deceptively simple lesson about perception. Morrone sometimes works on difficult replacement lines privately through headphones while the effects mixer is making a pass. The client then hears the finished line only once in context rather than listening to it repeated dozens of times during adjustment. Repetition directs attention towards the repair and teaches the listener exactly where to expect it. The same awareness shaped his humorous rule about never soloing loop group in front of a client. Background conversations that work perfectly as part of a scene may sound absurd when isolated and scrutinised. What matters is not whether every element survives examination on its own, but whether it performs its intended role in the scene.

    These examples reveal that Morrone is not simply mixing sound. He is managing attention, expectation and knowledge. Once listeners have been taught where an edit exists, they may hear it differently. Once an element has been isolated, they may judge it according to criteria that have little relevance to its actual purpose. Perception is shaped not only by acoustic information, but by what listeners have been encouraged to notice. Part of the mixer’s craft therefore lies in protecting the audience’s experience from unnecessary awareness of the mechanisms used to construct it.

    Yet group recording could also become a powerful storytelling tool. On Criminal Minds: Beyond Borders, episodes moved between international locations while much of the production remained based in Los Angeles. Carefully performed local-language group recordings, combined with music and other environmental elements, became essential to establishing each location convincingly. Here, group material could be brought forward rather than hidden. There was no universal rule governing how loudly an element should be mixed. Its appropriate level depended upon what the scene needed to communicate.

    By the end of the session, Morrone’s account of re-recording mixing had moved far beyond faders, plug-ins and delivery specifications. The technology matters enormously, as do templates, routing, restoration tools, monitoring and control surfaces, but those things serve a larger process. A mixer must understand performance, storytelling, perception, collaboration, translation and the subtle politics of working with clients under pressure. The session may contain hundreds of tracks, yet the audience should hear a coherent world rather than the complexity required to construct it. Morrone’s lecture revealed a craft built upon anticipation, contextual judgement and the careful management of attention. Preparation preserves the possibility of creativity under pressure. Technical knowledge allows technology to disappear from the conversation. Rules provide essential foundations, while experience reveals when the needs of a scene require them to bend. Great television sound is not created by making every element impressive or every recording perfect in isolation. It emerges from understanding what the audience needs to hear, recognising what they should never need to notice, and making hundreds of individual decisions feel like one continuous experience.

  • How Does Sound Affect Us? Julian Treasure on Listening, Wellbeing, and Designing with Our Ears

    Julian Treasure

    How does sound affect us?

    Most people think about sound only when it becomes a problem. We notice the neighbour’s loud music, the traffic outside a bedroom window, the distracting conversation in an open-plan office or the shrill alarm that interrupts an otherwise quiet day. Far less attention is paid to the countless sounds that quietly shape our emotions, influence our behaviour and affect our health from one moment to the next. During his online guest lecture for Edinburgh Napier University, Julian Treasure argued that this oversight represents one of the greatest shortcomings of modern design. Buildings, products and public spaces are often designed primarily for the eye, while the ear receives remarkably little attention. Yet sound continually influences the way people think, work, communicate and feel. Throughout the presentation, one message emerged repeatedly. If we wish to design better experiences, we must learn to design with our ears.

    Treasure began by asking what sound actually affects. The answer, he suggested, is surprisingly simple. Sound influences our happiness, our effectiveness and our wellbeing, along with those of everybody who shares the environments we create. This observation immediately shifts the discussion away from traditional concerns about noise control or acoustic specifications. Sound becomes a human issue rather than merely a technical one. The quality of an acoustic environment influences far more than whether a room sounds pleasant. It affects how effectively people communicate, how comfortably they work, how safely they respond to hazards and how they experience the spaces in which they spend their lives. Sound therefore deserves to be regarded as one of the fundamental materials of design rather than an afterthought considered once construction has already been completed.

    To explain why sound exerts such profound influence, Treasure described four principal ways in which it affects human beings. The first is physiological. Unlike vision, which depends upon the direction in which we happen to be looking, hearing continuously monitors the environment around us. Human beings cannot close their ears in the way they close their eyes. Throughout evolution, this has made hearing our primary warning system, continually searching for signs of danger beyond the limits of our vision. As a consequence, sound reaches deeply into the nervous system with remarkable immediacy. Sudden or unpleasant sounds trigger hormonal responses associated with stress and vigilance, while calmer acoustic environments encourage relaxation. Treasure illustrated this contrast using familiar examples. An unexpected loud noise immediately increases physiological arousal, whereas gentle natural sounds such as breaking waves often slow breathing and encourage a sense of calm. These responses are not matters of personal preference alone. Sound influences heart rate, hormone secretion, breathing patterns and even patterns of brain activity, quietly shaping the body’s internal rhythms throughout the day.

    The discussion then moved beyond physiology towards psychology. Music provides perhaps the most familiar illustration of this relationship. People instinctively choose particular music to celebrate, to concentrate, to relax or to reflect, recognising that different sounds evoke different emotional states. Treasure argued that natural sounds often produce similarly powerful responses. Birdsong, for example, tends to create feelings of safety and reassurance. Rather than being arbitrary preferences, these reactions may reflect deep evolutionary associations developed over thousands of generations. Birds sing when environmental conditions are relatively safe, allowing those sounds to become unconsciously associated with security. Although listeners rarely analyse these processes consciously, they nevertheless influence emotional experience in subtle yet persistent ways. For sound designers, this observation carries important implications. Designing an acoustic environment involves much more than controlling sound levels. It also requires understanding the emotional associations that different sounds naturally evoke.

    Treasure’s argument became even more relevant to contemporary workplaces when he turned to the cognitive effects of sound. Human attention is a limited resource. People often imagine that they can listen to several conversations simultaneously, though the reality proves rather different. Treasure observed that the human brain possesses only a limited capacity for processing speech, making it extremely difficult to concentrate when nearby conversations compete for attention. Open-plan offices provide a familiar example. Designers frequently value openness, flexibility and visual communication, yet the resulting soundscape often undermines the very productivity these environments seek to encourage. Relevant speech continually draws attention away from the task at hand, interrupting concentration and increasing mental effort. Research cited during the presentation suggests that productivity can fall dramatically under these conditions, illustrating that acoustic design contributes directly to cognitive performance rather than simply influencing comfort. Decisions about the sonic character of workplaces therefore become decisions about how effectively people can think.

    By this stage, a broader pattern had already become clear. Sound is not simply something that accompanies our activities. It shapes them. Physiological responses, emotional reactions and cognitive performance all depend, to varying degrees, upon the acoustic environments within which people live and work. This perspective challenges a long-standing tendency to regard sound as secondary to visual design. Treasure instead presented listening as a central consideration for architects, designers, engineers and sound professionals alike. Before deciding how a space should look, he suggested, we should also ask how it will sound, and how those sounds will influence the people who experience them every day. That question would remain at the heart of the remainder of the discussion.

    Having established that sound influences our physiology, psychology and cognition, Treasure turned to its effect upon behaviour. This influence often operates below the level of conscious awareness, making it particularly easy to overlook. Most people assume they make decisions independently of their acoustic surroundings, yet evidence suggests otherwise. Unpleasant environments encourage people to leave sooner, while attractive soundscapes invite them to remain longer. Treasure illustrated this with a striking study of consumer behaviour. In a supermarket displaying French and German wines with identical visual presentation, researchers changed nothing except the background music. On days when French music was played, French wine substantially outsold German wine. When German music replaced it, purchasing patterns reversed. Customers generally remained unaware that the music had influenced their choices, demonstrating that sound can shape behaviour without requiring conscious attention. For designers, retailers and architects alike, this example reinforced an important point. The acoustic environment is never simply a backdrop. It actively participates in shaping human decisions.

    The implications extend far beyond retail spaces. Treasure argued that every designed environment communicates through sound, whether intentionally or otherwise. A restaurant may create an atmosphere that encourages relaxed conversation, while another overwhelms diners with reverberation and competing voices. A hospital waiting room may reduce anxiety through carefully considered acoustics, or increase it through intrusive alarms and mechanical noise. An office may support concentration, or continually undermine it through poorly managed speech privacy. In each case, the acoustic environment becomes part of the overall design, influencing how people behave within the space. Designers therefore make decisions about human experience whenever they make decisions about sound, even if those decisions consist of ignoring it altogether.

    Treasure observed that this neglect reflects a broader imbalance within contemporary design practice. Buildings are routinely judged according to their appearance, products are evaluated through their visual form, and digital technologies devote enormous attention to graphical interfaces. Comparatively little thought is often given to how these same environments sound. This imbalance is surprising when one considers that hearing operates continuously. We can choose where to look, though we cannot simply decide to stop hearing the world around us. Sound therefore accompanies every activity, continually influencing perception in ways that visual design alone cannot achieve. Rather than treating acoustics as a specialist concern addressed late in a project, Treasure encouraged students to recognise listening as a fundamental design consideration from the very beginning.

    This perspective resonates strongly with professional sound design. Whether creating a film soundtrack, designing interface sounds, producing a virtual instrument or developing an interactive game, practitioners rarely add sound simply to occupy silence. Every sound communicates information, guides attention or influences emotional response. Treasure’s presentation broadened this principle beyond media production into everyday life. The same questions that sound designers ask while constructing a soundtrack also apply to architecture, product design and urban planning. What should the listener notice? Which sounds deserve emphasis? Which should remain unobtrusive? How can sound support rather than distract from the intended experience? The boundaries between sound design and environmental design begin to blur once listening itself becomes the central concern.

    Perhaps the most compelling aspect of Treasure’s argument lay in its optimism. If sound can undermine wellbeing, productivity and behaviour, it can equally improve them. Pleasant acoustic environments encourage relaxation, reduce physiological stress and support clearer thinking. Appropriate sound can strengthen communication, promote social interaction and make public spaces more welcoming. Rather than presenting acoustics as a matter of reducing unwanted noise, Treasure reframed the discussion in positive terms. The objective is not simply to remove bad sound, but to create environments in which good sound actively contributes to human wellbeing. This shift in perspective encourages designers to think creatively about the role sound can play rather than treating it solely as a problem to be controlled.

    These ideas naturally led towards a broader discussion of listening itself. If sound exerts such profound influence over human experience, then the ability to listen carefully becomes an essential professional skill rather than an incidental personal habit. Treasure suggested that hearing and listening are not the same activity. Hearing occurs automatically, while listening demands conscious attention, intention and practice. In an increasingly noisy world filled with competing sources of information, the ability to listen thoughtfully may be becoming more valuable rather than less. This distinction between passive hearing and active listening would ultimately form the foundation of his concluding message, not only for sound designers but for anyone responsible for creating environments in which other people live, work and communicate.

    Having demonstrated that sound influences physiology, emotion, cognition and behaviour, Treasure turned towards a more practical question. If sound has such profound effects upon human experience, what should designers actually do differently? His answer was strikingly optimistic. Rather than treating acoustics as a problem to be solved, he encouraged students to think of sound as a resource that can be shaped deliberately to improve people’s lives. Well-designed soundscapes do more than reduce unwanted noise. They encourage particular patterns of behaviour, support communication and create environments in which people feel healthier, calmer and more engaged. Designing with sound therefore becomes an act of positive intervention rather than damage limitation.

    Treasure illustrated this philosophy through a series of real-world projects. Airports, shopping centres and public spaces all benefited from carefully designed soundscapes that considered not only what people heard, but how those sounds influenced the way they behaved. Introducing natural sounds and thoughtfully composed musical environments increased customer satisfaction, encouraged visitors to remain longer and, in several cases, improved commercial performance. In one public space, the introduction of a biophilic soundscape was even associated with a measurable reduction in crime. These examples reinforced a central point running throughout the presentation. Sound does not merely accompany human activity. It shapes it. Decisions about the acoustic environment therefore become decisions about wellbeing, behaviour and social experience rather than simply matters of technical acoustics.

    Although these examples came from architecture and environmental design, their relevance extends directly to professional sound design. Every soundtrack contains foreground and background elements competing for the listener’s attention. Treasure encouraged students to think carefully about the role each sound should play within that wider acoustic picture. Not every sound deserves prominence, and not every moment benefits from additional music or greater complexity. Like a visual composition, an effective soundscape depends upon balance, hierarchy and clarity. He also encouraged designers to draw inspiration from natural environments, particularly through the thoughtful use of biophilic sound and adaptive or generative soundscapes that evolve over time rather than repeating mechanically. Different spaces support different activities, and their sonic character should reflect those differing purposes. Designing for concentration requires different acoustic decisions from designing for relaxation, learning or social interaction.

    The discussion naturally led back to listening itself. Treasure argued that hearing should never be confused with listening. Hearing is automatic. Listening is intentional. It requires attention, effort and continual practice. In an age characterised by constant distraction and increasingly complex acoustic environments, the ability to listen carefully becomes one of the most valuable professional skills a sound designer can develop. Technical expertise undoubtedly remains important, though it cannot substitute for careful listening. The most sophisticated recording equipment or software offers little value if the designer fails to recognise what listeners actually experience. Listening therefore becomes both a creative skill and an ethical responsibility. Before changing the sound of the world, designers must first learn to hear it properly.

    Treasure concluded by describing what he called the four foundations of effective listening: being conscious, committed, compassionate and curious. Conscious listening requires recognising that listening is an active process rather than a passive consequence of hearing. Commitment acknowledges that good listening demands time, attention and intention. Compassion encourages genuine understanding of other people through careful listening, particularly when viewpoints differ from our own. Curiosity reminds us that every sound and every conversation offers an opportunity to learn something new. Although these principles were presented in the context of listening, they also describe many of the qualities that distinguish thoughtful sound designers. Successful practitioners remain attentive, purposeful, empathetic and continually curious about how people experience the acoustic world around them.

    Treasure’s final appeal brought together everything that had preceded it. He encouraged students to become champions of listening and, above all, to “design with your ears.” This simple phrase encapsulated the wider philosophy running throughout the presentation. Sound should never be regarded as an afterthought added once visual design has been completed. It is one of the primary ways in which people experience the world. Every building, product, public space and interactive system possesses an acoustic identity that influences those who encounter it. Whether designing a film soundtrack, a hospital, a mobile application or a railway station, the same principle applies. The sounds we create shape the lives of the people who hear them.

    Taken together, Treasure’s presentation offered a compelling vision of contemporary sound design. It challenged the traditional tendency to regard sound as secondary to vision and instead positioned listening at the centre of human experience. Physiological responses, emotional wellbeing, cognitive performance, behaviour and communication all depend, to varying degrees, upon the acoustic environments we inhabit. For sound designers, this represents both an opportunity and a responsibility. Every decision about sound has consequences extending beyond aesthetics alone. Designing well therefore means more than creating compelling audio. It means understanding how people listen, recognising how profoundly sound affects everyday life and applying that knowledge to create environments in which individuals and communities can genuinely flourish.

  • What Can Sound Communicate That Words Cannot? Jim Metzner on Memory, Listening, and Going Places That Words Cannot Go

    Jim Metzner

    Jim Metzner began the lecture with a mystery.

    A sound was played. Students suggested possible explanations. Some heard machinery. Others heard something else entirely. For a few minutes the recording remained unresolved. Much of Metzner’s work inhabits that moment before a sound settles into a clear explanation. Before it becomes a bird, a vehicle, a voice, or a machine, it exists as an experience. During his online guest lecture for Edinburgh Napier University, discussions of field recording, travel, documentary production, family history, and memory repeatedly returned to this idea. How can sound communicate aspects of experience that are difficult to convey in any other way?

    Listening, in Metzner’s view, is not simply a way of gathering information. It is a way of encountering people, places, and experiences. Much of his work begins from a deceptively difficult question. How can a sound recording help somebody experience something they have never encountered for themselves?

    That challenge appeared repeatedly as students discussed their own recordings. Several described recording parks, public events, city streets, and everyday environments. Similar observations emerged from each example. Carrying a recorder changes the way people move through the world. Sounds that normally fade into the background suddenly become noticeable. Distant traffic acquires texture. Birds occupy distinct locations within a soundscape. Conversations, machinery, weather, and footsteps separate themselves into layers. The microphone becomes a reason to pay attention. One student described attempting to record ambience in a local park while aircraft repeatedly passed overhead. The interruptions were frustrating. Each time the environment seemed to settle, another aircraft arrived. Metzner responded with a story from his own work. While recording in the Great Swamp near a major airport, he encountered a similar situation. Waiting for silence would have meant waiting forever. Rather than treating the aircraft as a problem, he began treating it as part of the environment itself.

    Metzner’s answer reflected a recurring theme throughout the session. Recording is not always about removing the world. Sometimes it involves allowing the world to remain present. Sounds that initially appear intrusive may become important parts of the story. The aircraft was not simply interfering with the student’s recording. It was also shaping the student’s experience of being in that place. Standing in a park, looking upwards, waiting for the noise to pass, became part of the memory. In that sense, the aeroplane belonged to the story as much as the birds or the wind.

    The conversation then moved towards a problem that confronts many documentarians. The person who makes a recording remembers far more than the recording itself contains. They remember the weather, the location, the circumstances, and their own reactions. Future listeners possess none of this knowledge. How, then, can an experience be shared with somebody who was never there? A recording alone rarely provides a complete answer. Context becomes necessary. Yet explanation creates its own difficulties. Too little information leaves listeners uncertain about what they are hearing. Too much information can overwhelm the recording itself. Over the course of his career, Metzner has carried microphones through deserts, cities, forests, festivals, religious ceremonies, and countless other environments. Yet the purpose of these recordings has never been simply to build an archive of unusual sounds. Instead, they function as forms of communication. During the discussion, he compared recordings to postcards. A postcard never contains everything about a place. It presents only a fragment. Yet that fragment can still communicate something meaningful. Sound recordings operate in a similar way. They do not reproduce entire experiences. They provide partial access to them. Listeners complete the picture through imagination, memory, and interpretation.

    Listeners themselves become part of the process. No recording contains everything. Microphones record sound pressure variations. They do not record temperature, light, smell, movement, or the countless other details that contribute to an experience. Yet listeners rarely encounter recordings as collections of isolated sounds. They actively construct meaning from what they hear. A few seconds of ambience may be enough to suggest an entire environment. A familiar voice may evoke a person more vividly than a photograph. A distant church bell, footsteps in a corridor, or voices heard from another room can suggest a much larger world than the recording itself contains. Documentary production frequently relies upon this relationship between recording and imagination. Rather than attempting to communicate everything, the producer provides enough material for listeners to begin constructing their own understanding of a place, event, or experience. Recordings do not simply transmit information from one person to another. They create opportunities for participation. Listening becomes an active process through which people assemble impressions, associations, and memories from fragments of sound.

    Rain on a conservatory roof. Crickets during summer evenings. A vacuum cleaner moving through a family home. Songs sung by parents. Early computer games. Calls to prayer heard while travelling. When Metzner asked students to think about sounds they remembered from childhood, the answers arrived quickly. Few of the sounds were unusual. Their importance had little to do with acoustics. What mattered was everything attached to them. The examples revealed how deeply sound can become woven into personal history. Many of the memories were linked to recurring experiences rather than singular events. The sound of rain returning night after night. A family member singing repeatedly over many years. Household sounds that seemed insignificant at the time. Their importance emerged gradually through repetition. Long after specific conversations or individual days had been forgotten, the sounds remained. Several contributions also highlighted how difficult it can be to predict which sounds will become meaningful. People rarely decide in advance that a particular sound will become a lifelong memory. More often, significance emerges retrospectively. A sound that once seemed entirely ordinary acquires importance through later experience. Hearing a familiar sound years later can reactivate memories, emotions, and associations that extend far beyond the recording itself. What returns is rarely just the sound. People remember places, relationships, circumstances, and feelings connected to it. A recording therefore preserves more than an acoustic event. It can preserve pathways back towards experiences that might otherwise feel increasingly distant. A sound that appears entirely ordinary to one listener may carry decades of meaning for another. Hearing is rarely confined to the present moment. Certain sounds seem capable of collapsing time. A familiar voice, a piece of music, or an environmental sound can reconnect listeners with people, places, and relationships that might otherwise feel distant.

    While still in high school, Metzner began recording conversations with his grandfather. There was no documentary project in mind. He was not gathering material for publication. He simply wanted to preserve conversations with somebody he loved. Years later, those recordings became something entirely different. After his grandfather had died, the tapes acquired a significance that would have been impossible to recognise when they were first made. What had once seemed routine became irreplaceable. The story resonated with many listeners precisely because it involved no grand plan. Had Metzner waited until the recordings appeared important, it would already have been too late. Their value emerged from the simple decision to record ordinary conversations while the opportunity existed. From that experience came one of the clearest pieces of advice offered during the lecture. Record parents. Record grandparents. Record the people whose voices matter. Many recordings appear ordinary when they are made. Their value often becomes visible only later. The suggestion was not motivated by nostalgia alone. Voices contain forms of information that are difficult to preserve in any other way. Speech patterns, accents, pacing, humour, hesitation, and personality all become embedded within a recording. Written transcripts can preserve words. Recordings preserve presence. As Metzner reflected on these recordings, the discussion broadened into a larger point about time. Much of everyday life feels too ordinary to document. Conversations happen. People tell stories. Family members describe events that seem familiar and unremarkable. Yet these moments often become increasingly valuable as years pass. Recording provides a way of preserving details that might otherwise disappear unnoticed. Metzner’s reflections on these recordings returned repeatedly to the differences between memory and recording. Human memory is selective. Certain details remain while others disappear. Recordings preserve details indiscriminately. Accents. Hesitations. Laughter. Breathing. The rhythm of a voice. Background sounds that seemed unimportant at the time. Small details that might otherwise have been forgotten can later become deeply meaningful. A recording preserves more than information. It preserves traces of presence.

    Metzner has never been entirely comfortable with the phrase “capturing sounds”. The word suggests possession. It implies that a sound has somehow been seized and stored away. Throughout the discussion he returned to a different idea. Sounds are given rather than captured. Once a recording has been made, the challenge becomes helping somebody else experience what made that sound meaningful in the first place. Context matters. Stories matter. Yet explanation has limits. Documentary production often involves helping listeners approach an experience and then stepping aside so that the sounds can speak for themselves. A successful recording does not simply tell listeners what to think. It creates conditions in which they can form their own relationship with what they hear. The idea sits comfortably alongside much of his work. Recordings are not trophies collected from the world. They are invitations to listen more closely to it.

    Expensive microphones appeared surprisingly rarely in the lecture. Recording technology was never dismissed, though it was rarely placed at the centre of the discussion. Microphones matter. Recording techniques matter. Editing tools matter. Yet none of them can substitute for curiosity. A person who pays close attention to the world will often discover interesting sounds regardless of equipment. Conversely, expensive equipment cannot compensate for a lack of attention. Many of the examples discussed during the session pointed towards the same conclusion. Meaningful recordings often emerge from moments that other people would simply pass by. A sound heard while travelling. A conversation with a grandparent. Rain on a roof. An aircraft passing overhead. None of these experiences appear remarkable at first glance. Their significance emerges through listening.

    Near the end of the session, the lecture’s title, Going Places That Words Cannot Go, felt increasingly apt. Certain experiences resist straightforward description. The sound of rain on a roof. A grandparent’s voice. A crowded street in a distant city. A celebration, a conversation, or a moment of quiet. Words can describe such things. Sound can sometimes bring listeners closer to experiencing them. For Metzner, that possibility lies at the heart of listening. Sound does not simply tell us about the world. Under the right circumstances, it can preserve traces of people, places, and experiences long after the original moment has passed. More importantly, it can allow those experiences to be shared with somebody else. A recording offers only a fragment. A voice. A place. A conversation. A few seconds of sound preserved from a particular moment in time. Yet those fragments can remain meaningful for decades. They can reconnect people with memories, places, and relationships that might otherwise fade. Listening, as Metzner reminded students throughout the session, is not simply a way of gathering information about the world. It is one way of remaining connected to it.

  • How Do We Know What We Are Hearing? Professor Albert S. Bregman on Auditory Scene Analysis and Perceptual Organisation

    Albert Bregman

    How do we know what we are hearing?

    The question sounds simpler than it is. A voice is heard as a voice. A violin is heard as a violin. A passing vehicle is recognised almost immediately. Everyday listening creates the impression that sound sources reveal themselves directly. Most people rarely stop to consider how much processing has already taken place before recognition becomes possible. Professor Albert S. Bregman’s research begins from the observation that sound sources do not arrive at the ears. Acoustic mixtures do. By the time vibrations reach a listener, contributions from many different events have already combined. Voices, musical instruments, footsteps, ventilation systems, birdsong, machinery, and countless other sources may all contribute to the same signal. The auditory system must somehow determine which parts of that mixture belong together and which do not. Before Bregman’s work, hearing research had developed detailed accounts of pitch, loudness, masking, localisation, and frequency analysis. Considerably less attention had been paid to a more fundamental question. How does a listener determine what produced a sound?

    Bregman did not begin with a theory. He began with a puzzle. During memory experiments involving sequences of short sounds, he noticed that listeners often perceived groupings that were not physically present within the stimulus itself. Sounds sharing similar characteristics appeared to organise themselves into separate perceptual streams. The observation recalled ideas from Gestalt psychology, where visual elements combine into structures that cannot be understood simply by examining their individual parts. What began as an unexpected observation gradually became a larger problem. If listeners organise sounds into streams, how does that organisation occur? More importantly, what role does it play in perception itself? Bregman often approached the issue through analogy. Imagine standing beside a lake while observing only two floating markers moving up and down on the water’s surface. The movement provides evidence that something has happened, though many explanations remain possible. A boat may have passed nearby. Several boats may be moving in different directions. Wind may be disturbing the surface. Something may have fallen into the water. The available evidence does not identify the cause. Any conclusion depends upon inference. According to Bregman, hearing presents a similar challenge. The ears receive information about acoustic activity, though they do not receive direct information about the events that produced it. From patterns of pressure variation reaching two eardrums, listeners somehow infer the existence of voices, instruments, machines, animals, and other sound-producing events. Nothing in the signal arrives labelled. The auditory system must determine which acoustic components belong to the same source.

    Questions of speech perception, localisation, attention, and communication all depend upon this process. Before speech can be understood, before a melody can be followed, and before a sound source can be identified, the auditory system must first determine which acoustic components belong together. Organisation is therefore not one stage among many. It provides the conditions under which many other aspects of perception become possible. This perspective led Bregman towards what became known as auditory scene analysis. The term reflects an analogy with vision. Just as visual perception involves identifying objects within a visual scene, auditory perception involves identifying sound-producing events within an acoustic scene. The challenge lies in the fact that sound sources combine before reaching the listener. The auditory system therefore faces a decomposition problem. It must separate a complex mixture into components that plausibly belong to distinct events. A central claim running throughout the lecture was that perception involves more than detecting acoustic information. It also involves organising that information. Bregman’s demonstrations repeatedly returned to this point. Listeners often assume that qualities such as rhythm, melody, pitch, loudness, timbre, and location belong directly to sounds themselves. His examples suggested a more complicated picture.

    Auditory stream segregation provides one illustration. Under certain conditions, listeners stop hearing a single sequence of sounds and begin hearing multiple independent streams. Once this occurs, rhythms that were previously obvious may disappear. New rhythmic structures emerge. Melodic patterns change. The acoustic signal remains unchanged, though the perceptual outcome does not. Bregman’s demonstrations suggested that the consequences extend much further than rhythm or melody alone. Again and again, he returned to the idea that many perceptual properties depend upon how sounds are grouped. Listeners often assume that pitch, loudness, timbre, and spatial location belong directly to sounds themselves. Yet these properties can also be influenced by the way acoustic components are assembled into perceptual objects. When those groupings change, perception may change even when the underlying stimulus remains constant. This claim sits near the centre of auditory scene analysis. The framework is not simply concerned with separating one sound source from another. It is concerned with how perceptual objects are formed in the first place. Before listeners can judge the loudness of a sound, identify its pitch, recognise its timbre, or determine its location, the auditory system must first decide which components belong together. The resulting structure shapes many of the properties that listeners subsequently experience. From this perspective, perception becomes a problem of interpretation. Faced with an acoustic mixture, the auditory system must determine which explanation is most plausible. What listeners hear is not a direct copy of the physical world. It is the outcome of a process through which the auditory system attempts to reconstruct the events most likely to have produced the available evidence.

    Bregman argued that listeners exploit regularities commonly found in the physical world. Certain acoustic relationships provide evidence that components are likely to originate from the same source. Harmonicity offers one example. Many naturally occurring sounds contain frequency components related by simple numerical ratios. When such relationships are detected, the auditory system often groups those components together. Similar reasoning underlies what Bregman described as common fate. Components that begin together, change together, or move together over time frequently appear to belong to the same event. These principles do not guarantee correct interpretation. Rather, they provide strategies that usually correspond with the structure of the physical world. Auditory scene analysis is therefore concerned with probability rather than certainty. The auditory system rarely knows exactly what caused a sound. It generates interpretations that are likely to account for the available evidence. Most of the time those interpretations correspond closely enough to events in the environment that listeners remain unaware that any interpretation has occurred at all. Throughout the lecture, Bregman emphasised that these organisational processes usually pass unnoticed. Listeners rarely experience themselves as constructing interpretations. The world appears already divided into voices, instruments, footsteps, vehicles, and other familiar sources. Auditory scene analysis directs attention to the work required to produce that impression. The apparent simplicity of hearing may be one reason the problem remained difficult to recognise. Successful perception conceals many of the processes that make it possible.

    Music occupied an interesting position within the lecture. Bregman suggested that composers had discovered practical consequences of auditory organisation long before psychologists attempted to explain them. Counterpoint, orchestration, and performance practice frequently involve maintaining distinctions between perceptual streams or encouraging sounds to fuse into larger structures. Musical traditions therefore provide a long record of experimentation with the same organisational tendencies that auditory scene analysis later sought to describe. Music also offers situations in which these processes become unusually apparent. Changes in perceptual organisation can alter the melodies and rhythms listeners hear, making it possible to observe principles that often remain hidden during everyday listening. Bregman was not suggesting that composers were unconsciously applying psychological theory. Rather, centuries of musical practice had encountered many of the same perceptual constraints that later became objects of scientific investigation.

    Yet music represented only one instance of a broader phenomenon. Following a conversation in a crowded room, recognising a familiar voice over the telephone, locating a sound source in a busy environment, distinguishing one instrument from another, and understanding speech in noise may appear to involve different problems. Bregman’s framework suggested that each depends upon a prior act of organisation. Auditory scene analysis altered the relationship between many areas of hearing research by drawing attention to this common foundation. Rather than treating speech, music, localisation, and auditory attention as entirely separate domains, the framework highlighted organisational processes upon which they all depend. Seen in this way, auditory scene analysis is not merely a theory about particular auditory illusions or laboratory demonstrations. It addresses a question that sits beneath much of auditory perception research. How does a listener move from an undifferentiated acoustic mixture to a world populated by distinct events, objects, and sources?

    The framework also shifted attention away from sound as a purely physical phenomenon and towards perception as a process of inference. Earlier approaches often focused on the contents of the acoustic signal. Bregman drew attention to a prior question. Before a listener can recognise a voice, identify an instrument, understand speech, or respond to a warning signal, the auditory system must first decide what probably caused the sound.

    The answer is usually reached so quickly that the problem remains unnoticed. Voices appear as voices. Instruments appear as instruments. Bregman’s work suggests otherwise. Listening depends upon a continual process through which the auditory system constructs explanations from incomplete evidence. Most of the time those explanations correspond closely enough to the surrounding environment that hearing feels direct and effortless.