Overview | Mindmap | Literature review | Case study | Sounds | Video | Findings | Further questions
Under what conditions does organized sound generate pre-linguistic meaning?
A research project by Andreas Russo — ReSound Erasmus Mundus Joint MA, Universidade Lusófona, Lisbon, 2026.
Download the full paper.
Overview
Sound reaches us before we have the words for it. This project asks under what conditions organized sound — from music to soundscape — produces a felt response without requiring instruction, cultural knowledge, or conscious interpretation. That kind of meaning, registered in the body before it has been named or understood, is what this project calls pre-linguistic.
The investigation uses Russell’s model of affect as a framework for mapping felt states without cultural interpretation and draws on Jung and Schopenhauer, two thinkers who seem to be describing a similar psychic layer from different directions. Schopenhauer argued that music does not represent phenomena but expresses the inner nature of phenomena directly, bypassing cultural meaning entirely. Jung described archetypes as universal, inborn patterns of behaviour, present across all human experience, not culturally acquired. This project explores whether human predisposition to respond to organized sound operates at that same level.
© Vera Marmelo
Introduction
My practice as a sound designer and composer for visual media convinced me that sound reaches audiences at a level that precedes conscious interpretation. This idea turned into Stillness/Motion, a participatory quadraphonic live performance in which I mixed sound in real time for a room of people, using their movement as a score and my intuition as methodology. This project is the attempt to examine what the performance revealed about how organized sound generates felt meaning before cultural interpretation is possible.
Schopenhauer (1818/1969) argued that music accesses something prior to cultural representation. Jung (1959/2014a) described archetypes as universal, inborn patterns of behavior. This project asks whether human predisposition to respond to organized sound operates at that same level.
Sub-questions
1. How does organized sound reach the body before cultural interpretation occurs?
2. How do specific acoustic properties (spectral content, envelope shape, spatial movement) carry that response?
3. What does a practice-based, participatory sound performance reveal about the conditions under which this happens?
Case study summary
The central finding came near the end of the first of two performances. A rotating wind sound, accelerating through the quadraphonic speakers, pulled a room of people from different cultures into a collective circular movement, without instruction or cultural cue. When a technical failure abruptly cut the sound, the room stopped instantaneously. What was happening appeared to precede interpretation.
Mindmap
Literature review
Arthur Schopenhauer
Schopenhauer — The World as Will and Representation (1818)
In The World as Will and Representation, Schopenhauer claims that music, unlike every other art, is not a representation of phenomena, but rather an expression of their inner nature. Not this sorrow, this joy, this gaiety — but the essential nature of sorrow, joy, and gaiety in the abstract, stripped of any specific occasion. He argues that music “never expresses the phenomenon, but only the inner nature, the in-itself, of every phenomenon” — not any particular felt state, but the essential nature of felt states themselves. Where other arts represent their subject matter, music expresses it directly.
Carl G. Jung
Jung — Archetypes and the Collective Unconscious (1959)
Jung — The Structure and Dynamics of the Psyche (1960)
Jung’s theory of archetypes proposes that certain patterns of behaviour are inborn — present across all human cultures, not acquired through experience or cultural transmission. Archetypes are irrepresentable in themselves, but they produce what Jung calls archetypal images: culturally shaped manifestations that differ across traditions while pointing back to the same underlying predisposition. The Great Mother is not any particular mother or goddess but the underlying tendency that produces Isis, Mary, and Kali as its culturally specific forms. Similarly, the Underworld archetype produces Hades, Duat, and Hel across Greek, Egyptian, and Norse traditions.
These archetypes belong to what Jung calls the collective unconscious — the layer of the psyche shared across all human experience, prior to culture.
The fact that no known culture lacks music suggests a universal predisposition to organise and respond to sound. This is archetype-like. But organised sound is representable, transmissible, and performable — which excludes it from being an archetype. And it is not an archetypal image either, since it does not arrive as a culturally shaped product but as the medium through which archetypal responses are activated. It triggers a universal response without determining the cultural form that response takes.
The Jung/Schopenhauer connection
Schopenhauer describes what music does. Jung describes what might be the layer of the psyche it does it to. Neither thinker completed the connection.
Jung himself seems to have recognized the gap. In 1956, pianist Margaret Tilly gave him a private demonstration of music therapy. He reportedly said he avoided music because it deals with “such deep archetypal material” that most musicians don’t realise what they are working with. (McGuire & Hull, 1977/2020) After the session he said he would recommend music therapy in every analysis — but he died five years later without addressing it in his writings. That silence is part of what inspired this project.
© Vera Marmelo
James A. Russell
Russell — A Circumplex Model of Affect (1980)
Russell provides a framework for mapping the affective states that sound produces: all felt states can be plotted along two axes, valence (pleasant to unpleasant) and arousal (activated to deactivated).
Patrik N. Juslin & Daniel Västfjäll
Juslin & Västfjäll — Emotional Responses to Music (2008)
Juslin and Västfjäll identify several mechanisms through which music produces emotional responses. The most relevant here is the brain stem reflex: a physiological response triggered by acoustically salient events — sudden attacks, high frequencies, spatial movement — before cortical processing has occurred. The body responds to certain acoustic properties as signals before cultural knowledge is activated.
Spatial behaviour, spectral character, and temporal envelope carry pre-linguistic affect because they operate below the threshold at which cultural learning becomes relevant.
Pierre Schaeffer
Schaeffer — Treatise on Musical Objects (1966)
Pierre Schaeffer proposed reduced listening as a mode of attending to sound as a pure acoustic object that is free of context, stripped of its source, cause, and cultural associations.
The Stillness/Motion performance raised questions about this premise. Several sounds used in the piece (breath, non-tonal consonants, percussive mouth sounds, footstep-like tapping) carried embodied associations regardless of any instruction to hear them acoustically. If those associations are not cultural but physical — if they trigger a brain stem reflex before a learned response — then reduced listening may describe an ideal the body cannot actually reach. It may always respond before the instruction arrives.
Janet Cardiff
Cardiff’s installation is the closest practical precedent for the questions this research asks. Forty speakers, each carrying a single voice from a recording of Thomas Tallis’s Spem in Alium, are arranged in an oval around a room. Visitors move physically toward the voices that pull them — whether driven by curiosity or reflex, the movement precedes interpretation. Visitors navigate the sound spatially before they know what they are looking for.
The piece does not explain the mechanism. But it demonstrates that spatial sound produces embodied navigation before conceptual meaning is formed.
Case study
How the piece came to be
The original concept for this project was a participatory installation using algorithmic body tracking. Each participant would be assigned a unique tone; their proximity to others would generate harmonic relationships. After months of development, the technical layer proved too complex, and the panning too imprecise for people to identify which tone was theirs. The sonic consequences of their movement were inaudible.
On the advice of professor Adriana Sá, the piece became a live performance: manual fader riding in real time, with tape zones on the floor corresponding to sound layers. When someone occupied a zone, that layer became active.
© Vera Marmelo
The sound material
Four classmates (Autumn Bochart, Justin Enoch, Shyalina Muthumudalige, Irazema Vera) and I gathered around a microphone making non-tonal consonants, breath, and percussive mouth and hand sounds. These recordings, layered with field recordings of wind, rain, and sea waves, formed the texture of the piece.
All sounds were chosen for their potential to morph seamlessly into one another. No effects other than minimal equalisation. A Max/MSP patch allowed real-time mixing, placement, and rotation of each sound across the four speakers at variable speeds, with some material feeding only the LFE subwoofer.
What happened
The piece ran in two back-to-back performances of approximately eight minutes each, for fifteen to twenty people moving in and out of the room.
For most of the time, the relationship between the sound layers and the floor zones was illegible. Participants defaulted to passive listening. The instruction given at the door — that they were free to walk, dance, or stand still — communicated freedom without communicating consequence.
However, two moments broke through, near the end of the first performance:
Rotating wind – I began panning a wind sound through the four speakers in an accelerating circular motion. One participant started spinning in place. Then another began walking in a circle following the direction of the sound. Then all of them. As the rotation accelerated, so did the people. Within seconds everyone in the room was running in a circle, laughing.
Wooden box – While people were still circling, I introduced a granular, transient sound with the acoustic profile of footsteps. One participant began tapping her feet, still spinning.
Technical failure – The vibration from their stomping disconnected the USB cable from the audio interface. The sound cut completely. The room stopped instantaneously. People looked around in uncertainty, then clapped and laughed.
What it suggests
The circling moment is the central finding. The rotating wind sound carried no pre-established meaning (no narrative, no instruction, no cultural cue) yet participants followed its motion with their bodies. Spatial movement appears to have been decisive: it worked where static, undifferentiated layers did not.
What Juslin and Västfjäll would identify as a brain stem reflex appears to have preceded any conscious decision. The technical failure may reinforce this: when the sound cut, the collective movement stopped instantaneously, suggesting the response was being actively sustained by the sound rather than by social dynamics alone.
We don’t know what the spinning moment meant to each person. What is observable is that the response was collective, physical, and required no cultural knowledge: a room of people from different countries responded in kind to a rotating sound. This appears consistent with Jung’s framework — a layer of response that precedes cultural shaping — without confirming which archetype, if any, was activated.
Russell’s model is partially confirmed: arousal was clear, as the accelerating rotation produced a clear escalation in the room. Valence was not, since the piece never became legible enough as emotional territory for participants to navigate consciously.
Tensions between theory and practice
The failure in the performance was as informative as the success: organized sound generates pre-linguistic meaning only when the perceptual chain is immediate enough that the body responds before the mind asks what it means. Schaeffer’s premise also ran into trouble: sounds like breath, consonants, and footstep-like tapping carried embodied associations regardless of instruction. The conditions for pre-linguistic meaning appear more specific and more physical than the hypothesis anticipated.
Further questions
- What did participants actually experience? Reports of their felt experience collected immediately after the performance would be the obvious next step.
- Would a genuinely autonomous system produce the same results, or does the visible presence of the performer change the dynamic?
- How do these mechanisms operate in narrative film, where sound design intersects with visuals to tell a story?
© Vera Marmelo
AI disclosure
This research was developed in dialogue with Claude (Anthropic, claude-sonnet-4-5) across multiple sessions.
Interface: Text-based exchange. No persistent memory between sessions. Continuity was maintained by uploading an iterative context document at the start of each session. The user retains full agency over all content decisions. Claude reveals its reasoning inline but conceals training data and the basis for its suggestions. All theoretical references were independently verified by the author.
Sequence of use: Multiple sessions spanning the full research arc, from initial question formulation, through literature review drafting, to case study structuring and final paper editing. Each session began with an uploaded summary of previous sessions and outstanding tasks.
Critical assessment: The tool was most reliable for structural tasks (organising the argument’s architecture, identifying redundancies, catching citation errors) and less useful (if not detrimental) as a conceptual partner: Claude’s tendency toward agreement led the author into directions that went nowhere, required backtracking, and caused a lot of wasted time. The most productive use was treating it as an editor and mistake-finder rather than a co-thinker, which is how the author intends to use AI in the future.



