
When someone near you screams, it’s hard to know whether it’s fear, joy, or excitement. Another clue, this one visual, seems, on the contrary, almost immediate. Researchers, from Charles Darwin in 1872 to an article in
Psychology Today published in July 2025, describe the same thing: our brain does not process all emotional signals at the same speed. To decide quickly, it gives priority to a very specific channel, and it’s not the sound.
This channel is facial expressions, which we learn to decode from the first months. Work cited by Harriet Oster and Charles Darwin shows that babies respond to a happy or angry face long before they understand a sentence or tone of voice. A meta-analysis published in 2017 in the journal Social Cognitive and Affective Neurosciencebringing together 76 studies on the face and 34 on the voice, indicates that each channel mobilizes largely distinct brain networks. It all points to the same idea: by default, the brain first bets on what it sees.
Survival: when the face gives clear orders to the brain
In The Expression of the Emotions in Man and AnimalsCharles Darwin already described the face as a tool forged by evolution. A furrowed brow, clenched teeth or widened eyes gave our ancestors a clear instruction: flee, attack or submit.
Studies show that these codes are found in different cultures, for emotions such as joy, sadness, anger, fear, surprise or disgust. In a situation of danger, this visual language is better than a long sound deciphering. Certain micro-expressions, too brief to be controlled, make this language even more reliable.
Nonverbal sounds are fuzzier. An isolated cry can accompany terror, a sensational attraction or a sporting victory, and only the situation allows us to decide. The sound is quickly disrupted by distance, background noise or a cold that makes the voice tremble. Acoustic parameters such as pitch or volume overlap between several emotions, which makes the auditory canal less precise in guiding an on-the-spot reaction.
In the brain, shortcuts for the face, not for the voice
Neuroscience confirms this favoritism. The Fusiform Face Area (FFA) helps identify a face in a flash, while the Superior Temporal Sulcus (STS) tracks changes in expression.
The 2017 meta-analysis published in Social Cognitive and Affective Neuroscience also describes a strong involvement of the amygdala for faces, a key region for fear and emotional memory. The voice mainly recruits the superior temporal gyrus (Superior Temporal Gyrus (STG), specialized in complex sounds, which corresponds to more analytical and context-dependent processing.
When the voice takes the advantage to read emotions
The work of Michael Kraus, published in the journal American Psychologistremind us that voice matters: in his experiments, participants judged emotions better when they only had their voice, such as on the telephone or when the face is hidden.