Somewhere on the steppe below the Altai Mountains, a herder sits facing the open grassland and begins to sing. From his throat comes not one note but two: a deep, grainy drone and, floating above it like a separate instrument, a thin crystalline whistle tracing its own melody. No electronics, no trickery, no hidden second singer. If you did not know the physics, you would call it magic. The remarkable thing is that once you do know the physics, it does not stop sounding like magic.
One buzz, many notes
Every voice — yours included — is already a chord. When the vocal folds vibrate, they produce not a single frequency but a whole harmonic series: a fundamental plus overtones at two, three, four times that frequency and beyond. What we hear as "one note" is this bundle fused together, its inner proportions giving each voice its timbre. This is the source-filter picture of the voice that the Swedish acoustician Gunnar Fant formalized in 1960: the larynx supplies a harmonically rich buzz, and the throat, mouth and lips act as a filter, boosting some regions of the spectrum — the formants — and shading others.
An overtone singer learns to sharpen that filter to a knife's edge. By shaping tongue, lips and pharynx with unusual precision, the singer narrows a formant until it seizes on a single harmonic and lifts it out of the blend, loud enough for the ear to accept it as a note of its own. A 2020 study in the journal eLife put Tuvan singers in an MRI scanner and watched it happen: the singers merge two of the vocal tract's resonances into a single, unusually sharp peak that tracks one harmonic at a time. In the low kargyraa style, the ventricular folds above the vocal cords join in at half the frequency, dropping the drone by an octave and doubling the harmonics available for the filter to choose from.
Two notes from one voice is not a violation of acoustics. It is acoustics, played like an instrument.
A world map of the split voice
The heartland of the art is Inner Asia. In the Republic of Tuva, khoomei unfolds into a family of styles: sygyt ("whistle"), which isolates one piercing harmonic high above the drone; kargyraa, the sub-octave growl; ezengileer, which pulses like stirrups at a trot; borbangnadyr, a rolling shimmer of overtones. Across the border, Mongolian khöömii — inscribed by UNESCO on its list of the Intangible Cultural Heritage of Humanity in 2010 — is described by its own practitioners as an imitation of nature: wind across grass, water over stones, birdsong.
Then the map surprises you. On Sardinia, the cantu a tenore — recognized by UNESCO in 2005 — sets four male voices in a circle, two of them (bassu and contra) singing with overtone-rich guttural techniques that lay a harmonic bed under the melody. And in South Africa, Xhosa women practice umngqokolo, a form of overtone singing documented by the ethnomusicologist Dave Dargie only in the 1980s. Continents apart, with no plausible contact, people bent over the same physics and found the same door. Wherever human attention turns patiently to the voice, the harmonic series is waiting.
What the traditions use it for
The URL of this article contains the word healing, so let us take the question seriously — by asking who says what.
In Tuva, the ethnomusicologist Theodore Levin, who spent years recording there with the Tuvan scholar Valentina Süzükei (Where Rivers and Mountains Sing, 2006), describes throat singing as part of an animist world in which rivers, caves and mountains have spirit-masters, and sound is a way of addressing them. A herder singing sygyt toward a valley is not performing; in the tradition's own words, he is in conversation with the place. Within that world, sound also belongs to the shaman's toolkit — voice, drum and jaw harp are used in rituals of healing and protection. That is the tradition's claim, stated in the tradition's language, and it deserves to be heard as what it is: centuries of practice, described through a different lens than ours.
In Tibet, the monks of the Gyuto and Gyume monasteries cultivate an extraordinarily deep ceremonial chant in which a harmonic hovers audibly above the fundamental. When the scholar of religions Huston Smith first heard it in the 1960s he was startled enough to bring acousticians into the room; the result was a 1967 paper in the Journal of the Acoustical Society of America on "an unusual mode of chanting by certain Tibetan lamas" — a rare, lovely case of the primary source being both a monastery and a spectrogram. For the monks themselves, the chant is not a vocal feat but tantric liturgy: the voice is part of the ritual, not a performance about it.
What the other lenses see — and what they don't
The measurement lens can say a great deal here, and says it happily. The two notes are real, not illusory: they show up on any spectrogram. The vibration a singer feels in the skull and chest is real too — bone conduction, physics you can feel with a hand on the breastbone. The long, regulated exhalation the technique demands is the kind of slow, deliberate breathing that contemplative traditions from yoga to monastic chant have cultivated on purpose for millennia.
What the measurement lens cannot yet say is whether the practice does for the practitioner what the traditions say it does. Research on sound and the body — including the field known as vibroacoustics — is active, but its results so far are early-stage and mixed, and honesty requires saying so. The traditions describe care, ceremony and transformation in their vocabulary; the laboratory is still learning to phrase the question in its own. Those are two descriptions in two languages, and the comparison between them is precisely the interesting part — not something to be settled by shouting one language over the other.
The door is in your own voice
Here is the open secret: the harmonic series this whole story rests on is in your voice right now. Hum a comfortable note and slowly morph the vowel from ee to oo and back, keeping the pitch fixed. Somewhere in the glide, faint and silvery, you will hear a whistle sweep down and up — your own overtones, briefly stepping out of the blend. Herders, monks and mothers on three continents heard that whisper, followed it for generations, and built worlds of meaning around it.
Why a drone with one moving harmonic holds human attention the way it does — why this particular sound, of all sounds, became prayer in Tibet, conversation with landscape in Tuva, and ceremony in the Eastern Cape — is a question no spectrogram closes. It stays open. That is not a failure of the story. It is the story.