lullable

The Sleep Library · A definition

Boring voices to fall asleep to, and what makes one work

The short answer

A voice that is easy to fall asleep to is one that has given up trying to hold your attention. In practice that means a narrow pitch range rather than an expressive one, an even pace with no accelerations, a moderate and unvarying volume, and no performance — no rising inflection promising that something is about to happen. People search for boring voices to fall asleep to because boring is the correct specification, not a compromise. The qualities that make the best podcast voices — warmth, dynamics, timing, the withheld reveal — are precisely the qualities that keep a listener present.

What the ear is actually reacting to

There is a reasonably precise answer to which parts of a voice grab you, and it comes from work on emotional prosody rather than from anything about sleep.

In 2008 Wiethoff and colleagues played listeners words spoken in neutral, happy, erotic, angry and fearful intonations during a passive-listening fMRI experiment — nobody was asked to judge anything, they simply heard the voices. Responses in the right mid superior temporal gyrus were significantly stronger for all of the emotional intonations than for the neutral one. The interesting part is what they did next: they regressed those responses against the raw acoustics, and found the region tracked mean intensity, mean fundamental frequency, variability of fundamental frequency, and duration.

Translated out of the notation, that is loudness, how high the voice sits, how much the pitch moves around, and how long the utterance runs. And no single one of those explained the effect on its own — the stronger response disappeared only when they corrected for arousal, or for all of the acoustic parameters taken together.

So the specification is a conjunction rather than a single dial. A voice can be quiet and still be wrong, if the pitch keeps travelling. That is the entire failure mode of a good narrator reading a good book, and it is why podcasts keep people awake even at low volume.

The hearing does not clock off

The other half of the problem is that the voice has to keep behaving for the whole file, long after you have stopped consciously listening.

In 1999 Perrin and colleagues recorded auditory evoked potentials in ten adults, playing them their own first name mixed randomly among seven others, while awake and then during stage II and REM sleep. In stage II, all the names evoked K-complexes — but the early portion of that complex was selectively larger after the subject's own name. In REM, a late positive wave appeared for own names and not for the others.

Ten people, and a quarter of a century ago, so hold it loosely. But it establishes the thing that matters here: the sleeping brain is still sorting what it hears into more and less relevant, at least some of the time. A narrator who becomes suddenly emphatic at minute forty is speaking to a system that is still, in a reduced way, listening for exactly that.

Whispering is a different product

The nearest neighbour to this question is ASMR, and it is worth separating.

Poerio and colleagues ran two studies in 2018 — a large online experiment and a laboratory one — and found that ASMR videos increased pleasant affect only in people who already experienced the tingling response. The laboratory study found ASMR associated with reduced heart rate, and also with increased skin conductance, which is a marker pointing the other way. It is a real and physiologically rooted experience. It is not a universal one, and it is not straightforwardly low-arousal.

Whispering, in other words, is an intense intimate delivery that happens to be quiet. A flat voice at conversational volume is a different instrument, and the two get filed together only because both are unlike broadcast.

The delivery notes, plainly

- Pace: unhurried and constant. No accelerating into anything. - Pitch: narrow range, sentences falling at the end rather than lifting. - Volume: even. The peak matters more than the average, because the peak is what arrives. - Breath: audible and unedited. Removing it is a production instinct that costs you the most obviously human, least urgent sound in the file. - Structure: no cliffhangers, no held-back reveal, no question left open across a break.

None of this is difficult to do. It is difficult to want to do, because every instinct trained by narration work runs the other way.

Where we come in, honestly

Lullable's narration is built to this list, which is why our voices sound slightly wrong on first listen and correct on the tenth. The endings are given away early on purpose; the pace does not vary by design; nothing in the file is waiting to grab you. That is the whole of low-arousal audio, and it is what a sleep story is once you take the marketing out.

If pitch is the thing you notice most, the catalogue is read by a man and read by a woman, and it is worth trying both — the range that disappears for you is not the range that disappears for someone else.

Lullable reads material like this aloud — warmly, slowly, and quieter every minute — until you drift off somewhere around the fourth clause.

Get the app