LOADING 0%
// nav_menu.exe
Home Resume Blog Contact Order
English فارسی
~/blog / technology / dial-up-screech

What Was That Screech When Dial-Up Connected?

Dialing your internet provider always started normally: dial tone, a few clicks of dialing, a ring. But the moment the other end picked up, the home phone went feral — a few seconds of screeching, howling, nothing like any human language. Then sudden silence, and the familiar sign: you were online.

Most of us glanced at the modem and assumed it was "doing its thing." But that sound wasn't the sound of working — it was the conversation itself. Every word of it had meaning. And the strangest part: you didn't need a modem to hear it. The ordinary phone sitting on the table was enough.

What Was That Screech When Dial-Up Connected?

# Short answer: it was two computers talking

There was a modem on the other end too — your ISP's. Before a single byte of data could move, these two machines had to get acquainted: pick a shared language, test the phone line, and agree on the fastest speed that line could honestly deliver that night. The whole ritual is called a handshake. That screech was two machines shaking hands — the loudest, most talkative handshake in history.
Why was it audible on an ordinary handset? Because this conversation traveled through exactly the same band your own voice uses on the phone line. The far modem pushed its voice-shaped waves into the wire, and the wire has no idea whether the source is a human or a machine — it just carries waves. The regular phone was effectively a loudspeaker sitting in the middle of a machine-to-machine call. That's why kids back then could dial the ISP's number and listen to the "aliens" answer — with nothing but a phone. And get scared.
So let's listen to that conversation scene by scene. Because from the first beep to the last hiss, everything meant something:
1
Dial tone & dialing — the line is free; every key sends two pure tones
2
Ring & answer tone — the far modem introduces itself at 2100 Hz
3
Line probing & speed agreement — the screechy part: noise and echo get measured
4
Connected — the continuous hiss means data in transit

# Why was the internet riding on phone audio in the first place?

The telephone network was built for exactly one job: carrying human speech from one house to another. So the entire path — filters, amplifiers, the wiring itself — was tuned to let through only waves in roughly the 300 to 3400 Hz band, the range where human speech lives. Anything below or above it was simply thrown away.
And the network couldn't care less whether a wave was produced by a person's vocal cords or by the most carefully engineered electronics in the world — as long as it stayed inside the band. So the modem did something clever: it disguised data as sound. The bits that needed to arrive were encoded into the frequency, phase, and amplitude of waves that, to the phone network, looked perfectly voice-like. The phone thought it was moving voices; it was actually moving code.
This also answers the main question: why the screaming? Because dense data spreads across the entire audio band — no silence, no pauses, no rhythm. Human speech is made of orderly, predictable waves; a stream of bits is a nearly random wave across every frequency at once. And a nearly random sound, to human ears, is a screech. That sound didn't "sound bad" — it sounded exactly the way raw data has to sound.
Human voice

Orderly, rhythmic waves with silence and pauses — optimized for the human ear

Data sound

Nearly random, spread across the whole 300–3400 Hz band, without a single moment of silence — optimized for the receiving modem, not the ear

# Before the screech, the beeps: every key played two tones

The conversation had started with the very first beep; we just didn't know it. The dial tone meant "the line is free" — two pure waves laid on top of each other. And every key you pressed sent two tones: one from the row-frequency group, one from the column-frequency group. This scheme is called DTMF — dual-tone multi-frequency:
1209 Hz 1336 Hz 1477 Hz
697 Hz 1 2 3
770 Hz 4 5 6
852 Hz 7 8 9
941 Hz * 0 #
key 5 = 770 Hz + 1336 Hz
Each key = one row frequency + one column frequency; the network reads the digit from that pair
Why two tones instead of one? To resist lies. If every digit were a single tone, any single-frequency noise on the way could forge one. But a pair must arrive together, exactly matched — a rare accident. (And that "every key is a chord" design is also why you could play melodies on a phone keypad — anyone hammering star and hash in rhythm was essentially programming an instrument.)
After the last digit came the ringback tone. And this time, whoever picked up the phone wasn't human. It was a machine that hears a ring not with ears, but with a circuit.

# The 2100 Hz tone: "Hello, I am not human"

The ISP's modem answered not with "Yes?" but with a continuous 2100 Hz tone — the answer tone. It did two jobs at once. The first was identification: "there is a machine on this end." The second was more technical: your modem used this tone to start learning the line — measuring how its own sound came back, with how much delay and how much weakening, so it could later separate its own echoes from the other side's voice.
Then came a run of short double-frequency beeps — the capability menu. With these, each modem read out the list of languages it speaks, like two negotiators exchanging business cards before the meeting: "I can do all of these — what about you?" The two sides pick the best shared language and head into the main event. The one this article is named after.

# The main screech: probing the line and haggling over speed

Now the two modems begin sending patterns they agreed on in advance — that's the howling part. The logic is as clean as a lab trick: the receiving modem knows exactly what wave should arrive. Every difference between what was sent and what arrived is a measurement of the line itself: which frequencies the noise swallows, which get weakened the most, how many milliseconds the echo takes to come home. With those few seconds, the modem tunes itself for this line, this night, this exact moment.
The test result converts directly into speed. Clean line? Agreement at 56 kilobits per second — the ceiling of the V.90 standard. Noisy line? Speed steps down, because more speed means bits packed closer together, and a noisy line garbles them faster. So those few seconds of screeching were really this sentence: "let's find out how fast you can honestly talk tonight."
Clean line

Little noise → bits can be packed tighter → agreement at the ceiling: 56 kbit/s (V.90)

Noisy line

Heavy noise → packed bits collide → speed steps down to keep the transfer reliable

And after the agreement? A continuous, ever-present hiss: the data itself, flowing. No more clean tones or recognizable patterns — hundreds of bits at any instant, spread across the whole band. And anything that disturbed this stream — say, a handset lifted in another room — killed the connection. The most classic family fight of that era comes straight from here: "Get off the phone!" really meant "Don't corrupt the modems' agreement!"

# Let's rebuild that sound in Python

To hear with our own ears what "meaningful" sounds like, we'll rebuild the whole conversation phase by phase and save it as a WAV file: dial tone, dialing, ringback, answer tone, capability menu, line probing, and the connected hiss. One delightful detail before the code: the sampling rate is set to 8000 — precisely the rate the heart of the phone network itself worked at, because it never carried anything above 4 kHz. Our file lives under the same constraint the network imposed on modems:
dialup_synth.py
# dialup_synth.py - rebuild the screech, phase by phase
import wave, math, random, struct
 
SR = 8000 # the phone network's own sampling rate
 
DTMF = {'1':(697,1209), '2':(697,1336), '3':(697,1477),
        '4':(770,1209),  '5':(770,1336),  '6':(770,1477),
        '7':(852,1209),  '8':(852,1336),  '9':(852,1477),
        '*':(941,1209),  '0':(941,1336),  '#':(941,1477)}
 
def tone(freqs, sec, amp=0.30):
    out = []
    for i in range(int(SR*sec)):
        t = i / SR
        s = sum(math.sin(2*math.pi*f*t) for f in freqs)
        out.append(amp * s / len(freqs))
    return out
 
def silence(sec):
    return [0.0] * int(SR*sec)
 
def band_noise(sec, amp=0.30, lo=300, hi=3400):
    # random data spread over the voice band - that's why it screams
    out, y_hi, y_lo = [], 0.0, 0.0
    k_hi = math.exp(-2*math.pi*hi/SR)
    k_lo = math.exp(-2*math.pi*lo/SR)
    for _ in range(int(SR*sec)):
        x = random.uniform(-1, 1)
        y_hi = x + k_hi*(y_hi - x)
        y_lo = y_hi + k_lo*(y_lo - y_hi)
        out.append(amp * (y_hi - y_lo))
    return out
 
audio  = silence(0.3)
audio += tone((350, 440), 0.6)           # dial tone: an empty line
audio += silence(0.2)
for d in "808080":                   # dialing the ISP number
    audio += tone(DTMF[d], 0.08, 0.35)
    audio += silence(0.06)
audio += tone((440, 480), 0.9, 0.25)   # ringback tone
audio += silence(0.4)
audio += tone((2100,), 1.4)            # ANSWER tone: "a modem here"
audio += silence(0.2)
for f in [(1000,1600), (1100,1700), (1000,1700), (1100,1600)]:
    audio += tone(f, 0.06)            # V.8 menu bongs (simplified)
    audio += silence(0.02)
audio += silence(0.2)
audio += band_noise(1.6, 0.40)          # TRAINING: probing the line
audio += band_noise(3.0, 0.18)          # connected: the hiss IS the data
 
frames = b''.join(struct.pack('<h', int(max(-1, min(1, s)) * 32767))
                   for s in audio)
with wave.open('dialup.wav', 'w') as w:
    w.setnchannels(1); w.setsampwidth(2); w.setframerate(SR)
    w.writeframes(frames)
print(f"dialup.wav written - {len(audio)/SR:.1f} seconds of nostalgia")
This is a free reconstruction: the phase order and the key frequencies are real, but the fine details of the protocols are simplified so the script runs in seconds and produces a ten-second file. Run it and play the output — a ten-second flight to the past:
output.txt
dialup.wav written - 10.0 seconds of nostalgia
Now listen closely. Each chunk is one sentence: the flat first tone says "the line is free"; the short pairs say "eight, zero, eight..."; the sustained tone says "I am a machine"; the wild noise says "let's measure this line"; and the final hiss says "we are trading code now." The screech we heard for years as meaningless noise has subtitles now, scene by scene.

# Let's reverse it: pull the phone number back out of the sound

If our claim is right — that the sound was data, not noise — then it must be possible to read the data back out of the audio. In that same WAV file, we'll now recover the number we dialed in the code, purely by listening to the sound. The tool is the Goertzel algorithm: a tiny radar that asks each suspicious frequency, "how much of you is inside this chunk of audio?":
dialup_decode.py
# dialup_decode.py - the sound was data: read the number back
import wave, math, struct
 
SR, N = 8000, 400                  # 50 ms analysis windows
ROWS, COLS = (697,770,852,941), (1209,1336,1477)
KEY = {(697,1209):'1', (697,1336):'2', (697,1477):'3',
       (770,1209):'4', (770,1336):'5', (770,1477):'6',
       (852,1209):'7', (852,1336):'8', (852,1477):'9',
       (941,1209):'*', (941,1336):'0', (941,1477):'#'}
 
def goertzel(win, f):
    k = 2 * math.cos(2*math.pi*f/SR)
    s1 = s2 = 0.0
    for v in win:
        s0 = v + k*s1 - s2
        s2, s1 = s1, s0
    return s1*s1 + s2*s2 - k*s1*s2
 
with wave.open('dialup.wav') as w:
    raw = w.readframes(w.getnframes())
x = [v/32768 for v in struct.unpack('<%dh' % (len(raw)//2), raw)]
 
digits, run, current = [], 0, None
for i in range(0, len(x)-N, 80):   # slide by 10 ms
    win = x[i:i+N]
    e = sum(v*v for v in win)      # total energy of this window
    p = {f: goertzel(win, f) for f in ROWS+COLS}
    r = max(ROWS, key=lambda f: p[f])
    c = max(COLS, key=lambda f: p[f])
    # two pure tones dominating the spectrum = a key is held
    key = (r, c) if p[r]+p[c] > 0.1*N*e else None
    if key == current: run += 1
    else: current, run = key, 1
    if key and run == 2:           # stable across windows = real press
        digits.append(KEY[key])
 
print("dialed:", ' '.join(digits))
output.txt
dialed: 8 0 8 0 8 0
That's it. Out of ten seconds of audio, the program extracted — by listening alone — the same number we dialed in the code. The ISP's modem did exactly this, only with sharper filters, within milliseconds, for every digit of every customer. The screech was data from the very beginning; we just didn't speak its language.
ⓘ
Why Goertzel and not FFT? An FFT gives you the whole spectrum at once; but when — like a DTMF receiver — only seven specific frequencies matter, Goertzel computes just those, separately and far more cheaply. That's why real telephone chips used exactly this algorithm for DTMF detection.

# Lab: dial the number and hear the phases with subtitles

Your turn now. Press the button and this conversation will be synthesized and played right here in your browser — this time with live subtitles for every second. Once the connection is up, the "pick up the phone" button unlocks too; press it and see why that famous family fight was actually physics:
dialup_lab.exe — interactive
The audio is synthesized right here with Web Audio — the same phases, the same frequencies.
Pick up the phone and press the button…
ⓘ
Why did picking up the phone kill the connection? A handset adds a parallel circuit to the line: the impedance drops, the signal level and noise change, and the waveform the two modems agreed on no longer matches what arrives. The modem doesn't know "someone picked up the phone" — it just sees data that no longer makes sense. And like any relationship whose trust breaks down, it hangs up.

# And then, nobody screamed anymore

Dial-up ended when data no longer needed to disguise itself as voice: ADSL kept the same phone wire but pushed data through the frequencies above the voice band — territory the phone system simply wasn't listening to. The result: an always-on line, and a home phone that was always free.
But the ritual never ended; it just went silent. Modems still do this today — Wi-Fi, 4G, 5G, your home router — before any serious transfer they probe each other and agree on speed and a shared language. The only difference is that the negotiation moved into frequencies and channels neither your ear nor the phone next door can hear. Back then we heard the sound and missed the meaning; now the meaning is all around us every day, and the sound is gone.
takeaway.txt
The phone was built for humans to talk —
modems learned to scream.

# Related Posts