The Higgins Bet, Without The Snobbery
Imagine standing in Covent Garden on a wet London night.
A man with a notebook hears a flower seller speak and starts placing people on the map by sound alone. That man is Professor Henry Higgins, George Bernard Shaw's phonetics show-off from Pygmalion, later reborn in My Fair Lady. He boasts that he can place a stranger within six miles by speech, and in London sometimes within two streets.
Absurd? A bit.
But the ear does do this. Not perfectly. Not fairly. Not always kindly.
Still, every accent carries clues. A vowel moves. The r disappears. A t turns into a little catch in the throat. A sentence rises when another speaker would let it fall.
Accent Guessr is built around that old Higgins trick, minus the class worship. We hide the face, play a short clip of real English, and ask you to put the voice on the map.

What Are We Actually Hearing?
When we hear an accent, we are hearing a stack of small speech habits.
Vowels do the loudest work. The vowel in bath, coffee, about, price, or fish can place a speaker faster than the actual words they choose. Consonants help too: some accents keep the r in car, some drop it, some tap it, some roll it.
Then comes prosody — the music of speech. Prosody is stress, rhythm, pitch, and timing. It is why one voice feels clipped, another feels sing-song, and another feels as if every syllable gets equal weight.
Modern accent-recognition systems listen for the same families of clues. They track speech sounds, pitch, duration, energy, and the shape of sound in tiny slices. Not magic. Just a very patient Higgins with a lot more math.
Wait — is that fair? Not if we pretend an accent is a passport. People move, code-switch, learn English later, copy friends, act on camera, and soften or sharpen their speech depending on the room. A game map is a clue, not a verdict.

So What Makes One Voice Feel From Somewhere?
The short answer: patterns. The longer answer is more fun.
- Vowels: where the tongue sits, how open the mouth is, and whether one vowel slides into two.
- Consonants: whether r, t, h, th, l, and final sounds stay sharp, soften, or vanish.
- Melody: whether English moves in stressed beats, gives each syllable equal space, climbs, falls, or bounces.
- First-language transfer: when another language leaves fingerprints on English sounds.
English is not one accent with a few local stains. English is spoken natively, as a second language, and as a daily working language all over the world. The game treats that as the point.
We are not talking about correcting voices. We are talking about hearing the history inside them.

Where Do Britain And Ireland Get Loud?
RP / Posh English. Listen for non-rhotic speech: the r fades after vowels, the vowels stay clean, and the consonants arrive clipped. It sounds "educated" only because schools and broadcasters treated it that way for about a century. The accent did not descend from heaven in a waistcoat.
London / Cockney / Estuary. London often gives itself away in the middle of words. Bottle can pick up a glottal catch, th can turn toward f or v, and h may drop. Estuary English softens the edges; old Cockney pushes them forward.
Liverpool / Scouse. Scouse has a bright nasal edge and a rising tune. The k sound can scrape into a fricative, and the whole accent carries Liverpool's port history: Irish, Welsh, and sea traffic in one voice.
Birmingham / Brummie. Brummie often lands with a low, falling melody, flatter vowels, and a slight nasal tint. It is easy to mock if you are lazy. Harder to hear well.
Manchester / Mancunian. Manchester speech tends to keep vowels flat and endings drawn out. The u sound can feel short and blunt, and the rhythm can carry a little swagger. Oasis did not invent it. They just exported it.
Yorkshire. Yorkshire often sounds clipped, dry, and direct. Short flat vowels do a lot of the work, and the old dropped definite article — the famous t' — gives the rhythm its hard little hinge.
Geordie. In Newcastle and Tyneside, house can move toward hoose, and about can move toward aboot. That is not a cartoon. It preserves older northern vowel patterns that faded from many southern accents.
Edinburgh / East-coast Scottish. Edinburgh can sound clearer and more clipped than Glasgow, with a Scottish r and a tighter vowel set. It often gives you Scotland without the full punch of the west coast.
Glasgow / Glaswegian. Glasgow speech can be fast, dense, glottal, and strongly rhotic. The r may tap or roll, the vowels can feel broad, and the pace can make outsiders sweat.
Welsh English. Welsh English often sings. You hear it in the rise and fall, the clear vowels, and the tapped r. English is doing the words; Welsh rhythm is still in the room.
West Country / Bristol. The West Country keeps the r in car and hard, a feature many English accents lost. Add rounded vowels and a soft burr, and you hear one reason early American English did not all drop its r.
Southern Irish English. Southern Irish speech is usually rhotic and melodic, with vowels and t/d sounds that can make a sentence feel lifted. Dublin, Cork, Kerry, and Kildare do not sound the same, of course. No country is that tidy.
Northern Irish / Ulster English. Ulster speech often has flatter vowels, a firm r, and a rising sentence shape. It can sound Irish and Scottish at once, which is not a bug. It is history speaking through the mouth.
What Does North America Do To The Same Words?
US Generic / Midwest. This is the accent many people mistake for "no accent" because TV trained them to hear it that way. It is rhotic, clear, and low on obvious local markers. Invisible ink is still ink.
US Upper Midwest / Minnesota. The long o gets rounder, the voice can turn nasal, and the melody often feels Scandinavian-influenced. Boat and home carry more geography than they seem to.
Deep South. The Southern drawl stretches vowels until one sound becomes a small journey. Pin and pen can merge, the tempo may slow, and y'all is not decoration. It is a useful plural pronoun doing honest work.
Texas. Texas mixes Southern drawl with a sharper twang. Diphthongs stretch, vowels open, and phrases like fixing to carry local grammar as much as local flavor.
New York City. Classic New York speech is quick, dense, and vowel-rich. The r may drop after vowels, coffee can move toward cawfee, and the rhythm has elbows.
Boston / New England. Boston is famous for dropping the r after vowels, but the real clue is the whole vowel system around it. Park the car became the joke because the joke works.
LA / Valley Girl / California. California gives you fronted vowels, uptalk, vocal fry, and the social glue of like. The clue is not one sound. It is a posture, a rhythm, and a way of holding the end of a sentence in the air.
Canadian English. Canadian speech is usually rhotic and can sit close to General American. The clue is often Canadian raising: vowels in words like price and about shift before voiceless sounds. Subtle. That is why it is hard.
Jamaican / Caribbean English. Jamaican and Caribbean English often carry the rhythm of Creole speech, with different h and th patterns, pitch movement, and final consonant treatment. The hard part is code-switching: one speaker may move between broad Creole and international English in a minute.
Why Do Australia, New Zealand, And South Africa Fool People?
Australian English. Australian English is usually non-rhotic, with broad vowels, rising sentence ends, and a compressed mouth shape. In broader speech, the price vowel can drift toward something like proice. Tiny move. Huge signal.
New Zealand English. New Zealand sits close to Australia until the short i gives it away. The fish-and-chips joke exists because the KIT vowel centralizes, so fish can move toward fush. The joke is old. The clue still works.
South African English. South African English often sounds clipped and tense, with distinctive short vowels and a tapped or lightly rolled r in some speakers. Afrikaans, local English history, and many African languages all leave traces, depending on the speaker.
What Happens When English Grows Through Another First Language?
French-influenced English. French often leaves a uvular r, a softened or missing h, th shifting toward z or s, and more even stress. The mouth shape changes first. The accent follows.
German-influenced English. German often makes English consonants crisp. W can move toward v, th toward s or z, and final voiced sounds can devoice. The result is clean, firm, and easy to spot when the clip is good.
European Spanish-influenced English. Spanish brings pure vowels, a tapped or rolled r, b/v pressure, and sometimes an extra e before s-clusters. Spain can want to become espain. The ear notices.
Italian-influenced English. Italian gives English open vowels, musical timing, and stronger double consonants. Some speakers add a faint vowel after a final consonant, as if the word wants one more step before it stops.
Russian-influenced English. Russian often brings a dark l, a rolled r, final devoicing, and fewer articles because Russian does not use them the same way. The accent can sound heavy not because it is slow, but because consonants carry more weight.
Polish-influenced English. Polish can bring hissing sibilants, palatalized consonants, w moving toward v, and a pull toward penultimate stress. The clue is often the edge of the consonants rather than the vowels.
Scandinavian-influenced English. Scandinavian English is often highly fluent, so the accent hides in the melody. Clear vowels, a sing-song pitch shape, and clean consonants give it away.
Arabic-influenced English. Arabic can bring emphatic consonants, p and b pressure, throatier sounds, and vowel shifts. Many Arabic varieties do not use p as a native sound, so a tiny p-to-b move can become a large clue.
Chinese-influenced English. Mandarin and Cantonese do not create one accent, but both can shape English through tone transfer, clipped final consonants, l/r pressure, and th shifts. The rhythm may feel more syllable by syllable.
Japanese-influenced English. Japanese works in morae — small timing beats — so English clusters may pick up extra vowels. L and r move toward the Japanese flap, and final consonants often want a vowel to land on.
Korean-influenced English. Korean can add vowels around hard consonant clusters, soften or shift final sounds, and put pressure on f/p and l/r. The rhythm often comes in clear syllable-sized blocks.
Indian English. Indian English often has retroflex t and d sounds, where the tongue curls back, plus syllable-timed rhythm and w/v closeness. This is not broken English. It is one of the largest living Englishes on earth.
Singaporean / Singlish. Singaporean English can sound clipped, fast, and tone-shaped, with final consonants reduced and particles like lah, leh, and lor doing real social work. Switching between polished Standard Singapore English and Singlish is a skill, not a flaw.
Nigerian English. Nigerian English often has syllable-timed rhythm, crisp consonants, and a lively pitch shape. Nigeria is far too large for one voice, so the game listens for broad West African English clues, not a single national mask.
Kenyan / East African English. Kenyan English is often measured and clear, with Swahili-influenced rhythm and vowel shape. The clue is steadiness: each syllable gets room to stand up.
How Does The Game Use This Without Being Weird About It?
Accent Guessr is a game, not a citizenship office.
The speaker is hidden at first, so you cannot rely on face, fame, clothing, or background. You get the voice. Then you drop a pin. After the guess, the game shows who was speaking, where the accent belongs for this round, and how close your ear came.
The map labels are broad on purpose. RP overlaps with the South East. Australia and New Zealand can fool even good listeners. People move. Voices mix. A clip can be clear without being a perfect museum specimen.
So the rule is simple: respect the speaker, test the ear, enjoy the reveal.
What Do The Hints Do?
Sometimes the ear needs a nudge, not an answer. The hints are there to keep you listening when two places are fighting in your head.
50 / 50. This narrows the map to the correct region and one believable decoy. You still have to choose, but the game stops asking you to scan the whole world at once.
Unblur. This reveals the speaker before you guess. A face can give extra context, but it can also mislead, which is why the voice comes first.
Each hint is available once per game. Use them as training wheels for the ear, not as proof that a voice belongs in only one box.
What Did We Read Before Writing This?
- George Bernard Shaw's Pygmalion, Act I, for Professor Higgins and his phonetics boast.
- British Council and Manchester Metropolitan University research on UK accent history, RP, Geordie, Scouse, migration, and accent prejudice.
- The International Dialects of English Archive and the Speech Accent Archive for real-world English samples and accent comparison.
- Research on automatic accent identification, especially Najafian and Russell, 2020 and Yang et al., 2023, for phonemes, prosody, speaker embeddings, and speech-recognition error.
The in-game labels and playable accent set come from the Accent Guessr game description, accent IDs, and local accent manifests in this repo.
What Should The Ear Learn Next?
In 10 years, speech software will probably hear accents better than most of us. The better goal is not to erase them. It is to hear more of them, with less snobbery than Higgins and more curiosity than the room he walked into.