Minimal Pairs: 7 Proven ESL Drills
A student who has studied English for eight years can still say “I lost my key” and have it come out as “I los my key,” because Mandarin syllables rarely end in the consonant sounds English stacks onto the end of a word. Researchers who’ve studied how Mandarin speakers produce English describe this as a straightforward transfer problem rather than a lack of effort — Mandarin codas are mostly limited to vowels, /n/, and /ŋ/, so a sound like the final /t/ in “lost” has no equivalent slot to map onto.1 The fix isn’t more grammar drilling. It’s short, targeted listening and production practice built around the specific sound a student is missing, and that’s exactly what minimal pairs are built for.
What Minimal Pairs Actually Fix for Mandarin-Speaking Students
Before picking activities, it helps to know which minimal pairs are worth class time and which ones are a waste of it. Not every English sound contrast trips up every learner, and drilling a distinction a student already hears fine just burns a lesson slot that could go toward a real gap.
Four categories cover most of what shows up in a Taiwanese classroom. Final consonants are the biggest one — pairs like cap/cat/cab or bus/bug/bud force students to hold onto a sound they’d normally drop. The v/w merge comes next, since Mandarin has no /v/ at all, so vine and wine, or vest and west, collapse into the same word in a student’s mouth. Short-versus-long vowel pairs — ship/sheep, bit/beat, live/leave — matter because Mandarin vowel length doesn’t carry meaning the way English vowel quality does. And the “th” sounds, both voiced and voiceless, get replaced with /s/, /z/, or /d/, so think becomes sink and this becomes dis. The ASHA reference chart on Mandarin’s phonemic inventory is a useful gut check if you’re unsure whether a sound exists in a student’s first language before you plan a drill around it.2

Discrimination Listening Drills: Train the Ear First
Before a student can say the difference between “light” and “right,” they need to hear it reliably, and that’s a separate skill from production. A discrimination drill strips speaking out of the equation entirely: read a word aloud and have students point to, circle, or write which column it belongs to on a two-column handout.
The variation that does the most work is the odd-one-out version. Say three words — two identical and one from the contrasting pair, like cat, cat, cot — and ask which one didn’t match. This catches students who’ve started guessing based on context rather than actually hearing the sound, which happens more than teachers expect once a class gets comfortable with a word list. Run it cold, with no visual support, for the first pass, then let students check answers in pairs before you reveal the key. Five minutes of this at the start of class, run consistently for two or three weeks on the same sound pair, moves the needle more than a single 20-minute lesson ever will.

Minimal Pairs Bingo
Bingo turns a discrimination drill into something students will ask to play again, which matters more than it sounds like it should for a repetitive listening task. Build a 4×4 or 5×5 grid using words from two or three minimal pair sets — mixing ship/sheep, bit/beat, and live/leave on the same card works well because it forces students to stay locked onto vowel length rather than picking up on a single sound they’ve memorized.
Call words one at a time, at a normal speaking pace, and don’t repeat yourself — that repetition is exactly the crutch you’re trying to remove. Students who are still relying on lip-reading or guessing will lose fast, and that feedback loop corrects itself within a round or two without you having to say anything. For classes of eight or more, keep several different card layouts in rotation so students can’t win by memorizing a friend’s card instead of listening.

The Telephone Number Encoding Activity
This one is closer to real communication than most pronunciation drills get, which is why it tends to stick. Assign each digit, 0 through 9, a minimal pair word — for example 1 = “bit,” 2 = “beat,” 3 = “van,” 4 = “fan.” Students then “encode” a phone number, address, or ID number using the words instead of digits and read it to a partner, who writes down the number they hear and converts it back.
The activity fails loudly when it fails, which is the point. If a partner writes down the wrong number, both students immediately know a sound got dropped or swapped somewhere in the exchange, and they have to figure out which pair caused it. That built-in accountability does more for motivation than a worksheet ever will, because the stakes feel real even though the numbers are made up. Run it with actual classroom-relevant numbers — a made-up student ID, a locker combination — and the activity holds attention even in classes that usually check out during drills.

Mouth Position and Mirror Demonstrations
Some minimal pairs are listening problems and some are physical ones — a student can hear the difference between “van” and “fan” perfectly well and still can’t produce /v/ because Mandarin never asks the lower lip and upper teeth to vibrate together that way. For those pairs, showing the mouth matters more than repeating the word.
Hand out small mirrors, or have students use their phone cameras, and walk through the physical setup side by side: teeth on lip and voice on for /v/, teeth on lip and voice off for /f/. The University of Iowa’s Sounds of Speech tool is worth projecting on a screen for this — it animates the tongue, lips, and vocal folds for every English consonant and vowel, which saves you from trying to describe airflow with your hands.3 Students who’ve been told for years to “listen and repeat” often light up the first time someone actually shows them what their mouth is supposed to be doing.

Minimal Pair Sentences and Tongue Twisters
Single words in isolation are the easiest version of a minimal pair drill, and also the least useful for anything beyond the first week. Once students can reliably tell two words apart, move the contrast into a full sentence so they have to hold the distinction while also managing grammar and meaning at the same time — which is what actually happens in conversation.
Simple sentence frames work best: “I bought a new van/fan yesterday” or “Can you fill/feel this for me?” forces the listener to use context alongside the sound itself, which is a more realistic test than an isolated word ever is. For a lighter version, short tongue twisters built around one contrast — “Ted said Ed shed his red thread” for the /ð/ versus /d/ split — give stronger students something to compete over while everyone else keeps practicing the underlying sound. Pair this section with a look at sentence stress and rhythm once individual sounds are solid, since stress placement is the next layer students need for natural-sounding speech.
Running Dictation With Minimal Pairs
Running dictation gets students moving and listening at the same time, and it works especially well once a class has a decent bank of minimal pair sentences to draw from. Post five or six sentences around the room, each built around a different minimal pair. In pairs, one student runs to read a sentence and memorizes it, then runs back and dictates it to their partner, who writes it down exactly as heard.
The mistakes that show up in the written version are a direct readout of which sounds still aren’t landing — if “fill” keeps turning into “feel” on the page, that pair needs another round before you move on. Because the activity is timed and physical, students tend to forget they’re being tested on pronunciation at all, which lowers the anxiety that usually comes with speaking drills. It also doubles as a listening check for the discrimination work from earlier in the week, so you get two data points on the same sound pair without running two separate activities.

Record-and-Compare Self-Correction
The last step is the one most classrooms skip, and it’s the one that actually changes long-term habits: have students record themselves saying a set of minimal pair words or sentences on their phone, then play it back next to a model recording or your own voice.
Most students genuinely cannot hear their own pronunciation errors in real time while they’re speaking — the cognitive load of producing the sentence crowds out the ability to monitor it. A recording removes that pressure entirely. Have students mark which words sound closest to the model and which ones don’t, then re-record just the words that missed. Ten minutes of this every couple of weeks, saved and compared over a semester, gives students a concrete before-and-after they can hear for themselves — which tends to build more motivation than any grade you could put on a pronunciation quiz.

None of these seven drills need more than a printed word list and, at most, a few minutes of screen time. What they need is repetition on the same handful of sound contrasts across several weeks, not a different pair every day. Pick the two or three contrasts that show up most often in your own students’ speech, run the discrimination and bingo drills until listening is solid, then move to sentences, dictation, and recording once production catches up. If you’re building this into a broader speaking curriculum, it pairs naturally with the speaking fluency activities a phonics fundamentals covered elsewhere on this site.
Zdroje
- Production of English Syllable-Final /l/ by Mandarin Chinese Speakers, Journal of Language Teaching and Research — research on how Mandarin’s syllable structure affects English final-consonant production.
- Mandarin Phonemic Inventory, American Speech-Language-Hearing Association — reference chart of Mandarin consonant and vowel sounds compared to English.
- Sounds of Speech, University of Iowa — interactive animations showing tongue, lip, and vocal fold position for every English consonant and vowel.



