How to Write ESL Tests and Curriculum With AI: A Prompting Guide for Teachers
Ask an AI tool to “write me an ESL test” and you will get exactly what that request deserves: a generic worksheet, mis-leveled questions, and answer keys that quietly contradict themselves. Ask it the right way, with the right constraints, and the same model will hand you a targeted, level-appropriate assessment in under a minute. The difference is never the tool. It is the prompt. This guide is about the part of AI-assisted assessment that most teachers skip: learning to write the instructions that turn a general-purpose model into a reliable co-author for your tests and curriculum.
If you have already read about backward design, CEFR alignment, or distractor theory, treat this as the operational layer underneath all of them. Those articles tell you what good assessment looks like. This one tells you how to communicate that to a machine so it actually produces it.

Why “Write Me a Test” Fails Every Time
A language model predicts likely text. When you give it a thin request, it fills the gaps with the statistical average of everything it has seen, which is a bland B1-ish worksheet aimed at no one in particular. Your real class is not average. You have a specific level, a specific syllabus point you taught last week, a specific exam your students are training for, and a specific number of minutes in your period. Every one of those details is a constraint the model needs but cannot guess.
The mental shift that fixes most bad output is simple: stop thinking of the prompt as a question and start thinking of it as a specification. A good test-writing prompt reads less like “can you help me” and more like a work order you would hand a junior colleague who has never met your students.
The Five Ingredients Every Test Prompt Needs
Almost every weak result traces back to a missing ingredient. Before you press enter, check that your prompt names all five.
1. The learner profile
State the level in CEFR terms (A2, B1, B2), the age band, the first language if it matters for false friends, and the context. “Adult B1 learners in a Taiwan business-English evening class” produces different vocabulary and topics than “twelve-year-old A2 learners in a public school.” The model cannot see your room, so describe it.
2. The target construct
Name exactly what you are measuring. “Past simple versus present perfect,” “skimming for gist,” or “TOEIC Part 5 sentence completion” all point the model at a narrow, testable skill. Vague constructs like “grammar” or “reading” invite the model to test everything and measure nothing.
3. The item format and count
Specify the question types (multiple choice, gap-fill, matching, short written response), how many of each, and how many options per multiple-choice item. “Ten four-option multiple-choice items” leaves nothing to chance. If you skip this, you will get a random assortment you then have to reshape by hand.

4. The constraints and guardrails
This is where teachers add real value. Tell the model what to avoid: no vocabulary above the target level, no culturally narrow references, exactly one unambiguously correct answer per item, and distractors that are plausible but clearly wrong to a student who knows the point. Adding “every distractor must be grammatically parallel to the key” alone will lift the quality of a multiple-choice set dramatically.
5. The output shape
Tell it how you want the result laid out: student-facing questions first, then a separate answer key, then a one-line rationale for each correct answer. That rationale line is your fastest quality check, because it forces the model to justify each key and exposes the items where its own logic wobbles.
A Prompt You Can Copy and Adapt
Here is a complete specification-style prompt that folds all five ingredients together. Change the bracketed parts and reuse it as a template.
You are an experienced ESL assessment writer. Write a short quiz for B1 adult learners in Taiwan whose first language is Mandarin. Target construct: past simple versus present perfect. Format: 8 four-option multiple-choice items. Constraints: keep all vocabulary at B1 or below; exactly one correct answer per item; make all four options grammatically parallel; distractors should reflect the most common Mandarin-speaker errors with these tenses. Output: numbered student questions first, then an answer key, then a one-sentence rationale for each key explaining why the distractors are wrong.
Notice there is no cleverness here, only precision. The prompt tells the model who, what, how many, what to avoid, and how to present it. That single message will out-perform ten rounds of “make it harder” and “try again.”

Feeding the Model Your Own Source Material
The most reliable way to keep an AI test on-level and on-syllabus is to stop asking it to invent content and start asking it to work from yours. Paste in the reading passage you already use, the vocabulary list from this unit, or the transcript of the dialogue from your coursebook, and instruct the model to build items strictly from that text.
A prompt like “Using only the passage below, write five comprehension questions that cannot be answered without reading it, plus two inference questions” anchors every item to material you have already vetted. This single move eliminates most off-level vocabulary and most factual hallucination in one step, because the model is now editing a source rather than dreaming one up.
The same logic scales to curriculum. Paste your existing scope-and-sequence, then ask the model to draft a unit that fills a named gap: “Here is my ten-week B2 syllabus. Weeks 1 to 6 cover the tenses listed. Draft week 7 as a review-and-consolidation unit with objectives, three activities, and an exit assessment that recycles weeks 1 to 6.” You get continuity instead of a disconnected one-off.
Building Curriculum, Not Just Isolated Tests
Curriculum work rewards a different prompting rhythm than single tests. Rather than generating a whole course in one shot, build it in layers and keep the model’s earlier output in the conversation so each new layer stays consistent.
Start by asking for measurable learning outcomes in “can-do” form tied to a CEFR band. Once you approve those, ask for a unit map that sequences them. Then, unit by unit, ask for objectives, activities, and an assessment that measures the very outcomes you locked in at step one. Because the outcomes came first, the assessments the model writes at the end are automatically aligned to them. This is backward design executed as a conversation, and the model is far better at holding the thread when you build it this way than when you demand the finished course in a single prompt.

Iterating: The Second and Third Prompt Matter Most
Treat the first output as a draft, never a finished product. The highest-leverage prompting skill is knowing what to ask for in the follow-up. A few reliable moves:
- Stress-test the key. “For each item, argue that a second option could also be defended as correct. If you can, rewrite that item.” This catches the ambiguous items before your students do.
- Rebalance the difficulty. “Items 3 and 7 are noticeably easier than the rest. Rewrite them to match the difficulty of item 5.”
- Tighten the level. “Flag any word above B1 and replace it with a simpler equivalent.”
- Add a marking rubric. For open-response and writing tasks, ask for a band-based rubric with descriptors, then a sample answer at each band so your grading stays consistent.
Each of these follow-ups does work you would otherwise do manually, and it does it in the model’s own voice, so the revised items stay stylistically consistent with the originals.
The Human Checks You Cannot Delegate
No prompt, however careful, removes your responsibility for the final paper. Language models still produce answer keys that are wrong, comprehension questions that can be answered without the passage, and “B1” items stuffed with C1 vocabulary. Before anything reaches a student, read every item cold as if you were the test-taker, verify each key yourself, and confirm the vocabulary genuinely sits at your target level.

Pay special attention to fairness and cultural load. A model trained largely on English-language internet text will reach for references your international learners may not share. Scan for idioms, brand names, holidays, and assumptions that could disadvantage a student who knows the grammar perfectly well. This judgment is exactly what AI cannot supply, and exactly what makes you the assessment writer rather than the machine.
Data privacy deserves the same care. Avoid pasting identifiable student information, real grades, or anything sensitive into a public tool, and follow your school’s policy on where AI-assisted materials may be stored.
A Simple Workflow to Adopt This Week
Put the ideas together and the process is short. Write a specification-style prompt with all five ingredients. Feed the model your own passage or syllabus wherever possible. Ask for output split into questions, key, and rationales. Iterate with the stress-test and level-tightening follow-ups. Then do the human pass for accuracy, level, and fairness before it goes anywhere near a class.

Save your best prompts as templates. The B1 tense quiz you refined today becomes the A2 or B2 version tomorrow with three bracketed edits. Over a term, a small library of proven prompts turns a task that once ate your Sunday evening into a fifteen-minute draft-and-review. The AI is fast; your expertise is what makes it correct. Keep both in their proper roles and the tests will hold up.
Sources
- Council of Europe — Common European Framework of Reference for Languages (CEFR)
- British Council
- Cambridge English
- ETS (TOEIC and TOEFL)



