Introduction
Study & Recall
Better recall starts with better prompts.
“Turn your notes into questions” is decent advice until you try to do it.
You look at a page about DNS and write: “What is DNS?”
Technically, that is a question. It may even help you recall a definition. But it does not tell you whether you can trace a lookup, compare two server roles or diagnose why a name resolves on one network and not another.
Active recall is only as useful as the thing you ask your brain to produce. If every prompt tests a short definition, you can become excellent at short definitions while the real task still falls apart.
The aim is not to make every question difficult. It is to make the set representative.
Start with the performance you eventually need
Before writing a question, finish this sentence:
“After studying this, I need to be able to…”
The ending determines the prompt.
This is why copying every bold term from a textbook into a flashcard is not enough. The formatting tells you what the author emphasised. It does not automatically tell you what you must be able to do.
1. Core fact
Use this when a piece of information genuinely needs to be available.
Example: “Which DNS record type maps a hostname to an IPv4 address?”
Fact questions are useful, but a set made entirely of them can create isolated fragments. Follow an important fact with a question that makes you use it.
2. Why or how
Ask for the mechanism, reason or causal chain.
Example: “How can DNS caching reduce lookup time, and why can the same caching delay a record change?”
This prompt forces more than the phrase “DNS uses caching”. You have to connect the behaviour to two consequences.
3. Compare and contrast
Ask what changes, what stays the same and when the distinction matters.
Example: “How does a recursive resolver differ from an authoritative DNS server in the job it performs?”
Avoid questions such as “What is the difference between X and Y?” when the possible differences are endless. Name the dimension you care about: purpose, inputs, outputs, trade-offs or failure modes.
4. Apply to a new case
Give yourself a scenario that was not copied directly from the worked example.
Example: “A company moves a service to a new IP address, but some users still reach the old server. How could cached DNS data explain the split behaviour, and what would you check?”
Application questions help test whether the idea can travel beyond the sentence in which you first met it.
5. Reconstruct a process
Ask for the ordered steps and the decision points.
Example: “Starting with an empty local cache, trace the main steps from entering a domain name to receiving an IP address.”
Do not grade only the presence of keywords. Check the order, the role of each component and any branch that matters.
6. Diagnose an error
Present a wrong answer, broken output or flawed explanation and ask what is wrong.
Example: “A learner says the recursive resolver owns the definitive records for every domain it can resolve. What is wrong with that explanation?”
Diagnosis makes you distinguish a plausible statement from an accurate one. It is particularly useful for misconceptions you have already found in a Fail First attempt.
Too broad
Weak: “What is DNS?”
Stronger: “What problem does DNS solve for a client, and what does a successful lookup return?”
The stronger version defines the expected territory without printing the answer.
Too much of the answer in the question
Weak: “How does a recursive resolver contact other DNS servers to find the answer for the client?”
Stronger: “What does a recursive resolver do when it does not already have a usable cached answer?”
The weak version supplies most of the sequence as a cue.
Yes or no when production matters
Weak: “Does DNS caching improve speed?”
Stronger: “Explain one benefit and one risk of DNS caching.”
A yes-or-no answer can be recognised or guessed. An explanation has to be produced.
Vague self-assessment
Weak: “Do I understand TTL?”
Stronger: “What does a DNS record’s TTL control, and what practical effect can a long TTL have during a change?”
Confidence is useful evidence, but it is not a substitute for an answer you can inspect.
Trivia without a purpose
Weak: “What year was DNS invented?”
Stronger: “Which parts of the lookup path can cache a result, and how does that affect troubleshooting?”
A fact is not automatically valuable because it is easy to put on a card. Tie the question to the outcome you need.
Write the answer key at the same time
A prompt without a checking standard can become an argument with yourself.
For each question, write a compact answer key that contains:
The key does not need to be a polished essay. It needs to tell you whether your retrieval was accurate.
For the caching question, a useful key might be:
Now “I basically got it” has something concrete to meet.
Keep one main target per question
Short questions are easier to diagnose. If one prompt asks for five unrelated things and you miss two, the score tells you very little about where the gap is.
One main target per question is a good default:
Combine ideas deliberately when integration is the skill. The scenario about users reaching old and new servers intentionally joins caching, TTL and troubleshooting. It is useful because joining them is the task, not because longer questions are automatically better.
Avoid accidental cues
A retrieval prompt should tell you what to produce without smuggling the response into the wording.
Watch for:
You do not need to strip away all context. Real tasks contain context. Remove the clues that replace the thinking you meant to test.
Are multiple-choice questions bad for active recall?
No. They test a different form of retrieval and decision-making.
Well-written options can force you to distinguish close alternatives and diagnose misconceptions. That can be exactly what an exam or workplace decision requires. But options also provide cues, so a multiple-choice success does not prove that you could produce the explanation from scratch.
Use both where they match the goal. You might answer a scenario-based multiple-choice item, then explain why the chosen option fits and why the nearest distractor does not.
How many active recall questions should you make?
Enough to sample the important outcomes, not enough to create a second course about the first course.
Five to ten questions for a focused lesson is a practical starting point, not a universal number from a study. Adjust it to the size and variety of the material.
A compact set might include:
If the lesson has one narrow objective, use fewer. If you must perform several distinct tasks, use more. Delete duplicate questions that reveal the same gap.
How to mark answers without hiding the gap
Use labels that create a next action:
“Secure for now” matters. One successful answer is evidence from one retrieval, not a lifetime guarantee.
Turn the result into the next study action
Do not reread everything after every test.
For a missing answer, find the relevant explanation and rebuild the basic model.
For confusion, compare the concepts side by side and state the distinction in your own words.
For an application gap, study a worked example, then try a different case without looking.
For a guess, explain the reasoning and retrieve it again.
Then retest the weak question. If it becomes accurate, bring it back after a delay through spaced retrieval . That later attempt checks whether the answer remains available after the immediate familiarity has faded.
A 20-minute question-building session
If you want a simple workflow:
The timings are a practical constraint, not a scientific prescription. Their job is to stop question creation swallowing the study session.
You can also write the questions before studying and use them as a pretest . The first attempt shows the baseline; the same prompts after learning show what changed.
What the research does and does not say
Retrieval-practice research supports the broader act of trying to produce learned material rather than only studying it again. In experiments with prose passages, Roediger and Karpicke found a delayed-retention advantage for repeated testing over repeated study under the conditions they examined. Butler found that repeated testing could also support transfer to new inferential questions in his experiments.
That is good reason to include retrieval and application in a study routine. It does not prove that one universal prompt template, question count or marking system is best for every subject.
Use the evidence for the principle, then design the questions around the real performance.
A good question makes the gap visible
The test is not whether a prompt looks clever. It is whether the answer gives you useful evidence.
Can you recall the essential fact?
Can you explain why it works?
Can you distinguish it from the nearest alternative?
Can you apply it when the surface details change?
Can you reconstruct the process or spot the error?
Write questions that make you produce those things. Check them against a clear answer. Label the gap. Improve it. Then retrieve it again later.
That is active recall doing its actual job: showing you what your notes cannot show while they are open.
Research behind retrieval practice
These papers support the retrieval and transfer principles discussed above. They do not prescribe one universal question format.
• Roediger & Karpicke (2006), Test-enhanced learning: taking memory tests improves long-term retention
• Butler (2010), Repeated testing produces superior transfer of learning relative to repeated studying
• Karpicke & Blunt (2011), Retrieval practice produces more learning than elaborative studying with concept mapping
Build a question set that exposes the gaps
Start with the performance you need, answer without looking, mark the kind of gap and retest what you improve. Study & Recall turns that into a repeatable learning loop.