For tonight: about three, done properly. For the test: the same problems you were going to do anyway, spread across more days. That is not a motivational opinion. It is what happened when researchers actually counted.
Students ask "how many problems should I do?" expecting a number like 20 or 50, and the honest answer from the research is that the count is the least important variable in the question. What matters is when you do them, how they are mixed, and what you do with the ones you get wrong. Here is the evidence, and then the rule I give my own students.
The study that measured "more problems" directly
Doug Rohrer and Kelli Taylor ran the cleanest version of this experiment I know of, published in Applied Cognitive Psychology. Two hundred sixteen college students learned to solve a permutation problem, the kind you would practice on our permutation and combination calculator, and then followed different practice schedules.
Experiment 1: ten practice problems, either all in one sitting or split across two sessions a week apart. Same problems, same total effort. On a test four weeks later, the spread-out group's performance was roughly double the massed group's. And here is the detail that fools students: on a test just one week out, the two groups looked the same. Cramming does not feel costly, because its cost only shows up after the quiz you crammed for.
Experiment 2 attacked the count itself: three practice problems in a session versus nine. The extra six problems, the strategy researchers call overlearning, had no measurable effect on test scores one week later or four weeks later. Read that again with tonight's homework in mind. Problems four through nine, done in the same sitting after the skill was working, purchased approximately nothing.
The authors' conclusion was blunt: long-term retention was boosted by distributing practice and unaffected by overlearning. Twenty problems tonight is not twice as good as ten. Ten tonight and ten on Thursday is close to twice as good as twenty tonight.
The stopping rule, since "it depends" is useless at 9 pm
Here is the operating version I give students, which respects the evidence without requiring a stopwatch and a spreadsheet.
- Tonight, for a new skill: work problems until you have solved two in a row correctly, without notes and without peeking at an example. For most students on most skills that is somewhere between three and six problems. Then stop. You are done improving tonight; further identical reps are the overlearning that Experiment 2 priced at zero.
- Then schedule, don't add. The reps you were tempted to do tonight go on the calendar instead: a short set in two days, another early next week. Same total problem count as the marathon session, radically different retention. If a test is coming, this slots directly into the structure of a three-day study plan.
- Before a test, the target is total sessions, not total problems. Three short sessions of four problems beat one heroic session of twenty, at identical cost. The calendar is doing the memorizing, not the volume.
Notice what the stopping rule is measuring: performance, not effort. "Two in a row, unaided" is a completion condition you can verify. "I did a lot of problems" is a feeling.
What actually counts as one problem
The stopping rule only works if a "problem" means what the research means by it: a closed-book attempt, from the printed question to a final answer, followed by finding out whether you were right. Several popular activities feel like practice and are not:
- Rereading a worked example is not a rep. Recognition and production are different skills. Every worked example you follow feels clear, which is exactly why it measures nothing.
- Watching someone solve is not a rep. Not the teacher, not a video. It is useful before your attempts, worthless as a substitute for them.
- A problem you abandoned halfway and then read the solution to is worth something, but it does not count toward your two-in-a-row. Only completed, unaided, correct solves close the session.
- Checking counts as part of the rep. An unchecked problem is a coin flip you chose not to look at. Verify the answer before you count it, whether against the back of the book or by substituting it into the original.
What does a good mixed set look like in practice? Suppose the last three weeks covered factoring, radical equations, and systems. Tonight's six problems are two of each, shuffled so no two of a kind sit together, and none labeled. That set rehearses the exact skill the blocked textbook page can't: recognizing the problem type before executing it.
Make the reps you do keep count double
Two upgrades multiply the value of a fixed number of problems, both from the same research group.
Mix the problem types. In a follow-up experiment, students who practiced problem sets where types were shuffled together scored 63 percent a week later, against 20 percent for students who practiced the same problems blocked by type, even though the blocked group looked far better during practice itself. Blocked practice answers "how do I execute this method"; mixed practice answers the question tests actually ask, which is "which method does this problem want". Most textbook sections hand you pre-blocked sets, so mixing takes deliberate effort: pull two problems each from the last three sections instead of six from tonight's.
Keep testing what you already got right. A Karpicke and Roediger experiment on retrieval found that dropping mastered items from further testing roughly halved final recall, while dropping them from restudy cost nothing. Translation for a problem set: it is fine to stop rereading examples of the types you have down, but those types still belong in your mixed practice sets, one rep each, right up to the test.
Wrong answers are practice problems that pay double
The count conversation always assumes fresh problems, and that assumption wastes the best material you have. A problem you got wrong last week, redone cold from the printed question, is worth more than a new one, because it targets a demonstrated gap instead of a random spot. I walk through the full protocol in the 20-minute redo, but for counting purposes the rule is simple: your redo problems count toward tonight's total, and they go first.
One honesty note before the summary, because this blog does not oversell research. The overlearning experiments taught a single problem type to college students under lab conditions, and no study hands you your personal magic number. What the broader spacing literature adds is that the pattern is not fragile: the advantage of spreading practice over massing it is one of the most replicated results in learning research, across ages and materials. The two-in-a-row stopping rule is my translation of that evidence into something usable on a school night, not a figure copied from a table.
So the next time the assignment says "do the odd problems, 1 through 39", you have research-grade permission to renegotiate with yourself. Do enough tonight to hit two-in-a-row unaided. Bank the rest across the week in mixed sets, keep one rep of the mastered types in rotation, lead with redos. If you need fresh problems at a specific difficulty to fill those sets, MathSolver's practice quizzes generate them by topic and level, free and without an account. The students who improve are not the ones doing the most problems in a night. They are the ones whose problems are spread across the most nights.
