Assessment Should Improve Learning, Not Just Measure It

11 min read

Summary

A score describes a performance. Formative assessment earns its name only when evidence changes what the learner or teacher does next.

Assessment Should Improve Learning, Not Just Measure It

A score is a signal, not the finish line

Assessment is often treated as the end of learning: the lesson stops, the questions begin and a score enters a record. That sequence can certify performance. It can sort, select or report. But it does not automatically help anyone learn.

Formative assessment has a different purpose. It creates evidence about what a learner currently understands so that the learner, teacher or system can make a better next decision. The evidence becomes formative only when someone interprets it and acts. A quiz with no useful response is still a quiz. A dashboard no one can translate into instruction is still a dashboard. 1 2 3

This distinction matters wherever assessment becomes easy to automate. More questions, faster scores and richer analytics can increase the quantity of measurement while leaving the learning loop open. The central design question is not “How quickly can we mark this?” It is “What will change because we now know this?”

The assessment-to-action loop

A useful assessment loop has four stages.

  1. Elicit. Give the learner a task capable of revealing the intended knowledge or reasoning.
  2. Interpret. Use the response to locate a misconception, strategy, uncertainty or missing prerequisite.
  3. Respond. Provide a next step that is specific enough to use.
  4. Verify. Ask for revision, transfer or another attempt to see whether the response worked.
The assessment-to-action loop: Elicit, Interpret, Respond, Verify.
The assessment-to-action loop. Evidence becomes formative when it changes the next decision

Weakness at any stage breaks the loop. A multiple-choice item may be too ambiguous to reveal the target concept. A correct response may be a guess. A diagnostic label may be inferred from too little evidence. Feedback may arrive after the learner has mentally left the task. Or the course may provide no opportunity to use it.

The loop also makes clear why formative assessment is not a particular format. It can be a carefully chosen question, an annotated draft, a conversation, a worked example with a missing step, a concept sketch or a second attempt after feedback. What matters is the relationship between evidence and action.

Feedback helps on average, but sometimes harms

The broad feedback literature is encouraging but resistant to slogans. A meta-analysis covering 435 studies, 994 effects and roughly 61,000 learners estimated an average effect of d = 0.55. Yet 17% of effects were negative, and heterogeneity was very high. 4 “Feedback works” is therefore less useful than asking what the feedback says, when it arrives, who can use it and what happens next.

A 2023 meta-analysis focused on technology-rich learning environments synthesised 61 studies and 182 effects. It estimated an overall effect of g = 0.44, with explanation feedback the strongest coded type. 5 A 2024 synthesis of 46 articles and 116 digitally delivered feedback interventions reported a summary effect of 0.41 and found process-focused feedback contributed to performance in meta-regression. 7

These averages should not be casually compared or added. The studies used different outcomes, learners, designs and feedback categories. They do, however, support a common principle: information about a response is most useful when it helps the learner understand the task, the strategy or the next move, rather than merely report the learner's rank.

Feedback effects are meaningful, not automatic: Tech-rich settings, g = 0.44; Education overall, d = 0.55; Negative effects, 17%.
Feedback effects are meaningful, not automatic. Average effects conceal wide heterogeneity; content, timing, uptake and outcome alignment matter.

The content of feedback should match the intended learning. A 2024 writing meta-analysis drew on 95 papers, 200 treatment-control comparisons and 253 effects. It found that deep feedback directed at deep writing outcomes had an estimated effect of g = 0.80, while effects varied across learner groups and outcome types. 6 A grammar correction and an argument-level suggestion are not interchangeable; each should be judged against the change it is meant to support.

Diagnose before prescribing

Two learners can give the same wrong answer for different reasons. One has misunderstood the concept. Another knows it but misread the question. A third chose an unproductive strategy. A fourth made a transcription error. If the system sees only “incorrect,” it may prescribe the same explanation to everyone and call the result personalised.

Better diagnosis begins by designing tasks around interpretable evidence. Distractors can correspond to known misconceptions. An open response can request reasoning, not just a conclusion. A sequence of items can distinguish a missing prerequisite from a one-off error. Confidence can be collected as one fallible signal, never as proof.

Interpretation should remain proportionate to the evidence. One answer rarely supports a durable claim such as “this learner cannot reason algebraically.” A better statement is local: “this response is consistent with distributing the negative sign incorrectly; ask for one contrasting example.” That language protects the learner from a label and the educator from false precision.

Artificial intelligence can help classify responses, suggest feedback or identify patterns across a class. But a 2024 review of assessment for learning with AI found only 35 studies, and only 17 explored AI-related considerations and challenges. 8 That is a landscape of emerging practices, not mature evidence that automated diagnosis is consistently valid or beneficial.

Make feedback usable

Feedback is often written as if delivery completed the job. “Be more analytical.” “Check your work.” “Needs greater clarity.” Such comments may be accurate, yet still leave a learner without a route forward.

Actionable feedback connects three things:

  • the goal: what quality or knowledge the work is moving toward;
  • the current response: where it does and does not meet that goal;
  • the next move: a manageable action the learner can perform now.

This structure resembles influential feedback models that distinguish questions such as “Where am I going?”, “How am I going?” and “Where to next?” 2 It also aligns with reviews arguing that effective formative feedback should be specific, focused on the task or process, and presented in a way that does not overload or discourage the learner. 3

Timing depends on the task. Immediate feedback can prevent a basic error from propagating through practice. A delay can sometimes be useful when it preserves an independent attempt or invites self-evaluation. The right question is not whether immediate or delayed feedback wins in the abstract. It is which timing supports the target process for this learner and task.

Tone matters too, but kindness cannot substitute for information. Praise such as “Great job!” may support a relationship while saying nothing about what worked. Harsh precision can make useful information impossible to hear. Strong feedback is candid about the work, respectful toward the learner and specific about a possible revision.

The learner must do something next

Feedback literacy research reframes learners as participants rather than recipients. Learners need to understand standards, make judgements about quality, manage emotional responses and take action. A 2023 scoping review examined whether interventions can improve students’ feedback literacy, while an earlier systematic review organised the processes through which learners seek, interpret and use feedback. 11 12

This suggests an important product rule: do not design feedback as a dead-end message. Pair it with an action.

A learner might:

  • revise one sentence using a named criterion;
  • explain why a distractor is tempting but wrong;
  • compare their answer with an example and mark the difference;
  • solve a parallel problem with less support;
  • choose which feedback point to address first and explain why.

The response closes the loop. Without it, the system records that feedback was shown, not that it was understood or used.

Systematic reviews of formative assessment in higher education and K–12 reading broadly support assessment-and-feedback practices while showing that implementation, setting and study quality vary. 9 10 The mechanism remains plausible because evidence is connected to a decision. But the design must make that connection real.

Five tests for an assessment system

An evidence-minded assessment can be reviewed with five questions.

1. Target: What knowledge, reasoning or performance is the task intended to reveal? If the target is vague, the score will be difficult to interpret.

2. Evidence: Can the response distinguish likely patterns of understanding? Reliability matters, but so does validity: consistent measurement of the wrong construct is not useful.

3. Action: Is there a defined instructional response to each important pattern? Collecting data without a response plan creates analytics theatre.

4. Agency: Can the learner understand, question and use the feedback? Consequential inferences should not hide behind an opaque model score.

5. Recheck: Is there a meaningful opportunity to demonstrate improvement? A resubmission, parallel problem or delayed retrieval check turns advice into a testable intervention.

Five tests for useful assessment: Target, Evidence, Action, Agency, Recheck.
Five tests for useful assessment. Measurement closes a record. Formative assessment opens a better next attempt.

These tests apply beyond schools. Workplace learning, professional certification and self-directed study all suffer when assessment becomes a ledger rather than a loop.

Keep stakes proportional to the evidence

Low-stakes assessment can be deliberately incomplete. One exit question may be enough to decide whether tomorrow’s lesson needs another example. A short practice set may decide which hint appears next. Those uses tolerate uncertainty because the consequence is another learning opportunity.

High-stakes decisions require a different standard. Promotion, certification or access to an opportunity should not turn on a single opaque inference. Multiple tasks, appropriate reliability, documented validity, accessibility, review and an appeal route become more important as consequences grow. The same predictive model can be acceptable for recommending optional practice and unacceptable for deciding a learner’s future.

This proportionality is especially important with automated scoring. Efficiency does not reduce the obligation to test for systematic error, distribution shift or uneven performance across language and demographic groups. A responsible system records why an inference was made, how confident it is and when a human must intervene.

What the evidence does not show

The evidence does not show that more feedback is always better. Long comments can overload a novice; repeated automated prompts can become noise; detailed corrections can encourage dependence if learners never diagnose their own work. Nor does an average positive effect guarantee that a particular message helps a particular learner. The 17% negative effects in the broad feedback synthesis are a warning against treating delivery as benefit. 4

Current research does not establish that AI-generated feedback is consistently accurate, fair or instructionally superior. A model may produce an elegant explanation for a misdiagnosed problem. It may reward a familiar writing style, miss culturally varied expression or infer ability from incomplete data. Human review becomes more important as stakes rise. 8

The literature also does not collapse formative and summative purposes. Some decisions require a stable, standardised measure rather than coaching during the task. The mistake is not summative assessment itself. It is asking one score to certify achievement, diagnose misconceptions, motivate improvement and direct instruction without designing for those different jobs.

Finally, feedback cannot compensate for incoherent curriculum, inaccessible tasks or no time to revise. The response system is only one part of the learning environment.

Assessment earns its value through change

A score can tell us where a performance landed under specified conditions. That can be important. But learning advances through the decision after the score: revisit a prerequisite, alter an explanation, practise a strategy, revise the work, ask a better question, or increase the challenge.

The best assessment systems make that movement visible. They elicit interpretable evidence, keep inferences modest, give feedback a learner can use and check whether the next attempt improves. They measure less for the sake of measurement and respond more intelligently to what the evidence reveals.

Measurement closes a record. Formative assessment opens a better next attempt.

References

  1. Black, P., & Wiliam, D. “Assessment and Classroom Learning.” Assessment in Education: Principles, Policy & Practice, 5(1), 7–74. Source.
  2. Hattie, J., & Timperley, H. “The Power of Feedback.” Review of Educational Research, 77(1), 81–112. Source.
  3. Shute, V. J. “Focus on Formative Feedback.” Review of Educational Research, 78(1), 153–189. Source.
  4. Wisniewski, B., Zierer, K., & Hattie, J. “The Power of Feedback Revisited: A Meta-Analysis of Educational Feedback Research.” Frontiers in Psychology, 10, 3087. Source.
  5. Cai, Z., et al. “The effect of feedback on academic achievement in technology-rich learning environments: A meta-analytic review.” Educational Research Review, 39, 100521. Source.
  6. Scherer, R., et al. “How effective is feedback for L1, L2, and FL learners’ writing? A meta-analysis.” Learning and Instruction, 101961. Source.
  7. Dijks, M. A., et al. “A meta-analysis of the effects of context, content, and task factors of digitally delivered instructional feedback on learning performance.” Learning Environments Research, 27, 453–476. Source.
  8. Memarian, B., & Doleck, T. “A review of assessment for learning with artificial intelligence.” Computers in Human Behavior: Artificial Humans, 2, 100040. Source.
  9. Morris, R., Perry, T., & Wardle, L. “Formative assessment and feedback for learning in higher education: A systematic review.” Review of Education, 9, e3292. Source.
  10. Xuan, Q., Cheung, A., & Sun, D. “The effectiveness of formative assessment for enhancing reading achievement in K–12 classrooms: A meta-analysis.” Frontiers in Psychology, 13, 990196. Source.
  11. Little, T., Dawson, P., Boud, D., & Tai, J. “Can students’ feedback literacy be improved? A scoping review of interventions.” Assessment & Evaluation in Higher Education. Source.
  12. Winstone, N. E., Nash, R. A., Parker, M., & Rowntree, J. “Supporting learners’ agentic engagement with feedback: A systematic review and a taxonomy of recipience processes.” Educational Psychologist, 52(1), 17–37. Source.

Share this post

Spread the word

Comments

Ulearngo
Ulearngo provides study and exam preparation tools that help students learn effectively and prepare confidently for upcoming examinations.

Ulearngo is independent and is not affiliated with or endorsed by any examination board, government agency, university, or admissions body.

Products

Resources

Company

Legal

Copyright © 2026 Ulearngo. All rights reserved.