What Should an Assessment Actually Measure?

Every assessment makes a claim, whether that claim has been stated carefully or merely assumed. A vocabulary quiz may support a claim about recall. A worked case may support a claim about application. A simulation may support a claim about performance under specified conditions. Problems arise when the claim attached to the result is broader than the evidence the assessment has actually produced.

The central question is therefore not simply whether an assessment covers the right subject matter. It is whether the tasks require people to demonstrate the kind of capability that matters. Two assessments can address exactly the same content, use equally polished questions, and produce equally precise scores while measuring fundamentally different things.

A result is meaningful only when it completes a defensible sentence: on the basis of this observed performance, there is sufficient evidence to conclude that the person can do this, under these conditions, for this purpose.

The subject is not the capability

A subject describes the territory. It does not specify what a person should be able to do within it. Data protection, financial analysis, client communication, and project risk are content areas. Each can be assessed through tasks that demand anything from simple recognition to complex professional judgement.

This distinction matters because assessments do not directly reveal knowledge, judgement, or competence. They present a task, observe a response, and use that response to infer an underlying capability. The design of the task determines which aspects of capability have an opportunity to become visible.

Consider an assessment on incident response. Asking a candidate to name the stages of an incident-management process measures something different from asking the candidate to triage an ambiguous incident, identify missing information, choose an escalation route, and explain the decision to a senior stakeholder. Both tasks concern incident response. Only the second provides direct evidence of whether the candidate can use the process when several demands interact.

Different assessments measure different abilities

The distinctions below help clarify the evidence an assessment is expected to produce. They are not a rigid hierarchy, and they are not mutually exclusive. Complex tasks often draw on several abilities at once. The important question is which ability is central to the intended interpretation and which abilities merely support it.

Recall and recognition are not inherently weak forms of assessment. In many professions, immediate access to essential facts is necessary. A clinician may need to recognise a dangerous pattern without delay. A technician may need to recall an emergency sequence when reference material is unavailable.

The problem is not measuring foundational knowledge. It is treating foundational knowledge as sufficient evidence of capabilities that also require selection, adaptation, judgement, or execution.

The same subject can support very different assessments

Suppose several assessments all claim to cover data-incident management. The subject remains constant, but each change to the prompt changes the evidence produced.

None of these tasks is universally best. Each is appropriate for a different claim and decision. A short recognition test may be entirely adequate for checking awareness after a policy update. It would be inadequate as the sole basis for certifying that someone can lead a high-risk response.

Assessment quality begins with matching the strength of the evidence to the importance and breadth of the claim.

The format does not determine the level

A particular question format does not determine which ability is being measured. A carefully designed selected-response question can require diagnosis, comparison, or evaluation. An essay can do little more than ask for a memorised list. A simulation can be cognitively shallow if every decision has already been made and every step is heavily prompted.

Conversely, an assessment does not become authentic simply because it includes realistic documents, workplace imagery, or a detailed scenario. Surface realism can coexist with an artificial decision. What matters is the intellectual and practical work the candidate is required to perform.

Alignment means asking for the capability that was intended

A coherent learning system connects three things: the capability people are expected to develop, the experiences through which they develop it, and the evidence used to decide whether they have achieved it. This can be tested with three direct questions:

If the intended outcome is to resolve unfamiliar client problems, instruction must include more than explanations of the approved process. Learners need opportunities to interpret cases, decide which information matters, test alternatives, and receive feedback on their reasoning. The assessment must then require comparable intellectual work.

A quiz on process definitions may confirm that essential language is understood, but it cannot show that a learner can resolve a problem.

A useful test is to remove the course title and labels from the assessment and look only at what candidates must actually do. Would an independent reviewer infer the same capability described in the learning objective? If not, the system is probably rewarding something other than its stated priority.

Common mismatches

These mismatches have practical consequences. Learners direct effort toward what is rewarded. Instructors adjust emphasis toward what is tested. Decision-makers begin to treat scores as evidence of capabilities the assessment never elicited. Over time, the assessment can redefine a programme more powerfully than its stated objectives do.

Capability must be inferred from observable evidence

Knowledge, reasoning, and competence cannot be observed directly. What can be observed is a response: an answer selected, an explanation given, a product created, a procedure completed, a question asked, or a decision made.

Assessment is the structured process of moving from that observable evidence to a conclusion about an underlying capability. That process is strongest when four elements are explicit:

For example, “communicates risk effectively” is too broad to score consistently until the expected evidence is defined. A task might require a candidate to turn a complex technical analysis into a three-minute briefing for a non-specialist decision-maker and then answer questions.

Relevant evidence could include accurate prioritisation, adaptation to the audience, a clear statement of uncertainty, an actionable recommendation, and appropriate responses to challenge.

The observable evidence should be specified before candidate work is reviewed. Otherwise, assessors can be drawn toward whatever is easiest to notice: confidence, fluency, polish, length, or similarity to their own preferred style. Those features may matter in some roles, but they should influence the result only when they are part of the capability being assessed.

A plausible task is not always sufficient evidence

One well-designed task can still provide a narrow sample. A person may perform strongly because the context is familiar, the topic happens to match prior experience, or the task favours a particular approach. Where the decision is consequential, evidence should usually be sampled across more than one problem, context, or occasion.

Breadth and depth must be balanced. A large number of short questions can sample content broadly but reveal little about integrated performance. A single extended task can expose reasoning and execution in depth but leave large areas untouched. The appropriate balance depends on what the result will be used to decide.

Construct validity: does the evidence support the claim?

The capability an assessment is intended to represent is sometimes called its construct. Construct validity concerns whether the interpretation made from the assessment is adequately supported. In plain terms: are we justified in saying that this result means what we claim it means?

Validity is not secured by giving an assessment a professional title, covering the correct syllabus, or calculating a reliable total. It depends on the relationship among the intended capability, the task demands, the scoring approach, the people assessed, and the use made of the result.

One fundamental threat is leaving out important parts of the capability. An assessment of leadership that measures only knowledge of leadership models omits behaviour, judgement, influence, and adaptation. An assessment of written communication that uses only sentence-level grammar questions omits organisation, audience awareness, and purposeful expression.

Another threat is allowing irrelevant factors to influence the result. A technically demanding reading passage can turn a test of numerical reasoning into a partial test of reading proficiency. An unfamiliar interface can suppress the performance of otherwise capable users. Severe time limits can transform a test of careful analysis into a test of processing speed. Elaborate presentation requirements can contaminate an assessment intended to measure the quality of a recommendation.

The goal is not to remove every difficulty. Relevant difficulty is precisely what an assessment needs. The goal is to distinguish difficulty that belongs to the capability from difficulty introduced accidentally by the assessment.

Assessment conditions change what a result means

Access to tools, reference material, collaborators, and time should reflect the intended claim. If professionals normally use a calculator, template, code library, diagnostic instrument, or policy manual, prohibiting that resource may shift the assessment toward memory or manual execution.

Conversely, if immediate unaided recall is essential to safe performance, unrestricted reference access may prevent the assessment from eliciting that capability.

Assistance and prompting also matter. A candidate who succeeds after receiving a checklist, hints, and corrective feedback has provided evidence of supported performance. That may be valuable for a developmental purpose, but it is different from evidence of independent performance. The assessment should record the distinction rather than collapse both into the same score.

Accessibility, language demands, cultural assumptions, technology requirements, and testing conditions must also be treated as measurement issues. An assessment cannot support an accurate interpretation if avoidable barriers prevent some people from demonstrating the intended capability.

Authenticity is about representative demands

Authentic assessment is often understood too literally. A task does not become valid merely because it resembles an office, uses a realistic document, or tells a detailed story. A simplified case can provide stronger evidence if it preserves the reasoning, trade-offs, and standards that define competent practice.

Depending on the capability, the important features of real performance may include:

Authenticity is therefore selective. It preserves the features that make the capability what it is while controlling features that would add cost, risk, or irrelevant variation. A high-quality simulation is not a replica of an entire job. It is a deliberate sample of the decisions and actions that provide the strongest evidence.

Performance tasks also require disciplined scoring. Clear dimensions, observable criteria, examples of performance at different standards, and explicit treatment of critical errors help assessors distinguish evidence from impression.

Why a numerical score is not an explanation

A score compresses evidence. That can be useful: it supports comparison, summarises a body of responses, and enables consistent decision rules. But compression removes detail. A number cannot explain itself.

A score of 78 does not reveal, without further information, what was measured, how demanding the tasks were, which capabilities were strong or weak, how much uncertainty surrounds the result, or what standard the score represents. It may indicate 78 per cent of available marks, a scaled result, a position relative to other candidates, or performance against a defined standard. Those interpretations are not interchangeable.

Two candidates may receive the same overall score while presenting very different capability profiles. One may analyse evidence exceptionally well but communicate the conclusion poorly. The other may deliver a persuasive briefing built on weak analysis. If both dimensions are important, the shared total conceals a decision-relevant difference.

A defensible score report should make clear:

This is especially important for pass-or-fail decisions. A threshold creates a clear administrative outcome, but it does not eliminate uncertainty or turn small score differences into large capability differences. Someone just above a threshold is not necessarily meaningfully more capable than someone just below it.

Scoring rules should also reflect whether strengths can legitimately compensate for weaknesses. Strong client rapport may not compensate for unsafe technical advice. Excellent analytical reasoning may not compensate for failure to follow a mandatory control. Where a capability is essential, averaging it into a total can produce the wrong conclusion.

A practical test for assessment design

Before an assessment is built—or before an existing assessment is trusted for a new purpose—six questions provide a disciplined review.

These questions also expose when one assessment is being asked to serve incompatible purposes. A brief diagnostic intended to identify areas for further learning may not support certification. A certification assessment designed for consistent decisions may not provide sufficiently detailed developmental feedback. A recruitment screen may not predict performance across the full role.

Reusing an assessment is appropriate only when the new interpretation and use remain supported.

Measure what the decision needs to know

The quality of an assessment is not determined by the number of questions, the sophistication of the platform, the realism of the interface, or the precision of the score. It is determined by the strength of the reasoning that connects an intended capability to an observed performance and then to a proportionate conclusion.

That requires restraint as well as ambition. Not every assessment needs to reproduce authentic performance, and no finite assessment can capture an entire domain. The obligation is to select evidence that is fit for the intended decision, reduce plausible alternative explanations, and avoid claiming more than the evidence can support.

A sound assessment measures the capability that matters, under conditions that make that capability visible, and reports the result in terms that explain what the person can actually do.