Every assessment makes a claim, whether that claim has been stated carefully or merely assumed. A vocabulary quiz may support a claim about recall. A worked case may support a claim about application. A simulation may support a claim about performance under specified conditions. Problems arise when the claim attached to the result is broader than the evidence the assessment has actually produced.
The central question is therefore not simply whether an assessment covers the right subject matter. It is whether the tasks require people to demonstrate the kind of capability that matters. Two assessments can address exactly the same content, use equally polished questions, and produce equally precise scores while measuring fundamentally different things.
A result is meaningful only when it completes a defensible sentence: on the basis of this observed performance, there is sufficient evidence to conclude that the person can do this, under these conditions, for this purpose.
The subject is not the capability
A subject describes the territory. It does not specify what a person should be able to do within it. Data protection, financial analysis, client communication, and project risk are content areas. Each can be assessed through tasks that demand anything from simple recognition to complex professional judgement.
This distinction matters because assessments do not directly reveal knowledge, judgement, or competence. They present a task, observe a response, and use that response to infer an underlying capability. The design of the task determines which aspects of capability have an opportunity to become visible.
Consider an assessment on incident response. Asking a candidate to name the stages of an incident-management process measures something different from asking the candidate to triage an ambiguous incident, identify missing information, choose an escalation route, and explain the decision to a senior stakeholder. Both tasks concern incident response. Only the second provides direct evidence of whether the candidate can use the process when several demands interact.
Different assessments measure different abilities
The distinctions below help clarify the evidence an assessment is expected to produce. They are not a rigid hierarchy, and they are not mutually exclusive. Complex tasks often draw on several abilities at once. The important question is which ability is central to the intended interpretation and which abilities merely support it.
- Recall is the ability to retrieve facts, terms, rules, or steps from memory without being shown the answer. Evidence might include accurately stating a threshold, naming a principle, or reproducing a required sequence.
- Recognition is the ability to identify correct, relevant, or familiar information when possible answers or cues are provided. It may involve selecting the applicable policy, detecting an error, or distinguishing a valid example from a plausible alternative.
- Conceptual understanding involves relationships, causes, boundaries, and implications rather than isolated facts. It becomes visible when someone explains why a rule applies, compares approaches, predicts a consequence, or represents the same idea in another form.
- Procedural skill is the ability to carry out a defined method accurately, efficiently, and in the appropriate order. This might mean completing a calculation, configuring a system, conducting a check, or following a safety-critical sequence without omission.
- Application is the ability to select and use relevant knowledge or procedures in a contextualised situation. Evidence may include choosing an appropriate method for a case, applying a rule to supplied facts, or adapting a known process to ordinary variation.
- Reasoning involves drawing conclusions, diagnosing causes, weighing evidence, and justifying decisions where the answer is not supplied by a single remembered rule. Strong evidence includes explicit assumptions, evaluation of alternatives, and a conclusion supported by material facts.
- Transfer is the ability to use prior learning in a situation that differs materially from the examples encountered during instruction. It requires the person to recognise an underlying principle, adapt a strategy, and account for the features that make the new context different.
- Authentic performance integrates knowledge, skill, and judgement under conditions that preserve the important demands of real practice. It may involve producing a work product, conducting an interaction, or resolving a realistic problem using relevant tools, standards, constraints, and communication.
Recall and recognition are not inherently weak forms of assessment. In many professions, immediate access to essential facts is necessary. A clinician may need to recognise a dangerous pattern without delay. A technician may need to recall an emergency sequence when reference material is unavailable.
The problem is not measuring foundational knowledge. It is treating foundational knowledge as sufficient evidence of capabilities that also require selection, adaptation, judgement, or execution.
The same subject can support very different assessments
Suppose several assessments all claim to cover data-incident management. The subject remains constant, but each change to the prompt changes the evidence produced.
- Asking a candidate to state the organisation’s incident-severity levels measures recall of the classification framework. It does not establish that the candidate can classify a real incident correctly.
- Asking the candidate to select which event meets the definition of a reportable incident measures recognition in a bounded set of examples. It does not show whether the candidate can investigate incomplete or conflicting facts.
- Asking why a seemingly minor event could require escalation provides evidence of conceptual understanding. It does not show whether the candidate will escalate appropriately under operational pressure.
- Asking the candidate to complete a triage record using supplied information measures accurate execution of a standard procedure. It does not show what the candidate will do when necessary information is missing.
- Asking the candidate to classify a detailed incident and choose the next action measures application of the framework to a contextualised case.
- Asking the candidate to evaluate competing explanations, identify evidence gaps, and justify an escalation decision provides evidence of reasoning under ambiguity.
- Presenting an incident involving an unfamiliar technology or third-party relationship tests whether the candidate can transfer relevant principles to a materially different context.
- Requiring the candidate to manage a timed simulation, update the incident record, and brief a resistant senior stakeholder provides evidence of integrated professional performance under representative constraints.
None of these tasks is universally best. Each is appropriate for a different claim and decision. A short recognition test may be entirely adequate for checking awareness after a policy update. It would be inadequate as the sole basis for certifying that someone can lead a high-risk response.
Assessment quality begins with matching the strength of the evidence to the importance and breadth of the claim.
The format does not determine the level
A particular question format does not determine which ability is being measured. A carefully designed selected-response question can require diagnosis, comparison, or evaluation. An essay can do little more than ask for a memorised list. A simulation can be cognitively shallow if every decision has already been made and every step is heavily prompted.
Conversely, an assessment does not become authentic simply because it includes realistic documents, workplace imagery, or a detailed scenario. Surface realism can coexist with an artificial decision. What matters is the intellectual and practical work the candidate is required to perform.
Alignment means asking for the capability that was intended
A coherent learning system connects three things: the capability people are expected to develop, the experiences through which they develop it, and the evidence used to decide whether they have achieved it. This can be tested with three direct questions:
- What, specifically, should a person be able to do at the end?
- Did the learning experience provide a genuine opportunity to practise that kind of performance?
- Does the assessment require the person to demonstrate it?
If the intended outcome is to resolve unfamiliar client problems, instruction must include more than explanations of the approved process. Learners need opportunities to interpret cases, decide which information matters, test alternatives, and receive feedback on their reasoning. The assessment must then require comparable intellectual work.
A quiz on process definitions may confirm that essential language is understood, but it cannot show that a learner can resolve a problem.
A useful test is to remove the course title and labels from the assessment and look only at what candidates must actually do. Would an independent reviewer infer the same capability described in the learning objective? If not, the system is probably rewarding something other than its stated priority.
Common mismatches
- A programme intends to develop problem-solving capability but tests whether candidates can define problem-solving models and name their stages.
- Learners practise client communication, but the final assessment is a factual quiz on communication principles.
- A programme promises to build professional judgement, but every assessment question has one obvious answer and contains no competing considerations.
- Learners are expected to perform a technical procedure, but the assessment asks them only to write an explanation of how it should be performed.
- A course claims to prepare learners for unfamiliar cases, but the assessment reproduces worked examples with different names or numbers.
- Candidates are expected to produce evidence-based recommendations, but their presentations are graded mainly on visual polish and confidence.
- A programme intends to develop collaborative capability, but assigns a single mark to a group output without examining individual contributions.
These mismatches have practical consequences. Learners direct effort toward what is rewarded. Instructors adjust emphasis toward what is tested. Decision-makers begin to treat scores as evidence of capabilities the assessment never elicited. Over time, the assessment can redefine a programme more powerfully than its stated objectives do.
Capability must be inferred from observable evidence
Knowledge, reasoning, and competence cannot be observed directly. What can be observed is a response: an answer selected, an explanation given, a product created, a procedure completed, a question asked, or a decision made.
Assessment is the structured process of moving from that observable evidence to a conclusion about an underlying capability. That process is strongest when four elements are explicit:
- The claim: what the person should be able to know, decide, produce, or perform, including the relevant conditions.
- The task: a situation that gives the person a fair and sufficiently demanding opportunity to reveal the intended capability.
- The evidence: the actions, decisions, explanations, or qualities that would support or weaken the claim.
- The interpretation: the conclusion justified by the observed performance and the standards applied to it.
For example, “communicates risk effectively” is too broad to score consistently until the expected evidence is defined. A task might require a candidate to turn a complex technical analysis into a three-minute briefing for a non-specialist decision-maker and then answer questions.
Relevant evidence could include accurate prioritisation, adaptation to the audience, a clear statement of uncertainty, an actionable recommendation, and appropriate responses to challenge.
The observable evidence should be specified before candidate work is reviewed. Otherwise, assessors can be drawn toward whatever is easiest to notice: confidence, fluency, polish, length, or similarity to their own preferred style. Those features may matter in some roles, but they should influence the result only when they are part of the capability being assessed.
A plausible task is not always sufficient evidence
One well-designed task can still provide a narrow sample. A person may perform strongly because the context is familiar, the topic happens to match prior experience, or the task favours a particular approach. Where the decision is consequential, evidence should usually be sampled across more than one problem, context, or occasion.
Breadth and depth must be balanced. A large number of short questions can sample content broadly but reveal little about integrated performance. A single extended task can expose reasoning and execution in depth but leave large areas untouched. The appropriate balance depends on what the result will be used to decide.
Construct validity: does the evidence support the claim?
The capability an assessment is intended to represent is sometimes called its construct. Construct validity concerns whether the interpretation made from the assessment is adequately supported. In plain terms: are we justified in saying that this result means what we claim it means?
Validity is not secured by giving an assessment a professional title, covering the correct syllabus, or calculating a reliable total. It depends on the relationship among the intended capability, the task demands, the scoring approach, the people assessed, and the use made of the result.
One fundamental threat is leaving out important parts of the capability. An assessment of leadership that measures only knowledge of leadership models omits behaviour, judgement, influence, and adaptation. An assessment of written communication that uses only sentence-level grammar questions omits organisation, audience awareness, and purposeful expression.
Another threat is allowing irrelevant factors to influence the result. A technically demanding reading passage can turn a test of numerical reasoning into a partial test of reading proficiency. An unfamiliar interface can suppress the performance of otherwise capable users. Severe time limits can transform a test of careful analysis into a test of processing speed. Elaborate presentation requirements can contaminate an assessment intended to measure the quality of a recommendation.
The goal is not to remove every difficulty. Relevant difficulty is precisely what an assessment needs. The goal is to distinguish difficulty that belongs to the capability from difficulty introduced accidentally by the assessment.
Assessment conditions change what a result means
Access to tools, reference material, collaborators, and time should reflect the intended claim. If professionals normally use a calculator, template, code library, diagnostic instrument, or policy manual, prohibiting that resource may shift the assessment toward memory or manual execution.
Conversely, if immediate unaided recall is essential to safe performance, unrestricted reference access may prevent the assessment from eliciting that capability.
Assistance and prompting also matter. A candidate who succeeds after receiving a checklist, hints, and corrective feedback has provided evidence of supported performance. That may be valuable for a developmental purpose, but it is different from evidence of independent performance. The assessment should record the distinction rather than collapse both into the same score.
Accessibility, language demands, cultural assumptions, technology requirements, and testing conditions must also be treated as measurement issues. An assessment cannot support an accurate interpretation if avoidable barriers prevent some people from demonstrating the intended capability.
Authenticity is about representative demands
Authentic assessment is often understood too literally. A task does not become valid merely because it resembles an office, uses a realistic document, or tells a detailed story. A simplified case can provide stronger evidence if it preserves the reasoning, trade-offs, and standards that define competent practice.
Depending on the capability, the important features of real performance may include:
- the type and quality of information available at the point of decision;
- the need to identify missing, unreliable, or conflicting evidence;
- the tools and reference materials ordinarily used;
- time pressure or changing priorities where they are genuinely material;
- interactions with clients, colleagues, systems, or decision-makers;
- professional, legal, technical, or ethical standards;
- the need to explain, document, or defend a decision; and
- the consequences of critical errors, represented safely where necessary.
Authenticity is therefore selective. It preserves the features that make the capability what it is while controlling features that would add cost, risk, or irrelevant variation. A high-quality simulation is not a replica of an entire job. It is a deliberate sample of the decisions and actions that provide the strongest evidence.
Performance tasks also require disciplined scoring. Clear dimensions, observable criteria, examples of performance at different standards, and explicit treatment of critical errors help assessors distinguish evidence from impression.
Why a numerical score is not an explanation
A score compresses evidence. That can be useful: it supports comparison, summarises a body of responses, and enables consistent decision rules. But compression removes detail. A number cannot explain itself.
A score of 78 does not reveal, without further information, what was measured, how demanding the tasks were, which capabilities were strong or weak, how much uncertainty surrounds the result, or what standard the score represents. It may indicate 78 per cent of available marks, a scaled result, a position relative to other candidates, or performance against a defined standard. Those interpretations are not interchangeable.
Two candidates may receive the same overall score while presenting very different capability profiles. One may analyse evidence exceptionally well but communicate the conclusion poorly. The other may deliver a persuasive briefing built on weak analysis. If both dimensions are important, the shared total conceals a decision-relevant difference.
A defensible score report should make clear:
- which capabilities the assessment was designed to represent;
- which content, contexts, and task types were sampled;
- the conditions under which performance was observed;
- whether candidates are compared with a standard or with one another;
- how separate dimensions contributed to the total;
- whether critical requirements had to be met independently;
- how precise the result is, particularly near a decision threshold; and
- which conclusions the available evidence does and does not support.
This is especially important for pass-or-fail decisions. A threshold creates a clear administrative outcome, but it does not eliminate uncertainty or turn small score differences into large capability differences. Someone just above a threshold is not necessarily meaningfully more capable than someone just below it.
Scoring rules should also reflect whether strengths can legitimately compensate for weaknesses. Strong client rapport may not compensate for unsafe technical advice. Excellent analytical reasoning may not compensate for failure to follow a mandatory control. Where a capability is essential, averaging it into a total can produce the wrong conclusion.
A practical test for assessment design
Before an assessment is built—or before an existing assessment is trusted for a new purpose—six questions provide a disciplined review.
- What exact conclusion should the result support? Replace broad labels such as “understands compliance” with a statement of what a capable person should be able to identify, decide, explain, or perform.
- What would count as convincing evidence? Specify the decisions, actions, explanations, and qualities that would support the claim. Identify unacceptable errors and distinguish essential features from desirable refinements.
- Does the task make that evidence possible? A task cannot reveal judgement if all judgement has been removed, communication if no message is produced, or transfer if the context reproduces the learning example.
- What else could influence performance? Examine reading load, speed, technology, prior familiarity, cultural knowledge, physical conditions, prompting, and assessor expectations.
- Is the sample broad enough for the decision? Consider whether one question, one case, or one occasion is sufficient. Increase sampling when the claim is broad, performance is variable, or the consequence of error is high.
- How will the result be interpreted and communicated? Define what totals, profiles, thresholds, and performance levels mean. State limitations with the same care used to state conclusions.
These questions also expose when one assessment is being asked to serve incompatible purposes. A brief diagnostic intended to identify areas for further learning may not support certification. A certification assessment designed for consistent decisions may not provide sufficiently detailed developmental feedback. A recruitment screen may not predict performance across the full role.
Reusing an assessment is appropriate only when the new interpretation and use remain supported.
Measure what the decision needs to know
The quality of an assessment is not determined by the number of questions, the sophistication of the platform, the realism of the interface, or the precision of the score. It is determined by the strength of the reasoning that connects an intended capability to an observed performance and then to a proportionate conclusion.
That requires restraint as well as ambition. Not every assessment needs to reproduce authentic performance, and no finite assessment can capture an entire domain. The obligation is to select evidence that is fit for the intended decision, reduce plausible alternative explanations, and avoid claiming more than the evidence can support.
A sound assessment measures the capability that matters, under conditions that make that capability visible, and reports the result in terms that explain what the person can actually do.