Definition

An assessment concept defining how evidence of learning is collected, interpreted, and used for decisions. It governs measurement design, scoring, feedback, and reporting used to evaluate progress and attainment. It does not provide fair conclusions without alignment to objectives, consistent scoring, and attention to measurement error. It supports improvement and accountability by making performance observable and trackable over time. The concept is generally stable, though tools and standards for measurement evolve over time.

Principle

Principle
Gather converging evidence (content, response processes, internal structure, relation to other variables, consequences) that score interpretations are appropriate for their intended decisions.

Demonstration

Demonstration
A reading comprehension test validated for grade placement by showing item alignment to curriculum standards, consistent scoring behavior across populations, and correlation with external criteria like classroom grades.

Misapplication

Misapplication
Claiming an assessment is 'valid' from a single correlation or because it is widely used, without examining construct relevance, item functioning, or the consequences of decisions.

Consequence

Consequence
High-quality validity evidence increases confidence that decisions based on scores (placement, certification, remediation) are defensible and reduce unintended harms.

Reversal

Reversal
Invalidity: interpreting scores beyond their supported uses or for constructs the assessment does not capture.

Boundary

Boundary
Validity is claim- and use-specific; evidence supporting one interpretation (e.g., diagnostic use) does not automatically extend to other high-stakes interpretations (e.g., certification).

Semantic Tension

Semantic Tension
Tension exists between pragmatic validity (usefulness for decisions) and narrow psychometric indicators (e.g., a single reliability coefficient) that are necessary but not sufficient.

Synthesis

Synthesis
Validity synthesizes multiple evidence strands to justify specific score interpretations and decisions; it is not a property of the test alone but of the test in context and for explicit purposes.