Definition
An assessment concept defining how evidence of learning is collected, interpreted, and used for decisions. It governs measurement design, scoring, feedback, and reporting used to evaluate progress and attainment. It does not provide fair conclusions without alignment to objectives, consistent scoring, and attention to measurement error. It supports improvement and accountability by making performance observable and trackable over time. The concept is generally stable, though tools and standards for measurement evolve over time.
Principle
Principle
Gather converging evidence (content, response processes, internal structure, relation to other variables, consequences) that score interpretations are appropriate for their intended decisions.
Demonstration
Demonstration
A reading comprehension test validated for grade placement by showing item alignment to curriculum standards, consistent scoring behavior across populations, and correlation with external criteria like classroom grades.
Misapplication
Misapplication
Claiming an assessment is 'valid' from a single correlation or because it is widely used, without examining construct relevance, item functioning, or the consequences of decisions.
Consequence
Consequence
High-quality validity evidence increases confidence that decisions based on scores (placement, certification, remediation) are defensible and reduce unintended harms.
Reversal
Reversal
Invalidity: interpreting scores beyond their supported uses or for constructs the assessment does not capture.
Boundary
Boundary
Validity is claim- and use-specific; evidence supporting one interpretation (e.g., diagnostic use) does not automatically extend to other high-stakes interpretations (e.g., certification).
Semantic Tension
Semantic Tension
Tension exists between pragmatic validity (usefulness for decisions) and narrow psychometric indicators (e.g., a single reliability coefficient) that are necessary but not sufficient.
Synthesis
Synthesis
Validity synthesizes multiple evidence strands to justify specific score interpretations and decisions; it is not a property of the test alone but of the test in context and for explicit purposes.