Definition
An assessment concept defining how evidence of learning is collected, interpreted, and used for decisions. It governs measurement design, scoring, feedback, and reporting used to evaluate progress and attainment. It does not provide fair conclusions without alignment to objectives, consistent scoring, and attention to measurement error. It supports improvement and accountability by making performance observable and trackable over time. The concept is generally stable, though tools and standards for measurement evolve over time.
Principle
Principle
Make evaluative criteria explicit and operationalize performance levels so multiple raters can make consistent judgments and stakeholders can understand the basis for scores.
Demonstration
Demonstration
A rubric for program fidelity with criteria such as 'Adherence to Curriculum', 'Dosage', and 'Staff Preparation', each defined across four levels (1 = Not Implemented, 2 = Partial, 3 = Mostly, 4 = Full) with observable indicators (e.g., lesson frequency per week, presence of lesson plans, teacher training hours).
Misapplication
Misapplication
Creating a rubric with vague descriptors, too many indistinguishable levels, or using it without rater calibration, which yields unreliable scores and misleading comparisons across sites.
Consequence
Consequence
A clear rubric increases inter-rater reliability, enables aggregation of qualitative assessments into quantitative summaries, facilitates comparison across sites or time, and makes judgment criteria transparent to stakeholders.
Reversal
Reversal
No rubric or an informal checklist leads to idiosyncratic judgments, poor comparability, and difficulty in explaining why a program was rated in a certain way.
Boundary
Boundary
Appropriate for scoring artifacts, observations, and documents in an evaluation; it is not a measurement instrument for self-reported attitudes (unless specifically designed) and does not replace statistical validity checks for quantitative outcomes.
Semantic Tension
Semantic Tension
Tension exists between rubrics (qualitative-to-structured judgment) and simple numeric rating scales or checklists; rubrics demand definitional work but provide deeper interpretive guidance, whereas scales are faster but less diagnostic.
Synthesis
Synthesis
A Program Evaluation Rubric converts evaluation dimensions into explicit criteria and observable indicators across performance levels so teams can consistently assess program quality, implementation, and outcomes.