Definition
An education research concept defining methods used to evaluate interventions, programs, and policy impacts. It governs study design, measurement, and interpretation practices used to estimate effects and assess implementation quality. It does not establish causation without appropriate design choices and careful handling of bias and uncertainty. It supports improvement by identifying what works and under what implementation conditions. The concept is generally stable, though methods and reporting standards evolve over time.
Principle
Principle
Provide a consistent, observable set of criteria so different evaluators can measure presence or absence of required elements and reduce omission errors.
Demonstration
Demonstration
In an after-school literacy program evaluation, the checklist lists stakeholder interviews completed, baseline and follow-up assessments recorded, consent forms filed, data-cleaning steps logged, and fidelity checks observed.
Misapplication
Misapplication
Treating the checklist as an exhaustive substitute for professional judgment or context-sensitive analysis, or checking items superficially without validating underlying quality.
Consequence
Consequence
When used properly, evaluations become more comparable across sites and time, administrative oversights are reduced, and preparation for reporting and replication improves.
Reversal
Reversal
An open narrative audit without a checklist emphasizes holistic interpretation and emergent findings but risks inconsistent coverage of essential tasks and harder comparability.
Boundary
Boundary
Applies to verifying procedural and documentary elements of evaluation; it does not itself generate causal inference, replace analytic frameworks, or substitute for stakeholder consultation.
Semantic Tension
Semantic Tension
Tension exists between a checklist (binary presence/absence) and a rubric (graded judgment); checklists favor completeness and reliability, rubrics favor nuanced quality assessment.
Synthesis
Synthesis
A Program Evaluation Checklist is a reliability-focused tool that enumerates required activities and artifacts to ensure consistent, auditable execution of evaluation tasks while remaining subordinate to analytic judgment.