2026–2027 · Honest & Growth-Oriented · For Leaders & Evaluators

Teacher Evaluation
& Feedback Toolkit

The Widget Effect, the reform era's results, the accountability-vs-growth tension, observation frameworks, multiple measures, the value-added-measures controversy, feedback that develops, evaluator reliability and bias, keeping evaluation distinct from coaching, and difficult conversations. Center growth over ranking — accountability alone doesn't improve teaching. Plus 100 tips. From K12 Academics, free and with no login.

📋
Section 01

Welcome

Welcome to the K12academics Teacher Evaluation & Feedback Toolkit — a practical, honest guide to one of the hardest, most fraught parts of school leadership. Teacher evaluation matters enormously — yet its history is littered with failures, from systems that rated everyone 'satisfactory' to billion-dollar reforms that didn't work. This toolkit charts a better path: credible, growth-oriented, feedback-rich evaluation. It's built for leaders and evaluators.

How to use this toolkit
Who it's for
  • Principals and assistant principals
  • District leaders and evaluators
  • Instructional leaders and coaches
  • Anyone who observes and evaluates teachers
The stance this takes
  • Center GROWTH over ranking — accountability alone fails
  • Use multiple measures, not one observation or test score
  • The best 'feedback' develops teachers, not just rates them
  • Learn from a fraught history — including what didn't work

Teacher evaluation sits at a genuinely difficult intersection — of accountability and growth, of fairness and judgment, of a contentious history and an uncertain best practice. This toolkit is honest about all of it: what failed and why, what the research actually shows, and how to make evaluation a force for better teaching rather than a compliance ritual or a 'gotcha.' It's free and educational; adapt it to your context and policies.

🎯
Section 02

Why Teacher Evaluation Matters (and Why It’s Hard)

Teacher evaluation matters because teaching matters — but it's genuinely one of the hardest things school leaders do, for reasons worth naming honestly up front.

99%
teachers rated 'satisfactory' under old systems — no differentiation, and rarely useful feedback
The Widget Effect · TNTP
Little effect
the billion-dollar high-stakes evaluation reforms of the 2010s had little effect on student outcomes
reform research
Growth
evaluation works best as a DEVELOPMENTAL tool — accountability metrics alone don't improve teaching
evaluation research
Two purposes
evaluation serves accountability AND growth — and the two don't always work together smoothly
enduring tension

The case for taking evaluation seriously is strong: teachers are the most important in-school factor in student learning, and ensuring — and developing — good teaching is among the most consequential things a school does. But evaluation is genuinely hard, for reasons this toolkit doesn't gloss over: there's little consensus on how to define 'good teaching,' evaluation must somehow serve two purposes (accountability and growth) that don't always align, observing and judging teaching well requires real skill that evaluators often lack, and the field carries a fraught history of systems that failed. Doing evaluation well means facing all of this honestly.

Evaluation is high-stakes, deeply human, and genuinely difficult

It would be easy to treat teacher evaluation as a straightforward technical task — observe, rate, done. It isn't. It sits at the intersection of instructional quality, accountability, professional growth, employment decisions, and human relationships, and it's freighted with a history of approaches that either failed to differentiate anyone or overreached and backfired. Even defining 'good teaching' — the thing you're supposedly measuring — turns out to be contested. This toolkit takes the difficulty seriously rather than pretending it away, because the leaders who evaluate teachers well are the ones who understand what's genuinely hard about it — and who center evaluation on developing teachers rather than merely rating them.

📦
Section 03

The Widget Effect: Why Old Systems Failed

To understand teacher evaluation today, start with the landmark report that exposed why traditional systems failed so badly: 'The Widget Effect.' Its findings reshaped the entire field.

What the Widget Effect found
  • 99% of teachers rated 'satisfactory' — no differentiation
  • Evaluation was infrequent and low-standard
  • Teachers rarely got feedback that helped them improve
  • Excellence unrecognized; poor performance unaddressed
The core problem
  • Teachers treated as interchangeable 'widgets'
  • Not as professionals with real differences
  • The system was indifferent to performance
  • It disrespected teachers AND failed students
The old systems treated all teachers as interchangeable widgets

In 2009, TNTP's landmark report 'The Widget Effect' — studying 12 districts and surveying ~15,000 teachers — found that more than 99% of teachers were rated 'satisfactory.' Evaluation happened rarely, against low standards, and produced almost no useful feedback and almost no consequences: excellence went unrecognized and poor performance went unaddressed. The report's memorable framing: if teachers are so important, why do we treat them like interchangeable widgets rather than professionals with real, meaningful differences? This 'indifference to performance,' it argued, disrespected teachers and gambled with students' futures. The report was extraordinarily influential — it directly shaped the Race to the Top reforms (§04) — and it remains the essential starting point: the problem it named (evaluation that neither differentiates nor develops) is the one good evaluation must solve. See TNTP.

🔁
Section 04

The Reform Era: What We Learned (Honestly)

The Widget Effect sparked a decade of ambitious, expensive evaluation reform. Being honest about what happened — including its disappointing results — is essential to doing better now.

The reform era (2010s)
  • Race to the Top pushed 'rigorous' evaluation systems
  • Test-based value-added measures (VAM) were added
  • Complex multi-component systems replaced checklists
  • Billions of dollars were invested
The honest results
  • The reforms had little positive effect on student outcomes
  • High-stakes, test-heavy systems bred backlash
  • Many states have since rolled back VAM
  • The lesson: accountability alone doesn't improve teaching
The billion-dollar reforms largely didn't work — and that's a crucial lesson

Honesty about the reform era matters. In response to the Widget Effect, Race to the Top pushed states (with billions in federal grants) to adopt 'rigorous' evaluation systems incorporating student test-based value-added measures, and the Gates Foundation funded major research (the MET project). The systems became far more complex — multiple weighted components, test data, elaborate observation rubrics. And then the uncomfortable finding: rigorous research (including a major RAND study of the Gates-funded reforms) concluded these expensive, high-stakes reforms had little positive effect on student outcomes — and they generated significant backlash, with many states since rolling back the test-based components. The lesson isn't that evaluation doesn't matter; it's that accountability-by-measurement alone doesn't improve teaching. That reframing — from measuring teachers to developing them — shapes everything that follows.

⚖️
Section 05

Two Purposes: Accountability vs. Growth

At the heart of every tension in teacher evaluation lies a single fact: evaluation is asked to serve two purposes — accountability and growth — that don't always work together.

The two purposes
  • Accountability — rate performance; inform personnel decisions
  • Growth — develop teachers through feedback & support
  • Both are legitimate and important
  • But they can pull in different directions
Center growth
  • Accountability metrics ALONE don't improve teaching
  • Evaluation works best as a developmental tool
  • A screening device rarely develops anyone
  • Balance both — but center growth
Evaluation serves two purposes — and works best when it centers growth

The deepest tension in teacher evaluation is that it's asked to do two different jobs at once: accountability (judging performance against a standard, informing decisions about employment, pay, and tenure) and growth (giving teachers the feedback and support to improve). Both matter — but they can pull in opposite directions, because the same conversation can't easily be a supportive coaching dialogue and a high-stakes judgment (a teacher won't be vulnerable about their struggles with the person deciding their future, §13). The research points clearly one way: evaluation systems fail to foster improvement through accountability metrics alone, but multiple observations and detailed feedback can genuinely develop teachers — meaning evaluation works best as a developmental tool, not a screening device. Accountability still matters, but the wise emphasis is on growth.

📐
Section 06

Observation Frameworks

Most evaluation systems use a framework — a shared rubric describing effective teaching. The best-known are Danielson's and Marzano's. Used well, they give a common language; used poorly, they become checklists.

Common frameworks
  • Danielson's Framework for Teaching — the most widely used
  • 4 domains: planning, environment, instruction, professionalism
  • Marzano's teacher evaluation model
  • A shared language for what good teaching looks like
Use frameworks wisely
  • A framework is a tool for growth, not a checklist
  • Even Danielson warns hers has been misused
  • 'Good teaching' is genuinely hard to define & judge
  • Frameworks help — but don't do the thinking for you
Frameworks give a shared language — but they're tools, not answers

Most evaluation systems are built on an observation framework — a detailed rubric describing dimensions of effective teaching. The most widely used is Charlotte Danielson's Framework for Teaching (four domains: planning and preparation, classroom environment, instruction, and professional responsibilities); Robert Marzano's model is another. At their best, frameworks provide a shared language for good teaching and clearer expectations. But two honest cautions: even Danielson herself has become critical of how her framework is used — misapplied as a rigid checklist, and used by evaluators who often 'don't have the skill to differentiate great teaching from merely good or mediocre.' And there's genuinely 'little consensus on how the profession should define good teaching.' A framework is a valuable tool — but it can't substitute for skilled, thoughtful judgment. See Danielson Group and Marzano Resources.

📊
Section 07

Multiple Measures

No single measure captures teaching well. Effective evaluation draws on multiple measures — so results aren't distorted by one observation, one class, or one test.

Use multiple measures
  • Multiple observations, not a single moment
  • Student learning evidence (growth, formative trends)
  • Professional practice artifacts (plans, assessments)
  • Self-assessment and reflection
Why it matters
  • No single measure captures teaching well
  • One observation reflects one class on one day
  • Balance guards against distortion and bias
  • Research: you can't differentiate on observation alone
No single measure is enough — balance several

A core lesson of evaluation research (including the MET project): no single measure captures teaching well, so effective evaluation uses multiple measures. A single observation reflects one class, on one day, seen by one observer — far too thin a basis for a high-stakes judgment (and research found it's 'nearly impossible to differentiate performance using classroom observations alone'). Better systems combine multiple observations over time with student learning evidence (growth data, formative trends), professional-practice artifacts (lesson plans, assessments), and the teacher's own self-assessment and reflection. Combining measures produces a fairer, fuller picture, guards against the distortion any single measure introduces, and — crucially — provides richer information for growth. The point of multiple measures isn't more precise ranking; it's a truer, more useful understanding of teaching.

🔍
Section 08

The Observation Process

Observation is the part of evaluation nearly every teacher experiences — so getting it right matters enormously. Done well, it's a rich source of growth; done poorly, it's a distorted snapshot.

A strong observation process
  • Multiple observations over time, not one
  • A cycle: pre-conference, observe, post-conference
  • Trained, calibrated observers (§12)
  • Focus on evidence — what actually happened
Do it well
  • Gather objective evidence, not snap judgments
  • Look at learning, not just teacher performance
  • Mix announced and unannounced observations
  • Make the post-conference about growth, not just a score
Observation is the heart of evaluation — do it as a growth cycle, not a drive-by

Since classroom observation is the one part of evaluation nearly every teacher experiences, its quality largely determines whether evaluation helps or harms. A strong observation process is a cycle, not a one-time drive-by: a pre-conference (understanding the lesson and the teacher's goals), the observation itself (gathering objective evidence of what happens — including what students are doing and learning, not just the teacher's performance), and a post-conference centered on reflection and growth (not just delivering a score). It should involve multiple observations over time (one class on one day is a thin sample), trained and calibrated observers (§12), and a focus on evidence over snap judgments. Make the observation a source of genuine, growth-oriented feedback (§10) — that's where its value lies.

🧮
Section 09

Value-Added Measures: The Controversy

Few topics in evaluation are more contested than value-added measures (VAM) — statistical estimates of a teacher's contribution to test-score growth. It's important to understand both the appeal and the serious problems.

What VAM is
  • A statistical estimate of a teacher's impact on test scores
  • Factors in students' expected growth
  • Seen by some as a more 'objective' measure
  • Central to the reform era's systems
The serious problems
  • It relies on standardized tests (a narrow slice)
  • Scores can be unstable and are only estimates
  • It can be unfair, especially for teachers of disadvantaged students
  • It can't identify what to improve — just a number
VAM is controversial — and it can't tell a teacher how to improve

Value-added measures (VAM) use statistics to estimate how much a teacher contributed to their students' test-score growth, factoring in expected gains. Proponents saw them as a more 'objective' counterweight to inflated observation ratings — and they were central to the reform era. But the problems are serious and well-documented: VAM relies on standardized tests (a narrow slice of learning), scores can be unstable and are only estimates, they can be unfair (especially to teachers of disadvantaged students), and they don't even apply to the many teachers in untested subjects and grades. Perhaps most tellingly for a growth-oriented system: VAM cannot identify specific areas for instructional improvement — it yields a number, not guidance a teacher can act on. Given the mixed evidence and the widespread rollback of test-based systems, treat VAM with real caution — as, at most, one limited measure among many, never the centerpiece.

💬
Section 10

Feedback That Develops

If evaluation is to improve teaching, its most important product isn't a rating — it's feedback. And not just any feedback: specific, clear, actionable, credible feedback delivered as a collaboration.

Feedback that develops
  • Specific, clear, and actionable — not vague
  • Constructive and growth-oriented
  • Delivered collaboratively, not handed down
  • Credible — the evaluator's expertise matters
What teachers value most
  • Working collaboratively with the evaluator
  • Clear, specific, constructive feedback
  • Feedback paired with coaching & support (§13)
  • Feed forward — toward the next step, not a verdict
The most valuable thing evaluation produces is developmental feedback — not a rating

Ask what actually improves teaching through evaluation, and the answer isn't the score — it's the feedback. Research is clear about what makes feedback developmental: teachers report that the most helpful aspects of evaluation are working collaboratively with their evaluator and receiving constructive, clear, and specific feedback — and that feedback lands better when it comes from an evaluator with genuine credibility and expertise. Feedback should be actionable (a teacher can do something with it), collaborative (a two-way conversation, not a verdict handed down), and forward-looking (aimed at the next step). Crucially, feedback works best when paired with coaching and support to act on it — a randomized study found that individualized coaching in response to observation feedback improved classroom practice. The goal of feedback is improvement, not judgment. See our Instructional Coaching toolkit.

🌱
Section 11

Growth-Oriented Evaluation

Pulling it together: the whole point of evaluation should be to improve teaching, not merely to rank it. Growth-oriented evaluation turns a dreaded ritual into a genuine engine of improvement.

Emphasize growth over ranking
  • The goal is better teaching, not a score
  • Formative feedback throughout, not just an annual rating
  • Connect evaluation to targeted PD & support
  • Turn evaluation from anxiety into partnership
Formative + summative
  • Formative — ongoing, low-stakes, developmental (the emphasis)
  • Summative — a final rating (necessary, but not the point)
  • Center the formative growth work
  • Rankings inform; feedback improves
Growth over ranking — the point is better teaching, not a better number

The single most important shift in effective evaluation is from ranking to growth: the purpose is to help teaching get better, not merely to sort teachers or produce a rating. This means emphasizing the formative dimension (ongoing, low-stakes, developmental feedback throughout the year — 'the coaching that happens during practice') over the summative (the end-of-cycle rating, which is necessary for accountability but isn't where improvement happens). It means connecting what evaluation reveals to targeted professional development (if teachers consistently struggle with, say, formative assessment, provide learning there — not generic training). And it means treating teachers as professionals in a partnership, which turns evaluation 'from a moment of anxiety into' a genuine tool for growth. Rankings can inform decisions; only feedback and support actually improve teaching.

🎯
Section 12

Evaluator Reliability & Bias

An uncomfortable truth: classroom observation is often less reliable than the VAM everyone worries about. Evaluators vary wildly, and bias creeps in. Calibration and self-awareness are essential.

The reliability problem
  • Observation is a 'statistical Wild West' — highly variable
  • Different observers rate the same lesson differently
  • Ratings often inflated or inconsistent
  • Bias (halo, similarity, leniency) creeps in
Improve reliability & fairness
  • Train and calibrate observers (norm to the rubric)
  • Check inter-rater reliability
  • Guard against your own biases
  • But note: scoring skill ≠ skill at developing teachers
Observation is often less reliable than VAM — calibrate, and watch your bias

Here's a truth that surprises many: for all the (justified) worry about the variability of value-added scores, classroom observation is arguably less reliable in practice — one expert called it a 'statistical Wild West.' Different observers watching the same lesson often assign very different ratings; scores drift high (inflation) and inconsistent; and evaluator bias (the halo effect, favoring teachers similar to oneself, general leniency) quietly distorts judgments. The remedies: rigorously train and calibrate observers (norm them to the rubric with shared exemplars), check inter-rater reliability, and actively guard against your own biases. But one crucial caveat: an observer can get better at applying a rubric without getting better at helping a teacher improve — the gap between calibrated scoring and meaningful developmental feedback is often large. Pursue both reliable scoring and genuine feedback skill.

🔀
Section 13

Evaluation vs. Coaching: Keep Them Distinct

A crucial design principle: evaluation (judging) and coaching (developing) are different jobs, and blending them undermines both. Where possible, keep the roles — and the people — distinct.

Two different roles
  • Evaluation — judges performance (accountability)
  • Coaching — develops practice (non-evaluative)
  • Teachers won't be vulnerable with their judge
  • Ideally, different people play the two roles
Why separation matters
  • A teacher hides struggles from the person who rates them
  • Coaching needs safety; evaluation brings judgment
  • Blurring the roles produces superficial feedback
  • Even growth-oriented evaluation isn't coaching
Judging and developing are different jobs — keep them distinct where you can

A principle that connects this toolkit to our Instructional Coaching toolkit: evaluation and coaching are fundamentally different roles. Evaluation judges performance for accountability; coaching develops practice through a safe, non-evaluative partnership. The problem with blending them is well-documented: a teacher will not honestly reveal their struggles or take real risks with the person who controls their employment — so when the evaluator is also the 'coach,' trust collapses and the developmental conversation becomes superficial. That's why, where possible, the two roles should be played by different people (the principal evaluates; a non-evaluative coach develops). This doesn't mean evaluation can't be growth-oriented — it should be (§11) — but even the most supportive evaluation isn't the same as coaching, and pretending otherwise shortchanges both. See our Instructional Coaching toolkit.

🪞
Section 14

Self-Evaluation & Goal-Setting

Teachers aren't passive objects of evaluation — the most effective systems make them active participants through self-assessment, reflection, and goal-setting. Agency drives growth.

Engage teachers actively
  • Include teacher self-assessment and reflection
  • Have teachers set their own growth goals
  • Make it a partnership, not something done to them
  • Ownership fuels genuine improvement
Support teacher agency
  • Teachers reflecting on their own practice grow more
  • Goal-setting focuses development meaningfully
  • Combine self-assessment with observation & feedback
  • Professionals own their growth
Make teachers active participants — not passive objects of evaluation

Evaluation done to teachers rarely develops them; evaluation done with them does. The most effective systems engage teachers as active participants through self-assessment (teachers reflecting honestly on their own practice against the framework), reflection, and goal-setting (teachers identifying, with support, what they want to work on). This isn't just pleasant — it's more effective: professionals who take ownership of their growth, reflect on their own teaching, and pursue goals they helped set improve more than those who simply receive a rating. Pairing teacher self-assessment and goals with observation and feedback creates a genuine partnership, treats teachers as the professionals they are, and puts the energy of the whole process where it belongs — on the teacher's own commitment to getting better.

🗣️
Section 15

Difficult Conversations & Underperformance

Sometimes evaluation surfaces real problems that must be addressed honestly. Avoiding difficult conversations — the very failure the Widget Effect exposed — helps no one. Have them well.

Address real problems honestly
  • Don't avoid hard feedback — that's the Widget Effect
  • Be specific and evidence-based, not vague
  • Combine honesty with genuine support
  • Create a clear improvement plan and timeline
Have the conversation well
  • Be direct and respectful — clarity is kindness
  • Ground it in evidence, not opinion
  • Offer real support to improve
  • Document fairly and follow a fair process
Avoiding hard feedback is the Widget Effect — have the conversation, honestly and supportively

When evaluation reveals genuine underperformance, the temptation is to avoid the hard conversation — to soften, to rate 'satisfactory,' to move on. But that avoidance is precisely the Widget Effect that failed everyone: it denies the teacher the honest feedback they need to improve, denies students an effective teacher, and is ultimately disrespectful. Difficult conversations, done well, combine honesty with support: be direct and respectful (clarity is a form of kindness — vagueness leaves people confused and stuck), ground everything in specific evidence rather than opinion, offer real support and resources to improve, establish a clear improvement plan with a reasonable timeline, and document fairly throughout. The goal is genuine improvement wherever possible. Facing real problems honestly and supportively is hard — and it's exactly what the job requires.

🚪
Section 16

When a Teacher Isn’t Effective

The hardest reality of evaluation: sometimes, despite genuine support, a teacher isn't effective and can't become so. Handling this with both fairness and resolve is among leadership's most difficult duties.

The hard reality
  • Some teachers, despite support, aren't effective
  • Students deserve effective teachers
  • Ignoring this is the Widget Effect (§03)
  • Dismissal is a last resort — after genuine support
Handle it fairly
  • Provide real support and a fair improvement process first
  • Document thoroughly and consistently
  • Follow due process and your policies/contracts
  • Balance care for the teacher with duty to students
Some teachers can't become effective — act with both fairness and resolve

This is the part of evaluation no one enjoys but leadership can't dodge: occasionally, despite genuine support, feedback, and time, a teacher isn't effective and isn't going to become so — and students, who get only one shot at each grade, deserve better. Failing to act (the Widget Effect) protects an adult at children's expense. But acting fairly is non-negotiable: dismissal should be a last resort, reached only after providing real support and a genuine chance to improve (a fair improvement plan, resources, time), with thorough and consistent documentation, full due process, and strict adherence to district policies and contracts. Hold the tension between genuine care for the teacher as a person and your duty to students — both are real. This is among the hardest things a school leader does; do it with integrity, fairness, and resolve. See our School Leadership toolkit.

🎚️
Section 17

Low-Stakes vs. High-Stakes: The Debate

A central design question: should evaluation be high-stakes (tied to pay, tenure, dismissal) or low-stakes (focused on growth)? The evidence increasingly favors keeping the developmental work low-stakes.

The trade-off
  • High-stakes — strong accountability, but breeds defensiveness
  • Low-stakes — more honesty, reflection & growth
  • High stakes make teachers hide problems, not fix them
  • Satisfaction & trust drop as stakes rise
Lean toward growth
  • Compliance-driven, high-stakes evaluation triggers defensiveness
  • When teachers feel it's unfair, they disengage from feedback
  • The trend is away from high-stakes, test-based systems
  • Keep the developmental work as low-stakes as possible
High stakes make teachers defensive — keep the growth work low-stakes

A key design tension: the higher the stakes attached to evaluation (pay, tenure, dismissal), the stronger the accountability pressure — but also the more it backfires for growth. Research shows that satisfaction with evaluation declines as it becomes heavily compliance-driven and high-stakes, and that when teachers perceive the process as unfair, they disengage from the feedback itself — triggering defensiveness rather than reflection, and making professional growth harder, not easier. High stakes also incentivize hiding problems rather than surfacing them. This is a big part of why the field has moved away from the high-stakes, test-based systems of the reform era. The wise approach keeps the developmental work as low-stakes as possible (so teachers can be honest and reflective), while reserving genuine accountability consequences for the relatively rare cases that truly require them (§15–§16).

🚧
Section 18

Common Pitfalls

Teacher evaluation fails in predictable ways — most of them versions of either the Widget Effect (too soft, no growth) or the reform era's overreach (too high-stakes, distorting). Knowing them helps you steer between them.

Common pitfalls
  • A compliance checkbox — going through the motions
  • The Widget Effect — inflating ratings, avoiding truth
  • Gotcha evaluation — high-stakes, punitive, distorting
  • Ratings without growth — a score but no development
...and more
  • Relying on a single observation or measure
  • Uncalibrated, inconsistent, biased observers (§12)
  • Over-relying on test-based VAM (§09)
  • Getting good at scoring but not at developing teachers
Most failures are either the Widget Effect or the reform era's overreach — steer between them

Teacher evaluation tends to fail in two opposite directions, and both are traps. On one side, the Widget Effect: evaluation as a soft compliance ritual — inflated ratings, avoided hard conversations, no meaningful feedback, no growth, no accountability. On the other, the reform era's overreach: high-stakes, test-heavy, 'gotcha' evaluation that breeds defensiveness, gaming, and distrust while doing little to improve teaching. The other classic pitfalls cluster around these: relying on a single observation or measure, using uncalibrated and biased observers, over-relying on shaky VAM, and — subtly — producing ratings without development (the whole point). The path between the traps is the theme of this toolkit: credible, multiple-measure, growth-oriented evaluation that's honest about performance and genuinely helps teachers improve.

🧭
Section 19

For Leaders: Evaluating Teachers Well

Evaluation lives or dies by how leaders conduct it. Evaluating teachers well means centering growth, building credibility, and holding the accountability-and-development tension with skill and integrity.

Evaluate for growth
  • Center GROWTH over ranking (§11)
  • Give specific, actionable, credible feedback (§10)
  • Use multiple measures, not one (§07)
  • Connect evaluation to coaching and targeted PD
Do it with credibility & integrity
  • Build your own skill at observing & giving feedback
  • Calibrate; guard against bias (§12)
  • Keep the developmental work low-stakes (§17)
  • Face underperformance honestly and fairly (§15–§16)
Evaluate to develop teachers — with credibility, fairness, and honesty

Leaders determine whether evaluation improves teaching or becomes a hollow ritual (or a feared weapon). Evaluating teachers well means: centering growth over ranking; giving specific, actionable, and credible feedback (which requires building your own skill at observing teaching and coaching — a skill many evaluators lack); using multiple measures rather than a single observation or test score; connecting evaluation to coaching and targeted support so feedback becomes improvement; keeping the developmental work as low-stakes as possible; calibrating your judgments and guarding against bias; and, when necessary, facing underperformance honestly and fairly. Above all, treat teachers as professionals and the whole process as a partnership for better teaching. Get this right, and evaluation becomes one of your most powerful tools for improving learning. See our School Leadership, Instructional Coaching, and Professional Growth toolkits.

🔎
Section 20

Resources & K12academics

Teacher evaluation has strong (and contested) research and resources. Here's where to go deeper, plus K12academics for the wider world of education.

Frameworks & research
  • Danielson Group — the Framework for Teaching
  • Marzano Resources — the Marzano evaluation model
  • TNTP — 'The Widget Effect' & evaluation research
  • The MET project & reform-era research (read critically)
Practical & K12academics
  • Learning Forward, ASCD & Edutopia — evaluation strategies
  • IES — evaluation & professional growth research
  • Our Instructional Coaching, Professional Growth & Leadership toolkits
  • K12academics — the wider world of education
Evaluate to develop — center growth, use multiple measures, and be honest

For frameworks, the Danielson Group and Marzano Resources offer the leading observation models; TNTP's 'Widget Effect' is the essential (and reform-era research the essential critical) reading. Learning Forward, ASCD, and Edutopia offer practical strategies. Keep the through-line in view: center growth over ranking, use multiple measures, give credible developmental feedback, keep coaching distinct from evaluation, and be honest about performance. Pair this with our Instructional Coaching, Professional Growth & PD, and School Leadership toolkits. Start at K12academics.com.

Section 21

Toolkit Checklists

Six checklists for credible, growth-oriented teacher evaluation. Click any box to check it off; your progress stays for this session. Tap one to open it.

Learn From the History
Balance the Two Purposes
Use Sound Measures
Observe & Give Feedback Well
Keep Roles & People Right
Handle the Hard Parts Honestly
📥
Section 22

Downloads & Templates

Templates and guides referenced throughout this toolkit, ready to use in your evaluation practice.

Understand
  • The Widget Effect & reform-era summary
  • Accountability-vs-growth explainer
  • Multiple-measures framework one-pager
  • Value-added-measures (cautions) guide
Observe
  • Observation cycle (pre/observe/post) template
  • Evidence-gathering (not judgment) guide
  • Observer calibration & bias-check tool
  • Self-assessment & goal-setting template
Feedback & growth
  • Feedback-that-develops guide
  • Growth-oriented (formative + summative) framework
  • Evaluation-vs-coaching role clarity guide
  • Evaluation-to-PD connection tool
Hard parts & lead
  • Difficult-conversations guide
  • Improvement-plan & due-process template
  • Common-pitfalls (avoid these) checklist
  • Evaluating-teachers-well (leaders) checklist
Get the editable versions

Editable versions of these guides are available on request — see §26, Stay Connected.

👥
Section 23

Communities & Resources

Teacher evaluation has a substantial (and contested) research base. These are trusted places to learn and go deeper — read across perspectives.

Frameworks & models
The research & critique
  • TNTP — 'The Widget Effect'
  • RAND — evaluation of the Gates reforms
  • Research on VAM's limits & instability
  • Danielson's own reconsideration of her framework
Practical & feedback
  • ASCD & Edutopia — evaluation & feedback strategies
  • IES — evaluation & professional growth research
  • Observation & feedback protocols
  • Your evaluation team & mentors
Go deeper (companion toolkits)
📱
Section 24

QR Resource Hub

Scan any code below with your phone camera — perfect for a printed copy of this toolkit. The first codes go to leading teacher-evaluation resources.

Danielson Group QR code
Danielson Group

The Framework for Teaching.

TNTP QR code
TNTP

'The Widget Effect' & evaluation research.

Learning Forward QR code
Learning Forward

Standards for Professional Learning.

Explore K12academics QR code
Explore K12academics

Education resources & directories.

This Week in Education QR code
This Week in Education

Our weekly roundup for educators.

Join the Newsletter QR code
Join the Newsletter

Education news and resources.

State of Education Reports QR code
State of Education Reports

Free 2026 research reports.

Contact Us QR code
Contact Us

Questions or ideas for the next edition.

🛠️
Section 25

K12academics Resource Center

Beyond this toolkit, here's what K12academics offers educators, leaders, and families — much of it free.

📞
Section 26

Stay Connected

Ways to stay in touch with K12academics — and to help shape the next edition of this toolkit.

Subscribe
Contribute
  • Nominate a resource for a future edition
  • Request the editable guides from §22
  • Tell us what leaders need
  • Contact us
Follow
📚
Section 27

Sources & Further Reading

The findings in this toolkit come from teacher-evaluation research and a genuinely contested reform history. Read across perspectives, keep the focus on growth, and adapt everything to your context and policies.

The history & research
  • TNTP (2009) — 'The Widget Effect'
  • The MET project (Gates) — measures of effective teaching
  • RAND — evaluation of the Gates Intensive Partnerships (little effect)
  • The accountability-vs-growth literature (10-year review)
Frameworks & measures
Feedback & growth
  • IES — evaluation & professional growth research
  • Research: feedback + coaching improves practice (RCT)
  • Learning Forward · ASCD · Edutopia
  • Formative vs. summative evaluation
Go deeper (companion toolkits)

Teacher evaluation is a contested area with a fraught reform history; this toolkit presents that history honestly (including what didn't work) and centers growth-oriented, multiple-measure practice. Findings reflect the research above (the Widget Effect, MET, the accountability-growth literature) as of the 2026–2027 school year. This toolkit is an evidence-informed professional resource, not prescriptive — adapt it to your context, policies, and contracts.

🚀
Signature Feature

100 Teacher-Evaluation Tips

Everything above, distilled into 100 quick, practical reminders for evaluators and leaders. Twenty categories, five tips each.

Why Evaluation Matters
  1. Teaching is the biggest in-school factor.
  2. Ensuring and developing good teaching matters.
  3. But evaluation is genuinely hard.
  4. 'Good teaching' is hard to define.
  5. Take the difficulty seriously.
The Widget Effect
  1. 99% of teachers were rated 'satisfactory.'
  2. Old systems didn't differentiate anyone.
  3. Evaluation was infrequent and low-standard.
  4. Excellence unrecognized; poor performance unaddressed.
  5. Don't treat teachers like widgets.
The Reform Era
  1. Race to the Top pushed high-stakes VAM.
  2. Billions were spent.
  3. The reforms had little effect on outcomes.
  4. Many states rolled VAM back.
  5. Accountability alone doesn't improve teaching.
Accountability vs. Growth
  1. Evaluation serves two purposes.
  2. They don't always work together.
  3. Accountability metrics alone fail to improve teaching.
  4. Evaluation works best as a developmental tool.
  5. Balance both — but center growth.
Observation Frameworks
  1. Danielson's Framework is the most widely used.
  2. Marzano's is another.
  3. A framework is a shared language, not an answer.
  4. Even Danielson warns hers is misused.
  5. Don't use it as a rigid checklist.
Multiple Measures
  1. No single measure captures teaching.
  2. Use multiple observations over time.
  3. Add artifacts, self-assessment, reflection.
  4. You can't differentiate on observation alone.
  5. Balance guards against distortion.
The Observation Process
  1. Run it as a growth cycle: pre, observe, post.
  2. Gather evidence, not snap judgments.
  3. Observe multiple times.
  4. Look at learning, not just performance.
  5. Make the post-conference about growth.
Value-Added Measures
  1. VAM is a statistical estimate.
  2. It relies on standardized tests.
  3. Scores can be unstable and unfair.
  4. It can't tell a teacher how to improve.
  5. Treat it with caution — one limited measure at most.
Feedback That Develops
  1. The point is feedback, not a rating.
  2. Be specific, clear, and actionable.
  3. Make it collaborative, not handed down.
  4. Credibility matters — build your skill.
  5. Pair feedback with coaching.
Growth-Oriented Evaluation
  1. Emphasize growth over ranking.
  2. Center formative feedback throughout.
  3. Connect evaluation to targeted PD.
  4. Turn evaluation from anxiety into partnership.
  5. Feedback improves; rankings only inform.
Reliability & Bias
  1. Observation is often less reliable than VAM.
  2. Different observers rate the same lesson differently.
  3. Calibrate observers to the rubric.
  4. Guard against halo, similarity, leniency bias.
  5. Scoring skill isn't the same as developing teachers.
Evaluation vs. Coaching
  1. Judging and developing are different jobs.
  2. Teachers hide struggles from their judge.
  3. Use different people where possible.
  4. Coaching needs safety; evaluation brings judgment.
  5. Even growth-oriented evaluation isn't coaching.
Self-Evaluation & Goals
  1. Engage teachers as active participants.
  2. Include self-assessment and reflection.
  3. Have teachers set their own goals.
  4. Make it a partnership, not done-to.
  5. Ownership fuels growth.
Difficult Conversations
  1. Avoiding hard feedback is the Widget Effect.
  2. Be specific and evidence-based.
  3. Combine honesty with support.
  4. Create a clear improvement plan.
  5. Clarity is a form of kindness.
When a Teacher Isn't Effective
  1. Some teachers, despite support, aren't effective.
  2. Students deserve effective teachers.
  3. Dismissal is a last resort.
  4. Provide support and a fair process first.
  5. Balance care with duty to students.
Low-Stakes vs. High-Stakes
  1. High stakes breed defensiveness.
  2. Satisfaction drops as stakes rise.
  3. Unfairness makes teachers disengage from feedback.
  4. The trend is away from high-stakes.
  5. Keep the growth work low-stakes.
Common Pitfalls
  1. The Widget Effect: too soft, no growth.
  2. The reform era: too high-stakes, distorting.
  3. A single observation or measure.
  4. Ratings without development.
  5. Getting good at scoring, not developing.
For Leaders
  1. Center growth over ranking.
  2. Give specific, credible feedback.
  3. Use multiple measures.
  4. Connect evaluation to coaching and PD.
  5. Face underperformance honestly and fairly.
The Teacher Experience
  1. How teachers experience evaluation matters.
  2. Fairness determines whether they engage.
  3. Unfair processes trigger defensiveness.
  4. Treat teachers as professionals.
  5. Make it a partnership for better teaching.
Mindset
  1. Center growth over ranking.
  2. Use multiple measures.
  3. Give developmental feedback.
  4. Keep coaching distinct from evaluation.
  5. Be honest about performance.