
How can AI help administrators write Kim Marshall teacher evaluations?
How can AI help administrators write Kim Marshall teacher evaluations?

The Marshall rubric was built around a particular habit: get into classrooms often, for short unannounced visits, and follow each one with a quick conversation. Marshall is explicit that the year-end rubric should never be filled in from a single formal observation — it is meant to reflect a year's worth of mini-observations. That habit is what makes the rubric powerful, and it is also what generates the paperwork, because observing often means writing often. The part an evaluator writes most is the mini-observation write-up against the criteria. That's where AI can genuinely help: drafting the write-up, mapping evidence to the right criterion, and putting it in the rubric's own language. What it can't do is produce your district's summative rating. This post walks through where AI helps with Marshall, where it doesn't, and what to look for.
What the Marshall rubric is
The Kim Marshall Teacher Evaluation Rubric was created by Kim Marshall and is published through the Marshall Memo. It is designed to give teachers a year-end assessment of where they stand across every part of the job, informed by frequent mini-observations rather than a single visit. It is open source, and schools and districts are free to adapt it to fit their context.
The six domains
The rubric organizes teaching into six domains, with fifty-four criteria in all:
Planning & Preparation for Learning covers content expertise, standards-aligned goals and units, assessments, anticipating misconceptions, and an organized environment.
Classroom Management covers expectations, relationships, respect, routines, prevention of problems, and productive use of class time.
Delivery of Instruction covers clear objectives, engaging and rigorous lessons, connections, effective questioning, and a repertoire of strategies.
Monitoring, Assessment & Follow-Up covers checking for understanding, timely feedback, recognition, and using data to adjust and support every student.
Family & Community Outreach covers respect for families, clear communication, involving families in learning, and responsiveness to concerns.
Professional Responsibilities covers reliability, professional judgment, teamwork, openness to feedback, and ongoing growth and leadership.
The rating scale
Each criterion is rated on a four-level scale: Does Not Meet Standards, Improvement Necessary, Effective, and Highly Effective. Marshall is deliberate about the middle of that scale. Effective is solid, expected professional performance, and he is clear that teachers should feel good about it. Highly Effective is reserved for truly outstanding practice against very demanding criteria. Treating Effective as the real bar, rather than a consolation prize below the top, is central to using the rubric well.
Mini-observations: the habit the rubric is built on
The rubric's defining feature is the observation cadence behind it. Instead of one long formal visit, Marshall calls for short, unannounced mini-observations every few weeks, each followed by a face-to-face conversation. Over a year, those visits build a rich, accurate picture that a single observation could never provide. The catch is documentation: observing every teacher often means writing far more often, and that volume is exactly what wears principals down.
Where AI helps with the Marshall write-up
The strongest fit between current AI and the Marshall rubric is that high-volume, criteria-based documentation.
AI handles three parts of it reliably. First, it turns fragmentary mini-observation notes into coherent, evidence-anchored prose in the rubric's language. Second, it maps evidence to the right criterion across all six domains, including evidence that speaks to more than one. Third, it holds consistency across a caseload and across many visits, so the same quality of evidence lands in similar territory from teacher to teacher and week to week.
For Marshall specifically, the value is speed at volume. The rubric only works if you observe often, and the documentation load is what makes "often" hard to sustain. Cutting the write-up time per visit is what keeps the mini-observation habit alive.
Where AI doesn't help — and what stays with the evaluator
The honest scope follows from how the rubric is used.
AI cannot produce the summative rating, and it cannot supply any student-growth measures your district folds in. The Marshall rubric is a year-end judgment assembled from many visits and conversations, and how it is scored and combined is a district decision that sits outside any drafting tool.
AI cannot supply an evaluator's professional judgment about what a pattern across many visits really shows, or whether practice has reached Effective or only Improvement Necessary. That judgment is the point of the rubric, and it stays with the evaluator, who remains the last set of eyes on every rating and comment.
And general-purpose AI has no built-in understanding of Marshall. Paste notes into a consumer chatbot and it will invent criteria, ignore the mini-observation cadence, and hand out Highly Effective as if it simply meant good teaching. For a rubric this specific, those errors show.
What to look for in an AI tool for Marshall
A few questions worth asking before committing a tool to this work.
Does the tool actually know the Marshall rubric — its six domains and its four-level scale — or is it a generic writing assistant with Marshall vocabulary sprinkled in?
Is it built for speed at the mini-observation cadence, so documenting a short visit takes a minute rather than derailing your afternoon?
Does it treat Effective as the expected level rather than defaulting to the top of the scale?
Where does your observation data live — on the vendor's servers, or used to train models?
How EvalScribe handles the Marshall rubric
EvalScribe is built around the Kim Marshall Teacher Evaluation Rubric — all six domains, their criteria, and the four-level scale — and it is built for the mini-observation workflow the rubric depends on. An evaluator captures notes by typing, dictating, or photographing handwriting (Smart Scan OCR converts it to text), EvalScribe maps that evidence to the criterion it supports, and drafts a best-fit rating and evidence-anchored feedback in the rubric's own language. Every rating and comment is fully editable before you finalize it, and the evidence stays traceable to the note it came from.
Two scope notes, because the rubric calls for them. First, EvalScribe drafts the observation-based write-up. The overall summative rating and any student-growth measures are set through your district's process, not the app. Second, EvalScribe is an independent tool; it is not affiliated with or endorsed by Kim Marshall or the Marshall Memo, and it supports administrators who already use the rubric rather than replacing an official scoring platform.
Beta testers report saving 30 to 60 minutes per evaluation versus writing the documentation by hand. For a rubric built on frequent visits, that time compounds — every mini-observation you document faster is one more you can afford to do. More detail on the Marshall workflow is available at evalscribe.com/marshall.
Frequently asked questions about the Marshall rubric and AI
Does EvalScribe support the Kim Marshall rubric? Yes — all six domains, their criteria, and the four-level scale, drafting in the rubric's own language.
What are the six domains? Planning & Preparation for Learning, Classroom Management, Delivery of Instruction, Monitoring/Assessment & Follow-Up, Family & Community Outreach, and Professional Responsibilities — fifty-four criteria in all.
What are mini-observations? Short, unannounced visits every few weeks, each followed by a quick conversation. The year-end rubric is filled in from that accumulated evidence.
What are the rating levels? Does Not Meet Standards, Improvement Necessary, Effective, and Highly Effective. Effective is the expected level.
Is EvalScribe affiliated with Kim Marshall? No. It is an independent drafting tool that supports administrators who already use the rubric.
If your district uses the Marshall rubric, see how EvalScribe drafts across its six domains at evalscribe.com/marshall. Questions, or a school or district license? Reach the team at [email protected].
References
Kim Marshall, The Marshall Memo (the Teacher Evaluation Rubric, its rationale, and implementation guidance)
Kim Marshall, Rethinking Teacher Supervision and Evaluation (the case for frequent mini-observations)
District and state-hosted copies of the current open-source Marshall rubric (public domain, criteria, and scale documentation)
Related articles
Page maintained by Anthony D. Neely, Ph.D. — practicing K-12 educator with nearly 20 years in the classroom, 2025–2026 Walker County Distinguished Teacher of the Year, and co-founder of EvalScribe. Framework details verified against Kim Marshall's published Teacher Evaluation Rubric.
Last updated on August 6, 2026.
