AI in teacher evaluations — the observation (watching a classroom) versus the write-up (a rubric-aligned evaluation drafted on a tablet).

AI in Teacher Evaluations: The Line Between Watching and Writing

July 22, 20269 min read

AI in Teacher Evaluations: The Line Between Watching and Writing

AI in teacher evaluations — the observation (watching a classroom) versus the write-up (a rubric-aligned evaluation drafted on a tablet).

By Anthony Neely, Ph.D. — co-founder, EvalScribe

A Bronx assistant principal recently described what AI has done for her evaluations in a single sentence to GovTech: it "probably cuts my observation time in half, or the write-up part in half at this point."

Read that sentence again, because the small word doing all the work is or. Cutting observation time in half and cutting the write-up in half are not the same accomplishment. They are two different jobs that happen to share one word — "evaluation" — and they carry almost opposite risk profiles. For any district leader deciding how AI belongs in their evaluation process, the entire question lives in that or.

Two jobs hiding inside one word

Every teacher evaluation is really two acts stitched together.

The first is the observation: a trained human sitting in a room, watching instruction unfold, and exercising professional judgment about what is and isn't happening — the read of a class that went sideways for reasons that never appear in a transcript, the distinction between a quiet room that's disengaged and a quiet room that's deep in thought. This is judgment built on context, and context is exactly what a model sitting outside the room does not have.

The second is the write-up: taking the evidence a human already gathered and translating it into the language of a rubric — Danielson, your state's framework, whatever instrument your district has adopted — so the feedback is specific, aligned, and defensible. This is translation work. It is slow, it eats evenings, and it is the part most evaluators would happily hand off.

When you hear that AI "cuts the evaluation in half," the honest follow-up is always: which half?

Why the distinction is the whole ethical question

The American Federation of Teachers' Rob Weil, quoted in the same GovTech piece, cautions that leaning on AI here can be premature — his concern is precisely that these tools struggle to capture classroom context and that districts owe teachers transparency about how they're used. He's right, and the way to take his caution seriously is not to ban AI from evaluation — it's to refuse to let it do the watching.

Point AI at the observation, and you've outsourced the one part of the process that depends on a human being present and accountable. No teacher should learn that a rating of their craft was generated by a system that never saw their students. That's the version of "AI in evaluations" that deserves the skepticism.

Point AI at the write-up, and none of that is true. The judgment stays human. The evidence stays human. What changes is that the translation from "here's what I saw" to "here's how it maps to Domain 3b" stops costing an administrator their Sunday.

But the tool you choose for the write-up matters just as much

Here's where a lot of well-meaning administrators quietly take on risk. In that same reporting, the way the write-up gets done is by uploading a Danielson rubric into ChatGPT and asking it for suggestions. It works — but a general-purpose chatbot is the wrong instrument for a job this specific, for three concrete reasons.

You re-do the alignment every single time. A general chatbot doesn't know your framework. You have to hand it the rubric, explain the scoring, and hope it holds the thread — every observation, all over again. That's prompt engineering, and you've quietly made every evaluator in your building a part-time prompt engineer. A purpose-built tool has the framework built in. EvalScribe supports the frameworks already in use across all 50 states — Danielson, Marzano, Stronge, CEL 5D+, NIET, and state-specific instruments like Georgia's TKES and Tennessee's TEAM  (and Project Coach)— with the rubric hardcoded into the app, so there's no prompt engineering required.

The output drifts. Ask a general model the same thing twice and you can get two different answers, two different reading levels, two different scoring rationales. Across a stack of write-ups and a whole district, that inconsistency is exactly what makes an evaluation process feel arbitrary — and it's what a grievance latches onto. A tool built for this applies consistent, rubric-based scoring with evidence attached, so the language your first-year assistant principal produces and the language your thirty-year veteran produces are speaking the same instrument.

It was never built for the room. EvalScribe lets an evaluator capture the way they actually work — type, dictate, or scan handwritten notes — on the iPhone, iPad, or Mac they already carry into the classroom, then turns fragmented notes into evaluation-ready language and even drafts post-conference coaching scripts. The result administrators report is a sixty-minute write-up dropping to ten or less — on the order of 750 hours a year given back across a typical district.

The part nobody thinks about until it's a problem: where the notes go

Pause on what actually gets pasted into a general chatbot during an evaluation: candid, identifying notes about a named teacher's performance. Consumer AI tools, by default, may retain those inputs and use them to improve future models. For personnel records, that should stop you cold.

This is where EvalScribe's posture is built for districts rather than retrofitted. In its own words, on the site and in its privacy policy:

  • Local-first. "All data resides physically on your device. We do not maintain a central cloud database."

  • Transient processing. "Your data is sent securely to the engine, processed in memory, and returned."

  • Nothing kept on the server. "Neither the notes nor the evaluation are stored on our servers."

  • Training opt-out. "Our Azure instance is configured to opt out of all model training."

And the certifications your IT director will ask about belong to the infrastructure it's built on — which EvalScribe is upfront about. It runs on Microsoft Azure OpenAI Service, the same enterprise-grade infrastructure used by major banks and healthcare systems, which carries SOC 1, 2 & 3, ISO 27001, FedRAMP High, and HIPAA compliance. If your district needs paperwork, a Data Processing Agreement and a signed FERPA-aligned attestation are available on request. On FERPA itself: the tool is built to evaluate teachers — adults — not students, and does not maintain student records, with a standing recommendation to keep full student names out of notes.

That's not a marketing checklist. It's the difference between a workflow you can defend to your board and your teachers' association, and one you'd rather no one asked about.

Built by evaluators, vouched for by them

EvalScribe wasn't built by a general AI company that discovered schools as a market. It was co-founded by Anthony Neely, Ph.D., a district Teacher of the Year, and Andrea Neely, Ph.D., with a combined thirty years in classrooms — people who have written these evaluations by hand and know exactly which half of the job is worth automating.

And the people signing off are the people doing the work. A district assistant superintendent with more than forty years in education calls it a "game changer"; a principal says every public-school principal should be using it. For districts, that confidence is backed by a rollout designed to be painless — no IT integration, no rostering, one access code — so a decision doesn't turn into an implementation project. You can watch the demo and, if you want to feel it on a real observation, try three evaluations free — no card required.

The permission worth naming

Here's the reframe underneath all of this. You are allowed to use AI for the paperwork. You are not obligated to pretend it can do the watching — and you are not required to do the paperwork half inside a consumer chatbot that was never built for personnel records.

Most of the anxiety around "AI in teacher evaluations" comes from collapsing two decisions into one: should AI touch this, and which tool should do it. Pull them apart and the answers get clear fast. Point AI at the write-up, not the watching. Then use a tool that has your framework built in, keeps its outputs consistent, and keeps your teachers' data off the training pile.

The Bronx assistant principal's or wasn't a hedge. It was the whole answer. The districts that get this right will be the ones that heard it — and picked the instrument to match.

See the framework library for your instrument at evalscribe.com/frameworks, size the time your team is spending with the district capacity calculator, or reach the team directly at [email protected].

Frequently asked questions

Is it ethical to use AI for teacher evaluations? It depends entirely on which part of the evaluation the AI performs. Using AI to render professional judgment about a classroom observation is hard to defend, because the model lacks the context a present, accountable human has. Using AI to translate an evaluator's own gathered evidence into rubric-aligned language is a clerical time-saver that leaves the judgment fully human.

How is EvalScribe different from just using ChatGPT? A general chatbot has no built-in knowledge of your evaluation framework, so you re-supply the rubric and scoring every time, and its outputs vary from one run to the next. EvalScribe has the frameworks for all 50 states hardcoded in, applies consistent rubric-based scoring with evidence, needs no prompt engineering, and is built around how evaluators actually capture notes — typed, dictated, or scanned — on iPhone, iPad, and Mac.

Is EvalScribe secure? What happens to my evaluation data? EvalScribe is local-first: all data resides on your device, not in a central cloud database, and neither the notes nor the generated evaluation are stored on its servers. Observations are processed in memory on Microsoft Azure and returned, and its Azure instance is configured to opt out of all model training. The compliance certifications (SOC 1/2/3, ISO 27001, FedRAMP High, HIPAA) belong to the underlying Azure OpenAI infrastructure. A Data Processing Agreement and a signed FERPA-aligned attestation are available on request.

Does AI replace the administrator in a teacher evaluation? No. In an honestly scoped process, the administrator still observes, exercises judgment, and signs the evaluation. AI assists with the write-up — aligning evidence to the framework and keeping language consistent — not with the rating itself.

Can AI write teacher evaluations aligned to the Danielson Framework? Yes — aligning gathered evidence to a specific framework like Danielson is exactly the write-up task purpose-built AI is well suited to, provided the evaluator's own evidence and judgment come first.


Source: "AI for Teacher Evaluations: Major Time-Saver, or Premature?", GovTech, March 21, 2026.

Anthony D. Neely, Ph.D.

Anthony D. Neely, Ph.D.

Anthony Neely is the Founder of EvalScribe, a veteran educator, an AI integration consultant for teaching & learning, researcher, & author.

Back to Blog