The Stronge Teacher Effectiveness Performance Evaluation System uses seven performance standards, rated Ineffective to Highly Effective on a preponderance of evidence. Where AI helps with the write-up, where it doesn't, and what to look for.

How can AI help administrators write Stronge teacher evaluations?

August 09, 20267 min read

How can AI help administrators write Stronge teacher evaluations?

The Stronge Teacher Effectiveness Performance Evaluation System uses seven performance standards, rated Ineffective to Highly Effective on a preponderance of evidence. Where AI helps with the write-up, where it doesn't, and what to look for.

The Stronge model is built on a simple idea with a demanding consequence: an evaluator's job is to assemble a fair performance portrait of a teacher across the whole year, drawing on many sources of evidence rather than a single visit. That is what makes it rigorous, and it is also what makes the documentation heavy. The part an evaluator writes most is the observation-based write-up against the model's standards. That's where AI can genuinely help: drafting the write-up, mapping evidence to the right standard, and putting it in the model's own language. What it can't do is supply the Student Progress data or produce the summative rating. This post walks through where AI helps with Stronge, where it doesn't, and what to look for.

What the Stronge model is

The Stronge Teacher Effectiveness Performance Evaluation System, or TEPES, was developed by Dr. James H. Stronge of Stronge & Associates. It is a research-based framework used by districts and states across the country, and it uses a two-tiered structure: seven performance standards, each supported by multiple performance indicators that describe what the standard looks like in practice.

The seven performance standards

Six of the standards are research-based professional standards, and the seventh is results-based:

Professional Knowledge — understanding of the curriculum, subject content, pedagogy, and the developmental needs of students.

Instructional Planning — planning with standards, curriculum, effective strategies, resources, and data to meet learning needs.

Instructional Delivery — engaging students through a variety of instructional strategies that meet individual learning needs.

Assessment of/for Learning — gathering, analyzing, and using data to measure progress and guide instruction.

Learning Environment — creating a safe, respectful, and well-managed environment that supports learning.

Professionalism — ethical conduct, collaboration, communication with families, and ongoing professional growth.

Student Progress — the results-based standard: the work of teaching should result in acceptable, measurable student academic progress.

The rating scale

Teachers are rated on each standard using a four-level scale: Ineffective, Partially Effective, Effective, and Highly Effective. Effective is the expected level of performance — consistent practice, aligned with the school's mission, with a positive impact on student learning. Highly Effective reflects sustained, high-impact practice that can serve as a model for peers. Knowing that Effective, not Highly Effective, is the bar matters for rating fairly.

Preponderance of evidence

A defining feature of the Stronge model is how a rating is reached. Rather than scoring a single moment, the evaluator gathers evidence across the year from multiple sources — observations, artifacts, goal setting and student learning objectives, and sometimes surveys — and determines where the preponderance of evidence places the teacher on each standard. The observation write-up is a major part of that evidence, but it is one part of a larger portrait.

After the observation: where the work piles up

The observation itself is familiar work for an evaluator trained in Stronge. The write-up that follows is where the hours go: organizing fragmentary notes as evidence against the right standard, weighing that evidence against the indicators, settling on a defensible level, and writing feedback a teacher can act on, in the model's language, across a full caseload.

Where AI helps with the Stronge write-up

The strongest fit between current AI and the Stronge model is that standards-based documentation.

AI handles three parts of it reliably. First, it turns fragmentary notes into coherent, evidence-anchored prose in the model's language, tied to the indicators a standard describes. Second, it maps evidence to the right standard, including evidence that speaks to more than one. Third, it holds consistency across a caseload, so the same quality of evidence lands in similar territory from teacher to teacher.

For Stronge specifically, the value is drafting toward the standards and the four levels in a way that respects the preponderance-of-evidence approach — assembling what you observed into the standard it supports, rather than reacting to a single moment.

Where AI doesn't help — and what stays with the evaluator

The honest scope follows from how the model is built.

AI cannot supply the Student Progress data behind Standard 7, and it cannot produce the summative rating. Student Progress is set through the district's goal-setting and student-growth process, and the summative rating combines all seven standards. Those sit outside any drafting tool.

AI cannot supply an evaluator's professional judgment — what the body of evidence really shows, and where the preponderance of it places a teacher on a standard. That judgment is the heart of the model, and it stays with the evaluator, who remains the last set of eyes on every rating and comment.

And general-purpose AI has no built-in understanding of Stronge. Paste notes into a consumer chatbot and it will invent standards, blur the observation write-up together with Student Progress, and hand out Highly Effective as if it simply meant very good. For a model tied to a formal summative rating, those errors surface at the worst possible moment.

What to look for in an AI tool for Stronge

A few questions worth asking before committing a tool to this work.

Does the tool actually know the Stronge model — its seven standards and its four-level scale — or is it a generic writing assistant with Stronge vocabulary sprinkled in?

Does it treat Effective as the expected level rather than defaulting to the top of the scale, and does it keep ratings anchored to the evidence you captured?

Does it stay in its lane, drafting the observation-based standards and leaving Student Progress data and the summative rating to your district?

Where does your observation data live — on the vendor's servers, or used to train models?

How EvalScribe handles the Stronge model

EvalScribe is built around the Stronge Teacher Effectiveness Performance Evaluation System — all seven performance standards and the four-level scale. An evaluator captures notes by typing, dictating, or photographing handwriting (Smart Scan OCR converts it to text), EvalScribe maps that evidence to the standard it supports, and drafts a best-fit rating and evidence-anchored feedback in the model's own language. Every rating and comment is fully editable before you finalize it, and the evidence stays traceable to the note it came from.

Two scope notes, because the model calls for them. First, EvalScribe drafts the observation-based write-up. Student Progress (Standard 7) data and the overall summative rating are set through your district's process, not the app. Second, EvalScribe is an independent tool; it is not affiliated with Dr. James Stronge or Stronge & Associates, and it supports administrators who already use the model rather than replacing an official scoring platform.

Beta testers report saving 30 to 60 minutes per evaluation versus writing the documentation by hand. Across seven standards and a full caseload, that adds up to dozens of hours back — hours that can go to the coaching conversation the model is built to support. More detail on the Stronge workflow is available at evalscribe.com/stronge.

Frequently asked questions about the Stronge model and AI

Does EvalScribe support the Stronge model? Yes — all seven performance standards and the four-level Ineffective to Highly Effective scale, drafting in the model's own language.

What are the seven standards? Professional Knowledge, Instructional Planning, Instructional Delivery, Assessment of/for Learning, Learning Environment, Professionalism, and Student Progress.

What are the rating levels? Ineffective, Partially Effective, Effective, and Highly Effective. Effective is the expected level of performance.

Does EvalScribe produce the summative rating? No. Student Progress data and the final rating are set by your district; EvalScribe drafts the observation write-up.

Is EvalScribe affiliated with Stronge & Associates? No. It is an independent drafting tool that supports administrators who already use the model.

If your district uses the Stronge model, see how EvalScribe drafts across its seven standards at evalscribe.com/stronge. Questions, or a school or district license? Reach the team at [email protected].

References

  • Stronge & Associates Educational Consulting, Teacher and Leader Effectiveness Evaluation Systems (the seven performance standards and four-level rating approach)

  • James H. Stronge, Qualities of Effective Teachers (the research foundation for the model)

  • State and district Stronge implementation handbooks, including the Virginia performance standards built on the Stronge model (public standards, indicators, and rubrics)

Related articles


Page maintained by Anthony D. Neely, Ph.D. — practicing K-12 educator with nearly 20 years in the classroom, 2025–2026 Walker County Distinguished Teacher of the Year, and co-founder of EvalScribe. Framework details verified against Stronge & Associates Educational Consulting's published model overview.

Last updated on August 6, 2026.

Anthony D. Neely, Ph.D.

Anthony D. Neely, Ph.D.

Anthony Neely is the Founder of EvalScribe, a veteran educator, an AI integration consultant for teaching & learning, researcher, & author.

Back to Blog