
How can AI help Pennsylvania evaluators write teacher evaluations under Act 13?
How can AI help Pennsylvania evaluators write teacher evaluations under Act 13?

Pennsylvania evaluates teachers through the Act 13 Educator Effectiveness system, and the honest answer about AI is the same as it is everywhere: it helps with one specific part of the process. For a classroom teacher, the Observation and Practice component — the Danielson-based, four-domain piece — is where an evaluator spends the documentation hours, and it's where AI can genuinely help by mapping evidence to the right domain and drafting evidence-anchored feedback. What AI does not do is calculate your overall summative rating, decide which rating tool your district adopted, or resolve the one thing that trips up Pennsylvania evaluators more than any other: the fact that the state uses two different four-level rating scales that look alike but aren't.
What Act 13 actually is
For years, Pennsylvania's teacher evaluation ran under Act 82 of 2012. Act 13 of 2020, signed by Governor Wolf in March 2020, revised that Educator Effectiveness process, and the revised system took effect in the 2021-22 school year under Chapter 19 of Title 22 of the Pennsylvania Code. If you're evaluating teachers in Pennsylvania today, you're working under Act 13, not Act 82 — the distinction matters, because a fair amount of older guidance still floating around describes the Act 82 version.
Act 13 kept much of the existing structure but made some meaningful changes. The one most evaluators feel is the weighting. For a tenured classroom teacher, 70% of the summative rating is based on Observation and Practice, and the remaining 30% comes from building-level and teacher-specific data. That's a deliberate reduction in the weight of student-performance measures compared with the earlier system, and it puts even more importance on getting the observation documentation right.
The four Observation and Practice domains
The Observation and Practice component uses Pennsylvania's Framework for the Evaluation of Classroom Teachers, which the Department of Education adapted from Charlotte Danielson's Framework for Teaching (with additions for trauma-informed practice, cultural relevance, remote teaching, career readiness, equity and inclusion, and social-emotional wellness). It organizes practice into four domains, and an evaluator records a single rating in each:
Domain 1 — Planning and Preparation. Knowledge of content and of students, coherent instructional design, and assessment planning aligned to standards. The work that happens before students arrive.
Domain 2 — Classroom Environment. A culture of respect and rapport, clear expectations, managed routines and procedures, and a physical space that supports learning.
Domain 3 — Instruction. Clear communication, strong questioning and discussion, engaging students in the work of learning, and using assessment to adjust in the moment.
Domain 4 — Professional Responsibilities. Reflection on practice, accurate record-keeping, communication with families, professional growth, and contribution to the school community.
Pennsylvania's framework describes the state's vision of effective practice but does not dictate the exact data-collection tools an evaluator must use. The requirement is that an evaluator gather sufficient evidence to record a defensible rating in each of the four domains.
The part that trips everyone up: two four-level scales
Here is the single most important thing to understand about documenting a Pennsylvania evaluation, and the thing general AI tools get wrong constantly. Pennsylvania uses two different sets of four-level labels, and they are easy to confuse because they overlap.
The Observation and Practice domains are rated using Danielson's descriptors:
Unsatisfactory · Basic · Proficient · Distinguished
The overall summative rating a teacher ultimately receives uses different labels:
Failing · Needs Improvement · Proficient · Distinguished
Both scales run on the same 0-to-3 numeric spine, and the top two names are identical — Proficient and Distinguished. But the bottom two diverge. What reads as Basic at the component level corresponds to Needs Improvement in the overall rating, and Unsatisfactory corresponds to Failing. So a teacher can have domains described in Danielson's "Basic/Proficient" language while the summative document reports "Needs Improvement/Proficient." Same teacher, same evidence, two vocabularies.
This isn't pedantry. The two scales carry different consequences. A summative rating of Distinguished or Proficient is satisfactory, and so is a first Needs Improvement. A Failing rating, or a second Needs Improvement within the statutory window, becomes an unsatisfactory rating with real employment weight. Documentation that blurs the component descriptors into the summative labels — or vice versa — is documentation that won't hold up when it matters.
After the observation: where the work piles up
The observation itself is familiar work for an experienced Pennsylvania evaluator. The write-up that follows is where the hours go: taking fragmentary observation notes and turning them into evidence organized against the four domains, settling on a defensible rating for each in the Danielson descriptors, and drafting feedback a teacher can actually act on — then doing it again across a caseload, multiple times a year. The 70% weight on Observation and Practice means this documentation isn't a formality; it's most of the rating.
Where AI helps with the Pennsylvania write-up
The strongest fit between current AI and Act 13 work is the mapping-and-translation problem at the heart of the Observation and Practice component.
AI handles three parts of that reliably. First, it turns fragmentary notes into coherent, evidence-anchored prose in the framework's language — the difference between "kids explained their thinking to each other" and a sentence tied to Domain 3 on engaging students in constructing understanding. Second, it maps evidence to the right domain, including evidence that supports more than one — a well-run discussion can speak to both Domain 3 (instruction) and Domain 2 (classroom environment). Third, it holds consistency across a caseload, so the same quality of evidence lands in similar territory from teacher to teacher rather than drifting with the hour of the night.
For Pennsylvania specifically, the added value is keeping the domain ratings in the correct Danielson vocabulary and not silently sliding into the summative labels — the discipline the two-scale structure demands.
Where AI doesn't help — and what stays with the evaluator
The honest scope follows directly from how Act 13 is built.
AI cannot calculate your summative rating. EvalScribe and tools like it draft the Observation and Practice piece — the 70%. The overall rating combines that with the 30% building-level and teacher-specific data, and that roll-up into Distinguished, Proficient, Needs Improvement, or Failing happens in your district's rating tool, not in a drafting app.
AI cannot decide which rating tool your district uses. Under 24 P.S. §11-1138.6, a Pennsylvania LEA may use PDE's published practice models or adopt its own PDE-approved alternative rating tool. The four state domains and the Danielson descriptors are the fixed, shared foundation; the specific forms are a local fact to confirm.
AI cannot supply an evaluator's professional and local judgment — how a teacher's year unfolded, what the evidence really shows, how a rating fits the district's context. That stays with the evaluator, who remains the last set of eyes on every rating and comment.
And general-purpose AI in particular has no built-in sense of the two-scale structure. Paste notes into a consumer chatbot and it will happily blend the component descriptors and the summative labels, hand out "Distinguished" freely, and produce comments that could describe any teacher. For a document with contract weight, the gap between "sounds polished" and "actually defensible" is wider than the polish suggests.
Common mistakes under Act 13
A few patterns are worth naming.
The first is conflating the two scales — reporting a domain as "Needs Improvement" (a summative label) or a summative rating as "Basic" (a component label). Keep the Danielson descriptors for the domains and the Distinguished-to-Failing labels for the overall rating.
The second is treating the observation as the whole rating. It's 70% of it. The summative rating combines it with data measures under the district's tool.
The third is relying on stale Act 82 guidance — older weighting, older building-level data categories. Act 13 changed both.
The fourth is thin evidence. The framework requires a defensible rating in each domain supported by evidence noted in the record. A rating without traceable evidence behind it is the kind of thing that unravels in a challenge.
How EvalScribe handles Act 13
EvalScribe is built around the four Observation and Practice domains. An evaluator captures notes by typing, dictating, or photographing handwriting (Smart Scan OCR converts it to text), EvalScribe maps that evidence to the domain it supports, and drafts a rating and evidence-anchored feedback for each — using the framework's Danielson-based descriptors: Unsatisfactory, Basic, Proficient, Distinguished. Every rating and comment is fully editable before export, and the evidence stays traceable to the note it came from.
Two scope notes, because Pennsylvania's structure demands them. First, EvalScribe drafts the Observation and Practice piece — the 70%. It does not calculate the overall summative rating; that roll-up, which combines the observation score with the 30% building-level and teacher-specific data and produces the Distinguished-to-Failing label, happens in your district's rating tool. Second, because LEAs may use PDE's practice models or a PDE-approved alternative tool, EvalScribe is built on the four state domains and the Danielson descriptors the approaches share; confirm your district's adopted forms.
Beta testers report saving 30 to 60 minutes per evaluation versus writing the documentation by hand. Across a full caseload and multiple observation cycles, that adds up to dozens of hours back — hours that can go to the feedback conversation the ratings are meant to support. More detail on the Pennsylvania workflow is available at evalscribe.com/pennsylvania.
Frequently asked questions about Act 13 and AI
Is this Act 13 or Act 82? Act 13 of 2020 revised the earlier Act 82 of 2012 process, effective 2021-22. The current system — four Observation and Practice domains, 70/30 weighting for tenured classroom teachers, and the Distinguished-to-Failing summative rating — is Act 13.
What are the two rating scales, and why do they differ? The Observation and Practice domains use Danielson's descriptors (Unsatisfactory, Basic, Proficient, Distinguished). The overall summative rating uses Distinguished, Proficient, Needs Improvement, Failing. Both run 0-3 and the top two names match, but "Basic" corresponds to "Needs Improvement" and "Unsatisfactory" to "Failing."
Does EvalScribe calculate my summative rating? No. It drafts the Observation and Practice piece — 70% of a tenured classroom teacher's rating. The summative roll-up, which adds the 30% data measures and produces the final label, happens in your district's rating tool.
Can I use it with my district's own rating tool? Yes. Under 24 P.S. §11-1138.6, LEAs may use PDE's practice models or a PDE-approved alternative tool. EvalScribe is built on the four state domains and Danielson descriptors both approaches share.
Does AI replace evaluator judgment? No. The observation, the domain ratings, and the professional judgment are yours. AI translates your judgment into framework-aligned documentation.
Where can I read the official Act 13 sources? The Pennsylvania Department of Education's Educator Effectiveness page and the SAS Educator Effectiveness site, along with Act 13 of 2020 and Chapter 19 of Title 22 of the Pa. Code.
If you're evaluating teachers in Pennsylvania under Act 13, see how EvalScribe handles the four-domain Observation and Practice workflow at evalscribe.com/pennsylvania. Questions, or a school or district license? Reach the team at [email protected].
References
Pennsylvania Department of Education, Educator Effectiveness (Act 13 of 2020 revised the Act 82 process, effective 2021-22)
Standards Aligned System (SAS), Educator Effectiveness — Frameworks and Rating Tools
22 Pa. Code Chapter 19, §19.2a Classroom Teacher Evaluation (the four Observation and Practice domains)
Pennsylvania State Education Association, The Revised Educator Effectiveness System (70/30 weighting; summative ratings)
Related articles
Page maintained by Anthony D. Neely, Ph.D. — practicing K-12 educator with nearly 20 years in the classroom, 2025–2026 Walker County Distinguished Teacher of the Year, and co-founder of EvalScribe. Framework details verified against the Pennsylvania Department of Education's Educator Effectiveness resources, 22 Pa. Code Chapter 19, and Act 13 of 2020.
Last updated on July 28, 2026.
