
How can AI help Wisconsin evaluators write EE System teacher evaluations?
How can AI help Wisconsin evaluators write EE System teacher evaluations?

The Wisconsin Educator Effectiveness System (EE System) has two parts: an educator practice evaluation built on the 2022 Danielson Framework for Teaching — four domains, twenty-two components, a four-level rubric — and a student outcomes evaluation built around Student/School Learning Objectives (SLOs). AI can meaningfully help with the documentation burden of the practice portion: translating walkthrough notes into rubric-aligned language, mapping evidence to the right Danielson components, and producing draft ratings consistent with the framework's structure. AI does not replace evaluator judgment — particularly the Proficient-to-Distinguished call — and it does not author the SLO. Where AI lifts the most weight is the documentation itself.
What the Wisconsin EE System is
The Wisconsin Educator Effectiveness System is the framework Wisconsin uses for K-12 teacher and principal evaluation. It was established by Wisconsin Act 166 in the 2011-12 legislative session and codified at Wis. Stat. § 115.415. The system is administered by the Wisconsin Department of Public Instruction (DPI) and is required for teachers and principals in public school districts and independent charter schools.
The EE System is intentionally a two-part design. The first part is the educator practice evaluation, built on Charlotte Danielson's Framework for Teaching. The second part is the student outcomes evaluation, built around at least one Student/School Learning Objective (SLO) each educator develops and pursues annually. Both parts contribute to the educator's summary evaluation. The two-part design is intentional — it captures both the quality of professional practice and the evidence of student learning that practice produces.
Starting in the 2024-25 school year, Wisconsin DPI requires state model districts and independent charters to use the 2022 Danielson Framework for Teaching for the practice portion. The 2022 version updates the language and consolidates some elements of the 2013 framework, while preserving the four-domain, twenty-two-component structure.
The four domains and twenty-two components
The 2022 Danielson Framework for Teaching organizes practice into four observable domains.
Domain 1 (Planning and Preparation) covers six components: knowledge of content and pedagogy, knowledge of students, setting instructional outcomes, knowledge of resources, designing coherent instruction, and designing student assessments. Domain 2 (Classroom Environment) covers five: creating an environment of respect and rapport, establishing a culture for learning, managing classroom procedures, managing student behavior, and organizing physical space. Domain 3 (Instruction) covers five: communicating with students, questioning and discussion techniques, engaging students in learning, using assessment in instruction, and demonstrating flexibility and responsiveness. Domain 4 (Professional Responsibilities) covers six: reflecting on teaching, maintaining accurate records, communicating with families, participating in a professional community, growing and developing professionally, and showing professionalism.
That's twenty-two components in total, each carrying its own rubric language and its own evidence requirements.
The four-level rubric — and the Proficient-to-Distinguished judgment
The Wisconsin EE System uses four performance levels: Unsatisfactory, Basic, Proficient, Distinguished. Unsatisfactory signals practice that is below standard and harmful to learning. Basic signals practice that is inconsistent and remains teacher-centered. Proficient signals the professional standard — consistent, effective instruction. Distinguished signals exemplary practice where students assume strong ownership of learning and the teacher serves as a model for others.
The hardest call in the EE rubric — and the one that most divides evaluators in calibration sessions — is Proficient versus Distinguished. The two levels share a lot of surface features: organized lessons, engaged students, clear instructional outcomes. What separates them is who's doing the work. At Proficient, the teacher is driving the lesson well. At Distinguished, the students have taken over the heavy intellectual lifting — they formulate their own questions, sustain rich discussion, self-assess against criteria, and reinforce a respectful classroom culture. The teacher set the conditions; the learners produced the outcomes. That distinction is real, and it's the one most likely to get blurred when an evaluator is writing a summative quickly between observations.
The Wisconsin SLO requirement
Wisconsin's EE System differs from many other Danielson-foundation states in one important way: every educator completes at least one Student/School Learning Objective (SLO) annually as part of the student outcomes portion of the evaluation. SLOs are context-specific SMART goals that target the growth of a specific student population, based on baseline data and aligned to academic standards.
The SLO process runs across the evaluation cycle. The educator selects a focus area, gathers baseline data, sets a measurable growth target, designs and implements instructional moves intended to meet that target, monitors student progress, and reports on outcomes at the end of the SLO interval. SLO scoring is a holistic judgment about the quality of the process and the outcomes achieved.
This is meaningful because SLO authorship is professional judgment work — the educator and evaluator design and refine the SLO together, and AI shouldn't be authoring goals on behalf of the educator. What AI can help with is the documentation around the SLO: drafting the rationale, summarizing the baseline data narrative, and producing the SLO reflection at the end of the cycle. The goal-setting itself stays with the people doing the teaching.
After the walkthrough: where the work actually piles up
The observation itself is the part most evaluators trained for. What no one warned them about was the volume of writing that follows.
A single formal observation in the Wisconsin EE System can produce three to seven pages of evaluation documentation: rubric ratings across Danielson components, narrative justifications for each rating, evidence quotations from observation notes, suggested next steps, and language for the post-observation conference. Mini-observations and walkthroughs are shorter individually, but they accumulate — across a typical evaluator load over a Summary Year, the documentation pile is substantial.
This is where the day actually goes. The walkthrough takes fifteen minutes. The write-up takes forty-five.
Where AI helps with EE System write-ups
The strongest fit between current AI and Wisconsin EE evaluation work is the translation problem. An evaluator captures observation notes in shorthand — half-sentences, abbreviations, fragments — and then has to translate those notes into the formal language of the Danielson rubric. That translation is mechanical work. It's not where evaluator judgment lives; it's where evaluator time gets spent.
AI handles three pieces of that translation reliably. First, it can take fragmentary observation notes and produce coherent, evidence-grounded prose in the voice of the 2022 Danielson rubric. Second, it can map specific pieces of evidence to the relevant components across all four domains — telling the evaluator that the note about a teacher's wait time supports Component 3b (Questioning and Discussion Techniques), while the note about anchor charts supports Component 2b (Establishing a Culture for Learning). Third, it can produce internally consistent draft language across components in the same evaluation, so the summative reads as a coherent document rather than twenty-two disconnected paragraphs.
Where AI saves the most time is the third piece. Most evaluators don't struggle to write any one rubric justification. They struggle to write twenty-two of them in a single document while keeping voice and evidence-density consistent.
Where AI doesn't help — and what stays with the evaluator
The honest scope is narrower than the marketing on most AI tools suggests.
AI cannot make the borderline Proficient-to-Distinguished judgment for an evaluator. That call requires watching the lesson in real time, knowing the teacher's history and context, and judging whether the evidence rises to the standard of student ownership. AI can suggest a rating based on the notes provided, but the evaluator owns the final call.
AI cannot author the SLO. The SLO is the educator's professional growth artifact, developed in conversation with the evaluator and grounded in the specific student population being served. AI can help draft the SLO narrative or reflection language once the goal and approach are set, but the substantive design decisions — what to focus on, what baseline data to use, what growth target is meaningful, what instructional approach will move the needle — stay with the people who know the students.
AI cannot replace the relational work of the post-observation conference. The coaching conversation, the goal-setting around the SLO and educator practice goals, the harder conversations when performance is moving toward Basic — these are human work. AI can help draft talking points or summary language, but it cannot do the conversation itself.
And AI cannot do the contextual reasoning a long-tenured evaluator brings — knowing that this teacher had three new English-learners arrive midyear, that this lesson followed a fire drill, that this is the third week the heat hasn't worked in this hallway. That context shapes evaluator judgment in ways no chatbot will reconstruct from observation notes.
What AI does well, it does well. What it doesn't do, it shouldn't pretend to do. Pasting walkthrough notes into ChatGPT works fine as a glorified search engine — but for rubric-aligned EE documentation, the gap between "sounds polished" and "actually defensible under the Wisconsin EE System" is wider than the polish suggests.
Observable evidence: what each domain looks like in practice
Each domain has its own surface features — the things an evaluator actually sees, hears, and notes during an observation.
Domain 1 (Planning and Preparation) is largely evident before and after the lesson rather than during it. Look for lesson plans showing differentiation for IEPs and English learners, pre-observation conversations that mention specific student misconceptions the teacher is anticipating, materials prepared for multiple instructional pathways, and assessment design that aligns to the stated learning outcomes rather than just to activities.
Domain 2 (Classroom Environment) is most visible in the first five minutes. Look for whether the teacher knows students by name (including the quiet ones), whether the anchor charts on the wall are referenced during instruction or just decorative, how transitions between activities run (the thirty-second restart versus the five-minute one), and how the teacher addresses minor behavior without interrupting the lesson.
Domain 3 (Instruction) is the domain most evaluators feel most confident scoring. Look for wait time after questions (counted in seconds — eleven seconds for a Treaty of Versailles question is real evidence), the ratio of student-to-student dialogue versus teacher-to-student call-and-response, formative checks used in real time (thumbs up/down, exit tickets that actually change the lesson), and the teacher's responsiveness to confusion versus their drive to finish the lesson plan on schedule.
Domain 4 (Professional Responsibilities) is the domain most likely to be overrated, because much of the evidence lives outside the observation window. Look for reflective comments after the lesson that name specific moments rather than generic ones, documentation of parent communication and PLC participation, evidence of professional learning that's actually changed practice, and substantive growth across the educator's professional practice goals year over year.
Common scoring mistakes under the Wisconsin EE System
Even experienced evaluators slip into patterns worth naming.
The first is over-Distinguishing in Domain 4. Because Domain 4 lives partly outside the lesson, evaluators tend to credit "professional-feeling" teachers with high D4 ratings regardless of evidence. The fix is to require the same evidence rigor for Domain 4 as for Domain 3 — name the parent communication, name the PLC contribution, name the specific PD that changed practice.
The second is the Proficient default. When evidence is thin in either direction, evaluators slot Proficient because it feels safe. But the EE rubric requires evidence of Proficient practice, not just absence of Basic. A component with insufficient evidence is a coaching opportunity, not a default Proficient.
The third is component conflation. Adjacent components like 2a (Creating an Environment of Respect and Rapport) and 2b (Establishing a Culture for Learning) feel related, and evaluators often justify both with the same evidence. They're distinct claims — respect and rapport is about how the teacher and students treat each other; culture for learning is about whether the classroom communicates that learning matters.
The fourth is evidence-thin Distinguished. Distinguished requires evidence of student ownership, intentional design, or impact beyond compliance. "The teacher did it well" is not Distinguished — it's Proficient. Distinguished requires the move from teacher-driven excellence to student-driven excellence.
The fifth is the 1c/1e tangle. Setting Instructional Outcomes (1c) and Designing Coherent Instruction (1e) get mashed together routinely. They're separate questions: 1c asks whether the outcomes themselves are appropriately rigorous and aligned. 1e asks whether the lesson sequence will actually deliver them.
A sixth pattern is uniquely Wisconsin: separating the SLO conversation from the practice conversation. SLOs and Danielson ratings live in the same EE System, but they assess different things — the SLO is about student outcomes evidence, the practice rating is about observed instructional quality. Combining them into a single conversation, where the SLO outcome is used to justify a practice rating (or vice versa), blurs the system's intentional two-part design.
Wisconsin EE among other Danielson-foundation states
Wisconsin is one of several states whose statewide teacher evaluation framework is built on Danielson. Pennsylvania, Illinois, Kentucky, Arkansas, Maine, Maryland, Michigan, Minnesota, New Hampshire, North Dakota, South Dakota, Washington, and Hawaii all run Danielson-based systems with their own state-specific procedural rules and rubric language on top. Each state's variant differs in the details — number of tracks or cycles, summative cadence, growth-plan requirements, rating-scale labels — but the four-domain, twenty-two-component spine is consistent across all of them.
For Wisconsin EE evaluators, that means the rubric language and the underlying structure will feel familiar to evaluators trained in any other Danielson state. The Wisconsin-specific work is the procedural overlay: the SLO requirement, the Summary Year cycle, the equivalency option under PI 47 for districts that prefer an alternative practice model, and the DPI-specific documentation expectations.
How EvalScribe handles the Wisconsin EE System
EvalScribe is built around the actual EE System practice structure — all four Danielson domains, all twenty-two components, the four-level rubric — rather than approximating it with generic teaching-evaluation prose. The tool handles the translation problem natively. An evaluator captures observation notes by typing, dictating, or photographing handwritten notes (Smart Scan OCR converts handwriting to text), and EvalScribe maps that evidence to the relevant Danielson components and drafts ratings and comments in the rubric's actual voice.
Beta testers report saving 30 to 60 minutes per evaluation versus writing the documentation traditionally. Across a typical evaluator load — say thirty teachers and three to four observations per year — that math adds up to dozens of hours back over an evaluation cycle.
The evaluator stays in control of every rating, every comment, and every piece of mapped evidence. EvalScribe drafts; the evaluator decides. For districts using an alternative practice model approved under PI 47 equivalency, EvalScribe supports a range of models beyond Danielson. More detail on the Wisconsin EE workflow is available at evalscribe.com/wisconsin.
Frequently asked questions about Wisconsin EE and AI
Where can I read the official Wisconsin EE System guidance? The primary sources are maintained by the Wisconsin Department of Public Instruction (DPI): the Educator Effectiveness home page and the WI EE System Policy Guide. Statutory authority is Wis. Stat. § 115.415 and Wis. Admin. Code Ch. PI 47.
Does AI replace evaluator judgment under the EE System? No. AI drafts evidence-mapped evaluations and suggested ratings; the evaluator reviews, edits, and finalizes every document. The borderline judgment calls — particularly Proficient-to-Distinguished — stay with the practitioner who observed the lesson.
Can AI write my SLO for me? No, and it shouldn't. The SLO is the educator's professional growth artifact, developed in conversation with the evaluator and grounded in the specific student population being served. AI can help draft the SLO narrative or reflection language once the goal and approach are set, but the substantive design decisions stay with the people who know the students.
Does AI work for walkthroughs as well as formal observations? Yes. Brief walkthroughs, mini-observations, and full formal observations all flow through the same capture-and-draft pipeline, so the documentation across observation types stays consistent within the EE cycle.
Is Wisconsin EE the same as the Danielson framework? The Wisconsin EE System practice component is built on the 2022 Danielson Framework for Teaching. The four-domain, twenty-two-component structure and the four-level rubric come directly from Danielson. Wisconsin layers its own procedural rules — the SLO requirement, the Summary Year cycle, PI 47 equivalency for alternative models — on top.
With DPI ending the Frontline contract for 2025-26, how does an AI tool fit? Wisconsin DPI ended its contract with Frontline Education at the end of the 2024-25 school year, and now provides free forms and templates that meet EE process requirements. EvalScribe is a separate tool from the record-keeping platform — it focuses on the documentation translation work. The final exported evaluation can be saved into whichever record-keeping system your district has adopted.
How can I see what EvalScribe looks like for Wisconsin EE specifically? The Wisconsin EE page at evalscribe.com/wisconsin walks through the four-domain workflow, the Danielson rubric handling, and a sample exported evaluation. EvalScribe is available on iOS and macOS via the App Store, with three free evaluations to start.
If you're evaluating teachers under the Wisconsin Educator Effectiveness System, see how EvalScribe handles the full EE workflow at evalscribe.com/wisconsin.
References
Wisconsin Department of Public Instruction, Educator Effectiveness
Wisconsin Department of Public Instruction, WI EE System Policy Guide
Wisconsin Statutes § 115.415
Wisconsin Administrative Code Ch. PI 47
Wisconsin Act 166 (2011-12)
Charlotte Danielson, The Framework for Teaching (2022 Edition)
Related articles
Page maintained by Anthony D. Neely, Ph.D. — practicing K-12 educator with nearly 20 years in the classroom, 2025–2026 Walker County Distinguished Teacher of the Year, and co-founder of EvalScribe. Framework details verified against current Wisconsin DPI source documents.
Last reviewed: June 30, 2026.
