How can AI help Arkansas evaluators write TESS evaluations?

How can AI help Arkansas evaluators write TESS evaluations?

June 27, 202612 min read

How can AI help Arkansas evaluators write TESS evaluations?

How can AI help Arkansas evaluators write TESS evaluations?

Arkansas TESS — the Teacher Excellence and Support System — is built on Charlotte Danielson's Framework for Teaching, covering four domains and twenty-two components on a four-level rubric. AI can meaningfully help with the documentation burden of TESS evaluations: translating walkthrough notes into rubric-aligned language, mapping evidence to the right components, and producing draft ratings consistent with the framework's structure. AI does not replace evaluator judgment — particularly the Proficient-to-Distinguished call, which still belongs to the practitioner watching the lesson. Where AI lifts the most weight is the documentation itself.

What Arkansas TESS is

The Teacher Excellence and Support System is the statewide framework Arkansas uses for K-12 teacher observation and evaluation. It was established in 2011 under Act 1209 and codified at Ark. Code Ann. § 6-17-2801 et seq. The framework is administered by the Arkansas Department of Education's Division of Elementary and Secondary Education (DESE).

The structural backbone of TESS is Charlotte Danielson's Framework for Teaching — the same model used by Pennsylvania, Illinois, Kentucky, and a number of other states. Arkansas layers its own procedural rules, scoring requirements, and growth-plan structure on top, but the four-domain, twenty-two-component spine is Danielson. The TESS rubric was most recently updated in November 2023.

The four domains and twenty-two components

TESS organizes teaching practice into four observable domains.

Domain 1 (Planning and Preparation) covers six components: knowledge of content and pedagogy, knowledge of students, instructional outcomes, knowledge of resources, coherent instructional design, and assessment design. Domain 2 (Classroom Environment) covers five: respect and rapport, culture for learning, classroom procedures, student behavior, and physical space. Domain 3 (Instruction) covers five: communicating with students, questioning and discussion, engaging students in learning, using assessment in instruction, and flexibility and responsiveness. Domain 4 (Professional Responsibilities) covers six: reflecting on teaching, accurate records, communicating with families, participating in a professional community, growing and developing professionally, and showing professionalism.

That's twenty-two components in total, each carrying its own rubric language and its own evidence requirements.

The four-level rubric — and the Proficient-to-Distinguished judgment

TESS uses four performance levels: Unsatisfactory, Basic, Proficient, Distinguished. Unsatisfactory is below acceptable standards. Basic is developing but not yet meeting the standard. Proficient meets the standard with consistent, evidence-based practice. Distinguished exceeds the standard with student ownership, intentional design, and visible impact.

The hardest call in TESS — and the one that most divides evaluators in calibration sessions — is Proficient versus Distinguished. The two levels share a lot of surface features: organized lessons, engaged students, clear instructional outcomes. What separates them is who's doing the work. In Proficient, the teacher is driving the lesson well. In Distinguished, the students have taken over the heavy intellectual lifting. The teacher set the conditions; the learners produced the outcomes. That's a real distinction, and it's the one most likely to get blurred when an evaluator is writing a summative quickly between observations.

The TESS observation process

TESS runs on a track system. Probationary and novice teachers receive ongoing observation and feedback with no summative rating in their first year. Experienced teachers on Track 1 receive a summative evaluation at least once every four years, with formative observation in the intervening years. Track 2A is the interim appraisal — used between summatives to provide structured support. Track 2B is intensive support, triggered when a teacher receives an Unsatisfactory rating in any one domain or a rating of Unsatisfactory or Basic in the majority of components of a domain.

Within a given year, an evaluator typically conducts a combination of informal walkthroughs, formal observations with pre- and post-conferences, and a summative review that pulls evidence from observations and artifacts across the cycle. Each teacher also maintains a Professional Growth Plan (PGP), and teachers in intensive support follow an Intensive Growth Plan (IGP).

After the walkthrough: where the work actually piles up

The observation itself is the part most evaluators trained for. What no one warned them about was the volume of writing that follows.

A single formal observation can produce three to seven pages of evaluation documentation: rubric ratings across components, narrative justifications for each rating, evidence quotations from observation notes, suggested next steps, and language for the post-observation conference. A walkthrough is shorter but adds up — four walkthroughs across a teacher's load, three or four times a year, across thirty teachers, becomes a documentation pile no one budgeted time for.

This is where the day actually goes. The walkthrough takes fifteen minutes. The write-up takes forty-five.

Where AI helps with TESS evaluation write-ups

The strongest fit between current AI and TESS evaluation work is the translation problem. An evaluator captures observation notes in shorthand — half-sentences, abbreviations, fragments — and then has to translate those notes into the formal language of the rubric. That translation is mechanical work. It's not where evaluator judgment lives; it's where evaluator time gets spent.

AI handles three pieces of that translation reliably. First, it can take fragmentary observation notes and produce coherent, evidence-grounded prose in the voice of the TESS rubric. Second, it can map specific pieces of evidence to the relevant components across all four domains — telling the evaluator that the note about a teacher's wait time supports Component 3b, while the note about anchor charts supports Component 2b. Third, it can produce internally consistent draft language across components in the same evaluation, so the summative reads as a coherent document rather than twenty-two disconnected paragraphs.

Where AI saves the most time is the third piece. Most evaluators don't struggle to write any one rubric justification. They struggle to write twenty-two of them in a single document while keeping voice and evidence-density consistent.

Where AI doesn't help — and what stays with the evaluator

The honest scope is narrower than the marketing on most AI tools suggests.

AI cannot make the borderline Proficient-to-Distinguished judgment for an evaluator. That call requires watching the lesson in real time, knowing the teacher's history and context, and judging whether the evidence rises to the standard of student ownership. AI can suggest a rating based on the notes provided, but the evaluator owns the final call.

AI cannot replace the relational work of the post-observation conference. The coaching conversation, the goal-setting around the Professional Growth Plan, the harder conversations about Intensive Growth Plan placement — these are human work. AI can help draft talking points or summary language, but it cannot do the conversation itself.

And AI cannot do the contextual reasoning a long-tenured evaluator brings — knowing that this teacher had three new English-learners arrive midyear, that this lesson followed a fire drill, that this is the third week the heat hasn't worked in this hallway. That context shapes evaluator judgment in ways no chatbot will reconstruct from observation notes.

What AI does well, it does well. What it doesn't do, it shouldn't pretend to do. Pasting walkthrough notes into ChatGPT works fine as a glorified search engine — but for rubric-aligned summative documentation, the gap between "sounds polished" and "actually defensible under TESS" is wider than the polish suggests.

Observable evidence: what each TESS domain looks like in practice

Each TESS domain has its own surface features — the things an evaluator actually sees, hears, and notes during an observation.

Domain 1 (Planning and Preparation) is largely evident before and after the lesson rather than during it. Look for lesson plans showing differentiation for IEPs and English learners, pre-observation conversations that mention specific student misconceptions the teacher is anticipating, materials prepared for multiple instructional pathways, and assessment design that aligns to the stated learning outcomes rather than just to activities.

Domain 2 (Classroom Environment) is most visible in the first five minutes. Look for whether the teacher knows students by name (including the quiet ones), whether the anchor charts on the wall are referenced during instruction or just decorative, how transitions between activities run (the thirty-second restart versus the five-minute one), and how the teacher addresses minor behavior without interrupting the lesson.

Domain 3 (Instruction) is the domain most evaluators feel most confident scoring. Look for wait time after questions (counted in seconds — eleven seconds for a Treaty of Versailles question is real evidence), the ratio of student-to-student dialogue versus teacher-to-student call-and-response, formative checks used in real time (thumbs up/down, exit tickets that actually change the lesson), and the teacher's responsiveness to confusion versus their drive to finish the lesson plan on schedule.

Domain 4 (Professional Responsibilities) is the domain most likely to be overrated, because much of the evidence lives outside the observation window. Look for reflective comments after the lesson that name specific moments rather than generic ones, documentation of parent communication and PLC participation, evidence of professional learning that's actually changed practice, and growth in goals across the Professional Growth Plan year over year.

Common scoring mistakes under TESS

Even experienced evaluators slip into patterns worth naming.

The first is over-Distinguishing in Domain 4. Because Domain 4 lives partly outside the lesson, evaluators tend to credit "professional-feeling" teachers with high D4 ratings regardless of evidence. The fix is to require the same evidence rigor for Domain 4 as for Domain 3 — name the parent communication, name the PLC contribution, name the specific PD that changed practice.

The second is the Proficient default. When evidence is thin in either direction, evaluators slot Proficient because it feels safe. But TESS requires evidence of Proficient practice, not just absence of Basic. A component with insufficient evidence is a coaching opportunity, not a default Proficient.

The third is component conflation. Adjacent components like 2a (Environment of Respect and Rapport) and 2b (Establishing a Culture for Learning) feel related, and evaluators often justify both with the same evidence. They're distinct claims — respect and rapport is about how the teacher and students treat each other; culture for learning is about whether the classroom communicates that learning matters.

The fourth is evidence-thin Distinguished. Distinguished requires evidence of student ownership, intentional design, or impact beyond compliance. "The teacher did it well" is not Distinguished — it's Proficient. Distinguished requires the move from teacher-driven excellence to student-driven excellence.

The fifth is the 1c/1e tangle. Setting Instructional Outcomes (1c) and Designing Coherent Instruction (1e) get mashed together routinely. They're separate questions: 1c asks whether the outcomes themselves are appropriately rigorous and aligned. 1e asks whether the lesson sequence will actually deliver them.

TESS among other Danielson-foundation states

Arkansas is one of several states whose statewide teacher evaluation framework is built on Danielson. Pennsylvania, Illinois, Kentucky, Maine, Maryland, Michigan, Minnesota, New Hampshire, North Dakota, South Dakota, Washington, Wisconsin, and Hawaii all run Danielson-based systems with their own state-specific procedural rules and rubric language on top. Each state's variant differs in the details — number of tracks, summative cycle frequency, growth-plan requirements, rating-scale labels — but the four-domain, twenty-two-component spine is consistent across all of them.

For TESS evaluators, that means the rubric language and the underlying structure will feel familiar to evaluators trained in any other Danielson state. The Arkansas-specific work is the procedural overlay: PGP requirements, IGP triggers, Track 2A and 2B distinctions, and the DESE-specific documentation expectations.

How EvalScribe handles TESS specifically

EvalScribe is built around the actual TESS structure — all four domains, all twenty-two components, the four-level rubric — rather than approximating it with generic teaching-evaluation prose. The tool handles the translation problem natively. An evaluator captures observation notes by typing, dictating, or photographing handwritten notes (Smart Scan OCR converts handwriting to text), and EvalScribe maps that evidence to the relevant TESS components and drafts ratings and comments in the rubric's actual voice.

Beta testers report saving 30 to 60 minutes per evaluation versus writing the documentation traditionally. Across a typical evaluator load — say thirty teachers and three to four observations per year — that math adds up to dozens of hours back over an evaluation cycle.

The evaluator stays in control of every rating, every comment, and every piece of mapped evidence. EvalScribe drafts; the evaluator decides. More detail on the TESS workflow is available at evalscribe.com/arkansas-tess.

Frequently asked questions about TESS and AI

Where can I read the official Arkansas TESS framework and rubric? The primary sources are maintained by the Arkansas DESE: the TESS home page and the TESS Rubric Descriptors. Statutory authority is Ark. Code Ann. § 6-17-2801 et seq.

Does AI replace evaluator judgment under TESS? No. AI drafts evidence-mapped evaluations and suggested ratings; the evaluator reviews, edits, and finalizes every document. The borderline judgment calls — particularly Proficient-to-Distinguished — stay with the practitioner who observed the lesson.

Can AI help with the Professional Growth Plan or Intensive Growth Plan? AI can help draft language for PGP and IGP documentation based on observation evidence and the rubric language of the components flagged. The PGP conversation itself — the goal-setting, the coaching, the relational work — stays with the evaluator and the teacher.

Does AI work for walkthroughs as well as formal observations? Yes. Brief walkthroughs and full formal observations both flow through the same capture-and-draft pipeline, so the documentation across observation types stays consistent.

Is TESS the same as the Danielson framework? TESS is built on Danielson. The four-domain, twenty-two-component structure and the four-level rubric come directly from Charlotte Danielson's Framework for Teaching. Arkansas layers its own procedural rules — tracks, PGP/IGP, DESE-specific documentation — on top.

How can I see what EvalScribe looks like for TESS specifically? The Arkansas TESS page at evalscribe.com/arkansas-tess walks through the four-domain workflow, the rubric handling, and a sample exported evaluation. EvalScribe is available on iOS and macOS via the App Store, with three free evaluations to start.

If you're evaluating teachers under Arkansas TESS, see how EvalScribe handles the full TESS workflow at evalscribe.com/arkansas-tess.

References

Related articles


Page maintained by Anthony D. Neely, Ph.D. — practicing K-12 educator with nearly 20 years in the classroom, 2025–2026 Walker County Distinguished Teacher of the Year, and co-founder of EvalScribe. Framework details verified against current Arkansas DESE source documents.

Last reviewed: June 26, 2026.

Anthony D. Neely, Ph.D.

Anthony D. Neely, Ph.D.

Anthony Neely is the Founder of EvalScribe, a veteran educator, an AI integration consultant for teaching & learning, researcher, & author.

Back to Blog