Arizona teacher evaluation, Arizona Danielson framework, Danielson Framework for Teaching Arizona, ARS 15-537, ARS 15-341, Arizona Framework for Measuring Educator Effectiveness, Arizona local control evaluation, local control teacher evaluation, Arizona Professional Teaching Standards, AzTS, ADE evaluation repository, Ineffective Developing Effective Highly Effective, Danielson 2013 edition, Danielson 2022 edition, Danielson 22 components, four domains, Classroom Environment domain, Instruction domain, Planning and Preparation, Professional Responsibilities, observation domains, artifact evidence, Arizona teacher observation, two observations Arizona, continuing status teacher, probationary teacher, alternative evaluation cycle, Highly Effective three consecutive years, Mesa Public Schools evaluation, student data percentage evaluation, ARS 15-203, Arizona classroom walkthrough, AI teacher evaluation Arizona, teacher evaluation software Arizona, teacher evaluation app Arizona, observation write-up Arizona, Tucson Unified teacher evaluation, Chandler Unified evaluation, Gilbert Public Schools evaluation, Peoria Unified evaluation, Deer Valley Unified evaluation, Scottsdale Unified teacher evaluation, Dysart Unified evaluation, EvalScribe Arizona, how are teachers evaluated in Arizona, does Arizona have a state evaluation framework, which framework does Arizona use, does Arizona still use the Arizona Framework for Measuring Educator Effectiveness, which Danielson edition does my district use, what scale does Arizona use for teacher evaluation, what is the difference between effective and highly effective, how many observations does Arizona require, which Danielson domains come from observation, can AI write teacher evaluations, best AI tool for Arizona teacher evaluations, does EvalScribe support Danielson

How can AI help Arizona evaluators write teacher observations when the state has no single framework?

July 20, 202616 min read

How can AI help Arizona evaluators write teacher observations when the state has no single framework?

How can AI help Arizona evaluators write teacher observations when the state has no single framework?

Every other state I've written about hands you a framework. Alabama has the ATOT. West Virginia has Policy 5310. South Dakota adopted one Danielson edition statewide and wrote the rule around it. You may not like your state's instrument, but at least you can name it.

Arizona doesn't work that way, and the confusion that follows is the whole story.

If you search how Arizona evaluates teachers, you'll get answers about the Arizona Framework for Measuring Educator Effectiveness, delivered with total confidence, as though it were a live requirement. It isn't. And a general-purpose AI tool — trained on years of documents that describe that framework as mandatory — will make the same mistake, because the internet still thinks Arizona has a state framework it stopped requiring in 2019.

So before anything else, here is what actually happened, and what it means for the write-up on your desk tonight.

Arizona is a local-control state

In 2010, Arizona law required the State Board of Education to adopt a model framework for teacher evaluation, one that folded in quantitative student-progress data at a set percentage. That produced the Arizona Framework for Measuring Educator Effectiveness, and for most of the 2010s districts were required to align to it.

Then the 54th Legislature amended the statutes — ARS §15-537, §15-341, and §15-189.06 — and the requirement went away. In the Arizona Department of Education's own words, the changes mean districts are no longer required to align their evaluations to the Arizona Framework for Measuring Educator Effectiveness, emphasizing "more flexibility and local control in decision making."

What replaced a mandate was a menu. ADE now maintains a repository of compliant evaluation instruments, and it is careful to say the repository's purpose is not to require the use of any specific evaluation instrument or system. The Arizona Professional Teaching Standards — the InTASC-derived state model — sit in that menu as one option. Not the option. One.

So the honest answer to "which framework does Arizona use" is: whichever one your district adopted. There is no statewide rubric. There is your district's handbook, and it is the most important document in your evaluation life, because in a local-control state it's the authority that a statewide page can't be.

Most districts choose Danielson

Local control doesn't mean chaos. When districts are free to choose, most of them choose the same thing, and in Arizona that thing is overwhelmingly the Danielson Framework for Teaching.

That's not surprising. Danielson is the most widely adopted teaching framework in the country, it has decades of research behind it, and the Danielson Group sells training and materials that make adoption straightforward. If you're an Arizona district that just lost its state mandate and needs a defensible instrument fast, Danielson is the safe pick.

Which is why EvalScribe is built around it. Not because Arizona requires it — Arizona requires nothing — but because it's what you're most likely to actually be holding.

But which Danielson? This is where it gets sharp

Here's the part that trips up everyone, including the AI tools.

"Danielson" is not one instrument. There are two editions in active use in Arizona, and they are genuinely different.

The 2013 edition has four domains: Planning and Preparation, The Classroom Environment, Instruction, and Professional Responsibilities. Twenty-two components. This is the workhorse edition, and the large Arizona districts mostly run it — Mesa, the biggest district in the state, uses 2013 with all 22 components.

The 2022 edition kept four domains but renamed and reconceived them: Planning and Preparation, Learning Environments, Learning Experiences, and Principled Teaching. The language moved toward equity, cultural responsiveness, and student agency in ways the 2013 edition only gestured at.

These are not the same rubric with a fresh coat of paint. A component that exists in one may not map cleanly onto the other. If your evaluator is rating you on 2013's "Managing Classroom Procedures" and a tool drafts against 2022's "Maintaining Purposeful Environments," the words won't match the form, and the teacher will notice.

A general-purpose chatbot blends the two constantly, because both editions are in its training data and it has no way to know which one your district adopted. It will hand you a fluent paragraph that quietly mixes 2013 domain names with 2022 language, and it will look right until someone who knows the framework reads it.

Your handbook says which edition you're on. If you take one thing from this section, take that.

And which scale?

The same problem shows up in the rating levels, and it's worth its own moment.

Danielson's published editions use Unsatisfactory, Basic, Proficient, Distinguished. If you ask any AI tool for a Danielson rating, that's what it will give you, because that's what Danielson publishes.

But Arizona districts overwhelmingly don't use those words. They relabel to Ineffective, Developing, Effective, Highly Effective — the four levels tied to the performance-classification language in ARS §15-537. Mesa's handbook uses them. Most Arizona district handbooks use them. It's close enough to the Danielson scale to map one-to-one, but the words on your form are the state's words, not Danielson's.

EvalScribe ships the Ineffective/Developing/Effective/Highly Effective scale by default for exactly this reason. It's a small thing that matters, because a write-up that says "Proficient" on a form that reads "Effective" is a write-up that wasn't built for Arizona.

Where the evidence actually comes from

Danielson splits cleanly, and Arizona district handbooks state the split directly. Two of the four domains come from watching; two come from artifacts.

Domains 2 and 3 — The Classroom Environment and Instruction — are gathered primarily through classroom observation. This is what you see: rapport, culture, procedures, behavior, the physical space, communication, questioning, engagement, assessment in the moment, and flexibility when a lesson goes sideways. Ten components. This is the bulk of the write-up and the part that eats your evening.

Domains 1 and 4 — Planning and Preparation, and Professional Responsibilities — come from artifacts the teacher provides. Lesson plans, unit plans, assessment designs, communication logs, reflection, records, evidence of professional growth. The teacher assembles these and shares them with the evaluator; you don't rate them from a classroom visit, because most of them never appear in a classroom visit.

This matters for what a tool can honestly do. EvalScribe drafts Domains 2 and 3 from your observation notes, because that's where the evidence lives. You can also capture Domain 1 and 4 artifacts — photograph a lesson plan, scan a communication log — and draft from those. But the observation is where the tool earns its keep, and any tool that offers to rate all four domains from a single classroom observation is offering to invent evidence for the two domains it never saw.

How the year is shaped

Local control means the details vary by district, but ARS §15-537 sets a floor, and most Arizona districts build something recognizable on top of it.

At least two observations. That's the statutory minimum. Probationary teachers — typically the first three years — are usually evaluated on all four domains across a fall and spring cycle, with a self-assessment each time.

Continuing teachers get lighter cycles. A continuing-status teacher rated Effective or Highly Effective may be evaluated on Domains 2 and 3 only, and the second observation may be waived unless the teacher requests it or the evaluator documents a reason for one. There's usually a minimum stretch — 60 calendar days in Mesa — between formal observations, and observations aren't conducted right before major breaks.

The alternative evaluation cycle. Under ARS §15-537, a teacher who earns Highly Effective three consecutive years in the same district can move to an expedited review — formal walk-throughs with written feedback rather than the full observation cycle. It's the state's way of not making your best veterans jump through the same hoops every year.

Student data still counts, but the district sets how. The old statewide 33% mandate is gone with local control, but districts still fold in a student-data component and decide its weight and its source. Mesa, for instance, uses school letter grades at 20%. Your district decides this, and it's outside anything an observation tool touches.

Your handbook is the authority on all of it. I'll keep saying that because in Arizona it's true in a way it isn't elsewhere.

Where AI genuinely helps — and where it doesn't

It helps with the language. You watched the lesson. Rendering what you saw in Danielson's component vocabulary, at the right level, is the drafting work, and drafting is what these tools are for.

It helps hold the edition and scale straight. A tool configured to your district's edition and Arizona's scale won't drift into 2022 language on a 2013 form, or hand you "Distinguished" when your form says "Highly Effective." That consistency is worth more in Arizona than in a single-framework state, precisely because there's more to get wrong.

It helps with evaluator consistency. Same district, same framework, four evaluators, and four slightly different readings of "Effective." A tool anchored to the rubric narrows that spread — and when a district chose its own instrument, defending consistency across evaluators is the whole ballgame.

It does not know which framework you're on. This is the one place I want to be blunt. In a local-control state, a tool has to be told your district's instrument, edition, and scale. Anything that assumes it knows is guessing, and in Arizona the odds of guessing wrong are real.

It does not rate your artifacts for you. Domains 1 and 4 are the teacher's evidence, assembled and shared. A tool can help you draft from them once you have them; it can't manufacture them from a lesson.

It cannot give you the rating. It gives you a draft. You accept it, change it, or throw it out. The judgment is the job, and it stays with the person who was in the room.

What to write down

The draft can only work with what you captured. A few notes that make the Arizona write-up easier:

Write down which lets you see the level. The Danielson jump from Effective to Highly Effective is almost always the shift from teacher-directed to student-directed. Note the moment students took over — ran the discussion, corrected each other, drove the inquiry. That moment is the difference between a 3 and a 4, and you won't reconstruct it at 9pm.

Write down engagement versus compliance. The Developing-to-Effective line is compliance versus genuine engagement. "Students completed the worksheet" is compliance. "Students argued about the answer" is engagement. Capture which one you saw.

Write down the pivot. Component 3e — flexibility and responsiveness — only shows itself when something goes wrong. A student's confusion, a plan that stalls, and what the teacher did next. It's the most observation-dependent component in the framework and the easiest to forget to note.

Write down the questions, roughly verbatim. 3b lives or dies on whether questions were recall or reasoning. "What year was it?" and "Why do you think it happened then?" are different components' worth of evidence.

Write down what you couldn't see. If you were in the room 25 minutes, you saw a slice. The honest write-up names its own edges, and in a two-observation state that's not weakness — it's accuracy.

Common scoring mistakes in Arizona

Assuming there's a state framework. The root mistake, and the one the AI tools make. There isn't. There's your district's choice.

Using the wrong edition's language. 2013 and 2022 are different instruments. Rating in one against a form built on the other produces a write-up that doesn't match the rubric.

Defaulting to the Danielson scale. Your form probably says Ineffective/Developing/Effective/Highly Effective, not Unsatisfactory/Basic/Proficient/Distinguished. Small mismatch, real credibility cost.

Rating Domains 1 and 4 from the observation. That evidence is the teacher's artifacts, not what you saw in 25 minutes.

Reading Highly Effective as "really good teaching." It's leadership and student-directed learning, not polished delivery. A masterful teacher-led lesson is Effective.

Trusting a generic tool's Arizona knowledge. It will confidently describe a framework your district may not use, an edition it's guessing at, and a scale that isn't on your form.

Arizona in context

Most states made a choice and imposed it. Arizona made a choice and delegated it. Neither is obviously right — a statewide framework buys consistency at the cost of local fit, and local control buys fit at the cost of consistency. Arizona bet on fit.

The practical consequence is that "how do I write an Arizona evaluation" has no single answer, and anyone who gives you one confidently hasn't understood the state. What Arizona has instead is a strong default — Danielson, in one of two editions, on the state's four-level scale — surrounded by fifty-plus alternatives that some districts genuinely use. A tool built for Arizona has to handle the default well and be honest about the rest.

How EvalScribe handles Arizona

EvalScribe is an AI-powered evaluation tool built by a practicing K-12 teacher.

The Danielson Framework for Teaching is built in — all four domains, all 22 components — with language across the four levels, the look-fors an observer watches for, and the distinctions that separate adjacent ratings. EvalScribe draws the Developing/Effective line and the Effective/Highly Effective line where Danielson draws them: compliance to engagement, and teacher-directed to student-directed. It drafts in the framework's own vocabulary, on the Ineffective/Developing/Effective/Highly Effective scale Arizona districts actually apply.

You observe the way you already do. Capture notes by typing, dictating, or photographing your handwriting — Smart Scan OCR handles the handwriting, including the pedagogical vocabulary that trips up generic OCR. EvalScribe maps your evidence to the right components, drafts best-fit ratings with the evidence behind them, and you review, edit, and export a clean PDF.

Honest scope, because it matters more than the pitch.

We focus on the observation domains. EvalScribe drafts Domains 2 and 3 — the ten components a classroom observation actually reaches. Domains 1 and 4 are built from artifacts the teacher provides; you can capture and draft from those too, but the observation is where the tool does its real work, and we're not going to pretend a lesson tells us how a teacher's communication log looks.

We ship the common frameworks, and we'll add yours. Arizona's repository holds more than fifty compliant instruments. EvalScribe already ships Danielson in both the 2013 and 2022 structures, plus Marzano, Stronge, CEL 5D+, NIET, and Kim Marshall — which covers most of what Arizona districts adopt. If your district runs something else, email [email protected]. In a local-control state, telling us your instrument is how it gets built, and we'd rather add it than pretend one framework fits all of Arizona.

We don't know your district's choice, and we won't guess. EvalScribe has to be pointed at your edition and your scale. That's not a limitation to hide — it's the honest shape of building for a state that hands the decision to the district.

On privacy: observation notes are processed transiently on Microsoft Azure and are not retained on our servers. Evaluation records live locally on your device. Nothing is used to train public AI models.

Three evaluations are free. The Arizona page is at evalscribe.com/arizona. If you're leading a rollout across a district rather than a building — standardizing an edition, a scale, and evaluator consistency across everyone who rates — that's at evalscribe.com/district.

FAQ

Does Arizona have a state teacher evaluation framework? Not a mandated one. Since 2019, changes to ARS §15-537, §15-341, and §15-189.06 mean districts are no longer required to align to the Arizona Framework for Measuring Educator Effectiveness. Arizona is a local-control state, and ADE maintains a repository of compliant instruments rather than requiring one.

Then which framework do I use? Whichever your district adopted. Most Arizona districts choose the Danielson Framework for Teaching. Your district's evaluation handbook is the authority.

Which Danielson edition? Either 2013 or 2022 — they're both in use and they're different instruments. The 2013 edition uses Classroom Environment, Instruction, and Professional Responsibilities; the 2022 edition uses Learning Environments, Learning Experiences, and Principled Teaching. Your handbook says which.

What scale does Arizona use? Most districts use Ineffective, Developing, Effective, Highly Effective — tied to ARS §15-537's performance-classification language — rather than Danielson's published Unsatisfactory/Basic/Proficient/Distinguished.

Which domains does EvalScribe draft? Domains 2 and 3, the observation domains — ten components. Domains 1 and 4 are the teacher's artifacts; EvalScribe can help you draft from them once you've captured them, but it focuses on what the observation reaches.

How many observations does Arizona require? At least two under ARS §15-537. Probationary teachers are typically observed on all four domains; experienced continuing teachers may be observed on Domains 2 and 3 only, with a second observation sometimes waived. Highly Effective teachers may qualify for an alternative cycle after three consecutive years.

What's the difference between Effective and Highly Effective? Broadly, who drives the learning. Effective is consistent, skilled, professional practice. Highly Effective adds integration, leadership, and student-directed learning — a community of learners and influence beyond the classroom. A polished teacher-led lesson is Effective.

My district uses a framework you don't list. Can EvalScribe still help? Probably. EvalScribe ships Danielson, Marzano, Stronge, CEL 5D+, NIET, and Kim Marshall. If yours is different, email [email protected] — telling us is how it gets added.

Can AI write my evaluation for me? It can draft the write-up for the observation domains. It can't watch the lesson, know which framework your district picked, rate your teacher's artifacts, or make the judgment. Those are yours.

References

  • ARS §15-537, Evaluation of certificated teachers — the observation requirements, performance classifications, and alternative evaluation cycle.

  • Arizona Department of Education, Teacher/Principal Evaluation — the 2019 local-control changes and the state model instruments (Arizona Professional Teaching Standards).

  • Arizona Framework for Measuring Educator Effectiveness — the repository model and the compliant-instrument approach.

  • Charlotte Danielson, The Framework for Teaching (2013 and 2022 editions) — the framework most Arizona districts adopt.

  • Your district's teacher evaluation handbook — in a local-control state, the authority on your framework, edition, scale, and cycle.

Verify any citation against the current source before relying on it, and treat your district handbook as the final word on which instrument you're rated against.

Related articles

Arizona teacher evaluationArizona Danielson frameworkDanielson Framework for Teaching ArizonaARS 15-537ARS 15-341Arizona Framework for Measuring Educator EffectivenessArizona local control evaluationlocal control teacher evaluationArizona Professional Teaching StandardsAzTSADE evaluation repositoryIneffective Developing Effective Highly EffectiveDanielson 2013 editionDanielson 2022 editionDanielson 22 componentsfour domainsClassroom Environment domainInstruction domainPlanning and PreparationProfessional Responsibilitiesobservation domainsartifact evidenceArizona teacher observationtwo observations Arizonacontinuing status teacherprobationary teacheralternative evaluation cycleHighly Effective three consecutive yearsMesa Public Schools evaluationstudent data percentage evaluationARS 15-203Arizona classroom walkthroughAI teacher evaluation Arizonateacher evaluation software Arizonateacher evaluation app Arizonaobservation write-up ArizonaTucson Unified teacher evaluationChandler Unified evaluationGilbert Public Schools evaluationPeoria Unified evaluationDeer Valley Unified evaluationScottsdale Unified teacher evaluationDysart Unified evaluationEvalScribe Arizonahow are teachers evaluated in Arizonadoes Arizona have a state evaluation frameworkwhich framework does Arizona usedoes Arizona still use the Arizona Framework for Measuring Educator Effectivenesswhich Danielson edition does my district usewhat scale does Arizona use for teacher evaluationwhat is the difference between effective and highly effectivehow many observations does Arizona requirewhich Danielson domains come from observationcan AI write teacher evaluationsbest AI tool for Arizona teacher evaluationsdoes EvalScribe support Danielson
Anthony D. Neely, Ph.D.

Anthony D. Neely, Ph.D.

Anthony Neely is the Founder of EvalScribe, a veteran educator, an AI integration consultant for teaching & learning, researcher, & author.

Back to Blog