Massachusetts educator evaluation, MA Educator Evaluation Framework, DESE Model System, DESE Model Rubric, 603 CMR 35, Standards of Effective Teaching Practice, Massachusetts 5-step evaluation cycle, Educator Plan Massachusetts, Proficient Exemplary Massachusetts, AI for Massachusetts teacher evaluation, collective bargaining educator evaluation

How can AI help Massachusetts evaluators write educator evaluations?

July 01, 202618 min read

How can AI help Massachusetts evaluators write educator evaluations?

Massachusetts's Educator Evaluation Framework — governed by 603 CMR 35.00 and administered through DESE's Model System — is built on four Standards of Effective Teaching Practice with a four-level rating scale (Unsatisfactory, Needs Improvement, Proficient, Exemplary). Districts can adopt the DESE Model System, adapt it, or use their own "comparably rigorous and comprehensive" system — and because evaluation procedures are a mandatory subject of collective bargaining in Massachusetts, most district rubrics are locally-negotiated adaptations of the DESE Model. AI can meaningfully help with the documentation burden of Massachusetts evaluations: translating observation notes into rubric-aligned language, mapping evidence to the four Standards, and producing draft ratings consistent with the framework's structure. AI does not replace evaluator judgment — particularly the Proficient-to-Exemplary call, the integration of student feedback and MCAS growth data, or Educator Plan decisions. Where AI lifts the most weight is the write-up itself.

What Massachusetts's Educator Evaluation Framework is

The Massachusetts Educator Evaluation Framework is the statewide system for K-12 educator observation and evaluation. It is governed by 603 CMR 35.00 (Evaluation of Educators), administered by the Massachusetts Department of Elementary and Secondary Education (DESE), and implemented through the DESE Model System for Educator Evaluation. The Model System was most recently updated in 2019 and 2025, with role-specific rubrics piloting in the 2025-26 school year.

Unlike states with a single mandated statewide rubric, Massachusetts law provides significant district flexibility. School committees and school districts can adopt the DESE Model System, adapt it, or revise their own evaluation system as long as it aligns to 603 CMR 35.00 and is "comparably rigorous and comprehensive." Because evaluation procedures are a mandatory subject of collective bargaining in Massachusetts, the specific rubric each district uses is often the result of a locally-negotiated agreement between district leadership and the teachers' union — meaning district rubrics vary considerably in specific Indicator language, weighting, and evidence expectations, even though all are grounded in the four Standards defined in 603 CMR 35.03.

What is required of every Massachusetts educator evaluation: an evaluation cycle aligned to the Standards and Indicators, a rating on each Standard on the four-level scale, an overall summative rating, and multiple measures of student learning as evidence.

The four Standards of Effective Teaching Practice

Massachusetts organizes teaching practice into four Standards defined in 603 CMR 35.03, each subdivided into Indicators in the DESE Model Rubric.

Standard I (Curriculum, Planning, and Assessment) covers four Indicators: Curriculum and Planning (I-A), Assessment (I-B), Analysis (I-C), and Well-Structured Lessons. This Standard captures the pre-lesson work — the coherence of curriculum, the design of assessments aligned to standards, and the use of data to analyze student progress.

Standard II (Teaching All Students) covers four Indicators: Instruction (II-A), Learning Environment (II-B), Cultural Proficiency (II-C), and Expectations (II-D). This Standard captures what happens inside the classroom — the quality of instruction, the safety and productivity of the learning environment, culturally responsive teaching, and the maintenance of high expectations for all students.

Standard III (Family and Community Engagement) covers three Indicators: Engagement (III-A), Collaboration (III-B), and Communication (III-C). This Standard captures the teacher's work outside the classroom to build meaningful partnerships with families and community members in support of student learning.

Standard IV (Professional Culture) covers six Indicators — the most of any Standard — including Reflective Practice, Professional Growth, Collaboration, Decision-Making, Shared Responsibility, and Professional Responsibilities. This Standard captures the teacher's contribution to the broader school and district culture as a reflective, collaborative professional.

The DESE Model Classroom Teacher Rubric organizes these four Standards into twelve Indicators and twenty-five elements total. Because Massachusetts districts adapt the Model Rubric through collective bargaining, many locally-negotiated rubrics use a subset of Indicators or adjust element language to fit local context.

The four-level rating scale — and the Proficient-to-Exemplary judgment

Massachusetts uses four performance levels: Unsatisfactory, Needs Improvement, Proficient, Exemplary. The DESE framework is explicit about what each means. Unsatisfactory represents performance that has not significantly improved following a Needs Improvement rating, or performance consistently below the requirements of a Standard. Needs Improvement indicates performance below the Standard but not yet Unsatisfactory — improvement is expected. Proficient represents "the rigorous expected standard" — demanding but attainable, fully satisfactory professional practice. Exemplary is reserved for performance "of such a high level that it could serve as a model" for other educators.

The hardest call in Massachusetts — and the one that most divides evaluators in calibration sessions — is Proficient versus Exemplary. Both reflect strong teaching: organized lessons, engaged students, evidence-based practice, productive classroom culture. What separates them is whether the practice could serve as a model for others. Proficient is the expected standard; Exemplary is the standard by which others learn. The DESE framing is intentional: Exemplary is not "very good Proficient" — it's practice that could be filmed, shared, or replicated across a school. That distinction is real, and it's the one most likely to get blurred when an evaluator is writing summative documentation quickly between observations.

The Massachusetts 5-step evaluation cycle

Massachusetts evaluation runs on a distinctive 5-step cycle that shapes how documentation accumulates across a school year.

Step 1 is Self-Assessment. The educator reviews their own practice against the Standards and Indicators, identifies areas of strength and areas for growth, and prepares for goal setting. Step 2 is Goal Setting and Educator Plan Development, in which the educator and evaluator meet to set professional practice goals and student learning goals, and to determine the appropriate Educator Plan type — Self-Directed Growth, Directed Growth, Improvement, or Developing Educator Plan. Step 3 is Plan Implementation, the longest phase of the cycle, in which observations occur, evidence is collected, and progress toward goals is monitored. Step 4 is Formative Assessment or Formative Evaluation, a mid-cycle checkpoint that documents progress and adjusts the plan as needed. Step 5 is Summative Evaluation, the end-of-cycle rating on each Standard and overall, based on evidence collected across the cycle.

The 5-step cycle means Massachusetts evaluation documentation isn't a single summative event — it's an accumulating record of observations, formative check-ins, and reflection over months. The documentation load compounds accordingly.

The four Educator Plan types — and what they mean for documentation

Massachusetts assigns each educator to one of four Educator Plan types, and the plan type shapes both the evaluation cycle length and the documentation intensity.

The Self-Directed Growth Plan is for educators with Professional Teacher Status who received a Proficient or Exemplary summative rating in the prior evaluation cycle. It runs on a 2-year cycle with lighter observation and evidence requirements. The Directed Growth Plan is for educators who received a Needs Improvement summative rating — it runs on a 1-year cycle with more targeted evidence collection and increased evaluator involvement. The Improvement Plan is for educators who received an Unsatisfactory summative rating — a 30-day to 1-year plan with clear expected improvements and consequences for non-improvement. The Developing Educator Plan is for educators without Professional Teacher Status — a 1-year cycle for teachers in their first three years or in a new district.

For evaluators, the practical implication is that Educator Plan type drives how much documentation a given teacher requires in a given year. An evaluator managing a caseload of Self-Directed Growth Plan teachers has a very different documentation load than one managing a caseload of Directed Growth Plan teachers.

After the observation: where the work actually piles up

The observation itself is the part most evaluators trained for. What no one warned them about was the volume of writing that follows.

A single formal observation in a Massachusetts-aligned district can produce three to seven pages of evaluation documentation: rubric ratings across Indicators covered in the district's rubric, narrative justifications for each rating, evidence quotations from observation notes, alignment to the Educator Plan goals, and language for the post-observation conference. Over an evaluation cycle, a single teacher accumulates observation evidence, formative assessment documentation, and eventually a summative narrative that pulls it all together. The formative and summative documentation stages in particular require pulling together evidence from months of observations into a coherent, defensible narrative.

This is where the day actually goes. The observation takes twenty minutes. The write-up takes an hour.

Where AI helps with Massachusetts evaluation write-ups

The strongest fit between current AI and Massachusetts evaluation work is the translation problem. An evaluator captures observation notes in shorthand — half-sentences, abbreviations, fragments — and then has to translate those notes into the formal language of the Massachusetts rubric. That translation is mechanical work. It's not where evaluator judgment lives; it's where evaluator time gets spent.

AI handles three pieces of that translation reliably. First, it can take fragmentary observation notes and produce coherent, evidence-grounded prose in the voice of the Massachusetts framework. Second, it can map specific pieces of evidence to the relevant Indicators across the four Standards — telling the evaluator that the note about student discourse supports Indicator II-A (Instruction), while the note about anchor charts referencing prior learning supports Indicator I-A (Curriculum and Planning). Third, it can produce internally consistent draft language across Standards in the same evaluation, so the formative or summative reads as a coherent document rather than a stack of disconnected paragraphs.

Where AI saves the most time is the formative and summative aggregation piece. Most evaluators don't struggle to write any one observation write-up. They struggle to pull together three months of observations, evidence artifacts, and Educator Plan progress notes into a formative or summative document that reads as coherent professional judgment. That aggregation work is where AI can shave the most hours off an evaluation cycle.

Where AI doesn't help — and what stays with the evaluator

The honest scope is narrower than the marketing on most AI tools suggests.

AI cannot make the borderline Proficient-to-Exemplary judgment for an evaluator. That call requires watching the lesson in real time, knowing the teacher's history and context, and judging whether the evidence rises to the standard of "practice that could serve as a model." AI can suggest a rating based on the notes provided, but the evaluator owns the final call.

AI cannot integrate the required student feedback into evaluation judgments. Massachusetts requires student feedback as one of the evidence types (603 CMR 35.07), and DESE provides model surveys for grades 3-12. The interpretation of that student feedback — reading patterns across responses, weighing against other evidence sources, adjusting for cohort characteristics — is evaluator work AI doesn't do.

AI cannot handle the MCAS Student Growth Percentile integration for eligible teachers. For educators with 20 or more students who take statewide assessments, evaluators must consider SGP data as part of the summative rating. The DESE guidance defines "anticipated student learning gains" as a mean SGP between 35-65, with 65+ exceeding expected growth and 35 or below not meeting expected growth. Integrating growth data with other evidence sources is evaluator judgment work.

AI cannot navigate the Educator Plan work. The self-assessment conversation, the goal-setting meeting, the Educator Plan type determination, and the harder conversations when a plan moves from Self-Directed Growth to Directed Growth or Improvement — these are human work. AI can help draft language for plan documentation based on evidence, but it cannot do the conversation itself.

And AI cannot do the collective-bargaining-aware work that Massachusetts evaluation requires. Every district's rubric is a locally-negotiated document, and evaluators need to know the specific Indicators, weightings, and evidence expectations their district agreed to. AI trained on the DESE Model Rubric doesn't know that Framingham weighted Indicator II-A differently than the Model, or that Wellesley added a locally-negotiated element on trauma-informed practice. That contextual knowledge stays with the evaluator.

What AI does well, it does well. What it doesn't do, it shouldn't pretend to do. Pasting observation notes into ChatGPT works fine as a glorified search engine — but for rubric-aligned Massachusetts documentation, the gap between "sounds polished" and "actually defensible under your district's collectively-bargained evaluation system" is wider than the polish suggests.

Observable evidence: what each Massachusetts Standard looks like in practice

Each Standard has its own surface features — the things an evaluator actually sees, hears, and notes during an observation.

Standard I (Curriculum, Planning, and Assessment) is largely evident before and after the lesson rather than during it. Look for lesson plans showing alignment to the Massachusetts Curriculum Frameworks, evidence of anticipating student misconceptions, learning objectives that are measurable and communicated to students, and assessments aligned to the stated learning targets. In classrooms with SEI (Sheltered English Immersion) endorsements, look for Language Objectives alongside Content Objectives.

Standard II (Teaching All Students) is where the lesson lives. Look for students doing the "heavy lifting" of thinking and talking; a variety of grouping structures (pairs, small groups, whole class) used purposefully; higher-order questioning that goes beyond recall; tiered supports for students with IEPs and English Learners; safe and respectful environment where routines are seamless; and cultural responsiveness evident in materials, examples, and student voice choices.

Standard III (Family and Community Engagement) is largely evident outside the observation window. Look for two-way communication artifacts (translated communications for non-English-speaking families, class website updates, positive contacts not just deficit-reporting), evidence of family collaboration on student learning goals, and cultural responsiveness in family engagement approaches.

Standard IV (Professional Culture) is the Standard most likely to be overrated because most evidence lives outside the observation window. Look for specific evidence cited in self-reflection, active contributions in PLCs and collaborative planning, professional learning visibly applied to practice, and acceptance and action on feedback from prior evaluation cycles.

Common scoring mistakes under the Massachusetts framework

Even experienced evaluators slip into patterns worth naming.

The first is over-Exemplifying Standard IV. Because Standard IV lives partly outside the lesson, evaluators tend to credit "professional-feeling" teachers with high Standard IV ratings regardless of evidence. The fix is to require the same evidence rigor for Standard IV as for Standards I and II — name the PLC contribution, name the specific PD applied to practice, name the specific feedback the teacher acted on.

The second is the Proficient default. When evidence is thin in either direction, evaluators slot Proficient because it feels safe. But the DESE framing is explicit: Proficient requires evidence of "the rigorous expected standard" — demanding, but attainable. It's not a default. An Indicator with insufficient evidence is a coaching opportunity, not a default Proficient.

The third is Indicator conflation within Standard II. II-A (Instruction), II-B (Learning Environment), II-C (Cultural Proficiency), and II-D (Expectations) feel related, and evaluators often justify multiple with the same evidence. They're distinct claims. Instruction is about the pedagogical moves; Learning Environment is about the climate and structures; Cultural Proficiency is about responsive practice for diverse learners; Expectations is about the rigor and belief in student capacity.

The fourth is evidence-thin Exemplary. Exemplary in Massachusetts requires evidence of practice that could serve as a model. "The teacher did it well" is not Exemplary — it's Proficient. Exemplary requires evidence that the practice is at a level where it could be filmed, shared with peers, or replicated across the school as a professional development artifact.

The fifth is failing to integrate student feedback. Student feedback is a required evidence type in Massachusetts (603 CMR 35.07). Evaluators sometimes rate Standards on observation evidence alone without integrating what students said about their teacher's practice. That's not compliant with the framework's evidence requirements.

The sixth is uniquely Massachusetts: applying the DESE Model Rubric language when the district's collectively-bargained rubric uses different language. Evaluators trained on the DESE Model sometimes drift back to Model rubric language in write-ups even when the district's negotiated rubric uses different Indicator names or weightings. Teachers and their union representatives will notice — and it's a defensibility problem in any grievance process.

Massachusetts in context — the district-flexibility model

Massachusetts's approach to educator evaluation is genuinely different from many other states. States like Texas (T-TESS), Florida (FEAPs), Georgia (TKES), Arkansas (TESS), and Tennessee (TEAM) operate single statewide rubrics that every district uses. Massachusetts — like Indiana, Wyoming, and Missouri — provides a state model (the DESE Model System) and allows districts to adopt it, adapt it, or use their own comparably rigorous system.

The Massachusetts version of this flexibility model has a distinctive feature: because evaluation procedures are a mandatory subject of collective bargaining, the district-level adaptation happens through a formal negotiation between district leadership and the teachers' union. This produces rubric variation that is more structured than pure district-discretion states but less uniform than single-rubric states. Most Massachusetts districts operate with a locally-negotiated rubric that inherits the four Standards from 603 CMR 35.03 but adjusts Indicator language, weighting, or evidence expectations to fit local priorities and bargaining outcomes.

For Massachusetts evaluators, this means the specific rubric your district uses — the exact Indicators, the specific element language, the weighting choices — is your operational reality. Any tool that supports Massachusetts evaluation has to be flexible enough to accommodate these locally-bargained adaptations without losing the four-Standards structure that makes Massachusetts evaluation defensible.

How EvalScribe handles Massachusetts

EvalScribe is built around the four Massachusetts Standards of Effective Teaching Practice — with the four-level rating scale (Unsatisfactory, Needs Improvement, Proficient, Exemplary) applied through rubric-aware logic. The tool handles the translation problem natively: an evaluator captures observation notes by typing, dictating, or photographing handwritten notes (Smart Scan OCR converts handwriting to text), and EvalScribe maps that evidence to the four Standards and drafts ratings and comments in the framework's actual voice.

EvalScribe currently covers the four Standards through their core Indicators — Curriculum and Planning (I-A), Assessment (I-B), Instruction (II-A), Learning Environment (II-B), Family Engagement (Standard III), and Professional Culture (Standard IV). Expanded coverage of the full DESE Model Rubric — all twelve Indicators including Analysis, Cultural Proficiency, and Expectations, plus the twenty-five elements — is on our roadmap for a future update. For districts using a locally-negotiated rubric grounded in the four Standards, the underlying rubric-aware logic still applies.

Beta testers report saving 30 to 60 minutes per evaluation versus writing the documentation traditionally. Across a typical evaluator load — say thirty teachers on mixed Educator Plan types across an evaluation cycle — that math adds up to dozens of hours back over a year.

The evaluator stays in control of every rating, every comment, and every piece of mapped evidence. EvalScribe drafts; the evaluator decides. The Proficient-to-Exemplary judgment call, the integration of student feedback and MCAS SGP data, the Educator Plan decisions, and the harder coaching conversations remain with the practitioner who knows the teacher. More detail on the Massachusetts workflow is available at evalscribe.com/massachusetts.

Frequently asked questions about Massachusetts and AI

Where can I read the official Massachusetts Educator Evaluation regulations and rubric? The primary sources are maintained by the Massachusetts Department of Elementary and Secondary Education: the DESE Educator Evaluation home page, the 603 CMR 35.00 full regulation, the Massachusetts Model System for Educator Evaluation, and the DESE Model Classroom Teacher Rubric.

Does AI replace evaluator judgment in Massachusetts? No. AI drafts evidence-mapped evaluations and suggested ratings; the evaluator reviews, edits, and finalizes every document. The borderline judgment calls — particularly Proficient-to-Exemplary and the integration of multiple evidence sources — stay with the practitioner.

My district negotiated a rubric that's different from the DESE Model. Will AI work for us? In most cases yes, because most locally-negotiated rubrics still use the four Standards from 603 CMR 35.03 as their foundation. Tools built around the four-Standards structure with district-flexible logic are the better fit. Generic AI tools that assume one universal rubric miss the Massachusetts reality.

How does AI handle the 5-step evaluation cycle? AI is strongest at the observation-driven pieces of the cycle: capturing evidence during Plan Implementation, drafting formative assessment documentation, and pulling summative narratives together at the end of the cycle. The self-assessment, goal setting, and Educator Plan development steps stay with the educator and evaluator.

Does AI integrate student feedback into evaluations? Not directly. Student feedback is a required evidence type in Massachusetts, and its interpretation is evaluator judgment work. AI can draft the observation-based rubric portion of an evaluation; the evaluator integrates student feedback, MCAS SGP data, and other evidence into the summative judgment.

Does AI work for both walkthroughs and formal announced observations? Yes. Brief unannounced observations, longer announced observations, and summative evidence collection all flow through the same capture-and-draft pipeline, so the documentation across observation types stays consistent throughout the evaluation cycle.

How can I see what EvalScribe looks like for Massachusetts specifically? The Massachusetts page at evalscribe.com/massachusetts walks through the four-Standards workflow, the framework handling, and a sample exported evaluation. EvalScribe is available on iOS and macOS via the App Store, with three free evaluations to start.

If you're evaluating educators in Massachusetts using the DESE Model System or a locally-negotiated district rubric, see how EvalScribe handles the workflow at evalscribe.com/massachusetts.

References

Related articles


Page maintained by Anthony D. Neely, Ph.D. — practicing K-12 educator with nearly 20 years in the classroom, 2025–2026 Walker County Distinguished Teacher of the Year, and co-founder of EvalScribe. Framework details verified against current Massachusetts DESE source documents.

Last reviewed: July 1, 2026.

Anthony D. Neely, Ph.D.

Anthony D. Neely, Ph.D.

Anthony Neely is the Founder of EvalScribe, a veteran educator, an AI integration consultant for teaching & learning, researcher, & author.

Back to Blog