California teacher evaluation, CSTP, California Standards for the Teaching Profession, Stull Act, Commission on Teacher Credentialing, California induction program, AI for California teacher evaluation, California classroom observation

How can AI help California evaluators write teacher evaluations under the CSTP?

July 05, 202615 min read

How can AI help California evaluators write teacher evaluations under the CSTP?

EvalScribe

July 01, 2026 • 17 min read

California's teaching practice is described by the California Standards for the Teaching Profession (CSTP) — six standards and twenty-eight elements, revised in 2024 and adopted by the Commission on Teacher Credentialing that April. The CSTP scores practice on a five-level Developmental Continuum, where Applying — not the top level — is the professional standard. AI can meaningfully help with the observation write-up: mapping evidence to the right elements, drafting best-fit ratings across the continuum, and generating feedback in the standards' own language. What it can't do is determine how your district's formal evaluation program actually scores teachers, since California leaves that to local bargaining under the Stull Act.

What California's teacher evaluation landscape actually is

California is unusual among states in how it separates standards from evaluation. The CSTP is a professional standards document, developed by the Commission on Teacher Credentialing (CTC) in partnership with the California Department of Education. The 2024 revision — the first major update since 2009 — was built over a 2022–23 work group process and adopted at the Commission's April 2024 meeting. The CSTP describes what effective teaching looks like; it is foundational to teacher preparation and, especially, to the state's mandatory two-year induction program, where beginning teachers work with a mentor to self-assess against the CSTP and build an Individual Learning Plan (ILP).

Formal, consequential teacher evaluation, by contrast, is governed by the Stull Act (California Education Code §44660–44665), first enacted in 1971 and renumbered in 1976. The Stull Act requires every district to maintain an evaluation program and sets minimum frequency — but it leaves the actual rubric, scoring scale, and process to local bargaining between the district and its teachers' union. Some districts adopt the CSTP's own Developmental Continuum directly as their evaluation instrument. Others build a locally negotiated rubric that draws on CSTP and Danielson's Framework for Teaching but scores on a different scale entirely — Los Angeles Unified, for example, has historically used a four-point scale (ineffective, approaching effective, effective, accomplished) rather than the CSTP's five levels. Understanding which situation you're in matters for how you use any evaluation tool, including this one.

The six standards and twenty-eight elements

The CSTP organizes teaching practice into six standards:

Standard 1 — Engaging and Supporting All Students in Learning (4 elements): focus on students; knowledge of students; student backgrounds and family engagement; and diversity and equity.

Standard 2 — Creating and Maintaining Effective Environments for Student Learning (4 elements): the learning environment; student behavior; organizational and resource management; and an inclusive environment.

Standard 3 — Understanding and Organizing Subject Matter for Student Learning (5 elements): knowledge of subject matter and pedagogy; connecting subject matter to real-world contexts; curriculum and resources for specific students and groups; content and skills across subjects; and curriculum materials and resources.

Standard 4 — Planning Instruction and Designing Learning Experiences for All Students (4 elements): planning instruction; designing and developing instruction; facilitating instruction; and adapting instruction.

Standard 5 — Assessing Students for Learning (4 elements): understanding and using assessments; interpreting and using assessment data; communication of assessment and data; and assessment for continuous improvement.

Standard 6 — Developing as a Professional Educator (7 elements): reflection on practice; focused professional learning; collaboration with colleagues; collaboration with families, guardians, and the community; ethical conduct and professional responsibilities; activating access and equity; and personal growth and well-being.

That's twenty-eight elements in total. A defining feature of the 2024 revision is how thoroughly it weaves equity, cultural responsiveness, and social-emotional development into every standard, rather than treating them as a separate add-on — a deliberate design choice reflecting the CTC's stated goal of supporting California's full range of learners.

The five-level Developmental Continuum — and why Applying is the standard

The CSTP's scale runs Emerging, Exploring, Applying, Integrating, and Innovating. What makes it distinctive among state frameworks is its explicit framing as a developmental continuum rather than a scoreboard — the standards body itself describes it more like a map of a growth trajectory than a ranking.

Emerging describes a teacher still building understanding of the standard, leaning heavily on provided structures. Exploring describes experimentation — attempting practices, but inconsistently. Applying is the professional standard: the teacher consistently and effectively applies the practice for most students. Integrating describes practice that's become fluid and responsive, with real-time adjustment for individual students. Innovating is the top: exemplary, leadership-level practice, where students show strong agency and the teacher actively advances the profession.

The critical point for evaluators is that Applying is not a middling score. It is the level the standards are designed for a competent, effective teacher to consistently reach. Integrating and Innovating describe something beyond solid effective teaching — first greater fluidity and responsiveness, then genuine leadership and advancement of practice for others. Rating a strong, consistent teacher as Applying is not underrating them; it's rating them accurately. Reserving Innovating for real leadership-level evidence keeps the top of the scale meaningful.

Two separate processes: induction and formal evaluation

It helps to keep two California processes distinct, because both involve the CSTP but serve different purposes.

The induction program is the credentialing pathway every new California teacher completes over two years to move from a preliminary to a clear credential. It centers on the CSTP directly: with a trained mentor, the beginning teacher self-assesses against the standards, sets goals, and documents growth through an Individual Learning Plan. This is developmental and mentor-driven — not the same event as an administrator's formal evaluation.

Formal evaluation under the Stull Act is what determines employment status and satisfies the district's legal evaluation obligations. Frequency is set by statute (Ed. Code §44664): at least once each school year for probationary staff, at least every other year for staff with permanent status, and, for veteran permanent staff with ten or more years in the district who are highly qualified and were previously rated at or above standard, as infrequently as every five years by mutual agreement between the evaluator and the teacher. The Stull Act also requires that evaluations reasonably relate to pupil progress toward state and district standards — a requirement that came to public attention through the Doe v. Deasy litigation against Los Angeles Unified. Exactly how a district's formal rubric scores practice, and how it weighs pupil-progress measures, is set through local bargaining, not by the CSTP document itself.

After the observation: where the work piles up

Whether the context is induction reflection or a formal Stull Act observation, the underlying task is the same: turn what happened in a classroom into evidence-based, standards-aligned writing.

A single CSTP write-up asks the evaluator or mentor to take fragmentary notes and turn them into evidence mapped across up to twenty-eight elements in six standards, decide a best-fit level for each on a five-level continuum, and draft feedback in the standards' language that a teacher can act on — while keeping Applying anchored as the real professional standard rather than treating it as a disappointing middle score. Multiply that across a caseload of teachers and multiple observations a year, and the documentation becomes the largest time cost in the process.

The observation takes a class period. Turning it into a defensible, standards-aligned write-up takes far longer — and the things that make a CSTP write-up defensible are exactly the things that make it slow to produce by hand.

Where AI helps with the CSTP write-up

The strongest fit between current AI and CSTP documentation is the mapping-and-translation problem. An observer captures notes in fragments and needs to sort them across the standards and render them in the CSTP's own language.

AI handles three parts of that reliably. First, it turns fragmentary notes into coherent, evidence-anchored prose in the CSTP's voice — the difference between "kids helped each other figure out the word problem" and a sentence tied to the Standard 4 element on facilitating instruction. Second, it maps evidence to the relevant elements across all six standards, including evidence that supports more than one — a well-run, differentiated small-group task can touch Standard 1's knowledge of students, Standard 4's adapting instruction, and Standard 2's inclusive environment at once. Third, it maintains consistency across all twenty-eight elements so the write-up reads as one coherent picture of a teacher's practice.

For California specifically, the payoff is in holding the continuum's own logic steady: drafting Applying as the genuine professional standard it is, not padding ratings toward Integrating or Innovating without leadership-level evidence, and keeping the equity and cultural-responsiveness threads grounded in what was actually observed.

Where AI doesn't help — and what stays with the evaluator

The honest scope here is narrower than most AI marketing implies, and it's shaped by California's particular structure.

AI cannot tell you which rubric your district actually uses for formal evaluation. Because the Stull Act leaves that to local bargaining, a tool built around the state's CSTP continuum is not automatically the same as your district's negotiated evaluation instrument — that has to be confirmed locally.

AI cannot manage the induction program relationship. The mentor-teacher self-assessment process and the Individual Learning Plan are relational and developmental; a drafting tool can produce useful evidence-based writing, but it doesn't replace the mentoring conversation.

AI cannot incorporate pupil-progress measures into an overall rating. The Stull Act requires evaluations to reasonably relate to pupil progress; how that's measured and weighted is a separate, locally determined process outside a classroom-observation write-up.

AI cannot verify genuine student agency or leadership-level practice from a tidy paragraph. Innovating is meant to be rare and specific; confirming it requires the evaluator's own judgment about what was actually observed.

And AI cannot supply the context a seasoned California evaluator brings — that this school just adopted a new curriculum, that this teacher is new to a dual-immersion program, that a lesson fell the week before state testing. That context shapes a rating in ways no chatbot reconstructs.

What AI does well, it does well. Pasting notes into a general chatbot functions as a glorified search engine — but for CSTP documentation that maps evidence correctly, holds the developmental continuum's actual logic, and survives scrutiny in a Stull Act process, the gap between "sounds polished" and "actually defensible" is real.

Observable evidence: what each standard looks like in practice

Each standard surfaces in things an observer actually sees and notes.

Standard 1 shows in student-centered tasks that value diversity as an asset, visible knowledge of individual students' backgrounds and needs, genuine family partnership, and instruction that actively counters bias and inequitable access. Standard 2 shows in a respectful, identity-affirming climate, culturally responsive and restorative behavior systems, well-organized routines and resources, and inclusive structures built on student assets.

Standard 3 shows in accurate, coherent subject-matter instruction connected to real-world and culturally relevant contexts, curriculum adapted for specific learners and groups, cross-subject application, and well-selected, bias-reviewed materials. Standard 4 shows in student-informed planning, varied and sequenced instructional design, responsive facilitation with multiple ways to demonstrate learning, and genuine differentiation for a wide range of learners.

Standard 5 shows in multiple, purposeful assessment types, data that visibly informs instructional decisions, clear communication of assessment information to students and families, and ongoing improvement of assessment practice. Standard 6 shows in accurate self-reflection, applied professional learning, productive collaboration with colleagues and families, consistent ethical conduct, active equity advocacy, and sustained personal well-being. Across every standard, the signature of Integrating and Innovating is fluid, responsive practice that extends beyond solid, consistent Applying-level teaching.

Common scoring mistakes under the CSTP

Even experienced California evaluators and mentors fall into a few recognizable patterns.

The first, and most consequential, is treating Applying as a disappointing score. Because five-point scales often run "bad to great," it's tempting to read Applying as merely adequate. But the CSTP explicitly designs Applying as the professional standard. A teacher who is reliably, effectively applying the practice for most students is being rated accurately at Applying — not shortchanged.

The second is inflating toward Integrating or Innovating without the evidence. Innovating in particular is meant to describe leadership-level, profession-advancing practice. Awarding it for strong classroom teaching alone — without evidence of the fluid, real-time responsiveness Integrating requires, let alone the leadership Innovating requires — erodes what the top of the continuum signals.

The third is confusing the CSTP continuum with your district's formal rubric. Some districts use the CSTP directly; others use a differently structured, differently scaled locally bargained instrument. Assuming they're identical produces write-ups that don't match what the district's evaluation process actually requires.

The fourth is treating equity and SEL content as a separate checklist item. The 2024 CSTP weaves cultural responsiveness, equity, and social-emotional development into every standard, not as an add-on domain. Evidence for these threads should appear throughout a write-up, not in an isolated section.

The fifth is conflating induction reflection with formal evaluation. The mentor-guided ILP process under induction and a Stull Act observation serve different purposes and audiences; the tone and stakes of the writing should reflect which one is happening.

California in context — standards without a single mandated scoring engine

California sits at an unusual point on the national spectrum. Most states pair a set of teaching standards with a single state-designated (or at least state-endorsed) scoring instrument. California instead maintains a strong, well-developed standards document — the CSTP — while leaving the actual evaluation instrument almost entirely to local bargaining under the Stull Act. That produces real variation: some districts run the CSTP's own continuum as their evaluation tool, and others run locally negotiated rubrics with different scales and structures, sometimes hybridizing CSTP with Danielson's Framework for Teaching.

What's constant statewide is the CSTP's role in teacher preparation and induction — every new California teacher encounters these standards on the path to a clear credential, regardless of which rubric their eventual district uses for formal evaluation. That makes the CSTP the closest thing California has to a shared professional language for teaching practice, even where the summative evaluation mechanics diverge by district.

How EvalScribe handles the CSTP

EvalScribe is built around the actual 2024 CSTP — all six standards, all twenty-eight elements — rather than approximating them with generic teaching-evaluation prose. Every element carries the standards' own language across the five levels, the look-fors an observer watches for, and the distinctions that separate adjacent levels on the continuum. An evaluator or mentor captures notes by typing, dictating, or photographing handwriting (Smart Scan OCR converts it to text), and EvalScribe maps that evidence to the relevant elements and drafts best-fit ratings and feedback in the standards' voice — treating Applying as the genuine professional standard it's designed to be, and reserving Integrating and Innovating for practice that's actually more fluid, responsive, and leadership-level.

A few honest scope notes, because California's structure demands them. First, the CSTP is the state's professional standards continuum; EvalScribe is built around it directly. Whether that matches your district's own formal Stull Act evaluation instrument depends on local bargaining — many districts use the CSTP continuum as-is, while some use a differently structured local rubric, and we're glad to hear from you at [email protected] if yours differs. Second, EvalScribe supports the observation write-up; the induction program's mentor relationship and Individual Learning Plan, and any district's pupil-progress measures, remain separate processes outside the app.

Beta testers report saving 30 to 60 minutes per evaluation versus writing the documentation traditionally. Across a full evaluation cycle, that adds up to dozens of hours back. More detail on the California workflow is available at evalscribe.com/california.

Frequently asked questions about the CSTP and AI

Is the CSTP the same as my district's evaluation rubric? Not necessarily. The CSTP is the state's professional standards continuum; formal Stull Act evaluations are locally bargained and may use the CSTP directly or a different local instrument.

Why isn't Innovating the professional standard? Because the CSTP frames practice as a growth trajectory. Applying is the level a consistently effective teacher is designed to reach; Integrating and Innovating describe more fluid and, eventually, leadership-level practice beyond that.

Does AI replace evaluator or mentor judgment? No. EvalScribe drafts best-fit ratings and evidence-based feedback; the evaluator or mentor reviews, edits, and finalizes every rating.

Does EvalScribe handle the induction program's ILP? No. EvalScribe drafts observation write-ups that can support induction reflection, but the Individual Learning Plan and mentor relationship are separate, teacher-and-mentor-owned processes.

Does EvalScribe incorporate pupil-progress measures? No. The Stull Act requires evaluations to relate to pupil progress, but how that's measured and combined with the observation is a separate, district-determined process.

Where can I read the official CSTP sources? The Commission on Teacher Credentialing publishes the 2024 California Standards for the Teaching Profession; the statutory basis for formal evaluation is the Stull Act, Cal. Ed. Code §44660–44665.

If you're evaluating teachers in California on the CSTP, see how EvalScribe handles the six-standard, twenty-eight-element workflow at evalscribe.com/california.

References

Related articles


Page maintained by Anthony D. Neely, Ph.D. — practicing K-12 educator with nearly 20 years in the classroom, 2025–2026 Walker County Distinguished Teacher of the Year, and co-founder of EvalScribe. Framework details verified against the 2024 California Standards for the Teaching Profession and Commission on Teacher Credentialing sources.

Last reviewed: July 5, 2026.

Anthony D. Neely, Ph.D.

Anthony D. Neely, Ph.D.

Anthony Neely is the Founder of EvalScribe, a veteran educator, an AI integration consultant for teaching & learning, researcher, & author.

Back to Blog