
How can AI help Oklahoma evaluators write teacher evaluations under the TLE Tulsa Model?
How can AI help Oklahoma evaluators write teacher evaluations under the TLE Tulsa Model?
July 01, 2026 • 18 min read

Oklahoma's teacher evaluation runs on the Teacher and Leader Effectiveness (TLE) system, and the most widely used qualitative framework within it is the Tulsa Model — a Danielson-based rubric of five domains and twenty indicators, built by Tulsa Public Schools and validated by the MET Validation Engine and VARC studies. What makes the Tulsa Model distinctive isn't just its structure; it's how the pieces are scored. The five domains carry different weights, so Instructional Effectiveness alone accounts for half the composite. The rating scale runs Ineffective / Needs Improvement / Effective / Highly Effective / Superior, and any indicator scored off Effective carries a documentation requirement. AI can meaningfully help with the observation write-up — mapping notes to the right indicators and drafting rubric-aligned language. It does not replace the evaluator's judgment, the required narrative comments, the weighted composite the evaluator or district calculates, or the certified-observer decisions the state builds the system around.
What Oklahoma's teacher evaluation system is
Oklahoma's teacher evaluation operates under 70 O.S. § 6-101.16 and Okla. Admin. Code § 210:20-41-1, administered by the Oklahoma State Department of Education (OSDE) through its Teacher and Leader Effectiveness office. The TLE system sets a qualitative observation component and a set of state-approved frameworks that districts choose among. Unlike states with a single mandatory statewide rubric applied uniformly (Georgia TKES, Texas T-TESS, Florida FEAPs, Tennessee TEAM), Oklahoma lets each district select the OSDE-approved qualitative framework that fits its context.
Several frameworks carry OSDE approval: the Tulsa Model, the Marzano Focused Teacher Evaluation Model, Marzano 2014, the McREL framework, and the Oklahoma TAP Teaching Standards framework. A district picks one and applies it to all of its teachers. The Tulsa Model is the most widely adopted, and for good reason — it was developed collaboratively by Tulsa Public Schools teachers and administrators on a Charlotte Danielson foundation, and it's one of the few teacher-observation instruments independently studied for predictive validity. The MET Validation Engine and VARC research found that indicators in the Tulsa Model correlated with student-achievement gains, with the strongest predictive power in the indicators around monitoring student learning and adjusting instruction accordingly.
What ties an Oklahoma district's system together is that state backbone: the five-tier rating scale defined in administrative code, a chosen qualitative framework, and the certified-evaluator requirement that governs who may rate a teacher's practice at all.
The five weighted domains and twenty indicators
The Tulsa Model organizes practice into five domains and twenty indicators. The detail most evaluators underuse is that the domains are not weighted equally. Under the model's weighted-scoring design, the composite is a weighted average, and the weights are lopsided on purpose:
Domain 1 — Classroom Management (30%) covers six indicators: Preparation, Discipline, Building-Wide Climate Responsibility, Lesson Plans, Assessment Practices, and Student Relations.
Domain 2 — Instructional Effectiveness (50%) covers ten indicators: Literacy, Current State Standards, Involves All Learners, Explains Content, Clear Instruction & Directions, Models, Monitors, Adjusts Based Upon Monitoring, Establishes Closure, and Student Achievement.
Domain 3 — Professional Growth & Continuous Improvement (10%) covers two indicators: Professional Learning and Professional Accountability.
Domain 4 — Interpersonal Skills (5%) covers one indicator: Effective Interpersonal Skills.
Domain 5 — Leadership (5%) covers one indicator: Professional Involvement & Leadership.
Twenty indicators, but a weighted composite in which Instructional Effectiveness carries as much as the other four domains combined. A distinguished Leadership score barely moves the overall rating; a shaky Instructional Effectiveness score moves it a lot. That weighting is a design statement — the state is telling evaluators where the instructional heart of the observation sits — and it's the single most common thing generic tools get wrong, because they treat every indicator as if it counts the same.
The rating scale and the Effective-to-Superior climb
The Tulsa Model uses five performance levels: Ineffective (1), Needs Improvement (2), Effective (3), Highly Effective (4), Superior (5).
Ineffective means performance is harmful or absent. Needs Improvement means inconsistent performance that partially meets expectations and requires development. Effective is the professional standard — consistent practice aligned to state standards, the level a solid Oklahoma teacher lives at. Highly Effective exceeds expectations, and Superior is exemplary, leading-and-modeling practice.
The interesting line — the one that divides evaluators in calibration — is the climb from Effective through Highly Effective to Superior. The surface features overlap: organized lessons, engaged students, clear objectives, a respectful room. What separates the levels is who is doing the intellectual work. At Effective, the teacher is driving effective instruction. At Highly Effective, students begin to demonstrate self-discipline and connect new learning to what they already know. At Superior, students are driving and extending their own learning — formulating their own questions, monitoring their own progress, sustaining the classroom culture without the teacher managing it. Read the top-level rubric language across the indicators and the same pivot recurs: at Involves All Learners, Superior students "initiate or develop their own activities"; at Monitors, the Superior teacher recognizes when a strategy has stopped working and changes it live; at Establishes Closure, Superior students connect the lesson to real-world use they identify themselves. If those student-ownership markers aren't present, the correct rating is Effective — even for excellent teacher-directed instruction.
Two Oklahoma-specific mechanics sit on top of the scale. First, the model scores on best fit, not perfect fit: if a teacher shows evidence at the "3" level across the majority of an indicator's rubric language, the evaluator awards a 3 even if it isn't a flawless match, then uses the model's "push-pin" process to set a growth expectation. Second, and critically for defensibility, any rating off Effective triggers a documentation requirement: a 1 or 2 on any indicator requires a Personal Development Plan attached to the evaluation, and a 4 or 5 on any indicator requires specific supporting comments. Ratings of N/A and N/O exist for not-applicable and not-observed indicators, and they don't mathematically inflate the weight of the remaining indicators in a domain.
The Oklahoma evaluation cycle
TLE evaluation operates on a district-defined cycle that meets state requirements. Under the Tulsa Model's structure, career teachers must be evaluated at least once a year, and probationary teachers at least twice a year. Within a year, districts typically layer informal walkthroughs and mini-observations underneath the formal announced and unannounced observations, each with the pre- and post-conference structure the model expects.
Oklahoma evaluators are required to hold OSDE-certified evaluator training before they can rate a teacher under any TLE framework, and Tulsa Model evaluators complete a re-certification after their second year. That certification requirement is a defensibility anchor: the state decided that scoring a teacher's practice on a weighted, consequential rubric requires demonstrated competence, not just a login. It also means the human in the loop is not incidental to the system — it's the load-bearing element the whole design assumes.
Separate from the observation rubric, OSDE's PL Focus — the Professional Learning Focus — is the educator-driven professional-learning component that replaced the older "sit and get" model. It's the teacher's own work, conferenced on rather than scored inside the twenty indicators, and it sits alongside the observation-based evaluation rather than inside it.
After the walkthrough: where the work actually piles up
The observation itself is the part Oklahoma evaluators trained for. The writing that follows is the part nobody warned them about.
A single formal Tulsa Model observation can generate several pages of documentation: ratings across twenty indicators organized under five weighted domains, the weighted-composite calculation, narrative justifications, the required supporting comments for every 4 or 5, a Personal Development Plan for any 1 or 2, and language for the post-observation conference. Multiply that by a caseload — thirty teachers with annual evaluations for career staff and twice-yearly cycles for probationary teachers — and the write-up becomes the single largest time cost in the evaluator's year.
This is where the day actually goes. The observation takes thirty minutes. The write-up takes ninety. And the conferencing takes hours more across the year. The instrument that makes the Tulsa Model defensible — its specificity, its weighting, its documentation triggers — is the same thing that makes the paperwork heavy.
Where AI helps with Tulsa Model write-ups
The strongest fit between current AI and Oklahoma evaluation work is the translation problem. An evaluator captures notes in shorthand — fragments, abbreviations, half-sentences — and then has to translate them into the formal language of the Tulsa Model.
AI handles three parts of that reliably. First, it turns fragmentary notes into coherent, evidence-anchored prose in the voice of the rubric. Second, it maps specific evidence to the right indicators across all five domains — recognizing that the note about nine seconds of wait time supports Involves All Learners and Monitors, while the note about the exit ticket that changed the lesson supports Adjusts Based Upon Monitoring. Third, it maintains internal consistency across twenty indicators so the write-up reads as one coherent document rather than twenty disconnected paragraphs.
The Oklahoma-specific place AI earns its keep is the documentation triggers. A tool that understands the Tulsa Model flags — and drafts — the supporting comments a 4 or 5 requires and the Personal Development Plan a 1 or 2 requires. Those are exactly the pieces an evaluator forgets at 9 p.m. between observations, and exactly the pieces that get an evaluation questioned later. What AI should not do is the weighted-composite math: the weights are a real feature of the framework, but combining domain scores into the final composite is the evaluator's or the district's calculation, not something to hand to a chatbot.
Where AI doesn't help — and what stays with the evaluator
The honest scope is narrower than the marketing on most AI tools suggests.
AI cannot make the Effective-to-Highly-Effective-to-Superior judgment. That call requires watching the lesson in real time and knowing whether student ownership is genuinely present — whether students are formulating their own questions or just answering well, monitoring their own progress or just complying. AI can suggest a rating from the notes; the certified evaluator owns the final call.
AI cannot navigate the PL Focus. That's the teacher's own professional-learning work, conferenced on rather than scored, and it shouldn't be generated on the teacher's behalf.
AI cannot replace the relational work of the post-observation conference — the coaching conversation with a probationary teacher after a first formal observation, the harder conversation with a career teacher whose Instructional Effectiveness evidence is trending toward Needs Improvement. That's human work, and it's most of what the job actually is.
AI cannot substitute for OSDE's certified-evaluator requirement. That certification exists because the state decided rubric-based ratings demand specific competence. AI supports the certified evaluator's documentation; it cannot be the evaluator.
And AI cannot do the contextual reasoning a long-tenured Oklahoma evaluator brings — knowing this teacher absorbed two combined classes when the building lost a position, that this lesson fell the week of state testing, that this room has three newly-arrived multilingual students. Local context shapes judgment in ways no chatbot reconstructs from notes alone.
What AI does well, it does well. What it doesn't do, it shouldn't pretend to do. Pasting notes into a general chatbot works fine as a glorified search engine — but for Tulsa Model documentation that maps to the right indicators, carries the required narrative comments, and stands up when an evaluation is contested, the gap between "sounds polished" and "actually defensible" is wider than the polish suggests.
Observable evidence: what each domain looks like in practice
Each domain has its own surface features — the things an evaluator actually sees and notes.
Classroom Management (Domain 1, 30%) is most visible in the first minutes. Look for lesson objectives explicitly tied to a current Oklahoma standard and materials ready before students arrive (Preparation); "withitness" and corrections that stop misbehavior promptly while preserving student dignity (Discipline); a current substitute folder, rosters, seating charts, and emergency plans (Lesson Plans); assessment that guides instruction rather than only producing grades (Assessment Practices); and a teacher who knows every student by name and holds visibly high expectations (Student Relations).
Instructional Effectiveness (Domain 2, 50%) is where half the composite lives, so it deserves the most careful evidence. Look for literacy embedded as the vehicle for learning content — text-based instruction, students citing text (Literacy); active learning for roughly 80% of class time with Bloom's-scaffolded questioning and three-to-five seconds of wait time (Involves All Learners); at least two delivery modes with technology integrated meaningfully (Explains Content); explicit modeling that narrates the thinking, not just the answer (Models); a teacher moving to all areas of the room and using varied response techniques (Monitors); an observable mid-lesson adjustment when evidence warrants it (Adjusts Based Upon Monitoring); a deliberate closure that consolidates learning (Establishes Closure); and IEP accommodations actually implemented, not just documented (Student Achievement).
Professional Growth & Continuous Improvement (Domain 3, 10%) lives largely outside the lesson. Look for completed, applied PD hours that connect to instructional improvement (Professional Learning) and consistent reliability — punctuality, absence-notification procedures, reporting deadlines met (Professional Accountability).
Interpersonal Skills (Domain 4, 5%) shows in proactive, timely family communication that isn't only deficit-focused, and in active, open collaboration with colleagues.
Leadership (Domain 5, 5%) shows in meaningful participation in school and district initiatives, proactive advocacy for student needs, and — at the top level — leading colleagues and challenging biased or disrespectful practices that impede serving all students.
Common scoring mistakes under the Tulsa Model
Even experienced Oklahoma evaluators slip into patterns worth naming.
The first, and the one unique to this model, is straight-averaging instead of weighting. Treating the five domains as equal quietly rewards a teacher strong in the low-weight domains and understates one who's strong where it counts. Instructional Effectiveness is half the score; Interpersonal Skills and Leadership are a twentieth each. The composite has to reflect that.
The second is the Effective default. When evidence is thin either way, evaluators slot Effective because it feels safe. But the Tulsa Model requires evidence of Effective practice — consistent professional practice aligned to standards — not merely the absence of problems. A thinly-evidenced indicator is a coaching opportunity, not an automatic 3.
The third is the off-Effective documentation crack. Rate an indicator a 2 or a 4 and forget the required attachment, and the evaluation has a hole in it: a 1 or 2 needs a Personal Development Plan, a 4 or 5 needs specific supporting comments. This is precisely where a contested evaluation comes apart.
The fourth is over-scoring the low-observability indicators. Building-Wide Climate, Professional Learning, Professional Accountability, Interpersonal Skills, and Leadership all live mostly outside the lesson window, so evaluators tend to credit a "professional-feeling" teacher with high marks absent indicator-specific evidence. The fix is to demand the same evidence rigor there as in Instructional Effectiveness.
The fifth is perfect-fit paralysis. The model scores on best fit — award the level that matches the majority of the rubric language, then set a growth expectation through the push-pin process. Withholding a 3 because one sub-element isn't flawless misreads how the instrument is meant to work.
The sixth is indicator conflation. Monitors (13) and Adjusts Based Upon Monitoring (14) feel related, and evaluators justify both with the same note — but one is about checking for understanding and the other is about changing instruction because of what the check revealed. Explains Content (10) and Clear Instruction & Directions (11) get mashed together the same way. They're distinct claims that need distinct evidence.
Oklahoma in context — the district-choice, Danielson-based model
Oklahoma's approach differs from the mandatory-statewide-rubric states (Georgia TKES, Texas T-TESS, Florida FEAPs, Tennessee TEAM) and shares structural DNA with district-flexibility systems. Where Oklahoma is distinctive is that its flexibility operates at the framework level: districts don't just tailor a plan, they choose among several OSDE-approved qualitative frameworks — Tulsa, the Marzano models, McREL, Oklahoma TAP — and apply the chosen one to all teachers.
The Tulsa Model's Danielson foundation means it will feel familiar to evaluators moving from other Danielson-adopting states, while Oklahoma's specific choices set practical application apart: the five-tier Ineffective/Needs Improvement/Effective/Highly Effective/Superior scale with Superior at the ceiling, the lopsided weighted composite, the best-fit-and-push-pin scoring philosophy, the narrative-comment-and-PDP documentation triggers, and the PL Focus sitting outside the rubric. For an Oklahoma evaluator, the rubric structure is stable, but the exact cycle timing, observation cadence, and local expectations are set by the district within the state framework.
How EvalScribe handles the Tulsa Model
EvalScribe is built around the actual Tulsa Model structure — all five domains, all twenty indicators, the five-tier scale with Oklahoma's exact labels (Ineffective, Needs Improvement, Effective, Highly Effective, Superior) — rather than approximating it with generic teaching-evaluation prose. Every indicator carries the model's own rubric language across the five levels, the look-fors an evaluator watches for, and the distinctions that separate Needs Improvement from Effective and Effective from Highly Effective. The tool handles the translation problem natively: an evaluator captures notes by typing, dictating, or photographing handwriting (Smart Scan OCR converts it to text), and EvalScribe maps that evidence to the relevant indicators and drafts ratings and comments in the rubric's voice — including the supporting comments a 4 or 5 requires and the Personal Development Plan language a 1 or 2 requires.
One honest scope note on the weighting: the Tulsa Model composite is a weighted average, and EvalScribe surfaces each domain's weight for reference, but the final weighted-composite calculation is completed by the evaluator or the district's system — the app does not compute it. That boundary is deliberate. The weighting is a real feature of the framework; the arithmetic that rolls twenty indicator scores into a single composite is the evaluator's determination, and EvalScribe doesn't put its thumb on that scale.
Beta testers report saving 30 to 60 minutes per evaluation versus writing the documentation traditionally. Across a typical Oklahoma caseload — thirty teachers with annual evaluations for career staff and twice-yearly cycles for probationary teachers — that adds up to dozens of hours back over an evaluation cycle.
A further scope note: EvalScribe natively supports the Tulsa Model and the Marzano framework for Oklahoma. Districts using the other OSDE-approved qualitative components — McREL or Oklahoma TAP — would find Tulsa- and Marzano-specific coverage today; dedicated support for those two is on our roadmap for a future update. OSDE's PL Focus is the educator's own professional-learning component and stays outside the tool, which is exactly where the framework intends it. And the certified evaluator stays in control of every rating, every comment, and every piece of mapped evidence. EvalScribe drafts; the certified evaluator decides. More detail on the Oklahoma workflow is available at evalscribe.com/oklahoma.
Frequently asked questions about the Tulsa Model and AI
Where can I read the official TLE and Tulsa Model sources? The primary sources are maintained by OSDE and Tulsa Public Schools: the OSDE Teacher and Leader Effectiveness (TLE) pages and the TLE qualitative components library, which houses the Tulsa Model Observation and Evaluation Handbook. Statutory authority is 70 O.S. § 6-101.16 and Okla. Admin. Code § 210:20-41-1.
Does AI replace evaluator judgment in Oklahoma? No. Oklahoma requires OSDE-certified evaluator training before anyone can rate a teacher under a TLE framework, with Tulsa Model re-certification after the second year. AI drafts evidence-mapped write-ups and suggested ratings; the certified evaluator reviews, edits, and finalizes every document.
Why does the weighting matter so much? Because the Tulsa Model composite is a weighted average, not a straight one. Classroom Management is 30%, Instructional Effectiveness is 50%, Professional Growth is 10%, and Interpersonal Skills and Leadership are 5% each. A tool or evaluator that averages the domains equally will produce a composite the model doesn't recognize. To be clear about EvalScribe's scope: the app surfaces those weights for reference but does not calculate the composite — that final arithmetic is completed by the evaluator or the district's system.
What has to happen when I rate an indicator off Effective? A rating of 1 or 2 on any indicator requires a Personal Development Plan attached to the evaluation; a rating of 4 or 5 on any indicator requires specific supporting comments. Effective (3) is the level that needs no written justification. This documentation trigger is one of the Tulsa Model's core defensibility features.
Does EvalScribe support Marzano, McREL, or Oklahoma TAP? Marzano, yes — EvalScribe natively supports the Marzano framework alongside the Tulsa Model. McREL and Oklahoma TAP are not yet supported; dedicated support for those two is on the development roadmap and coming in a future update.
Does AI handle the PL Focus? No. OSDE's Professional Learning Focus is the educator's own professional-learning component, conferenced on rather than scored inside the twenty indicators. It stays teacher-authored. EvalScribe handles the observation-based evaluation.
How can I see what EvalScribe looks like for Oklahoma specifically? The Oklahoma page at evalscribe.com/oklahoma walks through the five-domain workflow, the twenty-indicator handling, the domain weighting shown for reference, and a sample exported evaluation. EvalScribe is available on iOS and macOS via the App Store, with three free evaluations to start.
If you're evaluating teachers in Oklahoma under the TLE Tulsa Model, see how EvalScribe handles the observation-based workflow at evalscribe.com/oklahoma.
References
Oklahoma State Department of Education, Teacher and Leader Effectiveness (TLE)
Oklahoma State Department of Education, TLE Qualitative Components
Tulsa Public Schools, TLE Observation and Evaluation Handbook for the Tulsa Model
Oklahoma Statutes, 70 O.S. § 6-101.16; Okla. Admin. Code § 210:20-41-1
Charlotte Danielson, The Framework for Teaching Evaluation Instrument
Related articles
Page maintained by Anthony D. Neely, Ph.D. — practicing K-12 educator with nearly 20 years in the classroom, 2025–2026 Walker County Distinguished Teacher of the Year, and co-founder of EvalScribe. Framework details verified against current OSDE and Tulsa Public Schools source documents.
Last reviewed: July 1, 2026.
