
AI Tools for T-TESS Teacher Evaluations: What Texas Admins Should Know
AI Tools for T-TESS Teacher Evaluations: What Texas Admins Should Know

If you're a Texas appraiser wondering whether AI can help with T-TESS evaluations, here's the honest answer: yes, but for one specific part of the process. AI can take the strategically-scripted evidence you gathered during an observation and turn it into formal, rubric-aligned documentation across the 4 domains and 16 dimensions. It does not replace your professional judgment as the appraiser, it does not touch the student growth measure or the self-assessment component, and it is not a shortcut around the T-TESS process the state requires. Used well, it gives you back the hours you currently spend translating what you saw in a classroom into the specific descriptor language each dimension and performance level requires. This post walks through where AI genuinely helps with T-TESS, where general tools fall short, and what to look for if you're considering one.
A quick refresher on what T-TESS actually asks of you
The Texas Teacher Evaluation and Support System is built around 4 observable domains: Planning, Instruction, Learning Environment, and Professional Practices and Responsibilities. Underneath those domains sit 16 dimensions in total, and each dimension is rated against 5 performance levels — Improvement Needed, Developing, Proficient, Accomplished, and Distinguished.
There's a feature of T-TESS that distinguishes it from most other state frameworks and worth naming explicitly: each domain is designed to focus on both teachers and students rather than separating them. The system was built around the idea that a constant feedback loop exists between teachers and students, and that evaluating the effectiveness of teachers requires consistent attention to how students are responding to the teacher's instructional practices. That framing isn't a stylistic choice. It shapes what evidence appraisers are supposed to gather, how dimensions are scored, and what good T-TESS documentation actually looks like.
The full T-TESS appraisal includes more than the observation rubric. It also includes a student growth measure and a teacher self-assessment that contributes to the overall rating. The rubric ratings, however, are where the bulk of an appraiser's time goes — and within those ratings, the bottleneck isn't the observation itself. It's translating strategically-scripted evidence into the specific language each of the 16 dimensions requires, distinguishing carefully between the 5 performance levels, and doing that for every teacher, multiple times a year.
A note on a recent T-TESS development worth flagging: in 2024-25, Texas adopted an optional Alternate Domain I rubric for districts using lesson internalization aligned with high-quality instructional materials (House Bill 1605). Alternate Domain I is optional — many districts continue using the standard T-TESS rubric — but appraisers in districts that have adopted it need to be trained on it specifically. EvalScribe currently supports the standard T-TESS rubric across all 4 domains and 16 dimensions. Alternate Domain I support is on the development roadmap and coming in a future update. More on this below.
The labor that takes appraisers away from classrooms isn't the observation. It's the translation step that follows.
Where AI genuinely helps in the T-TESS workflow
The honest, useful application of AI here is narrow and specific: you provide the evidence you gathered, and the tool helps produce T-TESS rubric-aligned narrative language and performance-level-informed scoring rationale across the 16 dimensions.
That's not hype. It's addressing the actual bottleneck. The cognitive work of observing a classroom and forming a professional judgment is yours and stays yours. What AI can take off your plate is the time-consuming work of rendering that judgment into the formal, rubric-aligned documentation T-TESS requires — distinguishing carefully between, say, Developing and Proficient on a given dimension, or between Accomplished and Distinguished, using the kind of descriptor language the rubric actually calls for.
The math is what makes it matter. T-TESS asks appraisers to make 16 dimension-level judgments per observation, each calibrated to 5 performance levels, with specific descriptor language for each. Multiply across a full faculty and multiple observation cycles, and you're looking at hours of documentation labor that used to be unavoidable. Those are hours you can spend in classrooms, having real post-conferences while the lesson is fresh, and doing the parts of instructional leadership that actually move the needle.
This is worth saying plainly, because a lot of appraisers carry guilt about it: using AI to handle the documentation translation is not cutting a corner. It's the right tool for a genuinely time-consuming task, freeing you to be present for the parts of the appraisal process that need a human. The job is still hard. The rubric still matters. The tool just handles the part that was never really about your expertise as an appraiser.
Where general generative AI breaks down for T-TESS
Most appraisers who try this reach for a general tool — ChatGPT, Claude, Gemini — because they're free, familiar, and right there. They can produce competent-sounding evaluation language. But for T-TESS specifically, three problems show up fast.
The tool doesn't know T-TESS. A general AI platform has no built-in understanding of the 4 domains, the 16 dimensions, the descriptors, or the 5 performance levels. To get aligned output, you have to feed all of that in — the rubric language, the dimension descriptors, the level distinctions — every single time, for every evaluation. The setup work you're trying to escape just moves to the front of the process.
The level-distinctions are where general tools fail hardest. This is a T-TESS-specific concern. Distinguishing between Developing and Proficient, or between Accomplished and Distinguished, requires understanding what those distinctions actually mean on a given dimension — and the descriptor language for those distinctions is specific. General tools default to generic "improving / good / great / excellent" framing that doesn't map to T-TESS's actual rubric character. You may get something that sounds plausible, but applied across a faculty, you'll find the level assignments drift in ways that are hard to defend.
The scoring isn't consistent. A general tool's output varies with the wording of your prompt, the day, and the model version. Two appraisers using the same tool to evaluate similar lessons can land on substantively different dimension-level ratings and substantively different rubric language. In a document that informs an overall appraisal rating with contract weight, that inconsistency is the kind of problem that surfaces in a personnel conversation — or worse, in a grievance.
There's also a quieter question worth raising: where does your teacher observation data go? When you paste classroom observations into a consumer-tier general AI tool, it's worth knowing whether that data is being retained or used to train future models. Policies differ by tool and they change over time, so the honest advice is to check the current data policy of whatever platform you're using before you put personnel-adjacent information into it.
If those limitations sound like dealbreakers, they're exactly why purpose-built tools exist. I wrote a fuller comparison of purpose-built versus general AI for teacher evaluations on the EvalScribe blog if you want to go deeper on that distinction.
What to look for in an AI tool for T-TESS
Whether you end up using EvalScribe or anything else, these are the questions worth asking before you commit to a tool for this work.
Are T-TESS's 4 domains and 16 dimensions built into the tool, or do you have to paste them in yourself every time? If you're providing the framework every time, the workflow is fragile.
Does the tool correctly distinguish between the 5 performance levels for each dimension — Improvement Needed through Distinguished — or does it produce generic high/medium/low scoring that doesn't map cleanly to the rubric?
Does the tool reflect T-TESS's "constant feedback loop" character — that each domain examines both teachers and students — or does it default to teacher-only language that misses what T-TESS is actually asking appraisers to assess?
Is the scoring consistent across appraisers and across teachers, so the same observation lands in similar territory regardless of who's running it through the tool?
Where does your observation data live? Is it stored on the vendor's servers? Is it used to train models?
Can your appraisers use it on day one without an additional training requirement? Texas appraisers already complete a state-required 3-day T-TESS appraiser training. The AI tool shouldn't add another certification burden on top of that.
How EvalScribe fits
EvalScribe is built for exactly the part of T-TESS we've been talking about — the rubric documentation step.
All 4 T-TESS domains and all 16 dimensions are built into the tool. The 5-level performance rubric is applied with rubric-aware logic that distinguishes between adjacent performance levels the way T-TESS actually requires — meaningful differentiation between Developing and Proficient, between Accomplished and Distinguished, calibrated to the descriptor language each dimension uses. That depth matters, because it's where general tools tend to flatten T-TESS into generic evaluation language that loses the system's actual character.
On the data question from earlier: EvalScribe runs on Microsoft Azure with no data storage on our end. Your teachers' observations don't sit on our servers.
And to be straight about scope, since trust matters more than reach here: EvalScribe handles the rubric documentation translation step. It doesn't run your pre- or post-observation conferences, calculate your student growth measure, conduct your self-assessment process, or replace your professional judgment as the appraiser. It translates the strategically-scripted evidence you gathered into formal, rubric-aligned T-TESS language. The appraiser remains the appraiser. The tool handles the translation work that was eating your evenings.
On Alternate Domain I specifically: EvalScribe currently supports the standard T-TESS rubric. Support for the Alternate Domain I rubric — for districts that have adopted it for lesson internalization aligned with HQIM — is on the development roadmap and coming in a future update. If your district uses standard T-TESS, EvalScribe handles it now. If you've adopted Alternate Domain I, it'll be supported soon.
A brief note on the framework alignment behind all of this: EvalScribe wasn't built by a vendor looking for a market. It was built by an educator with eighteen years in the classroom and PD-development experience at the national level. The framework alignment isn't surface-level adoption of state rubric language — it's pedagogical understanding of what the T-TESS rubric is actually asking appraisers to assess. That matters more for T-TESS than for most state frameworks, because T-TESS's "constant feedback loop" character requires the tool to think about teachers and students together, the way the rubric does, rather than separating them.
What that looks like in practice: some appraisers report completing an entire evaluation in under five minutes. Compared to writing one the traditional way, beta testers report saving thirty to sixty minutes per evaluation. Across a full faculty and multiple observation cycles, that adds up to your fall and spring looking very different.
The simplest next step
I built this because appraisers who love being in classrooms shouldn't have to dread the rubric documentation that follows — and because the teachers on the other end of these appraisals deserve thoughtful, rubric-aligned feedback, not whatever fell out of a rushed Sunday-night session.
If you want to see what your specific time savings might look like across your campus or district, there's a time-saver calculator at evalscribe.com where you can put in your numbers — number of appraisers, number of teachers, observations per year — and see how many hours EvalScribe would give back. Your actual situation, not a marketing claim.
If you want to try the tool itself, an individual appraiser license is $100 a year. That's deliberately set below most districts' procurement thresholds, which means you don't need a committee, a purchase order, or a six-week approval cycle to try it. You can decide for yourself.
If you have questions, or you'd like to talk about a school or district license, reach me at [email protected].
FAQ
Can I use ChatGPT to write T-TESS appraisals? You can, but the tool doesn't know T-TESS — you'll need to provide the 4 domains, the 16 dimensions, the descriptors, and the 5 performance levels every time, and the level-distinctions will vary depending on how you prompt it. For a document that informs an overall T-TESS rating with contract weight, that inconsistency is worth thinking carefully about.
Does AI replace the appraiser in T-TESS? No. The observation, the professional judgment, and the dimension-level rating decisions are yours. AI's legitimate role is helping translate your judgment into formal, rubric-aligned documentation — not making the appraisal for you.
What part of T-TESS can AI actually help with? The rubric documentation step — turning your strategically-scripted observation evidence into formal, dimension-aligned write-ups across all 4 domains and 16 dimensions. AI tools don't handle the student growth measure or the teacher self-assessment components, and they don't run your pre- or post-observation conferences.
Is AI-generated T-TESS documentation defensible? It depends on the tool. Consistent, rubric-aligned output from a purpose-built tool — with meaningful distinctions between the 5 performance levels — holds up better than output that varies by prompt wording or flattens the level distinctions into generic language. Consistency across appraisers is what makes documentation defensible.
Does EvalScribe support T-TESS? Yes. All 4 domains, all 16 dimensions, and the 5-level performance rubric (Improvement Needed through Distinguished) are built in, with rubric-aware logic that distinguishes between adjacent performance levels the way T-TESS actually requires.
Does EvalScribe support the Alternate Domain I rubric? Not yet. EvalScribe currently supports the standard T-TESS rubric. Alternate Domain I support — for districts that have adopted the lesson-internalization rubric aligned with HQIM — is on the development roadmap and coming in a future update.
How much time does it save on T-TESS evaluations? Some appraisers report completing an entire evaluation in under five minutes. Compared to writing them traditionally, beta testers report saving thirty to sixty minutes per evaluation. You can run your own numbers using the time-saver calculator at evalscribe.com.
Is my teacher observation data safe? EvalScribe runs on Microsoft Azure with no data storage on our end — your observations don't sit on our servers. For any general AI tool, check the current data policy before entering personnel-related information.
