Texas school districts are questioning automated scoring of STAAR written responses after more than 13,000 scores were corrected through human reviews. The controversy raises questions about assessment accuracy, accountability, and responsible educational technology.
Editorial Note
This article examines automated scoring of constructed-response items on Texas STAAR reading assessments. The scoring system should not be confused with generative AI chatbots, and the number of corrected scores should not be interpreted as the overall error rate for all STAAR examinations.
The available reporting establishes increased rescore requests and score adjustments, but it does not demonstrate that every initial automated decision was incorrect or that all score changes resulted exclusively from automation rather than other scoring factors.
Thousands of Texas Test Scores Change After Human Review
Texas's use of automated scoring for written standardized-test responses is facing renewed scrutiny after school districts requested tens of thousands of score reviews and more than 13,000 results were corrected.
According to the Texas Tribune, the number of requests to rescore open-ended responses on STAAR reading assessments nearly tripled from the previous year.
Data discussed by the Association of Texas Professional Educators indicate that the Texas Education Agency received 64,489 rescore requests during spring 2026, compared with 21,620 the previous year.
More than 13,300 scores were subsequently corrected after human review.
The findings raise important questions about the reliability of computer-assisted assessment, the opportunities available to challenge results, and the extent to which standardized scores should influence decisions affecting schools and students.
They also highlight the difference between using technology to improve educational efficiency and demonstrating that automated decisions meet acceptable accuracy standards.
How Does Texas's Automated Scoring System Work?
The State of Texas Assessments of Academic Readiness, commonly known as STAAR, measures student achievement in subjects including reading and mathematics.
Following revisions to the examination, Texas increased its use of questions requiring students to construct written responses rather than select answers from a list.
Evaluating these responses is more complicated than scoring conventional multiple-choice questions.
A student's written answer may demonstrate partial understanding, use unconventional language, or communicate an otherwise correct explanation differently from an expected example.
Texas began incorporating an automated scoring engine in December 2023 to assist with scoring constructed-response items.
The system uses scoring models developed from previously evaluated student responses.
Importantly, state officials distinguish this technology from generative artificial intelligence products such as conversational chatbots.
Texas also uses human scoring and review procedures within its broader assessment process.
The controversy is therefore not accurately described as a system in which every STAAR question is evaluated exclusively by AI.
Why Did So Many Districts Request Reviews?
School districts may request rescoring when educators believe a student's written work was not evaluated appropriately.
The Texas Tribune reported that requests increased substantially as educators questioned results assigned to open-ended responses.
More than 250 districts and public charter networks submitted requests during the latest reporting period.
The number of requests does not automatically equal the number of errors because a review may confirm the original score.
However, the substantial increase indicates that many educators perceived enough uncertainty to seek additional verification.
When a score is corrected following human review, it raises a reasonable question about whether the original evaluation adequately reflected the student's demonstrated understanding.
A reliable system should be able to explain how such differences occur and whether they reveal recurring weaknesses in assessment design, scoring rules, or implementation.
What Do the 13,300 Corrections Actually Mean?
More than 13,300 score corrections is a substantial figure, particularly when decisions based on standardized assessments can influence perceptions of student achievement and school performance.
However, the percentage requires a clearly defined denominator.
The corrected results represented roughly one-fifth of the rescore requests submitted, but those requests were not a random sample of all examined responses.
Schools typically seek reviews because they suspect an error, meaning the reviewed group may be more likely to contain scoring disagreements than the overall testing population.
The Texas Education Agency has emphasized that fewer than 0.5% of total open-ended responses on reading examinations received score changes.
Both observations can be true simultaneously.
A relatively small overall correction rate can coexist with a substantial number of changed results among the responses specifically selected for review.
The central issue is whether the scoring system provides sufficiently accurate, consistent, and transparent results for the purposes for which they are used.
Why Accuracy Matters for Students
Assessment results can influence educational decisions in several ways.
Teachers use assessment information to identify learning needs and determine whether students require additional instruction.
Parents may rely on scores to understand academic progress, while administrators use aggregated results to evaluate programs and intervention strategies.
Standardized performance measures can also influence institutional accountability ratings.
When a student receives a score that does not accurately reflect their work, teachers and families may draw inappropriate conclusions about what the student understands.
An underestimated result could suggest that a learner needs intervention in an area where they are already proficient.
An overestimated result could obscure a weakness that deserves attention.
Neither outcome is desirable.
The purpose of assessment should be to provide information that improves educational decisions rather than create an appearance of precision that the scoring process cannot consistently support.
The Difference Between Automated Scoring and Generative AI
The term artificial intelligence is frequently applied to technologies with very different capabilities.
Automated scoring engines may rely on trained statistical models or machine-learning methods designed to compare student responses with established scoring criteria.
Generative AI systems are designed for broader content generation and interaction.
The fact that a system uses machine learning does not mean it functions like ChatGPT or continuously modifies its behavior after evaluating each student response.
This distinction matters because different technologies introduce different technical risks.
Automated scoring can still struggle with unusual wording, alternative explanations, or features of student writing that are difficult to interpret through a particular model.
Those issues should be evaluated through rigorous validation, monitoring, and comparison with qualified human judgment.
Responsible educational technology requires understanding what a system does, not simply labeling it AI.
Why Texas Uses Automated Scoring
State officials have defended automated scoring partly on efficiency grounds.
According to the Texas Tribune, Texas Education Agency representatives estimated that relying entirely on human scoring would require substantially more personnel and increase annual expenses by approximately $15 million to $20 million.
Manual scoring could also delay the release of results.
These considerations are relevant because large state assessment programs process millions of examinations.
A system that reduces costs while delivering results more quickly may free resources for other educational priorities.
However, efficiency must be evaluated alongside accuracy, fairness, and the availability of effective appeals.
A cheaper scoring process is not necessarily beneficial if avoidable errors lead to additional review expenses, incorrect educational decisions, or declining trust.
The proper question is whether automation provides an acceptable balance of reliability, speed, cost, and accountability.
Should Human Review Be Required?
Human review can provide an important safeguard when computer-generated scores are disputed.
Qualified reviewers may recognize valid explanations or unusual writing patterns that an automated model evaluates incorrectly.
However, human evaluators can also disagree and make mistakes.
For that reason, simply replacing every automated decision with a human judgment does not automatically guarantee perfect accuracy.
A stronger assessment system may require multiple safeguards, including carefully designed scoring rubrics, representative training data, regular reliability studies, and procedures for identifying anomalous results.
Independent evaluations could examine whether scoring accuracy varies across grade levels, writing styles, English-language proficiency groups, or other relevant characteristics.
Clear appeals procedures are equally important.
Families and educators should understand when a review is available, what evidence is required, how long it takes, and whether a corrected score changes any consequential decision.
Implications for Educational AI Development
The Texas controversy offers a broader lesson for educational organizations developing artificial intelligence tools.
Technology can support instruction, assessment, and academic intervention, but its capabilities should be evaluated using measurable outcomes.
For example, an AI-assisted tutor might identify a student's weaknesses and recommend targeted practice.
That recommendation could be helpful, but it should not automatically be treated as a definitive determination of academic mastery.
Educators should remain able to review important conclusions, examine student work, and correct inaccurate results.
The same principle applies to automated feedback, placement recommendations, and diagnostic assessments.
Educational technology is most valuable when it increases the accuracy and usefulness of information available to teachers and students.
It becomes problematic when users are expected to trust complex decisions without meaningful transparency or opportunities for review.
For platforms such as New To Education, these issues are relevant to the responsible development of tutoring and AI-assisted learning services.
The objective should be improving instruction while keeping qualified educators and learners meaningfully involved in decisions.
What Happens Next?
Texas officials may face additional questions about the reliability of automated scoring, procedures for human review, and the financial implications of rescoring requests.
School districts and professional organizations may continue examining whether the existing model provides sufficient protections.
Further analysis could investigate differences between automated and human scores, the types of responses most likely to be corrected, and whether current safeguards identify significant inconsistencies.
The most important evidence will be whether the system can demonstrate reliable performance across student populations and maintain meaningful procedures for correcting errors.
Key Takeaways
- Texas received 64,489 requests to rescore open-ended STAAR reading responses in spring 2026.
- The number of requests was nearly three times the previous year's total.
- More than 13,300 scores were corrected following human review.
- State officials report that fewer than 0.5% of all open-ended reading responses received score changes.
- The reviewed responses were not a random sample of all tests, so the correction percentage cannot be generalized to every exam.
- Texas uses a hybrid scoring process incorporating automated technology and human evaluation.
- The controversy highlights the need for transparency, reliability testing, and meaningful appeals in educational assessment.
Frequently Asked Questions
Does AI grade all STAAR examinations?
No. Texas uses automated scoring for certain constructed-response items, with human evaluation incorporated into the broader scoring process.
Were more than 13,000 scores corrected?
Yes. Reporting based on Texas Education Agency data indicates that more than 13,300 results were adjusted after human review.
Does that mean 20% of all STAAR scores were wrong?
No. The approximate percentage applies to the responses selected for rescoring, not the full statewide testing population.
Are automated scores always less accurate than human scores?
Not necessarily. Reliability depends on the assessment, scoring model, training, validation, and review procedures. Both automated and human scoring can produce disagreements.
Could this affect school accountability ratings?
Standardized assessment results contribute to accountability decisions, making scoring reliability important. Whether a specific correction changes a school's rating depends on the applicable calculation and other performance measures.
Final Thoughts
Texas's automated scoring controversy illustrates a challenge that education systems will increasingly encounter as advanced technology becomes part of everyday academic decision-making.
Automation can reduce administrative demands, accelerate feedback, and support large-scale assessment.
However, technological efficiency should not come at the expense of accuracy or meaningful human oversight.
More than 13,000 corrected results provide a reason to examine scoring reliability, while the relatively small share of total responses changed demonstrates why the findings must be interpreted carefully.
The most productive response is neither to reject automation categorically nor to assume that computer-generated scores are inherently objective.
Instead, educational leaders should insist on transparent validation, effective correction procedures, and clear evidence that assessment tools support fair and useful decisions.
Students deserve evaluations that reflect what they know, and educators deserve information they can trust.
Support New To Education
New To Education supports accessible learning, educational news, tutoring, and the responsible development of technology-assisted educational services.
If you value independent reporting on education policy and emerging learning technologies, consider supporting our work.
https://newtoeducation.com/support-us
Related Articles
California Assessment Convening in Los Angeles Highlights Accessibility for CAASPP and ELPAC
An examination of assessment accessibility and meaningful evaluation.
Sources
The Texas Tribune — STAAR Rescoring Requests Triple After More Texas Schools Question Computer Grading (October 7, 2026)
https://www.texastribune.org/2026/10/07/staar-scores-computer-grading-texas-schools/
Association of Texas Professional Educators — STAAR Rescoring Requests Raise Questions About Test Accuracy and Accountability (October 9, 2026)
Texas Education Agency — Scoring Process for STAAR Constructed Responses
https://tea.texas.gov/data-reports/staar/scoring-process-staar-constructed-response-1.pdf
New To Education | Independent Education News and Analysis
This article is intended for educational and informational purposes.