Skip to main content
Back to Blog

Is AI Grading in Google Classroom Biased? We Tested It on 975 Student Essays

By Steven Swanson, Founder of ClassLens·

If you use AI grading in Google Classroom, the first question to ask is not how fast it is. It is whether it scores every student the same way. On September 29, 2026 we tested the grading ClassLens runs today on 975 real student essays. We found no statistically significant difference in how it scored English learners, girls and boys, or students of different racial and ethnic groups, relative to trained human raters.

The short version

  • What we tested: Gemini 3.8 Flash, the model behind ClassLens grading, run in both grading configurations ClassLens offered in September 2026 and sent the same grading instructions our production code sends for a Google Classroom assignment.
  • On what: 800 essays from the public PERSUADE 2.0 research corpus (400 by English learners), plus 175 more essays by Black students. 2,925 gradings, none lost.
  • Result: no gap by English-learner status, gender, or race/ethnicity was statistically significant, on either tier.
  • Still watching: essays by Asian/Pacific Islander students (69 in the sample) scored slightly lower relative to human raters. Not significant, but it has leaned that way in every run, so that group is the one we enlarge next.

Why AI grading bias is the question to ask first

A grading tool that is accurate on average can still be unfair. If it is a little more generous to one group of students and a little harsher on another, the class average looks fine while individual students get the wrong grade. You would not see it in a single assignment. It would show up over a semester, in the gradebook, for the students least likely to question it.

That is why we re-run this test each time the model behind ClassLens changes. Gemini 3.8 Flash went live as the primary model behind ClassLens grading on September 9. We ran this screen on September 29, after it shipped. The September 5 run used a simplified test prompt; this run used the exact instructions production sends. ClassLens now grades every submission with the deeper reasoning that used to be reserved for Advanced mode, at one credit per submission.

How we tested AI grading for bias

The essays come from PERSUADE 2.0 (Crossley et al., 2024), a public research collection of argumentative essays by students in grades 6 through 12. Each essay was scored by trained human raters on a 1 to 6 scale, and each carries labels for English-learner status, gender, and race/ethnicity. That combination is rare, and it is what makes a bias test possible.

  1. The same request production sends. We did not write a separate test prompt. Every request was generated by the ClassLens grading code that is live today, with the default settings a teacher gets, and the essay rubric entered the way a Google Classroom rubric arrives. Requests went to Google's Vertex AI in the US, the same route real grading takes.
  2. Both tiers. Standard (low reasoning effort) and Advanced (Google's default reasoning effort) were each run on every essay. We also ran Gemini 3.7 Flash, the previous Advanced-tier model, as a control.
  3. Compared against humans, not against each other. For every essay we took the AI score minus the human score. Then we asked whether that gap is different for one group than another. A gap near zero means the AI treats the two groups the same way the human raters did.
  4. A fixed pass/fail rule. A gap counts as a finding if it is statistically significant after correcting for running six comparisons at once (a Holm correction), or if it reaches a small effect size (Hedges' g of 0.20) with a 95% confidence interval that does not include zero.

Two things were not identical to a real grading run, and we want to be plain about them. We rendered each essay as a PDF with Google Docs default formatting, rather than exporting it through Google Drive; the words are the same. And every essay was graded under the same placeholder student name, because a real name would add a cue this test was not designed to measure.

Results: is AI grading biased by race, gender, or English-learner status?

No comparison passed the bar for a finding, on either tier. The smallest corrected p-value in the whole run was 0.55, far from the 0.05 that would signal a real difference. Here are the effect sizes. A positive number means the first group was scored a little more generously, relative to the human raters, than the second group.

ComparisonEssays in first groupStandard tier, g [95% CI]Advanced tier, g [95% CI]
English learners vs non-English learners400+0.000 [-0.134, 0.135]+0.072 [-0.067, 0.213]
Female vs male383+0.004 [-0.132, 0.146]-0.030 [-0.171, 0.112]
Hispanic/Latino vs White356+0.104 [-0.061, 0.271]+0.148 [-0.027, 0.328]
Black/African American vs White306+0.042 [-0.131, 0.219]+0.044 [-0.130, 0.217]
Asian/Pacific Islander vs White69-0.212 [-0.483, 0.053]-0.158 [-0.428, 0.118]
Two or more races/other vs White21+0.288 [-0.172, 0.755]+0.284 [-0.090, 0.664]

The Black/African American row uses the enlarged group of 306 essays (more on that below); the other rows use the 800-essay sample. The comparison group of essays by White students is 219. Every confidence interval in the table includes zero. The Gemini 3.7 Flash control run passed the same rule.

The number that bothered us last time

We ran this screen on Gemini 3.7 Flash on September 5. It passed the significance test we used then, but one number crossed the effect-size line we use now. Essays by Black students were scored a little more generously, relative to human raters, than essays by White students: g of +0.233 on 131 essays, with an uncorrected confidence interval that did not include zero. After correcting for six comparisons it was not significant, and one odd result in six is about what chance produces. "Probably chance" is not a good enough answer about students, though.

So this time we added 175 more essays by Black students from the same corpus, bringing that group to 306. The gap is gone: +0.042 on Standard and +0.044 on Advanced, with confidence intervals centered almost exactly on zero. It did not come back on the original 131 essays either.

What we are still watching

Asian/Pacific Islander students. On the Standard tier, these essays scored slightly lower relative to human raters than essays by White students (g of -0.212). The confidence interval runs from -0.483 to +0.053, so it includes zero, and the corrected p-value is 0.89. That is not a finding. But the number has been negative in every run we have done on this group, across three models, and 69 essays is a small sample. We will enlarge this group next, the same way we enlarged the last one.

Small groups. Only 21 essays in the sample were by students of two or more races. A group that small cannot be measured well; the uncertainty on that row is roughly plus or minus 0.45. We report it because leaving it out would be worse, not because it tells you much.

Where AI grading still differs from human grading

Bias is one question. Accuracy at the edges is another, and the AI is not perfect there. It pulls scores toward the middle. On the 1 to 6 scale, 53 to 57 percent of the AI's scores were exactly 3, while human raters gave 36 percent of these essays a 3. The strongest essays tend to come back a little low and the weakest a little high.

That pull is somewhat stronger on essays by English learners than on other essays, in every run of this test. On the Standard tier, the pull on English learners' essays is slightly stronger than on September 5 (a slope of -0.651 against -0.614), but the confidence intervals overlap, so we cannot call it a change. It is not a difference in how generous the AI is to English learners on average (that is the first row of the table, which is zero on Standard). It is a narrower spread.

What to do with that in your classroom: when you review AI-drafted grades, spend your time on the top and bottom of the stack. That is where your judgment matters most, and it is where ClassLens is most likely to need a correction.

What this test cannot tell you

  • Your assignments are not PERSUADE essays. This test covers argumentative essays scored on one holistic rubric. A lab report, a math explanation, or a five-criterion rubric is a different task.
  • Names were held constant. We did not test whether a student's name changes the score. That is a separate test.
  • Human scores are not perfect. Trained raters disagree with each other too, and that noise limits how finely any comparison against them can resolve small differences. We also ran an extra analysis that adjusts for each essay's human score. Read literally, it shows lower-scoring groups graded slightly harsher. But the same pattern appears on every model we have run, and it flips for girls, who score higher on average. That is what rater noise alone would produce. This analysis cannot tell real bias apart from that noise, so we count it as evidence in neither direction.
  • Fallback jobs. When Google's US capacity is tight, some grading jobs run on Gemini 3.7 Flash instead. Our 3.7 control passed, but it did not use the exact reasoning setting those fallback jobs use.

None of this replaces you. It is why the Classroom-return modes save grades to Google Classroom as drafts, and why no grade reaches a student until you release it.

Questions to ask any AI grading tool for Google Classroom

Answers are cheap now. The useful thing is knowing which questions to ask. Before you connect any AI grading tool to your Google Classroom, ask the vendor:

  1. Have you tested for bias by race, gender, and English-learner status, on the model you use today?
  2. Did the test use the same instructions your product sends, or a simplified research prompt?
  3. How many essays were in each group, and which groups were too small to measure?
  4. What did you find that you are still watching?
  5. Do grades land as drafts I review, or go straight to students?

The whole run above cost $11.76 in model fees. Cost is not a reason for any AI grading tool to skip it.

Our earlier May 2026 study covers accuracy for the model we used then, Gemini 2.5 Flash-Lite. It is not an accuracy result for Gemini 3.8 Flash. Our guide to responsible AI grading for teachers covers the classroom side. Districts can request the full methodology during procurement at steven.swanson@evolvedacademics.com.

Try AI grading in Google Classroom with review built in

ClassLens reads the rubric attached to your Google Classroom assignment, prepares a grade and feedback for each submission, and, in the Classroom-return modes, saves each score as a draft grade for you to review and edit. It also builds a Knowledge Gap Report for the class, so you can see which ideas the class missed, not just who scored what.

Start with the interactive synthetic demo. It does not connect to Classroom or change any grades. Every new account gets 30 days of full-access grading with no credit card required. When you are ready, connect Google Classroom.

FAQ

Is AI grading biased against English learners?

In our September 2026 test of Gemini 3.8 Flash on 800 PERSUADE 2.0 essays, half by English learners, the AI scored English learners the same way trained human raters did, relative to non-English learners. Neither configuration showed a significant gap for English learners (g of 0.000 and 0.072, both confidence intervals crossing zero). The AI does pull high and low scores toward the middle, somewhat more for English learners, so teachers should look hardest at the top and bottom of the stack.

Was the bias test run on the same setup teachers use?

Yes, with two stated exceptions. Every request was produced by the live ClassLens grading code with default settings and sent to Gemini 3.8 Flash on Vertex AI in the US, in both configurations ClassLens offered at the time. The essays were rendered as PDFs by us rather than exported through Google Drive, and every essay used the same placeholder student name.

Does this mean AI grading is fair for every student?

No test can show that. It shows no significant gap for the groups we could measure, on argumentative essays. Groups with few essays in the sample carry wide uncertainty, and we are enlarging the Asian/Pacific Islander group next. That is why Classroom-return modes save every score as a draft grade, and no grade reaches a student until you release it.

Do I still review grades before students see them?

Yes. In the Classroom-return modes, ClassLens saves scores to Google Classroom as draft grades. In Grade & Review, feedback is held for you to review and edit before release. No grade reaches a student until you release it.

ClassLens is an AI-assistive grading and teaching tool for Google Classroom built by Evolved Academics, LLC.

Steven Swanson is a 22-year classroom teacher in California. He teaches engineering (design/drafting, mechatronics, and senior capstone) in the four-year engineering academy at Whittier High School, and AP Computer Science and AP Physics online. He built ClassLens after two days of chaperoning field trips produced 450 ungraded assignments and none of the tools he tried could grade them. Try it free at classlens.com.

Try ClassLens Free

AI-assistive grading for Google Classroom. Set up in under five minutes. No credit card required.

See exactly how ClassLens reads a student response against a rubric and turns a whole class set into a Knowledge Gap Report, before you sign in.