Skip to main content
Back to Blog

We moved our Advanced tier to Gemini 3.7 Flash. Here is the measurement that did it.

By Steven Swanson, Founder of ClassLens·

ClassLens now grades Advanced-tier work with Google's newer gemini-3.7-flash. We tested it against the model we were already using, and the results said more about the old model than the new one.

What the testing showed (August 2026)

  • We graded 200 essays that human raters had already scored. gemini-3.7-flash agreed with those humans more closely than gemini-2.5-flash-lite, the model Google is retiring: 0.639 against 0.603 on a scale where 1.0 is perfect agreement.
  • gemini-3.1-flash-lite, the model we just took off the Advanced tier, scored between 0.29 and 0.43 on essays across three separate tests. That is the worst result of any model we have tried.
  • The Standard tier did not change. It stays on gemini-2.5-flash-lite for now.

How we decide whether a model can grade

We take essays that real people have already scored, hand them to the model, and see how close it lands. If it usually lands on or near the human score, and rarely lands far away, it passes. If it drifts, it does not. The number we use for this is called Quadratic Weighted Kappa, and all it does is put that closeness on a scale from 0 to 1, counting a five-point miss as worse than a one-point miss.

That measures agreement, which is not the same as being right. A model that scores like your colleagues do is one you can argue with on familiar ground, and arguing with it is still your job.

The honest reading of the new number

0.639 against 0.603 is a real difference and a small one. Run the statistics properly and gemini-3.7-flash clears the bar for being no worse than the model Google is retiring, but the gap is too small on this sample to say it is confidently better. That is the honest version, and we would rather print it than the press-release one.

The comparison that actually decided this was the other one. Our Advanced tier had been running gemini-3.1-flash-lite, and three separate tests put it between 0.29 and 0.43 on essays, far below everything else we tried. It was handing back nearly the same score to everybody. Essays that humans scored anywhere from 6 to 30 came back bunched near the middle, so a strong essay and a weak one finished a few points apart. We tried two different sets of grading instructions on it and both did the same thing, so there was no wording fix available.

It was not bad at everything. On short-answer factual questions the same model scored in a normal range, around 0.70 to 0.73. The problem showed up on long work graded across a wide point spread, which is the exact work the Advanced tier exists for.

Why we are telling you our own model was the worst one

Because the alternative is worse. Advanced is our paid tier, so the model that measured worst was the one sitting behind it. We found it by running the test with the same grading instructions ClassLens sends in production instead of trusting the model card, and the same test is what told us which model to put there instead.

This is the whole argument for measuring. AI answers are cheap now. What is not cheap is knowing whether the feedback in front of you reflects what a student actually understood, or just what a model felt like saying that afternoon. A model earns a place in ClassLens by grading close to a human on work humans already scored, and it loses that place when a later test finds something better.

What this changes for you

If you grade on Advanced, your essays are now scored by a model that does not hand back the same number to everyone. Nothing else moves. You still approve the rubric before anything is graded, grades still arrive as drafts, and no mode sends a grade to a student until you release it. A better model does not change who signs off.

gemini-3.7-flash costs us more per submission than the model it replaces. Your price does not change, and Advanced still costs the same three credits per submission it always has. We are covering the difference. You should know when the thing grading your students changes underneath you, including when the change costs us money.

The Standard tier stays on gemini-2.5-flash-lite while we run the new model on Advanced first and watch how it behaves on real work. When Standard moves, we will test it the same way and publish that too, including the parts that do not flatter us.

FAQ

Does ClassLens switch to every new Gemini model Google releases?

No. Every candidate model is validated against human-scored academic datasets before it grades any student work. A newer model ships only after it clears that bar, and a model already in service is replaced when continued measurement shows a better option. We have published one test where the newer model lost and we kept the old one, and this one where the newer model won.

What does a score like 0.639 actually mean?

It is a measure of how closely the model's scores match scores from human raters, on a scale where 1.0 is perfect agreement and larger disagreements count more heavily against it. A higher number means the model lands near the human score more often and rarely lands far off. It describes agreement with human graders. It does not mean the model is a better judge of the work than you are.

Does the better model cost me more?

No. Advanced costs the same three credits per submission it always has, and plan pricing is unchanged. The newer model costs us more to run and we are covering that difference rather than passing it on.

Do I have to change anything to use the new model?

No. Choose the Advanced tier the way you always have and the current Advanced model is used automatically. Rubric approval, draft grades, and the step where you release grades yourself are all unchanged. Read more about how we test grading accuracy and bias.

Steven Swanson is a 22-year classroom teacher in California. He teaches engineering (design/drafting, mechatronics, and senior capstone) in the four-year engineering academy at Whittier High School, and AP Computer Science and AP Physics online. He built ClassLens after two days of chaperoning field trips produced 450 ungraded assignments and none of the tools he tried could grade them. Try it free at classlens.com.

Try ClassLens Free

AI-powered grading for Google Classroom. Set up in under five minutes. No credit card required.

See exactly how ClassLens reads a student response against a rubric and turns a whole class set into a Knowledge Gap Report, before you sign in.