New LLM model doesn’t mean it’s better than its predecessor.

Reddit r/ArtificialInteligence Models

Summary

A user tested Gemini-3.5-flash-lite on grading classwork and found it performed worse than its predecessor Gemini-2.5-flash-lite, suggesting newer models are not always better.

I tested the new models released by Gemini, and they might be trained on new data and have more reasoning. However I tested it on grading classwork and it did worse than its predecessor. The day of testing showed that Gemini-3.5-flash-lite didn’t do as well as Gemini-2.5-flash-lite Breakdown and details: https://www.classlens.com/blog/gemini-3-5-flash-lite-grading-test
Original Article

Similar Articles