I hereby declare: Gemini 3.7 Flash, released last night, is the first model in the Gemini family in six months that can actually be used for agent work. I have no idea what 3.1, 3.5, and 3.6 were doing.
Back at 3.5 I collected the Reddit field reports, and the verdict was three times the price, worse Vision, disastrous tool calling. Same line, three versions later, finally turned around.
People ask me what actually changed after 3.5. My answer is three big advantages:
Want speed? You get speed.
Want quality? You get speed.
Want autonomous execution? You get speed.
Against what’s on the table right now, it feels stronger than Sonnet 5, but nowhere near Sol. The main difference, I think, is that it often misreads a user’s vague instructions, whereas Sol is better at guessing what you meant.
The price is also half, it has native audio and video multimodality, and the quota feels like the free extra you get when you buy cloud storage. You can’t use it up. I don’t wince at leaving it running all night. I’ve already wired it into Claude Code, and it’s capable enough for real work.
One more thing: for single-shot multimodal tasks, Gemini is still king.
Postscript: Multimodal Is King, Except When It Isn’t
I deliberately screenshotted a photo of a non-chain restaurant twice to strip the location data, figuring the AI would never place it, since Bangkok is wall-to-wall with this kind of shabu. GPT matched it anyway. Gemini, the one sitting on Google Maps’ own imagery, just guessed wrong.