Here's quite an impressive follow-up. I have a tool which knows how to render Markdown documents with embedded SVG content - I use it for the pelican test.
Since this transcript has HTML in it, I decided to upgrade that tool to also render HTML.
I set Gemini 3.8 Flash the task, using my own VERY shonky coding agent tool (llm-coding-agent) - and it did a solid job.
That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated".
Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.
Focus on speed and being OK with temporarily being #3/4 in intelligence might be the counterintuitive approach which makes Google win long term (whether accidentally or strategically). Can't wait to try Gemini Pro later this year!
I honestly can't believe serious people are making this argument on a straight face.
Gemini 3.7 flash outputs so many tokens per answer it doesn't matter how fast its TPS is, sol will end up being both cheaper and faster than Gemini. So ppl are paying more for a given task, waiting longer and using a dumber intelligence because "TPS number shiny".
Gemini 3.8 outputs 11k more tokens PER TASK on average in AAII than 3.7 putting it dead last in output tokens per task in the leaderboard.
There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them
it's such a weird split how most AI companies are trying to be the best, but Google really has a different mission statement. they already have users. lots of users. they need to be working on building models they can deploy and use with the most number of people, as they already have the users.
i don't know if Gemini models per se are fully is in line with that purpose, but the results we see keep seeming to be in-line with that split-of-focus.
They probably saw that report years ago of copilot dumping out the fast inverse sqrt function, and assume that's all they can do. From experience most anti-LLM people have either never used them, or used them back in the 3.5-4 era and then never again, though you might have even more experience with those people than I do. :P
I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried:
- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.
- Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.
- Document parsing (extracting the relevant trip info from PDFs).
If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.
Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.
Same here. I see so many people obsessing over the latest most state of the art bleeding edge models and yelling at Google for not being there, but I feel like the vast majority of people don't actually need those models. Flash has just been super useful and incredibly fast in my experience.
I prefer luna for most development, especially when I am guiding the process. Sometimes terra. I have had terrible results coding with sol. It is way over-tuned on RL to make something that completes the task, no matter what. I end up with way too much code that does a lot of things I didn't ask for.
I love Luna too. An excellent model and still usually better value per dollar than Gemini if you pay for API tokens. Things may change with 3.8 - we'll know soon.
IME you're supposed to have Sol drive Luna sub-agents to do 90% of the work. Sol should primarily be the verifier and goal setter. Use omp.sh with Task Delegation -> Always to strongly encourage Sol to drive Lunas. Also Luna prefers to be talked to with English in XML.
Google One plans are quite a good value actually - for a few bucks you get more Gemini plus space in Drive and other extras. Even through API, $3.75 for nearly Sol-level quality isn't that bad.
And let's not forget you can use it for free in AI Studio, and in the user app (even free accounts get tons of usage, though it's still 3.6 there), and in Antygravity.
That's the thing. I am completely lost because there are so many redundant paths to get the same thing and I'm trying to figure out which one is the best deal
I believe Gemini Flash is smart enough to know when to ground with web search. Their app has been saying it’s running a web search on almost all of my queries since 3.6. And given that Google … is Google, I trust them with web search grounding more than anyone else.
I wasn't trying to be precise originally, I just tried to fit activities into "morning / evening" buckets. I did the whole itinerary with Opus first, but when I gave it to Gemini 3.7 Flash to review, it started correcting it with "this place will close 5PM" or "this place is closed for good".
It was right on every nit, so it was surprising how well the model knows these things. If I ever release this I'll probably need the SERP API or Google Maps SDK (which I've heard is very expensive now), but for a personal trip where I will verify manually, using the LLM is okay for now.
When you called the Gemini API, did you opt in to using search grounding:
tools=[{"type": "google_search"}]
I'm curious whether in fact you were getting answers from the model weights (which is what I had assumed) or whether your API calls were resulting in web search tool calls.
Using grounding in Gemini is indeed backed by the same canonical data source for business information (like opening hours) as Google Maps. This stuff is available in its own API for a GCP fee, but we’ve built tooling to connect it to the Gemini agentic ecosystem as well.
I've swapped over to it in the past two weeks, it's been really good. It does what I ask and doesn't think it knows better than me, which so far has made it the most pleasing experience I've had when slop-coding.
My only wish is it were somewhat cheaper, as it tends to balloon pretty quickly when I'm using it in Opencode. I'm currently trying to offload a lot of work to subagents to stop the context expanding so rapidly. But on the upside, I rarely have to correct it - I've spent far less time arguing with this than with anything else so far.
I started trying out 3.7 Flash this week and it is competitive with opus/fable and also FAST. It is getting work done that anthropic models were struggling with and the speed with which it does is quite a bit noticeably faster.
Beginning to think Google is a dark horse in this race and some of Anthropic's "everything feels janky and rushed" karma is going to catch up.
Hype that burned out pretty quickly, it's hard to speak to the size and significance of old hype, I never felt it.
Every time I personally tried Gemini models up until last week they simply couldn't do the long complex tasks I'd being doing with Anthropic models for many months.
For awhile now I've found Gemini will use Google search for pretty much any real world knowledge, which is a huge plus IMO. It's basically Google with a much better frontend and no ads/seo nonsense.
I think it already has them but it's much more subtle. Also useful. When I've made certain sorts of queries I've had the distinct impression that it was attempting to very gently steer the conversation with suggestions. But it was brief, still answered usefully, and didn't resist going in the direction I wanted. So a win-win tactic I guess.
For example find a beautiful landscape shot of a place that just so happens to be accessible to tourists and ask it something along the lines of identifying the location. IME it will noticably steer the conversation towards relevant commercial offerings and offer (entirely unprompted) to help plan a trip.
Or ask it about a certain category of product with some requirements and it will initially present (relevant) options that look like paid placement to my eye. But if you ask it's happy to go on to turn up lots of alternatives and enumerate tradeoffs.
Assuming I'm correct the subtlety is on par with product placement in movies. Certainly leagues better than the internet advertising we've suffered to date.
I asked Claude to fix the grammar of my comment, and it changed "I am using 3.7 for" to "I've been using Claude 3.7", so they sneaked their own name on it.
incredible. further evidence supporting my personal stance to never ever let an LLM write or edit my writing intended for another human being to read. this is all me, baby
Once you've written something, it's incredibly easy to overlook minute changes to the text.
See: why authors wait days, weeks, or even months before editing what they've written (or, if you're more interested: cognitive regression, inattentional blindness, and the effects of misdirected saccades).
Eh that one is on me, if I think too much about my HN comment I end up deleting before posting it. I rely on the 1 min `delay` set in the profile page to fix before it goes live, but for some reason this time it was set to 0.
They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.
The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.
Not just speed, also reliability. IME, Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic.
That's interesting to hear. I should have added that I use Gemini through Google AI Studio as my general chat model, which probably explains our wildly different experiences.
Accordit to reddit talk, Fable 5.1 is worse than Opus 4.6 and 8B models are smarter than Qwen 3.8 Max, I wouldn't take anything said there with any more reliability than an instagram short.
I've been using 3.7 Flash to audit the work of Opus High, and Flash finds lots of subtle and insidious defects even while all the unit tests are green.
Then I tell Opus to read the audit report and implement what it agrees with.
Flash is really good at this, and it is blazing fast in Antigravity CLI. Easily 10x faster than Opus.
Can't wait to try 3.8 Flash. If it's good enough, maybe I'll switch Flash to primary and make Opus the auditor.
Idk, was building/maintaining simple esp32 control program with antig/opus. After last update it defaulted to gflash3.7. I pasted an email requesting 2 changes into the chat prompt, it did one and took me 4 turns to get that one right.
I know everyone is benchmaxxing but this one feels one step too far. Doesn't DeepSWE have both public and private tasks? I'd love to see the diff here.
It looks more like Google execs losing their mind and pressuring researchers to put DeepSWE directly into the training set.
Will look forward to the "feel" of the model in real testing. But I agree that these benchmarks do get "dealt with" rapidly. That's a shame, but I guess it's the times we live in.
I had qwen 3.8 3bit model drop into chinese on long runs. I had to remind it to use english. Its still better than every gemma model I tried. Gemma deleted files on a harddrive to make space when there was over 2TB free. For long runs, gemma is useless.
We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.
anthropic really needs something to address the cheaper end of the market before they get left behind. Sonnet 5 sucks, and Haiku hasn't been updated in a year. meanwhile we've got gemini flash, luna, and GLM5.3 all delivering 90% of the performance for a small fraction of the cost. paying $25/mTok is going to start looking pretty silly soon.
Wait a week with your judgement - most likely, Google is just bench-maxing very hard.
If you look at the previous Flash models and the announcement on Google I/O, it was an absolute disaster. Reality diverged very much from the marketing (supposedly great benchmarks).
The most interesting thing about the Gemini models is still their multi-modal support: they accept audio and video input, OpenAI and Anthropic's flagships are still image-only.
Gemini Flash is also pretty cheap, so it's a great family for performing media analysis, like extracting structured data from images and video.
Interesting side note: although Opus is still image-only, you can still drag videos into Claude Code and it doesn't blink an eye; it just strips it down to a series of images to parse.
True multimodal support would be way better, but I have no issues pasting in full screen recordings while QA'ing games and having Claude identify and fix issues in the video.
Yesterday's transcript of the best version from Fable looked like Fable already knew exactly what it should do without "thinking". In other words, there were no passages like "on the one hand I could do this, on the other hand ...".
It saw the fish in the basket from some other previous attempt but completely missed the gap between the tires and the rims where the background shines through (now it knows after scraping this comment and watch the next transcript).
Watch the video. It's from then-Gemini-lead Jeff Dean and the video shows off an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine.
It bugs me a little that "fidelity" has connotations other than "faithfulness to an original"---fidelity should be basically the same as correctness here!
The rendering of the gullet is very poor, because its both behind the handlebars but in front of the bike frame (impossible geometry). Surprising because gemini is usually pretty good on geo spatial skills.
Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak.
I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.
People have been sleeping on Gemini lately but these last few Flash releases (which were very rapid) are damn good.
These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).
I highly recommend just getting out of Anthropic's (or anyone's) vendor lock-in. Use opencode or pi. You can still use your subscription pricing using a proxy. I switched to opencode and haven't looked back.
Antigravity has been also rapidly improving lately, and your can also use any of the open coding harnesses. But I mostly meant “harness” as in your workflow/loop setup.
A month is not enough time for any meaningful change in an organization the size of Deepmind/Google. These models were surely the result of work streams and teams that started under Demis. I think Demis can safely feel proud Deepmind is getting back on track.
I like Google's strategy here. These new Flash models of late (Flash 3.6, 3.7 and now 3.8) have obviously been distilled from a much larger unreleased model (Gemini 3.5 Pro, iirc from the rumors).
One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February.
Small models cataching up with their bigger siblings are fantastic news.
A couple larger GCP customers requested this for sometime, especially on the cybersecurity side.
A SOC/IR or AppSec team doesn't need a generalized model that knows when Chaucer lived but it absolutely needs a model that can efficiently, quickly, and accurately prioritize vulnerability severity or validate patches.
Also just want to let my appreciation here for 3.7 it’s cheap super fast super reliable incredible at information parsing eu host able (important for us) and perfectly integrated into gcp. Great job google!
Looks like the strategy of regular updates with incremental improvements is working out well. Interestingly, the biggest jump in Artificial Analysis Intelligence Index score is for reasoning level Medium ( 3.7 was 51, 53, 57 for Low, Medium and High, 3.8 is 52,57, 59 respectively). I think scores at lower reasoning levels are more indicative of model capability since higher reasoning levels are focussed on benchmaxxing. We use the lowest reasoning level in production with good results.
I'm not an expert but I agree with your statement on the lower reasoning levels.
Lots of models seem to just allow the model to "bloatmax" tokens in order to get bumps at high/max reasoning levels. Many of the max reasoning levels allow models to use up to double or more the tokens the next lowest reasoning level uses. Its basically only useful for people who have no cost or time stipulations on anything.
I think I actually preferred it when we had models that either had reasoning enabled or didn't.
Something maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability literally outperforms at least 2/3 Chinese students, no to mention those who speak Chinese. After all, the model speaks like a real humankind if you prompt it well. That's AGI guys
Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.
Totally agree. When this current wave of GenAI really started heating up, I guess 2020-2021, my analysis was very straightforward. What are the high level inputs to long-term success? I basically came up with a couple of criteria:
1. Data. Lots of data.
2. Money. Lots of money.
3. Access to necessary hardware.
4. Business alignment/will to do it.
5. Access to talent, current and future.
This is certainly incomplete/naive. In my mind, though, Google was the clear answer.
On a more personal level, I've been deep into the Google ecosystem since I got diederich@gmail.com in 2005. (I actually paid 50 cents on ebay to get a very early invite.) There was no question in my mind that Google's AI work would deeply integrate into their whole ecosystem in very powerful and productive ways. (Yes, I can join you to discuss, at length, the various ways that Google's dominance is problematic/scary.)
Having said all that, I'm quite happy that there is, at the moment, a very rich competitive landscape. Indeed, not too long ago, with Gemini Pro 3.1 languishing, I moved most of my deeper thinking work to ChatGPT, which was, for me at least, clearly outperforming Gemini.
While I certainly didn't anticipate it, Google's strategy of making their fast/relatively inexpensive models surprisingly powerful has been a welcomed surprise.
Its not good at not making mistakes, but what it produces is structurally quite nice, not over-engineered (looking at you Sol) and its personality isn’t annoying (looking at you Claude). A bit like Grok Code, but Grok is a better coder.
It's pretty good if you can actively steer it , its actually really really good , the antigravity free tier and pro tiers are generous as well . I'm shocked at how fast it generates tokens.
They're quite selective in benchmarks, c.f. notably only bad one is 10% on TerminalBench. It's a really addled model, one time I said "Hi" and it built out a 4 panel hello world app with (fake) weather, a todo list, and a couple other things I forgot. I wouldn't be comfortable saying "ignore the #s!" except when I complained it was trash and way overcooked on agentic coding yet not good at it, and a couple DeepMind ML people liked the tweet.
I don't know if Google is having the worst marketing fumble or the most genius marketing one. Their "flash" models are very comparable to other companies' "pro" or "flagship" models. It seems to be a quite counterintuitive naming convention as it undersells the models.
Unless they have an even more powerful Gemini Pro in the oven...?
not saying they do have a beefier pro, but even if they did, isn't the delta between flash vs pro models reduced quite a bit? (e.g glm 5.3 flash vs 5.3, v4 flash vs v4 pro, sonnet 5 vs opus 5)?
"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)."
Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.
I'm curious if the knowledge cutoff is important, when the interface (Gemini app) can search online for recent information. Is there a big advantage to having everything internal?
Not directly - but latest research advancements, cleaner / richer datasets, etc. still require fresh base models. Not everything can be fixed through post training alone (e.g. why GPT-5.5 "Spud" was such a big jump, and also why GPT-6 "Astra" is now supposedly another big leap). Ofc model size etc also plays a role, but my (admittedly limited) understanding is that new base models _can_ also lead to big jumps even keeping parameter counts constant.
very important actually. just try to generate code for fresher frameworks/libraries. gemini sucks so bad in real work usage, everything it suggests are outdated and mostly useless.
I can’t wait until waiting hours and spending a big chunk of your usage per task seems antiquated, and real-time iteration on massive code changes is the norm. This might just be the year of efficiency, that truly allows AI to be used to the heart’s content.
How generous is the Google subscription quotas compared to Anthropic and OpenAI? This sounds like a really good potential model for high volume due to its speed and cost effectiveness.
(By high volume I mean things like "main app just updated with XYZ commits, please scan XYZ plugins and surface any compatibility issues")
I'm on the Ultra plan and use it for chat, antigravity, and some other work automations (similar to your example). The only time I've ever hit my limit is when I use Deep Think (which usually eats up 4-5% of the 6-hour usage limit per response).
The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.
You need more safeguards for sure, but also it tends to fly off down rabbit holes, rebuilding things in dumb ways, hacking around things, making assumptions etc, it seems very eager to go 'ta da! I did it look how quick I was', sometimes it nails it other times it created a lot of tech debt.
Also if it ever says, "I've found the root cause of ..", it definitely has not found the root cause and is making a non evidence based guess as it has run out of ideas.
Still refuse to search internet for stuff it thinks does not exist lol.
And even when searching for internet, it still cannot suggest a up-to-date approach to the problem.
For example I'm using crystal, it recently revamped the concurrency/parallel model. Even using web search, gemini still does not aware of the new feature and still give the outdated code.
I'm sure my crystal usage is not the unique case here.
You're telling me for only 5x the cost and 1/10th the speed I can use a Chinese model which performs worse than Gemini 3.8 Cyber? And I get to do all the hosting and setup work myself instead of just using a model and framework which is already integrated with GCP? Dang!
I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options:
- Flash-Lite
- 3.6 Flash [new]
- 3.1 Pro
The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tomorrow I'll be playing with the new model from OAI/Anthropic.
It's such a weird attitude, especially considering that 1) it's readily available on AI Studio 2) Anthropic models were not always available the moment they got released either.
(It also shows that the internet isn't dead. Even people who are not aware of Google AI Studio can express their valuable opinions on LLMs!)
"The new Gemini model isn't available in Gemini, the Gemini App Gemini model is two versions behind and marked as new and the actual new model is in AI Studio" is the kind of problem only Google has though.
AI Studio? Seriously, the hell is that? Gemini, AI Studio, Antigravity - what is all that nonsense? The 3.8 Flash announcement says the model is available to Google AI Pro customers. Is it the same as Gemini Pro, or some sort of AI Studio Pro? Based on the comments, i see the model is available in the Gemini App, not available in the UI, not available to Workspace accounts but is available to some personal accounts, yet I'm not a Workspace user. Some people have already mentioned that they are paid customers, yet they don't see the new model.
I know Google loves asking graph problems during their tech interviews, but I can't wrap my head why the customers should solve these problems as well.
it would be a bad take if the webui had 3.7 flash available in it today, and they just hadn't fully rolled out the latest model when they posted the launch announcement.
but the webui is currently offering 3.6 flash. the previous model still hasn't actually rolled out to it yet.
Workspace always gets things slower than normal Gmail accounts. They do a lot more to isolate data related to those accounts, so that's likely the cause here.
Anytime anything gets added to Workspace, I think Google has a lot more contractual obligations about keeping it around for X amount of time, so they tend to be more careful about adding things.
> I'm a paid Gemini subscriber via Workspace Standard accounts and yet I also only have access to 3.6.
Same and I have found it extremely annoying. I actually really like the Gemini models for question/answer stuff and reach for it before Claude (the other model family I have purchased) but it's getting long in the tooth at this point and I'm finding my Gemini usage shrinking to nearly 0.
That looks like the options that get presented for Workspace users (like at my company). The personal Google accounts give more recent models, for some reason I don't understand.
As someone with a Pro subscription, I had access to 3.7 the day it came out. Expecting to have access to 3.8 now, too. It's only the free accounts that are behind.
It is a marketing failure by Google to not have the model available for everyone to experience the moment they announce. Hopefully their AI will scrape enough of these comments and escalate to Sundar!
It's available in antigravity which I started using again (for small things until I can trust gemini for coding again).
it's weird how the web ui doesn't show the latest flash options while the desktop/mobile apps update the same day as the release. I saw the model in the model selection (by coincidence) before seeing it show up on HN
yeah in typical google fashion, the best way to use the gemini models is by avoiding google's actual products. i've got a vision project where gemini flash is the best option by a long shot, and i just use openrouter so i don't have to navigate google's mess.
I have mixed feelings about Gemini 3.7 Flash. I used it for a personal project in Java and it was ok: it was crazy fast and it reached the correct result, but the code quality was barely passable.
I also used it for a an app for my Garmin watch, and it wasn't good. The code was compiling, but functionality was totally broken and even with a lot of steering it wasn't able to make it work. GLM 5.3-flash instead was up for it and the code wasn't bad at all. I am curious to see if 3.8 is an improvement in this use case.
I’m interested in a general knowledge model (closed or open weight) and not coding specific. I want to plan for travel and trip. Do you have one of your favorite HN crowd?
I think about Google is the value you get of their plans, for 5$ a month you get their ai plus model combined with 400gb you can share this with your family. The other ai companies don't provide family plans
Perhaps the model is able to evaluate that it's not done, and to keep pressing on in the face of mounting failures, until it eventually arrives at a solution. Where Fable can skip that.
If I had to pay per token I would probably consider using this (they seem to be on the pareto of performance) but not being able to use opencode with a subscription is not really something I'm realistically going to do when claude and codex are around. Also never gotten along well with gemini-cli / antigravity-cli.
For my application, I'm still happily using gemini-2.5-flash and the only problem is when it reports being overloaded. It's for interpreting a downscaled phone camera photo of a hand-written shopping list on a whiteboard, and it works stunningly well. My handwriting sucks, too.
(I guess the only relevance here is that if your problem matches a model's strengths, then you can do fine with a model that is several generations out of date.)
3.7 flash was by far the best model for image recognition tasks according to my benchmarks. 3.8 flash didn't regress any candidates and improved some specificity (positive ID of common name vs species name of exotic fruit, correct identification of cast/replica of artifact and statue) but is still relatively weaker (26/30) on esoteric public figures (Korean beatboxers). I'm going to have to make my benchmark harder.
I’m very curious about your esoteric public figures benchmark, do you ask it in English or Korean to identify the person? Does it change the result? I wonder if having data labeled in only a given language (or web sources in only a given language) change the output.
Is the google infra stable enough right now? At the start of the year, the flash model was unusable for a whole month via gemini CLI. They could not fix it for a whole month and I was a paid customer.
been absolutely loving 3.7 flash for coding. it feels very fast and quality is decent for implementing product features. usually use opus or sol for hardcore debugging.
i think it's better than sonnet 5, especially when you compare speeds. i have to work with the llm anyway, the faster i can turn it the better the outcome.
It's interesting that Deepseek models were missing in the comparison. I see Deepseek v4 Flash a direct competitor to Gemini Flash for text-based agentic work.
On my short tests: This model is amazing and the speed makes it feel like another sort of AI.
But it's bad at code reviews (maybe it's the harness agy cli?). Could not get it to same quality level on reviews like Opus, GPT 5.6, Grok. Even tried special code review skills but no luck.
After struggling with Gemini for months, I think the trick to getting the most out of the model is writing a really solid personal intelligence/instructions prompt. The results are night and day in terms of performance.
Funnily enough you really do need a great prompting and SKILLS setup to use antigravity effectively in contrast to other providers which actually started benefiting from less detailed prompts over time. But I like it this way, its more customizable and much cheaper especially with a sub.
agy is good for those cases where you are willing to put the effort into the harness specifically for a task or family of tasks. The full suite, with evals, monitoring, hooks, custom tools, custom verifiers, etc,. It is not good if you want a "general coding assistant" like codex or claudecode.
The reality is that if you optimise a harness for a family of tasks[1], then most of these models give successful output. And there, gemini flash's speed shines.
For general coding assistant, you want it to be well, general, and you use a harness without too much customisation to something specific. Here you need deeply post trained coding assistants and implementors like codex/sol or claude/opus. Gemini flash in its current form will be too happy-go-lucky if you try using it the way we all use codex and is better used in a constrained setting.
tl;dr gemini flash for "LLM-aided workflows in production" is super good today. Cheap as well.
slot machine addict thinks if he pushes buttons in a certain order the odds get better.
In all seriousness, gemini has the best interactive planning document/orchestration. Tell it to create a plan document and work through it with it and it will preform really well(in antigravity products). But this is the case with plan modes with every model, I just think the interactive document that antigravity uses is really well thought out.
I'm surprised the introductory 50% discount is good for 4 months. It seems like frontier models release new versions every 2-3 months, so raising prices in 4 months seems like a bad plan: you're effectively planning to charge users twice as much for a model that is no longer frontier.
IME 3.1 Pro still has better system-instruction following than Flash 3.7, esp. when there're many conditions and clauses. 3.1 also writes better prose for technical material than Flash 3.7.
Once the system prompt complexity goes up, Flash starts to write very dense english. it might be fine for tasks like coding, but not for user-facing text meant to be digested by the average person.
3.7-flash has been useless many times, specially when context gets bigger. 3.1 is the only Google model that has seen use from me. With extended thinking, 3.7-flash is kinda usable but not without many problems. I find myself falling back to 3.1 often. I don't believe in any benchmarks because whatever they are doing to award 85% to 3.7 on anything, they should seriously reconsider that test for anything.
In my experience? No. 3.7 is faster and it just seems to get things right more often. Only big architecture tasks and analysis make sense with 3.1, perhaps, but honestly just use the Opus 4.6 to generate a plan and then switch back to flash for the implementation
It is still going to be better at text work, skills, document review, deep reasoning, architecture review, etc. It is only 6 months old, it isn’t like its world knowledge and software knowledge is really out of date. Use it to churn on harder design problems.
I see benchmarks beating sol terra and sonnet. But is actually better? Has someone used it? I don't see actually much people that use Gemini for coding.
Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.
That was sunset and replaced by Antigravity. FWIW until I abandoned it knowing the sunsetting, I was able to get good behavior out of Gemini CLI with overriding the system prompt. The default prompt crippled the harness with very poor instructions, but there was a hidden ENV to override it. Replacing it with Claude Code like prompts based on the model selected, it ran at a much higher intelligence level full stack with significantly less errors.
Do they officially support you use their AI Pro subscription (or whatever the heck it's called this month, the one that gives you models in antigravity) in a 3rd party harness?
i always thought alphabet’s own youtube videos must be a comparatively good source of new training data. if slop and other garbage is reliably filtered out it should leave plenty of higher quality content.
Dear Google, Kindly make you chat window on the right side of vscode in antigravity extension, There is a reason others kept it like that. I can see the code and inspect the files changed while Agents keep working. its critical for me personally.
[1] For tone and instruction following, a positive percentage increase represents an improvement in the tone of the model on sensitive topics and the model’s ability to follow instructions while remaining safe compared to Gemini 3 Flash. We mark improvements in green and regressions in red.
Gemini 3 Flash?! So is Gemini 3.8 Flash less safe than 3.7 Flash in all areas besides Text to Text Safety (and identical on Image to Text Safety)?
Why bother with a column “Gemini 3.8 Flash vs. Gemini 3.7 Flash” when you’re going to disregard the label for 20% of it? Also is the “Tone” label short for “Tone and Instruction Following”?
We also had GLM-5.3 flash and Qwen 3.8 Flash Next, everyone's getting flashed and I think it's a good trend.
Almost suspect that the rate of improvement to post-training is so fast that small models have an advantage - it takes much more compute to train a bigger model, so the flash models are just running in circles (well, not exactly of course) around the larger models right now.
3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.
Nice surprise. In a few of my own tests it seems maybe a tad slower than 3.7 (but still way faster than any other LLM I've used) and even smarter. With 3.7 I felt I could just not use 3.1 Pro at all and 3.8 seems even better.
It's a shame Google crams it ham-fistedly into search results and that Google has some of the reputation it has because I actually really enjoy Gemini and I don't even use it for the reason people often list which is that you can cross-reference it to stuff in your Google account
Whatever they’re using within the Maps app is not good at all. I cannot just ask it for things conversationally like I do with ChatGPT. They really need to put a better model in there. I don’t even think it maintains context across two different queries within the same session. It’s not seamless and doesn’t just “get it” like ChatGPT does.
Yesterday I asked for food stop on my road trip 45 minutes from the current time and it gave me some options, but then I changed my mind and specifically asked for Asian restaurants and it completely forgot about the 45 minutes and gave me the closest Asian restaurant to me.
>"safety performance" - this starting to get long in the tooth. Gemini cut programming session 3 times for "safety reasons" yesterday for mentioning image generation (I need to generate bunch of those for infinite zoom virtual training app experience). After I got creative and managed to trick it to answer t was of course because "think of a children"
And in my other app I was debugging and using OpenAI to optimize some path it cut me off numerous times because it did not like JIT functionality (this is my commercial business rule evaluation engine that compiles rules to executable code inside the app to increase performance using asmjit library)
I am basically paying for them to waste my tokens and time on these 2 tasks
But a good agents.md, starting from a clean slate, and specifying which key files to look into and follow the standards allows me to build gigantic projects even I struggle to keep in my head structurally.
Seems maybe you’re keeping a forever-session and multiple independent tasks end up overstaying in context?
I would say either start new sessions for new tasks or limit the context to something smaller than 1M.
I usually start with research/planning session, this goes into a detailed implementation plan and then a new session for the actual implementation.
If it's complex problem maybe a review/adversarial step between plan and implementation.
Also with forever-session any time you take a longer break (depends on model and provider as to how long) you will push an entire big context again without caching even if you don't need it. With 1M context this gets expensive.
And yet again another failed launch from Google. I pay for their AI plus Google one package to get more cloud storage (have no interest in their AI bundle but you have to pay). and all I see in the Gemini app is 3.6-flash
Well it's in AIStudio and Antigravity always the same day even free tier. Chat is usually low priority, especially if you don't have subscription. Are you on Pro/Ultra?
I'm trying it now for token heavy coding tasks, it's capable for many tasks but in noway compares to Claude/Sol - requires more prompts and the output isn't as good.
So just another mid-tier flash model, nothing exciting, but Antigravity has very generous quotas so it's a good workhorse model when your Claude/OpenAI subs run out.
And whilst it's a fast model, having to baby sit through and approve prompts every few seconds ends up making it slower than the Auto approve modes of Claude/ChatGPT - they definitely need an auto approve mode.
Latest rumor is that 3.5 pro was struggling to be meaningfully better than flash, since iterations on flash were moving much faster than iterations on pro, likely due to model size (flash is estimated to be in the 200-400B range).
Gemini is getting less useful with each update. I could edit a pdf with the 3-pro model before but 3.1-pro couldn't edit the given pdf nor it could generate one for me.
If you want to "edit" a PDF, then Claude Sonnet works well, although what it's going to do is regenerate it from scratch trying to retain overall formatting. It can even do this for scanned PDFs and foreign language ones that need translating.
If you just need to create PDFs, not edit them, then Gemini notebook (notebook.google) works well and has Google's usual very high free usage limits.
AFAIK in general you can't really edit PDFs since it's not a reflowable format - even with Adobe tools all that editing does is modify the text within a text box - not reflow the document to adjust to any change in size of the text box.
One place where I find the Flash models surprisingly bad is Google Search's "AI Mode".
A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe.
Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was there is no way to unsubscribe through the account, so I just blockthe emails instead.
I've run into this pattern quite a few times. AI Mode seems to make up things all the time.
I think that's just a limitation on the size of the model. I'm pretty sure that they use a pretty small model in those summaries to save money, which naturally makes them a little less smart.
>We will not send marketing emails to a user who has opted out of receiving them. Any marketing communications we send will include an unsubscribe link at the end of the email.
I don't think this is AI's fault. This is Pearson's publishing incorrect information and the only way to really know they are a bunch of lying assholes is to have an account and try to unsubscribe from it.
It's obvious that the Google Search AI Mode encourages the model to give an answer without spending unnecessary cycles investigating deeply.
They also heavily encourage keeping the context short. For example, it will remove the option to start a new turn after a small number of turns, depending on the topic.
It definitely makes things up all the time, but it gets it right surprisingly often. I really like it.
Not to rain on anyone's parade but I find it strange how excited and giddy people on HN get for any new X.X model releases. Pumping it straight to the top, clamoring to use it, check and compare benchmarks, bragging about it being your "daily driver"?
Are you people truly this excited about this crap? I mean I guess if you work for Google or Anthropic or whatever I could see it??? Otherwise, are these just bot comments?
Gemini Flash is the one I get most excited about, because it's so fast and so good at real-world knowledge, and it's improving so fast - look at how much the benchmarks improved in ~1 month. It's just categorically different than anything else.
Also, I use it every day, and it just got ~10% better at coding, according to the benchmarks. How is that not exciting?
You know how the saying goes that you have to pick two out of three: cheap, fast or good? This is all of those. Pretty exciting.
I'll wait for Astra and Grok 4.7 announcements but probably getting at least one Ultra subscription.
Since testing 3.7 on Pro for last two weeks I'm realizing just how long I'm waiting on other models. I've been multitasking to compensate but it's exhausting so I'd rather not.
Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":
https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f
Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...
reply