I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par, to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
I'm using it to run overnight tasks, and that's it until my quota runs out.
Canceled my subscription.
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
Including going first for decompiling AGY binary instead of searching the web for documentation...
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
Nobody has a moat.
This is the kind of story that ones tells to investors to justify the huge amount of cash burn. :-)
I think many/most of the players will crash and burn, and the ones that are left will divide the world.
The net effect is that the most likely scenario is if one big lab fails, they will likely all fail. Their revenues are all correlated.
To go to your dotcom comparison, the winner will be the ones picking through the assets that were written down by orders of magnitude and trying new products with the technology until one sticks to the wall. My base case isn't a dramatic crash but a slow burn. Telsa is a good example, revenue has a dramatic growth period, then stalls out. Stock price remains at a point that's unrealistic given the lack of growth but it can stay there so long as the balance sheet doesn't deteriorate.
Concerning the leverage on energy and real estate: don't forget that the AI companies have quite a lot of choice where to build their data centers. So AI companies have lots of opportunities to play several parties off against each other (in particular also for real estate and energy).
Uhm, what? LOL.
People dont value firms based on balance sheets fella. Have you taken a basic valuation class?
Tesla is a nice stock for traders - they like the volatility. Nobody holds Tesla as stock for investing. If you were to truly value it on an intrinsic value basis you'd have to bring in failure risk.
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.
Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.
Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.
I would argue those who already rule the world, will continue to do so.
What happens to OAI and Anthropic? No idea, probs go bust. Google just has to offer a half-decent offering in the long run and have a cost-advantage and it'll eventually knock OAI and Anthropic out as firms figure out what combination of models they want to be best for their economics and generating returns. Enterprises trust google over OAI and Anthropic. A clear signal of this was the Apple deal.
Dont forget those sweet returns fellas! CEO's are hired to make the owners wealthier. That is not gone.
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)
One possible explanation: because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible - combined with some willingness to burn money so that they don't strongly fall behind on their AI models.
On the other hand, OpenAI and Anthropic at least formerly concentrated on building and providing the best models that they could with concerns about financial feasibility taking a backseat.
Just to be clear: I do have the impression that by now (likely because of pressure from investors) OpenAI and Anthropic take these financial concerns more seriously, but nevertheless Google's vs OpenAI's/Anthropic's "DNAs" concerning on what to focus on differ.
There’s no magic there. You get an account executive and a call with a systems architect to find out what you’re doing.
Clouds gonna cloud, this is the reason they rolled deepmind into gcp and arguably the inverse is true, the labs are trying to become clouds
That feels right. It's not as if they've been missing out on great profits.
Because it's not an existential battle for Google. If OAI or Anthropic disappear from the absolute frontier for ~8 months the news cycle and churn will diminish them to the second rate. Google is processing near 4 quadrillion tokens every month, that's - I'm sure - significantly more than OAI or Anthropic, because Google is interested more so in their flash models and getting these competitive, which they are.
Meanwhile, Anthropic/OpenAI will struggle to survive the next 24 months on their current trajectory.
Most Google products even use flash lite underneath, so their frontier model is mostly used for distillation.
A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
This is no contradiction to your other claims, quite the opposite: perhaps (or even likely) Google wants to avoid that their models become a commodity in some (agentic?) application where the middleman who actually writes this application gets a disproportionate of the money that the customer of the application pays for it.
I used it for a month over the summer, right before they were going through the migration to antigravity. It was a fine workhorse IMO, no complaints from me.
Data centers are important but a few others also has them: Amazon, Microsoft, Meta. SpaceX will likely be in/at the top I AI dedicated precessing power in 2027 as well.
I don't see the moat. I see a company with a lot of other commitments that is not the best at delivering consumer facing products. They have some good cards but so do others.
Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.
Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
If not now, then when will these companies be AI leaders?
Even Google, with its staggering advantages in cash, compute, real estate, training data, and having basically invented the field only manages to briefly claim a 1-2 week lead once or twice a year.
Google is already on gen 8 of its TPUs and is certainly already working on the next version or two.
I'm sure the thinking out there, and hence investment, is all about how to tether the user to the most addictive, network-effected, incredibly deep, server-side, moat-able version of AI possible.
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
Are you sure?
--
China Just Built What TSMC Said Was Impossible
https://www.youtube.com/watch?v=Pk-w279ESHg
--
China Just Built What ASML Feared Most
This is marketing from Google, not a competitive offering
---
maybe it's this Anthropic post on GLM?
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details
I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
Not if the genius level IQs take the market share.
Whoever builds the deathstar wins!
And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
So far, I don't think the models are capable of running away on their own. Of course, it would be playing with fire to not at least consider the risks of such a runaway scenario and build in safeguards against it. But, there is no model that can build a better model on its own, thus far, to the best of my knowledge (which is far more limited than the models, so maybe I should ask them).
At least they’re led by trustworthy and honest people or we’d need to take their claims with some dose of skepticism.
https://tvtropes.org/pmwiki/pmwiki.php/Main/GirlfriendInCana...
It means I am saying something that is not very believable.
The model is not available yet, so Google is essentially saying "trust me bro".
Gemini not beating the "can't release a model" allegations
ok bro thx
I already pay $300+ for subs. Please don't tempt me with another $100 sub just because I got curious if the benchmarks were right.
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
I was going to say I don't know what they'd do for C, since Carbon and Calcium are already things. But knowing Google, they'll probably call it Chromium.
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
Imagine they aren't even familiar with rust but are deeply familiar with the product.
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
They did that for Go and it seems to have worked out for them though.
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.
It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
I am just trying to understand - why Google haven't done and have no plans for it. They have done it for Android and have Gemma models too.
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
So no, Google is not being punished, nor are they the people behind this technique.
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
It's also kinda wild how the competition being at v6.1 makes 3.x feel ancient, at least saying you are at v4 now changes public perception a bit imo.
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
Its amazing. It really is.
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
[0] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2023/p27...
(Full disclosure: I am one of the coauthors)
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.
This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.
Good addition to the arsenal.
There was Deepseek v4, which then later Deepseek v4.1 came out and it went back down again.
Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...
Argon will launch at an introductory price [1] of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
[1] After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
===So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
Many other models get things wrong, but Gemini is the only one to go on the defensive.
https://www.fastcompany.com/91383271/googles-chatbot-apologi...
https://www.businessinsider.com/gemini-self-loathing-i-am-a-...
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
None of the other AI labs do this. Really frustrating.
https://www.bloomberg.com/news/articles/2026-09-30/google-gr...
I would guess it's like opus 5ish level from this
Even advanced AI can't prevent broken templates haha.
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
These numbers are meaningless. Shame on them.
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
I miss you, Gemini 2.5 Pro :(
For real though. If they've become commercially uninteresting, that would be a pretty cool move.
This would explain why benchmarks are seemingly meaningless.
Cant wait for this AI hype to be over, so I can Terence this shizz as old school too
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
Close but no cigar!
Just compare how much better presented the Astra announcement was compared to this one: https://openai.com/index/gpt-6-astra/
but that's a side issue; my main point is that you are underrating the impressiveness of getting a safe rust port of a highly optimised c++ library even nearly up to par with the original. the tradeoffs rust makes for memory safety cost it some of the raw speed of c++ even with all the zero cost abstractions and purely compile time guarantees they have. (tangentially i wonder if ats (https://www.cs.bu.edu/~hwxi/atslangweb/) would be a good candidate for LLM assisted ports; it seems way more advanced than rust and might actually get c-level performance with safety, but it's really hard to write.)
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
- https://paritybits.me/google-should-provide-a-technical-post...
- https://gemini.google.com/share/6d141b742a13 (last message)
what about input?
(Maybe I missed it)
And was 2M tokens IIRC after release.
There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
nothing about this announcement gives me confidence that google is back on track as a model provider.
> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
Jokes aside, looks like an impressive model!
I think that should be a really bad sign, but hope its great.
Get lost.
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
Goshdarnit they didn't see my suggestion: https://news.ycombinator.com/item?id=49899171
So Google is migrating codebases from C to Rust? That is interesting...
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
Behind how?
I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.
I have no interest in benchmarks.
If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?
With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.
- Person 1: X is garbage compared to Y!
- Person 2: Why?
- Person 1: Because I like Y.
And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.
I fundamentally don't understand LLM "brand loyalty".
All of the models are constantly leapfrogging each other and always have been.
Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.