upvote
3.8 Flash is just quite good, and so is the Antigravity harness.

I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

reply
I use Antigravity but for some reason, `agy` in the command line feels very bad/incapable of doing things. I can't quite explain it but the most common issue I run into it is just hanging on being unable to finish a tool call
reply
Even if agy was the best (it's not, and is missing basic features) you wouldn't rather have a choice?

I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.

reply
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.

I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.

reply
I have used it for little more than 6 hours or so in total but I'm pretty sure it doesn't have compaction?
reply
How else would it work? Less technical people don't even watch their context usage.
reply
In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model's window.

That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.

reply
It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
reply
Auto mode?
reply
It absolutely has auto mode.
reply
agy cli does not have auto mode. I've tried and tried and tried to work with agy cli sandbox-mode and just failed.

  agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.

gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.

IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.

reply
I think it only has "--dangerously-skip-permissions" Claude and codex auto mode will reject certain actions. No secondary check on gemini/agy AFAIK
reply
Via cli switch, but in process w/o fine graining? If so please tell
reply
Auto mode means that another model reviews tool calls to attempt to disallow less safe ones. It's different from bypass permissions mode which typically just doesn't filter at all.
reply
Does it have /goal feature similar to Codex?
reply
I use a variety of models for various subagents. I don't want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
reply
There is a pi plugin to use agy directly from it.
reply
You get banned if they catch you.
reply
Yes and I've seen reports of it being an ENTIRE GOOGLE ACCOUNT BAN.

I don't want to mess with antigravity because my google account is too entrenched in my life.

reply
I've been tinkering with Gemini for several months and I think it's great. The most complex things I've had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.
reply
Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
reply
Gemini for day-to-day and top-of-head queries and claude for the real beefy work
reply
For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.

And I saw it do this twice, once for Android 14 and once for Android 16.

I think this is just within 3.8 flash's capabilities.

reply
3.8 Flash (but also last two ones) have really strong preference for dissecting binaries with quick thrown-together bits of python in my experience.

Including going first for decompiling AGY binary instead of searching the web for documentation...

reply
Astra also really loves reverse engineering binaries. I guess it's one of those things that isn't that complicated but is super tedious, and tedium means nothing to AI.
reply
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
reply
Sol 6.1 is quite good, but damn is it slow.

I'm using it to run overnight tasks, and that's it until my quota runs out.

Canceled my subscription.

reply
WHY ARE ALL OPENAI MODELS SO CHATTY - i thought claude kept going on, then i literally put it in claude.md that summarize your thinking in 200 words or less and tell me in points what you did and what's next. Did the same for CODEX - nope still keeps effing going on and on and on
reply
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
reply
I'm also not using it for coding but I've found Flash 3.8 to generate much better HTML output than Sonnet or Opus.
reply
The web version of Gemini is awful at search but I don't think that's the models fault.
reply
Hallucination seems a very dated term.
reply
Why? It's the same concept and root cause it was when we first started using it.
reply
Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.
reply
what is so terrible with their harness? I've been using gemini cli, now use agy, Pi agent harness, and agent (cursor), and my only real issue with agy was the permission handling, but other than that, it was ok.
reply
Please tell me you published your findings even as an issue on the llama.cpp GitHub
reply
he is still closing his jaw
reply
Spoiler alert: the problem didn't actually get fixed despite the jaw on the floor.
reply
Why? Anyone can run that prompt.
reply
Not everybody has access to AI. More than that, every prompt uses insane amounts of natural resources. So why not share it.
reply
The resources per prompt aren’t that much .

Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.

reply
All your examples are private goods: excludable and rival. If one person uses a unit, that prevents others from using them.

Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.

reply
Yet you participate in a society.
reply
If action X takes a million times more resources than action Y, it's silly to focus on or highlight action Y. Seriously: if you are a regular meat eater, your choices use several orders of magnitude more water than even a heavy LLM user. A quip from a comic doesn't somehow erase that or make it irrelevant.
reply
Checkmate. /s
reply
why reinvent the wheel and spend tokens for a problem that has already been solved?
reply
I had a similar but less impressive experience recently with Muse Spark 1.3.

Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.

It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.

reply
Godot encryption is laughably easy to break, there's tons of packages available for it. It's a well known drawback of using godot
reply