upvote
To back this up we have a discord chat for the board game terraforming mars where the agent takes input and vibe codes an open source implementation of the game.

https://tfmbot.com is the link (discord and source links on the splash screen).

The results are fucking incredible to the point where people in discord are stating "I'm surprised this is working so well". I am too.

I feel like there's a group online that missed the boat. Anything negative towards AI capabilities is still upvoted but I've been in the industry for over 25years, highly respected and can't fathom the "AI dumb lololol" type of comments i see on HN. AI is superseeding all other ways to develop.

reply
Every time weird stuff happened this week it was because the model choice in VSCode got set back to “auto” and some other model was trying its best. 5.5 is what I just set it to. Even Fable feels worse for my use cases.
reply
Even Anthropic rates Fable lower than 5.5 on pretty much all benchmarks.

"Why does Fable even exist" is a very very reasonable question right now.

reply
I would also say: I think this works great because the goal is well-defined and measurable.

I’ve also used Opus 5.5 on some hill-climbing, and a lot more steering is required here, because … eval is hard.

reply
Agree Opus 5.5 is incredible and efficient with Claude Max Usage.What a jump after the writing slop you got from Opus 5. I had moved to using Fable for my orchestration workflow mainly due to the communication issue (I think Opus 5 was capable enough but spoke in riddles so you lost confidence quickly)..With 5.5 it needs less steering now and communicates well, and I have had the same CC session running for the last week (obviously compacting with durable plans etc as the post), with it PM'ing my home built agent orchestration of the other coding agents(antigravity, codex, pi etc) and making decent decisions and all the recommendations are normally usually good.
reply
It’s great, but you still need to know what you are doing, not just the goal/s. I built a platform years ago from scratch, and now I am remaking it with more features and more polished design, I know exactly what needs to be done to tiniest details. The first prompt was very well detailed about the architecture and how everything should work, after an hour work at Xhigh, it did create the blueprint artifacts that I asked for, then I spent 3 days reading every single thing and writing notes, turned out it made the system overly complicated without adding extra value, plus I can see how some of the architecture design will be potentially a security risk. So after few days I fed my notes, this time took 4hours and 1M tokens! Later I spent more few days reviewing and writing notes, it was closer to what I want but still made architecture errors, the third run took around an hour and finally made it how it supposed to be, although there are still more notes on non critical stuff. So I don’t think we are yet at the stage where sitting goals and some high level is enough to produce quality results.
reply
9 hours?!
reply
Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
reply
I dare not ask about the cost, having burned $60 on a task running for 1h 16min once.
reply
I’m on the $100/month subscription; this session took about $500 in token-equivalent costs.

(Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)

reply
God I hope the prices drop quick. Once they stop subsidizing it these kinds of workflows will be unaffordable for anyone who isn't already wealthy
reply
Does it resume automatically on higher subscriptions?

I'm on a $20 plan and it never auto resumes. I have to go back in and type out resume or click a button.

reply
deleted
reply
Instruct it to arm a monitor (every hour or so) to wake him up in case of quota or api issue.
reply
Meh... There's a reason Opus 4.6 is still an option.

Pros know these are lower cost models.

reply
I do like Opus 4.6, but I think 5.5 on medium or low is a better value. I have to steer 4.6 more and build more scaffolding around the tasks. 5.5 just does what I ask. Visual spatial reasoning greatly improved in 5.5 as well.
reply
Is "wall-clock" an actual term you used before Claude? I had never heard it before the model used it and I can't stand it.
reply
It's a pretty old term, to distinguish from e.g. CPU time. This was in common usage even 30 years ago.

Example from 15 years ago: https://stackoverflow.com/questions/7335920/what-specificall...

reply
Trying to make us feel old? Very common term among people around my age and higher (40+). Please don't turn "I've never heard that expression before" into "must be AI saying it".
reply
It's a standard term and has been for ages. It distinguishes end to end time vs eg the amount of CPU time an individual process uses (which excludes time spent waiting for the system or time when the process was otherwise not scheduled on the CPU).
reply
“Wall time” is a pretty common systems term (you see it when comparing total runtime to kernel time, for example). I wouldn’t have indexed on “wall clock time” or other variants as an LLMism.
reply
I've used wall clock for many years, normally when compared to CPU time when talking about parallelizing some process.

CPU time might go up while wall clock time goes down

reply
Wow...it's super common. For example, timing a program you can get user CPU time and then wall clock time, which can be two very different numbers.
reply
Lolwut? You seriously have never heard anyone say this before?
reply
>> It’s a really good model.

In the meantime, I have cancelled my Anthropic subscription...

I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.

I start with some code produced by an Anthropic SOTA model...let’s call that Code A. Then I get Code B and Code C for the same task from models by two other vendors.

Then I ask each model to review and critique the other proposals.

By the end, both the Anthropic model and I usually run out of arguments... against them and agree that proposals B and C are better.

Claude then always asks whether it can incorporate the code or ideas from B and C into its own solution...

reply
It's not even funny any more. Chinese model, Chinese vendor, Chinese, Chinese... Did I say Chinese? Chinese!

Nobody in his right mind will use a Chinese clone when you have models like Opus 5.5 for peanuts.

reply
Yes ...I was so impressed I cancelled my subscription. I could not stand all the winning. I offer cheap hourly rates of $1000 for debugging AI slop.

Contact me at : prompt.plumber@gmail.com

reply
> Nobody in his right mind will...

let a few valley elites decide how humanity can use this technology

open and transparent is the way, China is showing how

reply
I’m rooting for open models, but SOL 6.1 and Opus 5.5 are absolute workhorses on a $100/mo sub. I share your fears, though, and really hope an open model catches up and can somehow compete with the subscription prices of the big 2.
reply
I have workhorse models, spend far less, the model matters less than people like to claim

there's no money long term in being a token vendor

reply
> Nobody in his right mind will use

Nobody AMERICAN in his right mind will use... Wait, actually a lot of them will.

But for me, as a non american, non chinese person: I'll use whatever the fuck is the best and cheapest for my task, because that's how fucking Capitalism works.

If that means that a (proclaimed) "communist" country cleans the carpet with the self-proclaimed land of the free: so be it!

reply