upvote
Sounds like OpenAI are in the token-maxxing camp, so who knows what individual employees are doing to work their way up the leaderboard?

If you spend $8000 to generate an animated pelican riding a bike, then how much tracking does it really need?

Is the guy who spent $300,000 or so translating the FLT proof to Lean going to get a big Christmas bonus?

reply
That was Anthropic.
reply
End of day, output and results are top target of measurements, token consumption is the obvious number that they would like to disclose for their own business benefits and a simple metrics that correlate with the output.

Rest assured, capitalist appears irrational in wasting money, but they certainly care more about profit.

reply
Taking a profit means you have to show numbers and the sooner you show numbers the harder it is to take people’s money.
reply
Can you elaborate on this? Especially the tooling.

I tried something similar and I remember it was still pretty dodgy in February.

reply
my stack in a sentence: refine the docs/prompts/skills often, that's your biggest job, use both frontier labs models reviewing each other, don't solve individual problems only the systemic ones (set standards strategically, don't define tactics)

If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review. If I had $100k to spend next month I could probably get through it, I'm running $2500+-api-equivalent a week at this point and I feel very token limited. Will be time for a 2nd or 3rd subscription soon for both labs I think.

Fable was a revolution, still learning how best to use it, 5.1 felt like a notable upgrade. At this point I launch a workflow with 10-20 minutes of interactive setup (and even that I feel might be too much), it runs for hours, and the PR is trivially mergeable (I still review every line, but 95% are just merge, maybe 4% are feedback needed, 1% are thrown away and regenerated, which implies I'm being insufficiently ambitious)

reply
These researchers are paid millions of dollars for their work. I doubt trust is really an issue at that level.
reply
Yes, because no employee with million-dollar comp has ever been untrustworthy in the history of business.
reply
Imagine if one of the humans at OpenAI was misaligned! We should get the AI to research this possibility once they've been aligned.
reply
That would be the mother of all circular accounting: the main clients of OpenAI are OpenAI employees.
reply
> let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware)

How are you running jobs unattended 24/7 without hitting your token limits?

reply
I'm currently running two 24/7 semi-autonomous AI research projects using Fable 5.1. It's on track to burn through my weekly quota in about 3 days. I check progress in the morning and in the evening, and provide some light steering.
reply
My only experience in >24h agents is with economically sane models (one of GLM5.2, 5.3-flash for orchestration, DSV4-flash for implementation, and glm5.3|sol|kimi3 agents + subagents reviewing at the end)

Over 24h my token spend is <30$. Excluding tokens for review it's <10$. With the absurdly gigantic subscription subsidies and a reasonable workflow I suspect one could run parallel agents.

I'm not sure what the point would be though unless working on some kind of optimization problem -- it takes me days to review <24h of the agent's output. It's almost always near enough to correct to be shippable; though I do give it feedback and iterate until it's better than the code I would have written.

reply
This sounds like more work than just writing the code yourself. You'll say it isn't. I don't believe you.
reply
/loop ?
reply
I suspect the $8000/day figure is the equivalent in API costs. But I also suspect gross margin on their API rates are 80-90%
reply
[flagged]
reply