Don't get ahead of yourself. Harnesses are not exactly rocket science and will be a commodity.
The real value providers here are the hardware, then the LLM as a distant second, and at a much larger distance the harness.
Labs are now post-training models with Harness so that Harness now gets absorbed into the weights.
Solar goes all the way up => power is commodity.
Some hyperscaler goes bankrupt => hardware is commodity.
Models get real good => output is a commodity, no profitable problems to solve anymore.
Open source models get good => models are commodity.
I really though this comment was a satire ...
This entire forum is infested with shameless hype chasers and biological linkedin bots.
I’m thinking of how in cyberpunk, people are replacing their cybernetic enhancements all the time. You could alternatively bioengineer your own body towards the desired outcomes, but that’s more constrained by the trajectory your body has already taken, whereas the promise of cybernetic parts is that they are more independently replaceable. (Probably an illusion in practice, but I’m talking about the fictional ideal.)
As another analogy, monolithic software tends to quickly become hard to change significantly, whereas a plugin architecture tends to be more flexible and modular, and people can share and combine their various plugins.
They probably used an LLM to come up with this bizarre metaphor.
- They can already reason better than many humans and are still improving all the time
- Harnesses are improving all the time
- We're already exploring things like long term memory, long term goals, and other things that humans have which LLMs traditionally lack
- An AI agent can read and reason about every piece of AI research ever published, including looking for insights that humans may have missed. A team of humans could never do this even if they dedicated their whole lives to it.
- They can design and execute experiments on a mass scale to determine what does and doesn't work
- Large AI labs have more than sufficient resources and motivation to throw at the problem, and are in fact doing this.
If you provide what you'd consider AGI, we may not agree on that definition, but I and other skeptics could at least discuss with you whether A.) that seems reasonably achievable given LLMs inherent limitations and B.) whether any of what you'd listed is actually likely to get us there.
As it stands, neither is possible without knowing what you believe AGI to be, but for what it's worth, coming from someone who both does see LLMs as valuable tools but whose definition for AGI also contains, among other things, reliable self-assessment of factual uncertainty [1] and basic counting and grade school maths [0][2] without tools or eternally scaling training data, I have yet to read any evidence that LLMs can achieve my, rather strict, metric for AGI.
These models are amazing tools, their ability to leverage massive amounts of high quality training data to further sciences truly awe inspiring, but that does not mean intelligence, at least in my definition that requires some internals these models have never been proven to possess. It's nuts that solving Erdos problems can be done by a model which struggles to count or solve a sudoku without external tools, but that's where the technology has been for years now and no paper I have read has shown that LLMs can overcome that to any scalable degree. You can push further with training data, but the limitations remain, albeit less noticeable. Any externalities, be it tools (self-scripted or called by the model), external memory solutions of all shapes and sizes, etc. I personally also feel cannot be required for or lead towards AGI as intelligence may be better leveraged by such externalities, but should never require them, so much of your suggestion I feel shouldn't be considered even if one believes LLMs can yield intelligence. I will admit that I am very extreme here though, this is not a position held by everyone for good reason. At the end I will always point towards the "extraordinary claims require extraordinary proof" of it all and that LLMs, in the face of any doubt, should be viewed akin to how Stockfish can play better than any grandmaster, but that does not mean intelligence, at least in my world.
If your definition for intelligence does not require basic arithmetics or an understanding of ones own knowledge gaps, then maybe LLMs can achieve that, but I'd push back on that truly rising to the AGI moniker. Maybe a more comprehensive or even my definition of AGI is possible whilst keeping the autoregressive nature after all, but there is no evidence supporting that by itself and quite a few things that haven't even begun to be overcome before something of that magnitude could be honestly considered.
It's akin to "let's colonise Mars by 2020 or 2030 or 2040 for sure, then terraform it" proposals. If that were possible, wouldn't we see a lot of these methods applied on earth and in a moon base long before (as in, we'd have had a permanent moon base in the early 2000s)? Same with LLMs, if they can truly yield AGI, we'd see some of the major deficiencies dealt with long before. The fact that we neither are terraforming earth, nor have any permanent off world colonies, nor have solved some of the listed, inherent limitations with LLMs by their design, that's what informs my skepticism that both are reasonably achievable in the timelines some industry "experts" (read hype merchants) propose on the regular. You tend to see some progress, a path toward solving actionable problems long before full implementation, at least in the real world...
[0] https://logicalintelligence.com/blog/energy-based-model-sudo...
And no I came up with the metaphor all on my own, send me the chat of you getting the LLM to come up with it. Why not argue based on merit instead of strawman and ad hominem attacks?
Harnesses (and the concept of agents before them) presuppose competence in LLMs which simply doesn’t exist.
0. https://www.businessinsider.com/sam-altman-ai-utility-electr...
His idea of metering is predicated on the thing he’s selling being AGI, it is not, and all his predictions have turned to dust.
Also that isn’t how metaphors work - they illuminate by comparison, if the comparison is not close they are not useful.
I don’t believe in AGI, but that doesn’t mean I don’t find AI useful. I just understand that the correct harness can take them to the next level.
Then how do you explain the wild success at using them for development?
That doesn’t make them intelligent agents which think independently.
I have a system that entirely reverse engineers old arcade games. Creates semantic symbol mappings that were considered impossible just a couple years ago.
Granted, it took me a couple weeks to build the system.
From impossible to a couple weeks in just a couple years.
Would you like to see it or continue to pretend these things don't exist? Your call.
(It's finding the coolest stuff - the anti-tampering hacks they put into the old machines is fascinating.)
You can never tell if the goomba opinion of the forum will agree we have reached AGI (seen that happen on a few threads lately) or will readily call that a ludicrous proposition.
The words "once that settles" are doing historic levels of work here.
No human on earth has a clear idea whether model technology will settle tomorrow or 100 years from now.
There's every reason to expect architectural breakthroughs will keep being discovered and causing nuclear blasts of forward progress.
I do agree that harnesses are going to extend AI capabilities a lot in the next year, but after reading Pi's page I don't see anything that makes it particularly special in terms of functionality, other than being more provider-agnostic.
Many of my harnesses eventually turn into customized UIs around the chat interface.
The harness facilitates the work animal doing work for you.
Not climbing harnesses to keep you safe.
1. You can use the '/new-tool' and tell what kind of tool you want (including whether it should be task-scoped, workspace-scoped, or global), the model builds it, the harness runs validation and other tests until the tool is ready
2. The model decides that in such and such task, it would be helpful to have a tool like this, it can build a task-scoped tool.
In either scenario, the tool catalog is rebuilt, and the new tool is instantly available in the next turn.
What I can see is a world where we end up with a Chromium-shaped harness, a fully featured standard implementation everyone builds against, because doing every single thing yourself would be crazy.
The antithesis to Pi, if you will.
Also, having only a "standard implementation" makes no sense for a harness. A standard implementation would need to try to be as good as possible at all things. But you'd often want a specialised harness designed for exactly your use case.
Some standard solution will emerge, which will be amplified by models being trained specifically to work with it.
I primarily like how it manages sessions, and how agents can easily reference other sessions.
So if you do want to use it, use the Codex sub. Once you install it, run Pi and /login and you’ll get login with ChatGPT. From there, Pi can tweak it’s settings if you ask. Check out their extensions (or ask Pi) and that will take you most of the way there.
What hiccups were you having?
Not out of the box, but you can add agent sdk. I'm not sure how great the results will be though.
looking at the website. i can't really tell if they have benchmarks and measuremnts on how all that improves capablities over just using regular agent withtout all that
Don't understand what people see in them.
also i think its hard to build general harnesses if they were trained on specific harness architecture.
There’s evidence of harnesses making a smaller, weaker model perform better than SOTA and some benchmarks ban harnesses because it becomes too easy.