This is many people at many jobs. We continue to see this weird thing when AI comes into an industry that suddenly the product being produced by the people was perfect/amazing/whatever. Maybe that's true for the majority of HN who are lucky enough to work with experts in every field, but in the general population that's simply not the case.
Many doctors, lawyers, and programmers are not very good at their jobs - I've seen it first hand. AI gives the general population a way to steer around these people a bit, and maybe ask the right questions. Is it perfect? Nope, but it's often better than the people someone has access to.
So, on one hand you can get an expert take that might be wrong, but is linked to an actual person, with a reputation and some level of ownership. On the other hand you have an over confident LLM that might be wrong and has no reputation, no ownership. How does that actually improve things compared to the older status quo?
It is entirely clear that is it not the case when I compare agent generated apps with what contracted software teams have produced.
For medicine it is likely worse.
Doctors do care, but they have to give an advice based on their 15 year old knowledge - they simply can't read through 83 papers in a quick session.
I think it is a matter of time before we see the first insurance companies assign greater risk to human advice (legal, tech, medicine, etc.) than to agentic advice.
You likely have to change your idea about this. Heck, this view was wrong 6 months ago. Sticking with is becomes a hazard to patients.
That's not my argument... I didn't draw the conclusion that because of those 2 factors humans generally provide better answers. I personally have no idea if that's the case, and for sure wouldn't rely on my personal feelings to evaluate that
> So, on one hand you can get an expert take that might be wrong, but is linked to an actual person, with a reputation and some level of ownership.
Regardless, I agree.
The newest studies still works on llms that are two years old. It doesn't appear that proper medical harnesses with frontier models has been evaluated.
My slight intuition is that we will already now see results that are much better than average human doctors.
That's the whole point, yes. LLMs are inherently unreliable. Both human experts and LLMs are unreliable in their own ways. The human expert has a reputation and some level of responsibility, the LLM doesn't
I think this extends to the other areas that AI is also disrupting.
Software engineers are also absolutely terrible: In being sloppy with handling errors, doing coercions, etc. All because always engineering after best practice simply takes too much time and mental load.
The real disruption from AI is likely that it does not get mentally fatigue. It can keep on, launch adversarial audits, make sure that everything is up to all standards.
It's worse than mental fatigue. A friend had GLM 5.2 suddenly become dumb for hours. It started forgetting instructions, randomly inserted some bits in simplified chinese... Apparently throttled due to demand. there was tons of tokens left. (btw when your llm runs out of tokens is another example that's worse than fatigue)
A tired programmer usually knows he's tired (and can take a break) and fucks up in more expected ways
6 months ago everybody talked about how AI was not able to plan their work. They could only do minor changes.
The introspective property of knowing when you are "tired" and acting appropriately will likely be trivially observed in 6 months.
operators will never expose this guess to the user (who wants to pay for a tool that complains and refuses work?) so they probably will use it to load balance differently. and if they are out of resources then they can't load balance so they'll just hope you won't notice, so we're back to where we started. maybe it's already roughly how it works now...
what's most likely in 6 months is a big correction in pricing, once they get people hooked up on this enough. maybe there will be a model that manages to detect when it's stupid but it won't be available outside FAANG and DOD
Naturally, LLMs do not get fatigue. As such they also don't need to detect it. That is correct. But LLMs can see if their performance have fatigue-like deterioration and correct for it.
That case is even more trivial and has nothing fundamentally to do with LLMs. Capacity is added everyday.
I would not base my long term projects on what happens in the market based on the fact that a single provider needs to throttle to satisfy demand now.
Your assumption that the current pricing is not sustainable might be fair. Personally I believe the opposite, and I havde not seen any indication that inference should not continue to decline in price.
In particular, LLMs appear to be hyper commoditisable. So if Anthropic or OpenAI is doing pricing shenanigans, people will likely move on.
they are not human so human sensations don't apply. but I heard when you run out of tokens it's also kind of like fatigue, just less predictable
Let's face it, only thing lawyers are there for right now is to take liability. And even that may not be needed 5 years from now.
Middlemen that can be automated will be automated.
I hope the irony is lost on no one that "Attention Is All You Need" was the paper that kicked the LLM boom off.
Part of the distinction between a how good one lawyer will be vs another is exactly that: attention to detail. And, as a trade, a practice, it seems like a skill that they try to hammer into the people who pursue this career (to varying degrees of success clearly).
It makes the last 20 years of leet code hiring look downright wrong. Your ability to recall what sorting algo to use, or solve some brain teaser conveys nothing about your ability to read (massive amounts) of code, and think critically about it.
At the end of the day, I think these paralegals have been quite weak links, and nearly nobody will notice a difference in quality when it's all handed over to LLMs. But the paralegals will lose their jobs, and probably have nothing better to do than purchase angle grinders and cut down Flock camera poles, or crown the front hood of Waymos with a traffic pylon.