upvote
when are we going to stop pretending these benchmarks have any meaning?

anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

reply
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.

(I'm not happy about the above being true, but it's the reality I seem to inhabit.)

reply
And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.
reply
I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
reply
Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.

Compare that to Muse spark 1.3

$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)

It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.

reply
With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
reply
and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.

This technology is strictly an extractive parasite on the world. Use it, but don't be excited.

reply
My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.
reply
global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.

i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.

depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.

the work does nothing or causes net harm.

reply
You’d expect that, wouldn’t you? But, alas…
reply
reply
"In Marxist philosophy..."

Well, that's about the same validity as "In Western astrology..." or "in flat earth theory..."

reply
Would you care to discuss the topic, or just throw grenades? Surely you can come up with something more substantive than this
reply
[flagged]
reply
I'm using AI to build things I wouldn't (and/or couldn't) have built before.

That's the opposite of parasitic.

reply
[flagged]
reply
Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.

People already started using contributor API, and your input is irrelevant.

reply
Don’t you have some looms to break?
reply
And the unabomber has entered the chat.
reply
I’m retired so it won’t be replacing my labor :)
reply
The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?
reply
[flagged]
reply
But is the score really reflective of the quality or are both models benchmaxxing?
reply
Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.
reply
Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.

This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.

reply
how much of it is from reallocation of staff to ai training and labeling
reply