There were open models from EleutherAI (GPT-Neo), Google Brain (T5X, Bert), and HuggingFace was promoting open models (and others doing open work I haven't listed here) all prior to 2023 and the big Chatgpt moment.
https://github.com/EleutherAI/gpt-neo/releases
https://github.com/google-research/bert
https://github.com/google-research/t5x
If you're learning about AI models it's still worthwhile to review these models and codebases because they continue to be the basis of the technology that's being produced today! It'll give you a good perspective of how things have changed, similar to say learning about propeller planes before moving onto modern jet engines.
Understand why T5Gemma is no longer under Apache-2.0 and honestly, have not seen that much advantage when testing that vs T0 in my experiments either way, but still, there are good reasons why plain old T5 and its descendants are still popular, licensing being among them.
Gemma team also has very consistently interesting models, especially like DiffusionGemma. Ironic, as (beside 2.5 Pro), I have never warmed up to the Gemini series of models but rate Gemma models far higher than e.g. Qwen in direct competition. In any case, thanks to the teams behind these for making as much possible.
[0] As in proper consumer hardware, not a cluster of DGX Sparks or Mac Studios solely for experiments that sometimes are asserted as being consumer grade...
When is Gemma5 coming out? :-)
I'm just beginning to enjoy the release of Gemma E4B why are you giving me a cold shower here telling me there may be no Gemma 5?
Here's hoping the supreme leader of the US does a speech like Chairman Xi.
What ended up happening was fairly limited gating followed by a “leaked” magnet link and llama.cpp which really brought a whole revolution in open use and changed the conversation completely.
I have no idea what role Meta played here, it may have been nothing, but they certainly could have been more guarded if they were really worried about the leak. The result was a big change in the trajectory of personal and open source AI use and even the dialog about it. Whatever the exact intentions, they were a key player.
You definitely have some scams and some software flaws being exposed, but I think for a revolutionary leap in tech this is all quite muted. Revolutionary tech advances often come with some severe consequences. For instance, to this day (after a century of safety improvements) cars still kill more than a million people a year, to say nothing of wrecking the atmosphere, but that's considered a reasonable price to pay for being able to get between places faster.
The reason I think cars are a good example is because that's certainly vastly higher than any price we're paying for LLMs, or probably ever will, yet the overall 'positive' effect of LLMs will likely be far greater than cars. Gotta put 'positive' in quotes because the possibility for automation and the like is going to be.... nuanced.... in effect, but at least in the longrun it'll be a very good thing.
I assume they are referring to LLMs escaping confinement and hacking other companies systems unprompted
> considered a reasonable price to pay
That's not a universal opinion and perhaps, just like for LLMs, we should have listened to the experts rather than gobbling up everything the industry pushed down our throats (to keep your car simile: SUVs are now ubiqutous in all European cities, there is no logical reason that should be the case)
I don't know a single white collar person who isn't worried for their job with llms. Every. Single. One. Lawyers, doctors, any type of office workers I ask. The definition of white collar. Plus don't let me start on various automatable blue collar jobs like drivers, warehouse workers and so on.
What do we get in return? Better search (for now, its already getting riddled with ads which by definition twist truth to highest bidder), some questionable psychotherapist for some desperate folks. What else? Cars are not flying, heck they are not even driving autonomously in any usable way, society is in deep shit everywhere I look, wars, environment reaching bad places and heading for worse, mentally unstable people holding way too much power, destroying lives of millions on morning whims.
Everybody feels like this is the revolution, it should be, it must be right just look at the numbers. Like proverbial guy with hammer, a very shiny cool hammer, looking for what to do with it. We all saw how sociopathic management in more harsh/capitalistic companies looks for any sign to let people go en masse.
I could go on for a long time. There is a lot of things to hate for most people, and very few to be happy for. It seems llms have the ability to get the best and worst out of humans, and worst part seems to be in abundance. Some revolution that is, 0.1% will get richer while everybody else the opposite and 1984 seems milder and milder version of reality out there.
I think I can agree with that. Technology does seem to have a way of crystallizing our worst tendencies.
<< Everybody feels like this is the revolution
I smiled at the analogy, but I would caution you to not trivialize it. There is a reason executives are pushing that point. There is enough of a revolution in it to make things complicated -- as if it was not already.
<< destroying lives of millions on morning whims.
True, but I am not quite certain what can be done about it at a personal level.
<< I don't know a single white collar person who isn't worried for their job with llms.
Dunno. Next few years are probably going to be fine. Society managers likely can't upend everything in one go. They would lose too much. I can't say that I am worried exactly. I can see the potential impact, but I think the potential benefits are worth it as long as we don't limit it to summarizing emails..
A lot of the stuff you mentioned isn't really AI (the risk of job loss from which I agree is real), but just everyday stuff that humanity has shrugged off since the dawn of time.
And that would be closed-AI companies failing to keep their AI closed in a box, thus proving open-AI proponents right and the hysterical group wrong?
https://illuminatingfacts.com/the-train-speed-panic-why-peop...
Reason that hysteria remains apropos now is that then, and now, we're figuring out how to deal with something beyond merely incremental change.
Irony is roads and sidewalks should have changed much more. We never did get around to good controls leading to zero deaths*, we decided a 100 people dead per day is a reasonable cost of convenience.
How much do our cyber traffic and cyber pedestrian controls need to change for everyman to get to drive AI? Car seats? Seatbelts? Air bags? Speed limiters? Pedestrian only living spaces? Driverless cars?
Many practical responses, likely a mix of things we haven't thought of yet, just as horses and horse drawn carriages didn't require most of them to coexist. Controls develop like scar tissue more naturally than they appear in advance.**
And if the better analogy for LLMs turns out to have been less like cars, more like flammable gas blimps, we'll figure that out too and tell cautionary stories for generations... but long before the stories are forgotten we'll have already come up with something more practical, faster, and – oh well – perhaps even deadlier per mile.
---
* Still working on https://en.wikipedia.org/wiki/Vision_Zero
** Not saying that's as it should be.
On the other hand some kinds of hysteria occurs due to a divide between the public and subject matter experts on topics. Kind of like Dunning Kruger. Take for example the people saying 5G causes cancer. Then there are people who blanket dismiss their worries, because "it's non ionizing you dummy". Given sufficient power you can still fry someone with non ionizing radiation, for example in radio broadcasts. Of course a regular 5g antenna can't put out that kind of power but my point is that people too quickly raise or dismiss concerns without actually critically examining the entire topic. This will likely worsen from specialization and progress in all fields and regular Joes get left further and further behind.
Problem is I’m fairly sure the link is wildly overstating its case, and looks like a content farm. Might even be, ironically, AI slop itself! For example, it talks about “railway spine”: “Physicians of the time believed that the jarring motion and vibrations of train travel could shatter the nervous system, causing lasting mental and emotional distress.”
This is bullshit. As just one example of the low trustworthiness of the “article”, railway spine was the result of a train crash, a notably traumatic event, and the symptoms described are in part just PTSD, a real condition and a real concern (not normal rail travel). For that matter railway travel in those days was genuinely unpleasant (lots of vibrations and jarring movement, poorly ventilated cars, and so on) which ironically modern science would probably validate as being some kind of health risk.
I’d actually view the listed example of train paranoia a great example of historical ignorance. People of the past were not stupid, contrary to popular belief. History has some genuine examples of silly hysteria, but these are usually the exception not the rule.
Sorry for the particular link, I'd grabbed whatever.
If one takes a minute, one can find contemporary newspaper articles. IIRC, many were about how people could suffocate.
https://books.google.com/books?id=Lm03AQAAMAAJ&pg=PA337&dq=t...
Contemporary women's health:
https://books.google.com/books?id=9-FXAAAAMAAJ&pg=PA71&dq=ho...
And contemporary survey of fears:
https://books.google.com/books?id=-RdAAAAAcAAJ&pg=PA107&dq=%...
Aaand, a 2023 takedown of claiming claims about historical tech hysteria:
It does line up with "PTSD" only becoming a thing after WWI.
Llama has never been open source. It's source-available, but still proprietary, under terms that (among other things) say "no competing with us, you have to buy a license for that".
https://opensource.org/blog/metas-llama-license-is-still-not...
Ps: I work for meta, but not in AI related orgs.
I kid. I’m all for meta releasing more open weight models.
I mean… I’m also 1000% certain that China is going to undercut whatever they can do if even not a technology reason but a legal reason. I’ve seen some wild stuff posted that has been made with Minimax… things a US company could never allow to happen.
Such as?
You can generate a video of just about any thought in your head at all.
IP holders and politicians will not like that.
Granted, human artists have always been capable of this. It's just never been worth it to bring most ideas into fruition. Now there's minimal cost to do so.
Early Sora, Grok Imagine, and Seedance 2.0 had few filters. Disney, Nintendo, et al. eventually stopped them all from using their IP, which made the appeal fade for a lot of users.
Kling and Nano Banana can still generate IP oddly enough.
It arguably didn't really get leaked, and they had the .edu req mainly for fair use education exemption when legality of models was much more uncertain.
OTOH, to the extent that the original model trainers rely on training on certain data not requiring a license from the copyright holder, there is at least an argument that with an open source licenses for the weights and training and inference code, a transparent training corpus to which the original trainer has relied on no special permissions not granted to the general public to train on it, to the extent that the legal theory behind the original trainer believing that it is free to train on the data is correct, provides all of the essential features of open source.
At the same time, there are things portrayed as open weights where training data is undisclosed and the weights have a license which limits purpose of use and other aspects of use; the models are free-of-charge (for limited uses) but not meaningfully open.
For example, if I was being pedantic:
> nobody seems to care about using the right words in only this context.
"nobody" would include you.
And then I might complain about we use the word "weight" for something massless, or how "bugs me" is *ento*mologically incorrect: https://xkcd.com/1012/
This is of course not a good use of time. I wonder if illustrating the point about how language is dynamic and meanings are descriptive not proscriptive, was a good use?
"Communicating": https://xkcd.com/1860/
Within tech circles, open weight != open source. Outside them, they’re synonyms.
The correct battle would have been weights + regime. But extremists insisted on data, too, which left Meta as the only other real voice arguing with anything practical. They had open weights. I think eventually open use was negotiated and that closed the case except for the folks still arguing about how to pronounce GIF.
They also kick off a ton of other nefarious things we are still paying for.
The internet is this vast, intellectual (in the academic, university, .edu sense), cypherpunk, government/activist... thing. It should be interesting. Instead the best we have for social networking is Mark "they trust me; dumb fucks" Zuckerberg and you getting banned from the site at any time for any reason
There's also Mark "company over country" Zuckerberg.
Google is incentivised to collect a bunch of data (like FB) to improve their ad serving etc. Apple does not have that incentive as they don't make as much off advertising as hardware / app revenue.
NVIDIA is friendly with open source as they want to commoditise the model layer and take the gains in the hardware / data center layer.
I am trying to think what the incentive is here.
- Juicing demand for AI workloads for which they can lease out Meta data centers to establish themselves in the infrastructure as a service space
- Preventing extreme capital concentration at any single competitor
- Fostering a market of small innovative companies they can then acquire
Tbh React might be as unforgivable as the Myanmar genocide
When we critiqe these sorts of institutions it's important not to prescribe value judgements on them and examine the circumstances of their condition materially.
To be clear, I'm not saying being an employee of Meta is comparable to being a member of the aforementioned groups. But I don't think it is always a fallcy to apply a value judgement to an institution.
"Meta did...kick off the origin of the open source race back in 2023"
ignores the majority of open source software history[1].
[1] https://en.wikipedia.org/wiki/History_of_free_and_open-sourc...
This is why bad people keep inheriting the earth. We keep forgiving them and they keep doing their crap.
I will make an exception for Musk and DODGE
Even if he is openly courting a Bond villain image and talking up his "robot army".
Essential employees were fired before the government had even established that it could safely do without them, and then...
After brandishing the chainsaw of efficiency in public, in the most cowardly way, went and then invoked all legal protections from being deposed on DODGE actions, personally and not answering under oath about the key decisions in that dismantling.
The model was "accidentally" released.
Meta was giving it to approved researchers only until someone leaked a torrent. Whether that was a researcher, an insider, or Meta's plan all along, we don't know.
Meta has withheld its best models, as have a lot of other "open weights" Chinese companies. When an "open weights" company gets ahead in one domain or modality, they tend to start withholding their releases. Tencent, for instance, began withholding their Hunyuan models once they became competitive. Alibaba has done the same.
The "open weights" strategy for the majority of players is this: open source when you're not in first place. Use the ecosystem to poison your rival's margins and play catch up on distribution.
In the West, it tends to take on yet another hook: "shareware weights until you pass $1M ARR, then you must license." See Flux, K2, etc.
The only way for open weights to make sense financially is if you have another income stream and are dumping on the market to destroy competition and/or can get people into using your inference infra / product ecosystem / tooling. Nobody's cracked this yet.
Let me be clear: with electric cars and solar, elon musk has done more net good for the world then almost everyone on HN. Is elon musk perfect? No. Far from it. It's within the imperfectness that people find reasons to attack him.
> It's the most ludicrous statement ever.
That is also polarized to the extreme, to the point that it's ludicrous.
Oh, please. Feel free to explain.
Facebook, Meta. 15+ years of emotional, child and human exploitation. Perverted glassware that spies on folk, lobbyists for age verification and who knows what else. They release an open model and all is fine and dandy? Nah.
Please get your priorities straight.
What do you think this Open LLM model is doing if not processing data from their murky sources?
“Enough” for what? Surely not enough to offset all the bad they’ve inflicted and allowed on the world.
https://en.wikipedia.org/wiki/Criticism_of_Facebook
> I think it's worth giving them some reasonable doubt.
Zuckerberg has shown through repeated action that he does not deserve any benefit of the doubt. This is the guy who called people “dumb fucks” for trusting him.
This is not even a case of “fool me once” anymore. If you continue to believe Zuckerberg, you’ve been fooled dozens, hundreds of times, and shame is definitely on you.