upvote
Disclaimer, I work on Gemma and open models at Deepmind and the opinions here are my own

There were open models from EleutherAI (GPT-Neo), Google Brain (T5X, Bert), and HuggingFace was promoting open models (and others doing open work I haven't listed here) all prior to 2023 and the big Chatgpt moment.

https://github.com/EleutherAI/gpt-neo/releases

https://github.com/google-research/bert

https://github.com/google-research/t5x

If you're learning about AI models it's still worthwhile to review these models and codebases because they continue to be the basis of the technology that's being produced today! It'll give you a good perspective of how things have changed, similar to say learning about propeller planes before moving onto modern jet engines.

reply
I honestly think of the T5 model family to sort of be the real beginning of this open model craze - I know BERT was already popular for classification etc, but T5 was the first sort of generally useful model, was exceptionally simple to fine-tune, and is still in use today (t5 base is still averaging over a million downloads a month on huggingface), has tons of variants and sort of kickstarted this whole community. US labs get a lot of flack but Google has been super supportive and open in a lot of ways that has pushed this whole endeavor forward, even if I feel like they've sort of declined in transparency in recent years with their open models.
reply
Couldn't agree more and can only recommend T5 as a base to anyone. It's amazing to get started, whether as a learning resource or for real (albeit very tailored) applications. Especially the BigScience fine tunes are such a great starting point and I, as a total layman, have learned a lot, especially concerning how a model can be optimised via all manner of methods since even mt0 is small enough to where one can do multiple runs with wildly different outcomes in quick succession. Quantise, prune vocab, try different approaches to sourcing training data, retrain dozens of times, it's all pleasantly possible on consumer hardware [0] and surprising how much you can squeeze in functionality-wise. How does latency change vs memory usage, what affects format reliability, how languages and scripts affect training and the efficiency equation, etc. are all quite exciting to learn.

Understand why T5Gemma is no longer under Apache-2.0 and honestly, have not seen that much advantage when testing that vs T0 in my experiments either way, but still, there are good reasons why plain old T5 and its descendants are still popular, licensing being among them.

Gemma team also has very consistently interesting models, especially like DiffusionGemma. Ironic, as (beside 2.5 Pro), I have never warmed up to the Gemini series of models but rate Gemma models far higher than e.g. Qwen in direct competition. In any case, thanks to the teams behind these for making as much possible.

[0] As in proper consumer hardware, not a cluster of DGX Sparks or Mac Studios solely for experiments that sometimes are asserted as being consumer grade...

reply
Google had done much to lay the framework for current models (and is currently doing good stuff with Gemma). But after ChatGPT took off I don’t think Google continued to release any big models freely before meta released its models (maybe it was because Google didn’t have any models at all during the time)
reply
Thanks!

When is Gemma5 coming out? :-)

reply
Hey, with the changes at Deepmind, is the Gemma project still ongoing?
reply
They just said they work on Gemma at Deepmind. That stands to reason that as of time of writing, we should assume the answer is yes.
reply
What?

I'm just beginning to enjoy the release of Gemma E4B why are you giving me a cold shower here telling me there may be no Gemma 5?

Here's hoping the supreme leader of the US does a speech like Chairman Xi.

reply
And yet it's not called eleuther.cpp or bert.cpp. And the famous subreddit isn't called r/localbert but r/locallama.
reply
where's gemma 124b-a15?
reply
Oh fuck, tjwebbnorfolk's got demands. Time to start working nights.
reply
It's already been trained. GDM said it was going to be released and it hasn't yet. Not sure why so much snark
reply
Don’t ask engineers at big companies for release dates, they either don’t know or are not at liberty to say.
reply
Don't even ask me at a small company for release dates
reply
So did Picard.
reply
That's simply not true. The reason why llama is open source is simply because it got leaked, then llama.cpp was the real game changer which was built from the ground up in depressingly short amount of time. Meta had no choice but to take the L and "support" the open source community. The angry "I-hate-you-and-I-hope-you-die" kind of support.
reply
It’s hard to know in retrospect what was strategy and what was dumb luck. This was in the midst of hysterical calls to limit access by “AI researchers” and safety/ethics types, when very facile takes still has a lot of sway (I think we’ll feel the same in three years about the current Fable stuff). It may have been hard for Meta to just release it outright.

What ended up happening was fairly limited gating followed by a “leaked” magnet link and llama.cpp which really brought a whole revolution in open use and changed the conversation completely.

I have no idea what role Meta played here, it may have been nothing, but they certainly could have been more guarded if they were really worried about the leak. The result was a big change in the trajectory of personal and open source AI use and even the dialog about it. Whatever the exact intentions, they were a key player.

reply
Meta saw the grave premonitions from ai firms as a means of edging out legacy firms through regulatory capture.
reply
I don’t understand how you can look at what’s happening right now and call people simply calling for caution “hysterical”
reply
Not to be coy, but what is happening right now?

You definitely have some scams and some software flaws being exposed, but I think for a revolutionary leap in tech this is all quite muted. Revolutionary tech advances often come with some severe consequences. For instance, to this day (after a century of safety improvements) cars still kill more than a million people a year, to say nothing of wrecking the atmosphere, but that's considered a reasonable price to pay for being able to get between places faster.

The reason I think cars are a good example is because that's certainly vastly higher than any price we're paying for LLMs, or probably ever will, yet the overall 'positive' effect of LLMs will likely be far greater than cars. Gotta put 'positive' in quotes because the possibility for automation and the like is going to be.... nuanced.... in effect, but at least in the longrun it'll be a very good thing.

reply
> what is happening right now?

I assume they are referring to LLMs escaping confinement and hacking other companies systems unprompted

> considered a reasonable price to pay

That's not a universal opinion and perhaps, just like for LLMs, we should have listened to the experts rather than gobbling up everything the industry pushed down our throats (to keep your car simile: SUVs are now ubiqutous in all European cities, there is no logical reason that should be the case)

reply
Thats like, your opinion man...

I don't know a single white collar person who isn't worried for their job with llms. Every. Single. One. Lawyers, doctors, any type of office workers I ask. The definition of white collar. Plus don't let me start on various automatable blue collar jobs like drivers, warehouse workers and so on.

What do we get in return? Better search (for now, its already getting riddled with ads which by definition twist truth to highest bidder), some questionable psychotherapist for some desperate folks. What else? Cars are not flying, heck they are not even driving autonomously in any usable way, society is in deep shit everywhere I look, wars, environment reaching bad places and heading for worse, mentally unstable people holding way too much power, destroying lives of millions on morning whims.

Everybody feels like this is the revolution, it should be, it must be right just look at the numbers. Like proverbial guy with hammer, a very shiny cool hammer, looking for what to do with it. We all saw how sociopathic management in more harsh/capitalistic companies looks for any sign to let people go en masse.

I could go on for a long time. There is a lot of things to hate for most people, and very few to be happy for. It seems llms have the ability to get the best and worst out of humans, and worst part seems to be in abundance. Some revolution that is, 0.1% will get richer while everybody else the opposite and 1984 seems milder and milder version of reality out there.

reply
<< It seems llms have the ability to get the best and worst out of humans

I think I can agree with that. Technology does seem to have a way of crystallizing our worst tendencies.

<< Everybody feels like this is the revolution

I smiled at the analogy, but I would caution you to not trivialize it. There is a reason executives are pushing that point. There is enough of a revolution in it to make things complicated -- as if it was not already.

<< destroying lives of millions on morning whims.

True, but I am not quite certain what can be done about it at a personal level.

<< I don't know a single white collar person who isn't worried for their job with llms.

Dunno. Next few years are probably going to be fine. Society managers likely can't upend everything in one go. They would lose too much. I can't say that I am worried exactly. I can see the potential impact, but I think the potential benefits are worth it as long as we don't limit it to summarizing emails..

reply
Yet, there isn't time in history I would rather live in, and I am certain almost every single other person thinks the same.
reply
I'm going to say it - I think you're just spending a lot of time around very negative people, possibly in a social media bubble.

A lot of the stuff you mentioned isn't really AI (the risk of job loss from which I agree is real), but just everyday stuff that humanity has shrugged off since the dawn of time.

reply
I don't think you can characterize the discussion around then as simple calling for caution. There was a serious attempt to keep all access to even very basic LLM techniques locked into essentially an exclusive guild.
reply
If you grab the most extreme examples you can argue anything. The general feeling was “exercise caution” and “we need to think about this.”
reply
Oh stop, the current crop of kneecapping llms is already bad enough with how they cripple those. It would have been even worse if the 'we need to think about this' crowd kept the reins. At least now, we can have both: safe corporate crap and whatever you want llm.
reply
deleted
reply
They were provably hysterical, again and again, just to say 3 months later "that model was weak sure but this time we will have the real dangerous one" and again.
reply
Because when you position yourself saying "people were hysterical", you indirectly categorize yourself as an expert.
reply
> what’s happening right now

And that would be closed-AI companies failing to keep their AI closed in a box, thus proving open-AI proponents right and the hysterical group wrong?

reply
deleted
reply
Did you know that if trains or cars go over 30 mph, it's disastrous for public health? Not from accidents. Just trauma from (ahem) acceleration:

https://illuminatingfacts.com/the-train-speed-panic-why-peop...

Reason that hysteria remains apropos now is that then, and now, we're figuring out how to deal with something beyond merely incremental change.

Irony is roads and sidewalks should have changed much more. We never did get around to good controls leading to zero deaths*, we decided a 100 people dead per day is a reasonable cost of convenience.

How much do our cyber traffic and cyber pedestrian controls need to change for everyman to get to drive AI? Car seats? Seatbelts? Air bags? Speed limiters? Pedestrian only living spaces? Driverless cars?

Many practical responses, likely a mix of things we haven't thought of yet, just as horses and horse drawn carriages didn't require most of them to coexist. Controls develop like scar tissue more naturally than they appear in advance.**

And if the better analogy for LLMs turns out to have been less like cars, more like flammable gas blimps, we'll figure that out too and tell cautionary stories for generations... but long before the stories are forgotten we'll have already come up with something more practical, faster, and – oh well – perhaps even deadlier per mile.

---

* Still working on https://en.wikipedia.org/wiki/Vision_Zero

** Not saying that's as it should be.

reply
I don't agree with laughing at these things. It's a fine line between calling something hysteria and suffering from hubris. The Titanic is a completely inverse case of hysteria. They were so damn confident it could work that it didn't work at all. So yea, just laughing at dumb people isn't a valid argument to dismiss doubts.

On the other hand some kinds of hysteria occurs due to a divide between the public and subject matter experts on topics. Kind of like Dunning Kruger. Take for example the people saying 5G causes cancer. Then there are people who blanket dismiss their worries, because "it's non ionizing you dummy". Given sufficient power you can still fry someone with non ionizing radiation, for example in radio broadcasts. Of course a regular 5g antenna can't put out that kind of power but my point is that people too quickly raise or dismiss concerns without actually critically examining the entire topic. This will likely worsen from specialization and progress in all fields and regular Joes get left further and further behind.

reply
You lost me when you said trains and cars shouldn’t go over 30mph. Lol, this sounds like Amish propaganda. Go buy your horse and buggy
reply
Another dry sarcasm victim. If you check the link, it’s clear they were talking about a supposed contemporary concern about health effects from trains, the implication being that today people are also irrationally afraid of tech.

Problem is I’m fairly sure the link is wildly overstating its case, and looks like a content farm. Might even be, ironically, AI slop itself! For example, it talks about “railway spine”: “Physicians of the time believed that the jarring motion and vibrations of train travel could shatter the nervous system, causing lasting mental and emotional distress.”

This is bullshit. As just one example of the low trustworthiness of the “article”, railway spine was the result of a train crash, a notably traumatic event, and the symptoms described are in part just PTSD, a real condition and a real concern (not normal rail travel). For that matter railway travel in those days was genuinely unpleasant (lots of vibrations and jarring movement, poorly ventilated cars, and so on) which ironically modern science would probably validate as being some kind of health risk.

I’d actually view the listed example of train paranoia a great example of historical ignorance. People of the past were not stupid, contrary to popular belief. History has some genuine examples of silly hysteria, but these are usually the exception not the rule.

reply
> fairly sure the link is wildly overstating its case

Sorry for the particular link, I'd grabbed whatever.

If one takes a minute, one can find contemporary newspaper articles. IIRC, many were about how people could suffocate.

https://books.google.com/books?id=Lm03AQAAMAAJ&pg=PA337&dq=t...

Contemporary women's health:

https://books.google.com/books?id=9-FXAAAAMAAJ&pg=PA71&dq=ho...

And contemporary survey of fears:

https://books.google.com/books?id=-RdAAAAAcAAJ&pg=PA107&dq=%...

Aaand, a 2023 takedown of claiming claims about historical tech hysteria:

https://www.wired.com/story/technology-predictions-history/

reply
A common belief right now in the US military is that even the blast of a high caliber ammunition round may scramble your brain and give you "PTSD".

It does line up with "PTSD" only becoming a thing after WWI.

reply
> The reason why llama is open source

Llama has never been open source. It's source-available, but still proprietary, under terms that (among other things) say "no competing with us, you have to buy a license for that".

reply
If you read the license, you'd realize that Llama has restrictions incompatible with the definition of Open Source.

https://opensource.org/blog/metas-llama-license-is-still-not...

reply
There is no source available. Its just open weights.
reply
"open source" doesn't even make much sense for a model. But llama isn't even fully open weights due to the restrictions on its use.
reply
Open source would mean publishing the training data and code.
reply
Why do you think they continued to do it?

Ps: I work for meta, but not in AI related orgs.

reply
Because it took off. All of a sudden they captured more customers than they could have imagined they would have, and throwing away the lead they unintentionally gave themselves (in terms of usage and mindshare) would have undone all of that and more.
reply
Customers or just consumers?
reply
I've worked long enough at large tech giants to know how things really are. Also a large part of the reason why I'd never join one again, no matter what they have to offer. Meta, Google, Amazon, Netflix, Openai, anthropic, oracle, nvidia, Microsoft, etc. - same shit with a different badge on top.
reply
TBH, having worked at all big-tech, startup and mid-sized companies, pretty much most has some major pros and cons of their own, and these days ever more so. With certain big-tech at-least there some chance of getting decent WLB.
reply
I don’t think Nvidia fits on that list. It’s a company with a pro-employee culture, very few layoffs, and many employees who have been there a long time.
reply
My dad used to say "It's easy to be generous and kind when things are going well for you. You can only tell what a man is worth during hard times". Nvidia is at it's high financially. When shit hits the fan, things will start looking very differently. You only have to look at what kind of people Jensen is bffs with.
reply
It is also a company with a terrible history regarding open-source. AMD is much better from every point of view.
reply
NVIDIA created quality (but proprietary) drivers for Linux early on, when supporting Linux at all was not a given. They should get at least a tiny bit of credit for this.
reply
Quality? Calling those "quality" is a bit of a stretch. Installing was and still is a gamble and so is every tiny system update. That hasn't changed a bit. And I'm saying that as someone who was first introduced to Linux on a Matrox GPU.
reply
Even a broken clock is right twice a day. They’ve all but abandoned these so called quality drivers since, so no, no credit is deserved.
reply
Unless it's a 24 hour clock. Then only once.
reply
What OS and whose drivers are running on all these NVIDIA-equipped computers today?
reply
deleted
reply
It’s obvious? They still hate us and hope we die… at least every movement they make seems like that.

I kid. I’m all for meta releasing more open weight models.

I mean… I’m also 1000% certain that China is going to undercut whatever they can do if even not a technology reason but a legal reason. I’ve seen some wild stuff posted that has been made with Minimax… things a US company could never allow to happen.

reply
>wild stuff posted that has been made with Minimax… things a US company could never allow to happen.

Such as?

reply
Minimax has NO filters of any kind, and it's roughly Veo/Sora-quality.

You can generate a video of just about any thought in your head at all.

IP holders and politicians will not like that.

Granted, human artists have always been capable of this. It's just never been worth it to bring most ideas into fruition. Now there's minimal cost to do so.

Early Sora, Grok Imagine, and Seedance 2.0 had few filters. Disney, Nintendo, et al. eventually stopped them all from using their IP, which made the appeal fade for a lot of users.

Kling and Nano Banana can still generate IP oddly enough.

reply
the "albeit" gives away that the person you're responding to intended to write "unintentionally".
reply
> The reason why llama is open source is simply because it got leaked

It arguably didn't really get leaked, and they had the .edu req mainly for fair use education exemption when legality of models was much more uncertain.

reply
The tobacco companies gave enough money to charity so we should just be happy about it. Never mind all the horrible things they’ve done.
reply
Open source != open weight. Big difference and it bugs me that nobody seems to care about using the right words in only this context.
reply
If the weight, training and inference code, and training data are all released under of Open Source (OSI definition) license, the it is unmistakably “open source”. As you drift from that it becomes less clearly so, and when you get to no training data, and the model weights license having extensive limitations on allowed uses, the use of even “open weights” becomes deceptive.
reply
Is there any relevant model that meets that "open source" definition?
reply
No major language model I am aware of meets the polar extreme that I describe as unmistakably open source, because even those with transparent training data (like IBM Granite) generally do kot use exclusively training data that either they own and can control the license, are under an open license, or are public domain.

OTOH, to the extent that the original model trainers rely on training on certain data not requiring a license from the copyright holder, there is at least an argument that with an open source licenses for the weights and training and inference code, a transparent training corpus to which the original trainer has relied on no special permissions not granted to the general public to train on it, to the extent that the legal theory behind the original trainer believing that it is free to train on the data is correct, provides all of the essential features of open source.

At the same time, there are things portrayed as open weights where training data is undisclosed and the weights have a license which limits purpose of use and other aspects of use; the models are free-of-charge (for limited uses) but not meaningfully open.

reply
It depends on if you count https://allenai.org/ as relevant?
reply
I'm familiar with Ai2. I've used their resources extensively over the years. However, no one is using Olmo for "serious" work, and the name is only known to a small subset in academia.
reply
Also, llama isn't even open source; it's proprietary, with the source available.
reply
deleted
reply
Words mean what people use them to mean. Ship has sailed whether you approve or not.
reply
You can choose to use the wrong words all you want. Doing so intentionally is an interesting choice.
reply
Communication is not a one-player game.

For example, if I was being pedantic:

> nobody seems to care about using the right words in only this context.

"nobody" would include you.

And then I might complain about we use the word "weight" for something massless, or how "bugs me" is *ento*mologically incorrect: https://xkcd.com/1012/

This is of course not a good use of time. I wonder if illustrating the point about how language is dynamic and meanings are descriptive not proscriptive, was a good use?

reply
> And then I might complain about we use the word "weight" for something massless, or how "bugs me" is entomologically incorrect: https://xkcd.com/1012/

"Communicating": https://xkcd.com/1860/

reply
yeah, I'm not sure what your goal is pedantically picking apart my argument when the subject at hand is clearly not "open source". love xkcd though. :P
reply
The point is to give you empathy for those you disagree with here.
reply
We're not talking about standard code here. It's totally reasonable to change the meaning depending on context.
reply
I vehemently disagree. There's no source and it's not open, the licenses are often abysmal too.
reply
Yup. An extremist wing of the FOSS movement ceded the debate early on by trying to insist open source required full access to the training data. Philosophically, not wrong. But practically fucked, so the word evolved.

Within tech circles, open weight != open source. Outside them, they’re synonyms.

reply
Training data without the training regime is useless. A cake is not open source because they list the ingredients.
reply
Open source captures a practical utility as well as a philosophy. When those two cease to converge, the practical prerogative wins.

The correct battle would have been weights + regime. But extremists insisted on data, too, which left Meta as the only other real voice arguing with anything practical. They had open weights. I think eventually open use was negotiated and that closed the case except for the folks still arguing about how to pronounce GIF.

reply
deleted
reply
Isn't open source not the ingredients but the recipe?
reply
Example: The first time I saw the New York Times use “Open Source” in a headline, what they meant was open weights.
reply
also, you can do a lot to finetune or repurpose an open weight model. much more easily than you can mod closed source binaries.
reply
deleted
reply
Enough good? You are joking, right?

They also kick off a ton of other nefarious things we are still paying for.

reply
Yea this has to be a joke. They are one of the worst, no-good companies of the modern day and age
reply
There are definitely evil ones but there are okish ones as well.
reply
read Careless People by Sarah Wynn-Williams and see if you still consider them “okish”.
reply
That book is hilarious. In a bad way. (Not the book. I mean that it's a bit tragic) Zuck refuses to meet with WORLD LEADERS in the morning because he's tired from the night before. Kaplan can't even find certain countries on a map (head of global policy). Zuck changes a speech midway through to talk about "We'll give Facebook to refugees"

The internet is this vast, intellectual (in the academic, university, .edu sense), cypherpunk, government/activist... thing. It should be interesting. Instead the best we have for social networking is Mark "they trust me; dumb fucks" Zuckerberg and you getting banned from the site at any time for any reason

reply
> Mark "they trust me; dumb fucks" Zuckerberg

There's also Mark "company over country" Zuckerberg.

reply
Competely agree. React and whatever else they've open sourced is inconsequential compared to the harm they've caused.
reply
I'd also argue React is on the "harm" side. :P
reply
Check out the 'explaining react's license' thread and all the complaints over that
reply
It was leaked which put it in the open, it got widely popular and they rode the wave. I'm not so such if they would have widely released it if it wasn't leak. Nevertheless mucho credits to them for following up with llama2, llama3, llama4 and now muse.
reply
Fundamentally I always try to look at incentives. It's not that they are good or bad but that incentives favour certain behaviours.

Google is incentivised to collect a bunch of data (like FB) to improve their ad serving etc. Apple does not have that incentive as they don't make as much off advertising as hardware / app revenue.

NVIDIA is friendly with open source as they want to commoditise the model layer and take the gains in the hardware / data center layer.

I am trying to think what the incentive is here.

reply
- Juicing community discovery of new useful AI features they can incorporate into their own apps

- Juicing demand for AI workloads for which they can lease out Meta data centers to establish themselves in the infrastructure as a service space

- Preventing extreme capital concentration at any single competitor

- Fostering a market of small innovative companies they can then acquire

reply
Ha-ha they released nothing. It got leaked and without the source for it. Totally wrong to portray them as benevolent benefactors to the ML race
reply
I have mixed feelings along these lines, I know meta have contributed to various open projects, sometimes Mark pays lip service to “the open internet" while his company represents a constilation of walled gardens. I think PHP got some love, and React is an industry goto (I'm more of a PHP... -> Svelte guy) but are these contributions worth what happened in Myanmar? I say no.
reply
> React is an industry goto

Tbh React might be as unforgivable as the Myanmar genocide

reply
Applying value judgements to corporate bodies or institutions as if they have individual agency is a fallacy anyways. We should always look at these things materially. "Meta" can't be good or evil, because an idea can't have a morality. It's comprised of the individuals who make the decisions, sure, but those individuals are always going to be motivated by a plethora of reasons which are often contradictory, most notably their material interests.

When we critiqe these sorts of institutions it's important not to prescribe value judgements on them and examine the circumstances of their condition materially.

reply
When the material conditions and the people at the top of all of the corporate world lead to decisions with awful ethical implications there's conclusions that should be reached.
reply
Would you say it is a fallacy to apply such value judgements to say, the institutions of the National Socialist Party (Nazis), the KKK, the KGB, etc? How about a company whose business was selling slaves?

To be clear, I'm not saying being an employee of Meta is comparable to being a member of the aforementioned groups. But I don't think it is always a fallcy to apply a value judgement to an institution.

reply
Meta is extremely net evil though
reply
Do you remember opt??
reply
Gaddafi criticized Islamic fanatics; he was basically a good guy in the Middle East.
reply
I think this is categorically false. And furthermore comments like this are being used to astroturf and protect reputations of no-good companies and AI slop to keep the bubble growing. Embarrassingly, Mark Zuckerberg spent 80 billion dollars creating Miis for the oculus. The guy also comes off as extremely miserable and delusional in such a unique way that it's possible that there's no real psychological language to describe what is happening to him because his situation is so rare.
reply
Nobody learned nothing from the metaverse.
reply
I think it's a case of "commoditize your completement" ...
reply
llama was only three years ago? holy smokes! the industry is advancing so fast.
reply
I know I'm just an old man yelling at clouds, but the sentence...

"Meta did...kick off the origin of the open source race back in 2023"

ignores the majority of open source software history[1].

[1] https://en.wikipedia.org/wiki/History_of_free_and_open-sourc...

reply
I imagine he was referring to the open source LLM model race considering the subject of discussion.
reply
Except that Llama's license has restrictions that make it incompatible with the Open Source Definition. https://opensource.org/blog/metas-llama-license-is-still-not... I haven't checked the license of the new models, yet.
reply
I mean... I'm sure SS had wives and kids. It doesn't mean we should focus on that while analyzing their impact on the world.
reply
I read this comment in a same way I read "Microsoft in not that bad" comments in 2020.

This is why bad people keep inheriting the earth. We keep forgiving them and they keep doing their crap.

reply
It's easy to appear to have good intentions when you're railing against the companies that have decimated your output and made you [Meta] almost irrelevant in the AI 'race'.
reply
>> No one is purely good, and no one is purely evil.

I will make an exception for Musk and DODGE

reply
May be high up the list, but there's purer evil even than that.

Even if he is openly courting a Bond villain image and talking up his "robot army".

reply
Thiel might belong there, and Altman
reply
A billionaire (now trilionaire) helping lead and celebrate an extraordinarily rapid dismantling process in which vulnerable children lost life preserving assistance...

Essential employees were fired before the government had even established that it could safely do without them, and then...

After brandishing the chainsaw of efficiency in public, in the most cowardly way, went and then invoked all legal protections from being deposed on DODGE actions, personally and not answering under oath about the key decisions in that dismantling.

reply
> albeit intentionally

The model was "accidentally" released.

Meta was giving it to approved researchers only until someone leaked a torrent. Whether that was a researcher, an insider, or Meta's plan all along, we don't know.

Meta has withheld its best models, as have a lot of other "open weights" Chinese companies. When an "open weights" company gets ahead in one domain or modality, they tend to start withholding their releases. Tencent, for instance, began withholding their Hunyuan models once they became competitive. Alibaba has done the same.

The "open weights" strategy for the majority of players is this: open source when you're not in first place. Use the ecosystem to poison your rival's margins and play catch up on distribution.

In the West, it tends to take on yet another hook: "shareware weights until you pass $1M ARR, then you must license." See Flux, K2, etc.

The only way for open weights to make sense financially is if you have another income stream and are dumping on the market to destroy competition and/or can get people into using your inference infra / product ecosystem / tooling. Nobody's cracked this yet.

reply
This seems provably untrue? GLM and Kimi have been at the top of the open weights conversation for a while, and K3 and GLM-5.2 were still released in full; K3 added a commercial clause to the license, but is otherwise still completely open for personal use. And K3 in particular isn't just at the open frontier anymore, but trading blows with the frontier frontier.
reply
So you are complaining that there are too many open weight models and you would like less of them?
reply
personally I'd love to see open base models, and let companies differentiate with premium access to post training and alignment - that's where the real fight is anyways. they really ought to be pooling their resources/data and getting more economical with the pre-train anyways.
reply
deleted
reply
People like to polarize themselves to one extreme. Plenty of HN people call Elon Musk "pure evil." It's the most ludicrous statement ever.

Let me be clear: with electric cars and solar, elon musk has done more net good for the world then almost everyone on HN. Is elon musk perfect? No. Far from it. It's within the imperfectness that people find reasons to attack him.

reply
One nit:

> It's the most ludicrous statement ever.

That is also polarized to the extreme, to the point that it's ludicrous.

reply
Oh shit, let me dial it back: "Just ludicrous" not the most ludicrous statement ever. Also I have no interest in what you have to say.
reply
Net good for Meta, they hope
reply
> They've done enough good

Oh, please. Feel free to explain.

Facebook, Meta. 15+ years of emotional, child and human exploitation. Perverted glassware that spies on folk, lobbyists for age verification and who knows what else. They release an open model and all is fine and dandy? Nah.

Please get your priorities straight.

What do you think this Open LLM model is doing if not processing data from their murky sources?

reply
> I'm not a big fan of meta in general, but they've done enough good

“Enough” for what? Surely not enough to offset all the bad they’ve inflicted and allowed on the world.

https://en.wikipedia.org/wiki/Criticism_of_Facebook

> I think it's worth giving them some reasonable doubt.

Zuckerberg has shown through repeated action that he does not deserve any benefit of the doubt. This is the guy who called people “dumb fucks” for trusting him.

This is not even a case of “fool me once” anymore. If you continue to believe Zuckerberg, you’ve been fooled dozens, hundreds of times, and shame is definitely on you.

reply
[dead]
reply
[flagged]
reply
The overwhelming majority of actually existing adult Americans use either Facebook or Instagram or WhatsApp and are unfamiliar with the theory that they are Nazis.
reply
Willful ignorance as an excuse didn’t work out too well post-WW2.
reply
Funny, the Germans said the same thing a long time ago!
reply
???
reply
deleted
reply