upvote
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
reply
How do you cleanse the data at this scale?
reply
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.

We have a lot of details in the tech report if you want to go deeper.

reply
Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.

I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!

reply
Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?
reply
How exactly would an open project do that?
reply
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
reply
to be fair, this discussion has been had numerous times here and the industry has arrived on "open-weight" to describe the practice of releasing the post-training weights in an open manner but not releasing the data it was trained on.

That's what they call this and I think it's a pretty clear definition these days to people in the industry.

reply
How, is the training data public?
reply
I wish universities would take it upon themselves to curate the training sets for these models.
reply
How about the private companies already doing it make it public because they built it on our data (without our permission)?
reply
The paper mentioned is here: https://tej.as/blog/aleph-alpha-kolibri.

(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we're merging the threads.)

reply
This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.
reply
Yeah the pdf alone is awesome as a learning tool.
reply
Such a crazy change from the times of Luminous, when they published a three pager with a claim that the model is similar good as „gpt 3“ (which??) with some graphs without y axis.

Bravo team!

reply
Hopefully this becomes the new standard.

It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.

reply
Be cautious what you wish for. Tools don't tell you what to do with them.

Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.

Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.

Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.

reply
Of course openness is only worse than leaving everything under the control of a select cabal of you believe that cabal to be more ethical than the rest of us.
reply
The select cabal who believe in an eschaton they're actively trying to bring about, as well.
reply
Would you apply this reasoning to the proliferation of nuclear weapons?

edit: why is the parent rationale sensible for AI and not nuclear technology?

reply
Absurd and incomparable.
reply
Why? That seems like a shallow dismissal. Why is a world changing technology okay in the hands of a small cabal in the one case and not another?
reply
"Why is the wheel not okay to gatekeep but the nuclear bomb is?" Even the framing is manipulative from the beginning. By asking this question you're already assuming they're in any way comparable. They are not even remotely equivalent and the entire comparison is utterly absurd.
reply
AI is more like a nuke than a wheel. Do you disagree? Can you suggest a less manipulative framing? I am personally a proponent of open source AI and models but I found this cabal framing strange when we do indeed rely on this kind of control for other world-altering technologies. And AI is different in that it enables technological development in ways quite unlike the wheel in a general sense.
reply
Dude ChatGPT is not the equivalent of a device that can level a major city killing millions in a flash. It is self evident. This entire discussion is ridiculous.

Nuclear weapons are a wholly unique threat to mankind.

reply
The main thing that bugs me about the risks discussion around AI is the lack of specificity. Commenters here have a good grasp of the risks around finding vulnerabilities faster than they can be patched. That's good and it matches the applicability of LLMs to coding.

But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.

As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of LLMs. What is going to be novel and calls for our spiritual development in other domains?

reply
I find it odd that more people don't realize that we are talking about risk of elevated *general intellectual-domain capabilities* as a resource. To be clear, I'm not claiming that LLM + RF is necessarily THE technology that poses the risk, the risk is in recursive self-improvement and whatever technologies will result.

It's very clear to that any specifics couldn't capture the risks, because the capabilities, including the risks, are one level higher than any specific techonolgy. It is the process of advancing technology itself, in accelerating speed, that poses the risk.

reply
The reason I use coding and vulnerabilities as an example is that it is a concrete example. It is what people pay for now when they buy AI. And the risks are specific and can be examined in detail. Some threads on this board currently show that even these more concrete and specific risks are often overblown, with LLMs finding low risk bugs and sucking up resources to evaluate and fix them.

Here you are claiming that AI products are going to reach AGI or RSI in the foreseeable future. Of course you can't "capture the risks" with specifics because those are inherently unspecific futures. It's a bit like saying when we invent antigravity all hell will break loose.

I would believe those future risks more if there were a progression of risks. What other than finding vulns has those characteristics?

reply
What poses the risk is the combination of abilities past a certain point enabling you to do basically anything.

While being unable to judge whether you should in the first place.

reply
ai cant move atoms at unlimited rate and also have limited energy. so your claim that "basically anything" is a bit of a stretch.
reply
You're right, people weirdly lack imagination on what "higher intelligence" (minus ethics) actually affords you, let's have a look:

What do average people currently want? They're taught, the most important thing was being rich. So they will ask their AI to make them rich. Most real life ways to get there are "sketchy" to say the least, usually downright unethical and anti-social, but US society turns a blind eye when the "Wolf of Wall Street" comes out on top and the schemes don't easily fit into average people's abilities of moral judgement.

-> Large parts of US society suddenly engaging in all kinds of "semi-legal/hyper-illegal" fraud schemes, at the expense of already saturated environmental and societal resilience. Guaranteed collapse.

Or, let's get rid of those pesky neighbors/wrong-colored people/annoying opinions? Again, "legal" is a pretty squishy concept and only really applies when you don't have the legal expertise to get around it. Now you can.

Or, look at the basics: what is "real"? You only "know" because you trust certain people and institutions. Generative AI can help with that /s.

It's not only about "building weapons of mass destruction". It's about doing the same shit as usual, but a thousand times faster/amplified. Look up poly-/metacrisis for starters. Going faster with AI when there's a wall in front of you isn't the best idea.

reply
This is analogous to the problem of spam, which is a problem about five minutes younger than email. Before spam you had to buy ads in the back pages of magazines you think target vulnerable demographics. Meta already spews fraudulent ads in horrific volume.

In other words, ambitious frauds have already explored all of the angles and bought all the ads. At worst, LLMs will create a few more successful but less ingenious frauds.

reply
Whenever classic p(doom) sentiments are explained through a text with obvious LLM markers, I wonder whether I am looking at a superhuman persuasion attempt.

..or maybe the commenter did look at superhuman persuasion long enough to believe it would be best to channel those ever the same fear fantasies from the LLM through their account to the reader.

On a more serious note, just look at the doom premises here: "Large parts of US society suddenly going criminal" is from the movie "The Purge", I think. It is fiction.

The idea that generative AI takes away our ability to find out reality. ... I don't know. People write about that a lot, but it still seems very far fetched.

Maybe through some terminally online overconsumption, like with social media? I wouldn't know.

With new AI capabilities we will have to adjust, I am sure. Media, science, education and law are changing very visibly right now. Those p(doom) narrations just seem to be pre-IPO hype though.

It is just so so dangerous. That is why they want to go public and only want to care about optimizing for the next quarter ...right before breaking into AGI. /s

reply
Ah, didn't think of this. Some rogue militia group might try to use this LLM (or create their own LLM based on this work) to help them create biological weapons or to do a mass hacking the infrastructure of targeted country.

Wonder what safeguards Kolibri uses to prevent this? Or if they even can

reply
I don't think that AI is as much of a boost to bioweapons as people think. Lab work doesn't get easier just because the experiment design part does.
reply
Chemical and biological warfare is hard. A cult in Japan created a mass casualty event using nerve gas. Which is the only somewhat "successful" terrorist WMD attack I know of. Knowledge of how to create these weapons isn't new. Guns and bombs are the most widely used terror weapons for a reason.
reply
Enter LLMs to help through all those pesky hard parts.
reply
The hard part is probably finding the lab equipment and chemicals without being noticed, and choosing not to use it to manufacture drugs instead which would be much more profitable
reply
Intelligence was never the bottleneck tho
reply
For terrorists it is tho. Four lions is basically a documentary. The only clever and creative terror attack happened 25 years ago.
reply
Didn't Mexican cartels kidnap telco workers to build them separate infra? We expect terrorists to be less resourceful?
reply
What's stopping them from using claude today? What could anyone possibly do to stop them from using open models? This line of thought seems like corporate/political/pr pandering more than a meaningful concern.
reply
Oh definitely, and I’m in the cybersecurity space so I’m already on the “worst case scenario committee” hah.

But the alternative just seems… so much worse to me?

A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.

And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.

reply
Yes, a "select" (rather, self-selected) elite "controlling" AI according to their wishes, what could go wrong?

Clearly not a "better" scenario. The real problem though seems people feigning helplessness? You can't leave society "to its own". You are part of it and go where it goes. So better start steering.

When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that. Just like you shouldn't be allowed to drive a car or fly a plane or command a rocket without proper guardrails, safeguards, prerequisites, etc.

"General" intelligence isn't present in humans, why does it need to be in AI?

reply
> When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that.

What do you mean "cannot"? As in you are granted abilities that have no responsible use?

reply
> When access to AI gives you abilities you cannot use responsibly

I don't think this is realistically a problem at all. It just makes certain types of research cheaper and less time-consuming. And, again, this is also a problem with the american services.

reply
I think the unethical things are happening already with boutique firms. I admit I don't get the concern.
reply
> Humanity has the serious problem of being underdeveloped in the "spiritual" department.

As is shown to us by filthy rich people every day.

Or did you mean the burglar in the fawellas?

reply
His point about regulation and innovation is great and I wish more people thought like that.

One of humanity’s biggest problems here is we don’t know how to do moderation.

We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.

The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.

I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.

reply
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we've since merged the threads, so I've moved it into the subthread which is specifically about the paper being responded to.)
reply