upvote
I'm all for more open models, but talk is cheap and this is a rather pointless announcement without anything backing it up. Publish your weights and HF repo or shut up IMO.
reply
And also a "proprietary data set" hahaha... Probably just means they don't want to show it, and it is data, that either they shouldn't have, or that there is nothing special about their training data and it is just meant to sound like there is some secret ingredient, while there is none.
reply
Not sharing the data is pretty standard because 1) it tends to get the lawyers involved and 2) good data is critical for getting good results.

Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.

This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.

reply
Data has copyright issues, so one can't share it generally without getting permissions from all of the copyright holders. The data is not theirs to share, anyways. The derived (learned) weights are a different matter.
reply
this is very normal for frontier lab companies. you need good data either synthetic or labelled (all the chinese open source models have their own armies of data labelers)
reply
> Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model.

> We will release the weights, technical report, model card, and developer artifacts later this month.

reply