upvote
The ethicists who see their job as publishing papers and acting as philosophy professors, sure, but someone has to own the training and eval suite for how AI thinks about trolley problems.
reply
Once the AI is working, surely the corporation can just ask it to come up with an inexpensive accountability sink for trolley problems? Perhaps it will suggest hiring an impeccably credntialed AI ethicist to legally opine that it has the most ethical trolley problem solutions?
reply
Why? Are trolley problems making it hard to make money?
reply
Because leaving that to emergent behavior is silly. Why is the idea of training models on ethical reasoning so strange?
reply
Because it's not going to make the investors richer, which is what most alignment work focuses on -- how do you make its output palatable enough to avoid embarrassing customers. The entire focus of the AI industry is pumping up the dollars coming into the company, and if you're not contributing to the bottom line, your neck is on the line.

You haven't explained why investors make more money by spending money on people to solve trolley problems.

reply
Because 1) ethics aren't universal. It's perfectly easy to frame ethics for the outcome you want depending on your a priori assumptions. 2) it's blatantly obvious that the business "needs" will supersede any ethical concerns or the ethics will be framed in a way that aligns with the primary goal of gaining money and power.
reply
I don't see either of those as reasons not to train on it. Training isn't zero sum. You don't need to take something out to train ethics, and whether you explicitly train it or not, models will take ethical stances. Your second point actually argues in favor of training because that is still an ethical stance.

I think a lot of people confuse ethics training with "don't be evil"

reply
When you put it like that, it's actually probably worse to train it on a subset of ethics curated by the people doing the training, aligning with their world view and interests. If the model learned ethics just from raw training it would at least be statistically in line with it's training corpus.

Manually training on "approved" ethical stances sounds a lot like censorship.

reply
I mean, sure, its censorship to the extent that you can apply the term to an institution that is also the creator of the censored work, and that's exactly the point and the whole meaning of “alignment”.
reply
I would eat my shorts if we could get alignment in a HN thread on what the top 25 "don't be evil" required ethical items for a AI lab would even look like.

The first 10 maybe easy, but we'd see so much disagreement even without money nothing would get done.

reply
They were. Like, right before the big public-use-of-AI wave.

Those left are the ones not seen as troublesome then (or those who have entered the field under them.)

reply