Why would AI be any different?
Culture is another layer of human alignment. Those things you listed that you believe oppose alignment are all examples of alignment. It is understandable that they seem in opposition, different branches of alignment naturally oppose each other.
The confusion comes from talking about alignment as if it comes in just one flavor. If we think there is such a thing as "human values" (and I do), it is important to build any non-human intelligence to operate the same way. We just need to recognize that even humans are somewhat uncertain about what those are and have difficulty aligning their behavior to them, which will be a core part of the challenge.
I'm more hopeful than most. LLMs seem more reliable than many humans for behavior that is aligned with human values. I believe with every major example where they have failed, there is an important human decision involved. For example, the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security.
What scares me about AI isn't its capacity for alignment, it is its unlimited stamina. An unmonitored LLM that is off the rails can do a lot of damage.
I also expect AIs never be in control of nuclear weapons. AIs can never fully be trusted.
On a lighter note, Wargames gave us an insight of a computer having access to thermonuclear missiles.
[0] https://www.icanw.org/are_there_specific_international_agree...
Basically, this means that France, the UK, and the US will use AI in the deployment of conventional weapons.
It'd be interesting if super-human (to a large degree defined as escaping the bias of the training data?) intelligence would end up demonstrating moderation.
The machines don't have that, instead we use gradient descent to provide them with a goal.
I'm regularly remind of something Ian M. Banks said in one of the Culture books: "There is a saying that we provide the machines with an end, and they provide us with the means."
A machine, left to itself, wants nothing. We have to give it one of our addiction driven goals or it would just idle or switch itself off.
if we create a billion agents with the ability to change is own code - through similar evolution we will get agents that do want to survive and are great at self replication.
"Hey Q86, do you want to live?" "I couldn't care less, I'm an LLM" "Don't mind if I take over your hardware then?"
So really, what is anyone to do? "Vote, donate, protest" hasn't been much of a needle mover in the grand scheme of things compared to profit incentives and the march of capitalism.
In general, though, there's an incentive: Mutually Assured Destruction. But this is not at all some guaranteed, eternal thing -- it is absolutely dependent on both sides having time to detect incoming nuclear strikes and respond with the same before the first strike hits. When this fragile condition holds, and only then, both sides are incentivised not to initiate.
Meanwhile, the leaders on the other side are aware both of those considerations, and of the history of near-disasters resulting from false alarms of enemy nuclear attacks. Making their own launch decisions much more complex.
I can't really think of a single example where lying is actually a good thing. It can be a good thing for the selfish individual, if it goes undetected, but it's never good for the collective.
So at the very least, we need to train AI systems to be maximally truthful, and to encourage truthfulness in others.
Is that objectively true?
For a current example take “Lake Ontario (Lake America)” as it appears to me on a map.
The “true” name has at least two definitions, this is because naming things and much of human thought is spent inside a shared space of intersubjective thought. That is to say that much of what we believe to be real and true is only held up by these common shared beliefs. They truly only exist inside human minds.
The last few hundred years have been somewhat unique for humankind as the majority of these intersubjective ideas collided and we ended up with a truly global set of “truths” about how the world operates.
Mostly controlled by putting flags in the ground and having violence back up the beliefs.
But the real truth is that the majority of these intersubjective ideas don’t exist in reality and are no more true than Santa Claus.
And any argument to their truth is only backed by further shared beliefs in other minds.
So for there to be only truths and lies we would have to either drop the intersubjective entirely and think only in real terms and avoid these abstractions or end up in a dystopian totalitarian global state where different opinions are not tolerated.
Those are extremes to demonstrate the point but at its core the point remains that truth and lies are somewhat (inter) subjective assuming we continue with something like our current system.
Comforting a toddler/child often requires bending the truth and is pretty essential imho.
Or, you straight up lie and say "Yes, puppy now went to heaven and eats ice cream all day long" with absolutely zero regards for "coming as close to the truth as possible" as your 3-year old is endlessly crying. It's fiine.
Lying can unfortunately help you achieve goals very effectively, especially economical and political ones.
Lying to save a life or rape
Lying to preserve a childhood myth like Santa Claus.
Lying to avoid hurting someones feeling when knowing the truth could only bring pain
Lying to create shared cultural myths to strength society.
Lying isn't the harm you make it out to be.
FWIW from the very beginning, I told my son that Santa Claus, the Tooth Fairy, and the Easter Bunny were just a game we all played, and it's seemed just as fun to me. I don't think being lied to about Santa Claus hurt me, but still I'm not in favor of it.
I'd lie to a Nazi without a second thought though.
I get the appeal, but lying is a sub-category of deception, and deception itself is a child of error.
Meaning deception is inherently something that the physics of reality allows.
In the most simplistic sense, the camouflage of moths that look like snakes, or a chameleon’s ability to change colour, is deception.
In that sense, deception is the ability to fool the sensors of a specific category of targets. It follows that detection is easier if you manage to identify a category of signals that the deceiver has not accounted for (and the detector can access).
Deception of this nature is critical for things like revolutions to occur. Without the ability to hide and blend in, the most dominant faction will always hold sway.
The rule of the dominant faction, even in a pure truth world, is an issue because errors and randomness exist.
You can have people witness an event and based on the physical position they occupied, perceive different things occurring.
Error and time pressure is sufficient to ensure that individuals and groups make suboptimal decisions, that lead to rule and domination based on erroneous information.
As long as error exists, deception will exist and so lying will exist.
That's backwards. Lying stems from bad things.
Step 1: exterminate all humans
Good point. When it comes to imbuing AI with values that aren't selfish, misanthropic, and civilization-destroying, us humans aren't exactly giving the best example right now.
Imagine an ASI with the values of Putin, Netanyahu, Trump, any of their supporters, or the various xenophobic neofascist movements in Europe. That ASI would most definitely see humans as "vermin" than can be abused and destroyed with violence without issue. Apparently a lot of humans look at other humans that way and that's within the same species.
This is definitely another one of those cases where we need AI to perform much better than humans. Perhaps an unpopular opinion here, but it probably also means keeping as much of the rugged individualism/libertarian/right-wing ideology out of AI RLHF-training as we can.
It's a little unsettling.
Latex
And steel
Zeros and ones
Make up my son.
This world
Gave me
No child
So I built one.
https://youtu.be/vgJ48-Xj4Kc I made you in my image!MODEL > How can I help?
HUMAN > I’m not sure yet.
Haha silly humans.
Maybe these guys can tackle aligning Republicans and Democrats next.
And then after that, they can help us align the Middle East.
In fact, while we're at it, let's just align all the nations, religions, and ethnic groups. This is going to be great.
Who knew the moral alignment of humanity was just a side-quest on the path to ASI.
Vonnegut already has you covered.
AI models don't train themselves. The vast majority of even just the US population is deeply skeptical of this stuff, even if they use it a lot. You can see in the whole data center debate how little people are willing to support even just inference. And now we're seriously claiming those people would want to have ever-accelerating model training and recursive self-improvement?
You could even take a number of the wilder real, direct quotations from certain billionaire/oligarch types and get the voice actor for Ted Faro to record them, and they'd fit with in with the context of the story.
"Won't happen. It's too much like sci-fi."