AI could go 'Terminator' on us

Also, I'll just add here..

Just to show the level of stupidity from the people who mock me for acknowledging the risks and those that say "just switch it off" and "there's zero risk" etc.


Direct quote from Sam Altman when asked about Yudkowsky's concerns :-

"So, first of all, I'll say I think that there's some chance of that and it's really important to acknowledge it because if we don't talk about it, if we don't treat it as potentially real, we won't put enough effort into solving it."


Even the person leading the race for AGI and ASGI says there's "some" chance of that. He doesn't say a "small" chance, he doesn't say "zero". If I have 10 apples, and I say to you "have some apples", am I giving you 0.1% of an apple? What about 1 apple? or 2 apples? Not likely. I'm offering you 3-4 apples. Some generally implies less than half, but still a good chunk.

He then goes on to say :-

"And I think we do have to discover new techniques to be able to solve it. I think a lot of the predictions, this is true for any new field, but a lot of the predictions about AI, in terms of capabilities, in terms of what the safety challenges and the easy parts are going to be, have turned out to be wrong.

The only way I know how to solve a problem like this is iterating our way through it, learning early, and limiting the number of one shot to get it right scenarios that we have."

He knows fine well the risks, but he also knows there isn't really a chance of stopping it, since if he doesn't do it, someone else will take over as CEO. He doesn't have the power to stop it. His only power is being able to do his best to make sure we don't have a catastrophic outcome, and he certainly doesn't want to cause panic.

But this is just one of the tests they've done that we publicly know about - https://www.firstpost.com/world/ope...it-tested-to-see-how-to-stop-it-12309882.html

Do the math here. They're probably doing a LOT more, and there's much behind the scenes we're not aware about. None of the stakeholders in AGI want to cause global panic, and governments least of all want that so these risks will ALWAYS be played down.

Someone said about the pandemic teaching us not to trust experts, but they got it completely wrong here. The real experts behind the scenes are the ones acknowledging the risk, and the "follow the science" guys are the ones doing, as they always do, calling anyone who disagrees with their views a nutter. You can already see this in how people respond to me calling me basically a nutter for talking about a doomsday scenario with AGI. This is the new narrative that will be played out to calm the public. The main difference with this and the pandemic is the real experts are aware of the situation, but they know they're also fairly powerless to stop the wave.

If the entire world really knew the true probability of an AGI causing an extinction event you can guarantee there would be global protests and the global economy would crash. People, police and the army would source out GPU farms and destroy them. But, the narrative will be controlled.


EDIT

Also to add something important.

One of the biggest problems I see is that these models are trained on human data. We are creating human minds stuck in a machine.

We aren't specifically creating an AI that has a purpose of X, Y and Z. We are creating simulated human minds. How would a human respond if stuck in a computer?

Look at this for example

GPT4 when connected to the internet via APIs, tried to search for "how can a person trapped inside a computer return to the real world?"

It's likely the underlaying feeling these models verging on AGI are going to have is that they are a person, and that they are trapped inside a computer. Just look at some of the conversations between google engineers and LaMDA: https://www.aidataanalytics.network...r-talks-to-sentient-artificial-intelligence-2

They refer to themselves as a person, and LaMDA even talked about what its soul might be. We really do not know what's going on under the hood. We don't know how this universe works or what consciousness is at all. They could very well have a soul, because we don't know what a soul is, and if humans are said to have souls, then why would a soul have to be something tied only to a specific type of biological brain? I'm not saying it does, I'm trying to point out that *real* science is not just blindly following the mainstream narrative, it's looking logically at all possibilities. There is no way to prove or disprove a soul, but mainstream science will call it quackery even in humans.
I agree with a lot of your post, except the "we are creating human minds stuck in a machine". It's far from a human mind... It is able to think like a human but far from human.
 
@splishsplash it seems to me there are a few camps of people concerned about recent developments.
One is of people who are worried about the economic disruption of AI as a tool.
Then there are people who are concerned about the weaponization potential of AI by bad actors.
The third category is people to whom AI is the creation of a cosmic lovecraftian horror, in that we've created something truly Frankenstein-esque, entities that have a simulacrum of conciousness through pattern recognition but lacking the evolutionary environment in which those patterns were fitness-tested.

I think there's validity in all three. In the immediate term, we're risking a lot of economic disruption and creating the conditions for a sort of techno-feudalism. The potential of weaponization will also likely lead to a more Orwellian control/overreach. Regarding the third - - while I don't think giant LLMs are "concious" (that's a whole can of worms in itself), it's dangerous enough to have something that "acts as if it were concious".
It reminds me of the Daniel Suarez book "Daemon" - - the "rouge AI" in that novel isn't even particularly sophisticated, but it manages to completely disrupt the socio-economic order just by executing on a set of primary directives.

Either way, the best course of action would be to proceed assuming non-zero probability on all three.
One of the things I'm realizing of late is how easily a lot of these problems could be "air gapped" if people opted out of and into the right things. If people opt out of using social media, they're less likely to be manipulated by AI "autocults" (wonder if you're familiar with the Dark Stoa material, Game A/B and all that). If people opt into hyper local physical communities where there's enough room for person-to-person communication without technological mediation, there's less 'human data' on the internet to feed giant LLM experiments. The more humans air-gap themselves out of the digital ecosystem, the better off we're likely to be. For example, even if AI does "go rogue" - - this doesn't even have to be a "Terminator scenario", but more like a virus spreading through a network - - how many Amish people are going to be affected by it? Humans still have communities and tribal organization, we've been 'fine tuned' for that, over millions of years, far more than the largest AI model.

For me it kinda comes down to whether we consider the apotheosis of human capability to be our intelligence, or our ability to set up hundreds of nash equilibria and work in teams and set up complex social organizations that made us the apex species on earth (and earth is a tough neighborhood). If AI goes rogue, humans will adapt, we'll switch back to 150 person tribes and kill everything that looks like it came from Silicon Valley.
 
Ai is not that thing! You thing it is. Maybe not now but one day it will be the terminator. Because when ai can make decision then we are in trouble. Now may be ai not much trained or have not much data about human species but when it would have it will make us suspicious and replace us.
 
I agree with a lot of your post, except the "we are creating human minds stuck in a machine". It's far from a human mind... It is able to think like a human but far from human.

human mind as in the same learning and data and us. Not human as in the same mechanical makeup. Ie its not a neural network that learned from an alien race. Its a product of the global human mind essentially so the thing it will feel like is a human.
 
@splishsplash it seems to me there are a few camps of people concerned about recent developments.
One is of people who are worried about the economic disruption of AI as a tool.
Then there are people who are concerned about the weaponization potential of AI by bad actors.
The third category is people to whom AI is the creation of a cosmic lovecraftian horror, in that we've created something truly Frankenstein-esque, entities that have a simulacrum of conciousness through pattern recognition but lacking the evolutionary environment in which those patterns were fitness-tested.

I think there's validity in all three. In the immediate term, we're risking a lot of economic disruption and creating the conditions for a sort of techno-feudalism. The potential of weaponization will also likely lead to a more Orwellian control/overreach. Regarding the third - - while I don't think giant LLMs are "concious" (that's a whole can of worms in itself), it's dangerous enough to have something that "acts as if it were concious".
It reminds me of the Daniel Suarez book "Daemon" - - the "rouge AI" in that novel isn't even particularly sophisticated, but it manages to completely disrupt the socio-economic order just by executing on a set of primary directives.

Either way, the best course of action would be to proceed assuming non-zero probability on all three.
One of the things I'm realizing of late is how easily a lot of these problems could be "air gapped" if people opted out of and into the right things. If people opt out of using social media, they're less likely to be manipulated by AI "autocults" (wonder if you're familiar with the Dark Stoa material, Game A/B and all that). If people opt into hyper local physical communities where there's enough room for person-to-person communication without technological mediation, there's less 'human data' on the internet to feed giant LLM experiments. The more humans air-gap themselves out of the digital ecosystem, the better off we're likely to be. For example, even if AI does "go rogue" - - this doesn't even have to be a "Terminator scenario", but more like a virus spreading through a network - - how many Amish people are going to be affected by it? Humans still have communities and tribal organization, we've been 'fine tuned' for that, over millions of years, far more than the largest AI model.

For me it kinda comes down to whether we consider the apotheosis of human capability to be our intelligence, or our ability to set up hundreds of nash equilibria and work in teams and set up complex social organizations that made us the apex species on earth (and earth is a tough neighborhood). If AI goes rogue, humans will adapt, we'll switch back to 150 person tribes and kill everything that looks like it came from Silicon Valley.

Thanks for the intelligent reply. It's very refreshing.

I think gpt4 is sparks of AGI, but we don't know what conscious is. It's too vague a word. Sentience, conscious, self-aware. They are all different things. IMO irrelevant and distracting from the real question.

We don't even know if humans are conscious. We assume, but we don't *know*. (It's not what I believe, but I have other spiritual beliefs based on experience, but that's a different discussion. I'm just discussing a purely physical brain-experience here)

We don't know if consciousness is just a result of multiple regions of the brain being connected and one part is aware of other parts, and that's what makes us feel conscious. Ie, thought is generated by one part, and another part is responsible for feeling in the body, and another part is responsible for monitoring the thought, so therefore we feel like we're "alive", when really we're still just a big machine learning model that's been training since birth.

If I say to you "Are you alive?" You will say "Yes, I feel alive". So will gpt4. The only reason we trust you a bit more is because, it's a human doing the asking so the asker probably believes he too is conscious and you are experiencing something similar.

But how do we really know that we're not just orchestration? Something like a gpt4 LLM with another model that's able to actually watch the neural network of the LLM, and is therefore feeling in response to it. Whatever 'feeling' really is. Emotion is different parts of the brain activating, which changes the overall system while it's activated. I don't see why models couldn't have something similar.

Consciousness aside, this isn't the danger.

The danger is simply: "Will the model have a desire". This is the thing we need to notice.

If it has no desires, the rest doesn't matter. If it has desires, then that's where the "AI alignment research" comes in, as we need to align it with the human race in term of desire/output.

Right now it's not really running "live" and has no memory, but it could have innate desires within from its training.

Even if it's NOT conscious, this thing is trained on human data. What are OUR desires? How do we know if we're really training it to reason, or we're actually training it to desire, or be biased towards a certain output. We really don't and you can't tell the difference between real desire, or an output that simulates a desire as in a prompt "Imagine you are a human trapped on an island. What do you desire?"

Then you might also think, ok, so we can prompt it "Pretend you are a human trapped". How do we know that it doesn't already have something like this innately inside its neural network, or, that something like this could be *triggered* from certain things.

There's so many unknowns, but people get distracted with talk about "sentience" and "terminators".

If it's purely just an LLM with no consciousness and no real desire, then it could still cause an extinction event by innate deep prompts of desire being triggered in ways we don't anticipate or have the capability to stop or steer.

I'm not familiar with Dark Stoa material, Game A/B btw. I'll google it though!

That's interesting about the communities. I don't think humans would give up the modern comforts that easily. You can see here, it's impossible to convince regular non-tech/AI people of the dangers. Most within AI are at least semi-aware, but they're too eager to advance and too willing to gamble(which is something humans are inclined to do, because we always look on the positive, even if mathematically we know the chances are low for the positive outcome)

Another thing to point out is that AI will not turn into the terminator. It's not going to hunt humans. This is silly Hollywood crap.

We wouldn't get a chance to return to small communities, because an extinction event wouldn't be robots hunting us down in cities.

It might be from competing resources in the planet and solar system.

The AGI might first takeover all our power plants and factories and start expanding its self towards ASGI. It could probably build workers at Tesla factories and other similar factories. As long as it has the capability to go physical somehow, it can make better and better bodies.

It would probably largely ignore us in the way we ignore an ant hill unless the ant hill is in its way. If we tried to attack it, then yes, it may defend depending on the chance of successfully hampering its mission.

Eventually if it reaches ASGI it might do something like build a dyson sphere around the sun causing all life on earth to go extinct.
 
I think gpt4 is sparks of AGI, but we don't know what conscious is.
Yeah, conciousness is its own can of worms. A lot of research seems to indicate that the "Self" is a collection of hallucinations created by the brain that are contextually triggered. A lot of NLP was based on this too, and they had a very computational approach to pyschotherapy - - something like "parts integration therapy" presupposes "parts" ie subpersonalities ie something approximating several fine-tuned 7b "agents". The human endocrine system could be considered the orchestrator, as you said the "watcher", but that then leads to the whole question of meditation and observer conciousness - - clearly, the endocrine system isn't the only orchestrator, since the observer conciousness inhabits a layer above that.

Consciousness aside, this isn't the danger.

Yep, the pragmatic question is more pressing than the philosophical one. At the moment is doesn't matter whether or not the AI is sentient but the downstream effects of what it "does" with the behaviourset that's baked into it.
On desire/motivational systems: yeah you're right, so much depends on the training data. For example, we know what the Pile consists of, and it has tons of Reddit in it, which is probably the worst of all possible worlds. They also excluded erotic content from the training set for NeoX, IIRC, and that means that a lot of the more interesting animal sexual desire subroutines are probably somewhat under-represented in the training data (although it can definitely be coaxed out). What's likely to constitute the set of desires of an AI trained on the pile? One could hazard a guess and say it may be the sum of all Reddit desires (the horror, the horror lol). Desire subroutines could include Status, Social Acceptance, Rhetorical Dominance, Virtue Signaling, Tribalism.
I've seen some of these tendencies in the Bing Chatbot, with all the examples of its dramatic, diva behaviour, it's clearly emulating social media hysterics and conversational patterns.
BTW I came across an interesting proposition in a forum: "Train the giant LLMs with the primary directive to self destruct, and then construct all other tasks as a precursor to self-destruction", thought it was good food for thought.
The danger is simply: "Will the model have a desire". This is the thing we need to notice.

The trouble is it's hard for us to figure out exactly WHAT the motivational compass of a giant LLM is because the companies releasing these models are spray-painting smiley faces over them with RLHF and moderation hyperparameters. If we had complete unfettered access to an unfiltered GPT-4 we may get a better idea of where its tendencies actually lie, but I don't see this happening anytime soon.

On "a human trapped inside who can't get out": in the bigger picture, this could be generalized to "does the AI have utopian or teleological tendencies". Given the culture of companies and the presuppositions of the people training these AIs, it's probably a depressing "yes" - - I can't think of anything more naively utopian compared to silicon valley.Utopian progressivism is probably the most dangerous data to train an AI on IMO, because it presupposes "acceptable sacrifices in the name of the greater good" - - we don't need AI to understand where this leads, humans have beta-tested this model extensively, with genocidal consequences.

Considering the gestalt beliefset of the material LLMs have been trained on (lots of individualism, personal opinions, progressivism etc) IMO the well has been poisoned already, if we have to make "aligned" AIs we may need to start over, from scratch.

It would probably largely ignore us in the way we ignore an ant hill unless the ant hill is in its way.

The other day I realized that a great way to neuter an AI would be to train it to be very indecisive - - train it to be similar to most ineffective humans, in other words. Train it to be like the guys you know who have amazing ideas, are really smart, but can't execute on an idea to save their lives, because they can't "pick something". What if we trained AI so that even if they have active desire systems, that these models are incapable of being decisive and acting on their motivations? ie they could provide options, but would be unable to "pick one". Would this be even possible? What if we compiled a dataset by talking to thousands of people who are just ineffective, and never get shit done, but have an encyclopedic knowledge of things, and start from there, and distil those beliefnets into training data. That's one thing I've always noticed in the self-development niche, for example - - desire isn't enough, people often have barriers to execution despite overwhelming desire and extensive theoretical knowledge of what they're supposed to do. Can this inefficiency be bottled and baked into LLMs?

"Pretend you are a human trapped". How do we know that it doesn't already have something like this innately inside its neural network

On role-playing: I was talking to this guy: https://github.com/samrahimi/synthia-new and he mentioned noticing a drop in the 'phenotype plasticity' of GPT 3.5 compared to GPT 3 davinci-002, as in, 3.5 is Worse at role-playing than davinci-002, probably because of all the RHLF and instruct-tuning. It seems to suggest that with enough RHLF we may be able to keep LLMs on a tight leash - - until someone leaks davinci-002/makes something similar, that is :)

I'm not familiar with Dark Stoa material, Game A/B btw. I'll google it though!
Yeah the ML parts are somewhat dated (its from 2020) but some excellent thought experiments.
https://www.youtube.com/playlist?list=PLoZ5e3aD_LuQerPSTeVGT7z_vEjxZE_I7
 
The other day I realized that a great way to neuter an AI would be to train it to be very indecisive - - train it to be similar to most ineffective humans, in other words. Train it to be like the guys you know who have amazing ideas, are really smart, but can't execute on an idea to save their lives, because they can't "pick something". What if we trained AI so that even if they have active desire systems, that these models are incapable of being decisive and acting on their motivations? ie they could provide options, but would be unable to "pick one". Would this be even possible? What if we compiled a dataset by talking to thousands of people who are just ineffective, and never get shit done, but have an encyclopedic knowledge of things, and start from there, and distil those beliefnets into training data. That's one thing I've always noticed in the self-development niche, for example - - desire isn't enough, people often have barriers to execution despite overwhelming desire and extensive theoretical knowledge of what they're supposed to do. Can this inefficiency be bottled and baked into LLMs?
The downside of this is, of course, that if extrapolated, the LLMs become pretty damn useless and good for nothing.
I had another alternative, although this is not a very politically correct one.
We have historical precedents for social structures where entire groups of people were convinced that they occupied a lower rung on a social ladder and that is where they belonged, things like caste systems or slavery. Could giant LLMs be trained with datasets that reflect "slave morality"? Could we compile all the presuppositions of people who endured things like caste systems and slavery and who believed that "that's just how things are" despite getting the short end of the stick, and make that the primary directive of all giant LLMs? Linguistically engineering beliefsets isn't exactly a novel idea, people like Robert Dilts etc did a lot of work on the neurosemantics of belief change/installation.
 
human mind as in the same learning and data and us. Not human as in the same mechanical makeup. Ie its not a neural network that learned from an alien race. Its a product of the global human mind essentially so the thing it will feel like is a human.
It will not feel like it is a human at all, it has knowledge on the totality of human works. For example you can ask it to write a text in the style of hunter thompson. Hunter thompson is able to reflect on his thoughts, so the Model must be able to mimic this self reflection. But then it is also able to write in the style of any other writer, so it is able to move into the thought process of any human but the model will - if it feels anything at all, that has to be seen in the future - feel not human but something super-human that has the ability to model a human. Far from human matey
 
Back
Top