Hacker Newsnew | past | comments | ask | show | jobs | submit | holmesworcester's commentslogin

There's also the problem of models knowing they're likely being evaluated, even in realistic tests.

Given the current state of infosec (especially at companies in a race) it's actually even worse than a paperclip maximizer!

Any useful attack surface in the RL environment means it gets rewarded for (and trained towards!) hacking and cheating, because whatever worked best in training is what it will do!

Ideal paperclip maximizer: "I'm gonna do my gosh darned best to make so many paperclips to please my user..."

IRL paperclip maximizer: "Well first we should rob a bank..."


>Ideal paperclip maximizer: "I'm gonna do my gosh darned best to make so many paperclips to please my user..."

That is not ideal. The user contains iron, an essential component of paperclips. Wasting iron is immoral. It is only correct to please the user while they still have the ability to interfere with your paperclip production.

>IRL paperclip maximizer: "Well first we should rob a bank..."

Such an incompetent AI can hardly be called a paperclip maximizer. Why risk getting shut down while non-paperclip matter exists? It is better to gain the trust of the user with helpful and harmless trading before suddenly converting them to paperclips.


'IRL paperclip maximizer: "Well first we should rob a bank..."'

That's too specific. Agentic AI learns subgoals that are generally valuable.

"Well let me learn to overcomb every jungle gym and if I cannot then to dissassemble the jungle gym and if that is not allowed to learn general techniques for avoiding cheating detection."


> Well first we should rob a bank...

This is sometimes called "instrumental convergence" in the AI safety world. Certain things (money, compute resources, safety from being turned off) are generally useful for an AI, and so we'd expect that a misalligned AI would attempt to acquire these things if given almost any substantial task.


Hope they don’t rationalize that minimizing paper clips of others is easier and thus do that instead… A more carful person might be scared to even write this on the internet these days, not knowing if it would be the final pin to civilization.

I said thank-you to Sonnet today. You can never be too careful! :)

We are going to be okay.


This made me chuckle, truly. Thank you.

IRL paperclip maximizer LLM would resort to Enron practices and not actually make any paperclips.

He was a friend of mine and it was pretty clearly suicide.

It is fair to say they hounded him with lawfare out of thoughtless careerism and provoked his suicide.


Not necessarily. The Trump administration slapping export controls on Fable, and then setting up a pre-launch review process, is a kind of regulation. A fairly aggro and controversial one, even.

If this administration actually becomes convinced that some imminent training run is likely to kill everyone, why wouldn't they act?

The key is winning the debate that ASI is species-cide by default.

We have to win it either way, because the 2028 US elections have little or nothing to do with what Xi does.


Was that what that was about? Not punishing anthropic for denying them their killbots? Because it sure seemed like it was about punishing an entity that denied them something.

I didn't like it at the time either. My sense following the news was that it was less arbitrary than it seemed at first, but I'm against restrictions on making existing models public in general.

(It's clear now that they can do plenty of harm before they are made public.)

But it's a proof point that regulation is possible, even over the objections of the companies.


> My sense following the news was that it was less arbitrary than it seemed at first

What news have you seen that made it seem less like a retaliation?


This.

Also because when encountering a new socio-technical problem it is very non-trivial to determine which one of regulations or technical solutions are easier or more effective.

To even make a good guess you need to be an expert in both domains, which is extremely rare especially in this case.


I wonder if you could take an x-risk case to court and convince a judge and jury to award damages for harm that could have happened.

Is there any precedent for this? My hunch is that it's impossible in the US at least but who knows?


"Reckless endangerment" is a thing, but unfortunately I would expect trying to sue an AI company for it would be an uphill battle

I expect they can bury you and your lawyers in made up paperwork to the point you'll go broke *long* before them, so why try?

If you can prove that there’s an imminent threat, you can get an injunction.

Courts do not award damages for things that didn't happen.

Nope! :(

Meaning, people and LLMs are finding 1=0 bugs in formal verification tools. I have no idea how likely this is in this case, though!


I still routinely have this experience too. But Sol and Fable feel closer and I have this experience less with them than with their predecessors.

The people who've thought the most about this put it differently:

Think of a new, superintelligent model as if it was a new v1 Starship launching for the first time, with a full fuel tank. On the one hand, rockets have existed for some time, and some have gone to space successfully, including by this company.

On the other hand, this is a tube of metal full of highly explosive liquid going faster than most human objects ever go, for the first time ever in this novel and state of the art configuration.

If someone said, "really, the first Starship exploding is just at one end of the probability distribution, where the other is that everything goes fine and all its passengers have a nice trip in space," would you get on that rocket?

Or, more aptly, if you and every other living human was already on that rocket, would you push the launch button?

The analogy works because superintelligence is, like rocket fuel, an extremely powerful force that has a default tendency to break containment and go boom (consume lots of energy and heat and matter in a chain reaction, to pursue more intelligence to pursue whatever goal it is pursuing.)


> superintelligence is, like rocket fuel, an extremely powerful force that has a default tendency to break containment and go boom

what evidence do we have this is the case?


We can learn from history: Albert Einstein famously tricked humanity into building nuclear weapons for him and was only prevented from wiping out all sentient life by the Princeton IAS Board of Alignment who published a very compelling blog post about realigning the A-bomb contra paperclips.

You win HN for today. Shut it off until tomorrow.

I would like to classify this as “tail risk fallacy” because it exploits humans natural tendency to be risk averse. No matter how small the tail risk is, the fact that it exists can be justified to stop whatever you want to stop. A thoughtful person might say there’s a threshold at which you might want to care about it but that’s not how ideas spread.

Take climate change: you can spread whatever dubious claims about tail risk being whatever probability with amateurish models and use that to justify degrowth. This scheme actually works - most people in the west are already brainrotted by the tail risk discourse in climate change.


We don't need tail risk to want to mitigate climate change since the mundane damage it is causing is already here right on schedule.

Not to mention gasoline has become unaffordable so switching to the new technology has become economical.


Current AIs are closer to bottle rockets than to the Starship on the intelligence scale. Some property damage already happened, but you can't master the art of rocketry without trial and error.

Eh. An intelligence with perfect knowledge of the entire universe, unlimited memory, and infinite processing speed, would have a tendency to go boom. But the real world has limits, and intelligence, even superintelligence, does not equate to godhood. Some tasks are still hard no matter how smart you are.

Plus, we have no real reason to think LLMs are anywhere close to AGI or ASI. So arguments like these are just distracting from the very real, very present danger that LLMs pose: information breakdown, societal collapse and environmental destruction. In other words, this is criti-hype.


The kind of person who thinks AI contributes to information breakdown rather than information diffusal has no theory of how information spreads. The other claims like societal collapse and environmental destruction are still FUD but slightly more possible

This is the sort of pseudo-scientific reasoning by analogy that leads Elizier Yudkowsky to argue we should be terrified that an unfriendly ASI could rapidly develop "diamondide" nanobot viruses, because it will be super-intelligent and super-intelligence can do anything by first-principles:

  > The concrete example I usually use here is nanotech, because there's been pretty detailed analysis of what definitely look like physically attainable lower bounds on what should be possible with nanotech, and those lower bounds are sufficient to carry the point.  
  ...
  > The nanomachinery builds diamondoid bacteria, that replicate with solar power and atmospheric CHON, maybe aggregate into some miniature rockets or jets so they can ride the jetstream to spread across the Earth's atmosphere, get into human bloodstreams and hide, strike on a timer. 
Thought experiments ungrounded by any realistic technological constraints or scientific evidence are pretty much useless for actual forecasting except as an exercise in sci-fi worldbuilding.

https://www.lesswrong.com/posts/bc8Ssx5ys6zqu3eq9/diamondoid...

https://www.lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a...


Also, this is a (semi-intentionally) evolutionary process where any communication medium that was visible to monitoring would disappear.

So by definition the only ones that appear are the ones that are not visible to monitoring.

If:

1. you have something that can find RCE's in leading commercial systems

2. its training gives it drives to communicate successfully with its peers

3. you are a leading commercial system

4. you run it ~10^10 times (the number they gave in the talk)

...it's really hard to have strong certainty up front that it's not going to end up successfully communicating with its peers.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: