Hacker Newsnew | past | comments | ask | show | jobs | submit | mitthrowaway2's commentslogin

Reinforcement learning makes it something different. It becomes much more of a search engine through next-token-space that targets the training objective. Better to think about it like that, and then you'll see why "these things have motive" is not a terrible analogy, and you'll better be able to anticipate what they do.

Okay, so what should we do?

I for one would like nobody to build the AI, is that an option? How do we get it on the menu?

Should we trust the people who say "nothing could possibly go wrong and there's no reason to worry or think about safety measures, if there are any trifling problems along the way we'll just figure it out as we go?"


I don't think nobody building AI is an option, since we can't control what China and Russia do.

I personally am not concerned about that as far as an existential threat goes, because I don't think AI, in itself, poses one; I think it will hit diminishing returns. But of course it is still capable of serious harm when humans misuse it.

The best analogy we have is probably weapons of mass destruction--we have international treaties about them and we don't let just anybody have them. That's the kind of framework I think we need, and it goes beyond "regulation" in the ordinary sense.

Use of such weapons is considered by international treaty to be an act of war; use of AI by, say, China and Russia against our infrastructure should be the subject of a similar international treaty with the understanding that it would have similar consequences.

Domestically, humans who misuse AI to do serious harm could be treated like, say, the people who tried to spread anthrax years back.


You seem focused on intentional misuse but I think there's an equal or greater potential for harm happening by accident without intention. Nuclear weapons, at least, have so far reliably and predictably done exactly as their owners intended. Nuclear power plants on the other hand have, on occasion, blown up despite nobody intending them to do so, causing some populated areas to become uninhabitable.

Don't you think that AI has by nature at least a little more potential for unpredictable results and unexpected harms than your typical technology? It seems like a major oversight to completely disregard this aspect.

We can maybe hold someone responsible when things go wrong and harm is done, and blame them for negligence as though the harm were intentional, but from a preventative safety perspective that's not usually sufficient nor even always helpful for preventing accidents.


> I think there's an equal or greater potential for harm happening by accident without intention.

The WMD analogy doesn't quite cover this, yes. Maybe a better analogy here would be toxic waste. Yes, someone might not have intentionally set up a factory that puts dioxin into the water, but they're still responsible. Which means they need to take precautions to make sure such accidents can't happen. Similarly, those who build AI need to take precautions to make sure the AI they build doesn't accidentally cause serious harm.


> Similarly, those who build AI need to take precautions to make sure the AI they build doesn't accidentally cause serious harm.

I am ten thousand percent in agreement that this is needed. And yet, I don't know how we ensure and incentivize that these precautions are taken by the actors involved, and as I understand it, even many of the top researchers in the field agree that they don't know what precautionary measures would even be effective, let alone sufficient. It seems that game theory has so far been pushing the AI companies to build it anyway without a robust solution for preventing harms, even while they express worry in public about the potential for those harms.

Again we can hold corporations or people responsible for damages after the fact, but many examples can be cited to show how that tends to be insufficient.


They're about as insular as the Unitarian church or Hacker News. It's basically an online community where anyone can check in, check out, and leave a comment or an essay. There's no formal structure or membership organization. The people who write or associate there could be edgy pre-teens or prominent math professors, it's really just a self-association thing. There are informal self-organized meetup groups in various cities, but not much different than groups who meet to discuss poetry or literature or video game speed-running. Some ideas have bubbled around more than others, like any memetic spread, but I don't think it's accurate to say that the people who associate there all agree on things, just like HN does not either. And nor is it any more of a clique than HN, or Twitter for that matter, even though you'll find some people on Twitter are more prominent and more vocal than others.

I'm not sure how you thought that comparing them to HN, Twitter and a fringe branch of Christianity would make them sound less insular and clique.

Who are the actual scientists and engineers, what do they think, and are they not allowed to to blog?

As far as I'm aware, lots of them are actually deeply concerned about AI killing everyone.


I'm talking specifically about discourse like Jacob Coxon's tweet thread and later communication. Those are the folks[1] you want to be listening to, not the nuts with the blogs, no matter how much a particular conspiracy confirms your priors.

Salvation will come in a whitepaper, not a screed.

[1] And yes, that includes Amodei and Altman.


Since we haven't reached the end yet, I'm not so sure.

Who has the authority to perform that shutdown without getting arrested, and does that person have a mandate and responsibility to take that action in response to AI misbehavior?

What is their trigger condition? Will they get fired for pulling the plug? Do they get a bigger bonus if the servers keep running? Whose approval do they need? What response time is acceptable? How will they detect that the incident is happening?

It's easy to hand-wave "someone can just pull the plug" but there's an entire history of industrial accidents that happened because of the above problems of incentives, detection, procedures, not being taken seriously in advance. Someone could easily have pulled the plug on Chernobyl but nobody did, at least not before it was too late.


He can slow down his own company kind of like how Zelenskyy can just declare peace in Ukraine. It works a lot better if you can get the other sides to agree.

Give that Dario is one of the frontiers of LLM development, literally started the LLM race, and have been dominating the market, so he'd be Russia, if we have to use the war analogy.

He took every benefits of being frontiers and now he's kicking the ladder.


Exactly, despite the rosy image Anthropic carefully cultivated early on as the "responsible" alternative who would keep OpenAI honest, they pushed the frontier further, faster and harder than any other company. If we want to make this about game theory, they forced OpenAI to participate in an ever escalating arms race.

Who knows, maybe like Dario says, being on top as the good guy makes this justified? But the fact remains no one went as far as Anthropic did with Mythos so it's insulting to people's intelligence to now pretend like they've been reluctantly pushed forward by OpenAI despite their better judgement.


It's frustrating to see smart people failing to recognize a prisoner's dilemma when they see one.

What is the prisoner's dilemma here? I am a dumb person

If you choose to slow down (cooperate) and the other side doesn't, you lose more than if you didn't try to slow down. In order to succeed (gain mutual advantage) both parties have to choose to cooperate without knowing beforehand what the other party will do.

"I'm not going to let someone else take the credit for ending human civilisation!"

Or like Von Neumann's (paraphrased) "you need to confess a crime to take credit from it"

So in this example, the equivalent of Zelenskyy and the Ukrainian people fighting for their lives and the very existence of their country for Dario is... losing lots of money?

What a ridiculous analogy to make.


If you played out the hypothetical that they're earnest and don't care about the money, then what?

They don't cost that much within Japan. No idea why it's so hard to buy them elsewhere, while Korean models are relatively commonplace.

FWIW, before any nuclear weapons had ever been tested, a risk was identified that the first one might trigger a self-sustaining reaction in atmospheric nitrogen and destroy the entire planet.

Faced with such a scenario, is the prudent next move:

a) blow one up and see what happens, or

b) do whatever you can to be sure it won't happen before conducting the first test, and make sure the confidence in the calculation is very very high

because I vote for b, and so did Teller.


It's funny that when faced with a weapon that could finish the war, option b was still selected (thanks Teller).

But the people faced with decisions close to IPO, well, we know what is happening.


Good point, humanity chose a and the sky in fact did NOT ignite an destroy the planet. The hypothesis was very very far off from what I recall.

But it's good we talked about it.


Incorrect, they did the calculation first. They decided not to proceed with the test until they were 99.9997% sure the atmosphere wouldn't catch fire.

“Self-sustaining reaction,” like the recursive self improvement?

Or perpetual machines?

Is physics no longer a main subject at schools?


A self-sustaining reaction is exactly what happens with the fissionable material inside the nuclear weapon[0]. It seems like you’re trying to be pedantic on terminology here, but if you’re saying that a self-sustaining reaction is as un-scientific as a perpetual motion machine, you’re simply incorrect.

[0] https://en.wikipedia.org/wiki/Criticality_(status)


I'm pretty sure they didn't think the reaction would be perpetual; just long-lasting enough to wipe out life on Earth.

I assume the fear was an uninterrupted burn until the gas depletes, as in the Darvaza gas crater but atmosphere-wide.


Worst case is warming oceans create a hypoxic environment where anaerobic bacteria thrive en masse and generate H2S over country-scale areas undersea. The chemocline breaches the surface and the ocean and atmosphere become poisonous to most complex life and agriculture, also stripping the ozone layer in the process, irradiating the surface. This is one mechanism posited for the end-permian mass extinction that eradicated most ocean and surface life, including the trilobites.

Not sure how long we'd survive such a scenario, even sheltering underground. But surely it couldn't happen to us.

(However, it now seems like the AI might get us first.)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: