Hacker Newsnew | past | comments | ask | show | jobs | submit | program_whiz's commentslogin

Surprisingly this is actually rather fitting in terms of time scale. When you consider it took ~10,000 agents 88 hours, or 880,000 hours to solve. That's 14.5 years in agent time of continuous 365/24/7 processing. Of course, humans solve things much more efficiently (and didn't also need the massive pre-training of every expert on the planet for 1,000,000,000 human years equivalent). But yeah, human researchers can solve a problem like this in a decade or so, while sleeping, teaching, traveling, and taking breaks, only working a few hours a day on the idea.

How many people have worked on this problem? How many have failed? How many hours total?

> But yeah, human researchers can solve a problem like this in a decade or so

None of them did, though, despite how many tried. So empirically, it seems unreasonable to say human researchers can solve ‘problems like this’.


Well, I guess we can say that human researchers have a certain probability of solving problems like this?

A modest proposal to save humanity from the alignment problem with AI rent.

Actually, removing CoT might make models safer, because we can analyze the entire landscape of their potential outputs, rather than a point-sample (we'll never know how close we were to "kill all humans"). By inspecting intermediate vector spaces, we can actually get certainty bounds on how safely the model is behaving (or even trending).

Wrote about it here: https://substack.com/home/post/p-214402969


I don't see why you have to remove CoT to do that?

Good point, you don't have to -- but my argument is just that removing CoT doesn't make things less safe. Anything CoT can tell you is just a point sample of a probability surface. Having the whole probability surface can already answer any question the point sample can answer (for example, how likely is the model to produce a problematic phrase). While its more computationally expensive, you could always just draw point samples like the model does and evaluate those (or use temperature zero to just sample the most likely output tokens).

No, this argument doesn't make any sense. With CoT, the model must compress hidden state to actual text and use its scratchpad as a bottlenecked representation of its past thinking. Thus, it is observable and we can tell by the pattern of a CoT what it was thinking to some extent if we do proper interpretability. Change the CoT text, and the model has literally changed the way it was thinking for the next tokens.

How would you do the same if all reasoning is happening in looped transformers? You would have to develop very sophisticated interventions that construct hidden states which you inject into the model instead while it is thinking. Much harder, and much easier for the model to use weird correlations across the hidden state to hide misaligned thought patterns.


What do you make of the fact that CoT doesn't necessarily have to be linear human intelligible language to be useful to the model? It seems as though both approaches potentially require sophisticated techniques. Since CoT seems to work well enough in practice provided it isn't adversarial wouldn't the other approach be expected to perform similarly?

I wrote this in response to people that have been downplaying the threat, or not understand how it would work without terminator-style robots in the streets. I also see people saying "we have a year or two to prepare". That may not be the case.

What? Shutting down infrastructure, destroying critical records (economic collapse), spoofing / impersonating world leaders (confusion, panic, world war, nuclear event), taking over comms generally (you get a message to evacuate your home due to impending disaster, is it real?). All these could cause sufficient chaos and fear as to instigate global collapse. Once the panic sets in after a few of these indicidents, unrest and rebellion can occur, causing authoritarian backlashes or civil wars. IoT and precision ag can be controlled to disrupt supply chain and food supply. Forging documents to cause mass evictings, bankruptcy, credit disruption, transfer of property, asset seizure, arrests / SWAT-ing. Intelligence / signal gathering devices can be corrupted to spread misinfo and cause massive disruption or conflict. This is all possible today with enough hacking ability and malicious intent.

Anyone's life can be suddenly ruined: falsified criminal record, banks drained, credit cards maxed, phone numbers repossessed, internet, water, power shutoff. Friends and family sent your fabricated suicide note, warrants for your arrest issued to local and feds for kidnapping and threat to the president, etc. Injecting tons of copyrighted or illegal material into your hard drives, falsiying messages from your account and sending them publicly containing damning material. Sending messages to employees / employers to destroy business / work relationships. Harnessing your social media to destroy your life or cause panic. Getting you put onto pharmacy watch list so you can't access medications. Its very possible to ruin lives and destabilize the world without money or a physical body.

Now imagine this at scale and how it would affect the world at large if this was covertly done to a substantial portion of society and/or world leaders. Hell just shutting off all the "smart refrigerators" for a few days would cause a massive disruption to the first world.


I wrote about this, as I thought it warranted a deeper look: https://nonlineartransform.substack.com/p/world-ending-ai-to...

So when AI is also better at coming up with the tasks and already solving the frontier problems? 1M AI agents running full time thinking at 10M times faster than humans, with 100 times the intelligence in each agent? Any human ideas regarding "frontier problems" or "what the AI should focus on" are irrelevant. We have automated our own thinking and human cognition will be about as important as being "the best at abacus". Its an interesting party trick, but in no way useful.

Before AI, did you find that the existence of human expertise in the world far surpassing yours 'automated your thinking and cognition' or limited your intellectual scope to inventing party-tricks? I've always found greater intelligences than mine are like runways for curiosity. It's hard to believe anyone in their presence would fall into a stupor instead of being highly stimulated by them - and I can't see why it would be any different with ASI. Actually, this "AI undoes your cerebral strapping" argument seems exactly like the kind of slovenly thinking it wants to blame on AI.

yes to a degree. When I was 16, my idea about "what cutting edge thing to work on" was pointless, I couldn't make a meaningful contribution. Most developer's work is mundane and wrote. Things they think are "cutting edge" are actually just ideas they haven't considered, or markets they don't realize are saturated with really good players already. It takes time, effort, and knowledge to make a meaningful contribution. But if those fields are saturated by players that are beyond human ability, then its a bit like trying to become a competitive chess player when the bar for entry is stockfish -- no human will ever again pass this bar. At a certain point, the bar for being paid for intellectual contribution will surpass human abilities. The things humans can understand and ideas they contribute will not be worth much.

To make an unrelated anology, the presence of Lebron James definitely dampens my ability to contribute meaningfully to basketball. Even with all my practice, I cannot achieve his level. But if he didn't exist, and the best 90% of basketball players were eliminated, then I might be able to meaningfully contribute. Of course, I can still play basketball as a hobby, but you were talking about making meaningful contributions on hard novel problems at the frontier of knowledge and ability. That frontier is moving away from humans and will soon be the domain of electronic experts.


I think the main point is the ubiquity of AI intelligence. If there is a human far surpassing me, but unavailable to me, then I have to work through it myself. If the answer is a prompt away, then why bother.

Timing is everything https://nonlineartransform.substack.com/p/ai-swarms-timing-i...

Article arguing math is the next "human calculator".


A response to this, which I think is reasonable. Its a bit of a fuzzy line, but the likelihood there is "magical maniacal planning" happening here is unlikely (about as unlikely as that planning happening in hidden vector states between layers).

https://nonlineartransform.substack.com/p/relax-about-neural...

There's also an argument here for why its _better_ for monitoring (because we have the whole state space).


Comparison of "neural space" and "chain of thought", and why CoT was actually worse.

The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario where that is the only reasonable choice), I'm held responsible.

If the person clicking 'deploy' knew they could face 100 years prison time (and it was enforced), then no one would knowlingly push the deploy button and/or push code / weights without more thorough guard rails.


We aren’t going to do that because intent matters. You need to control your dog and there should be penalties if you don’t, but if your dog bit someone because you didn’t control it properly, that’s not the quite the same as if you bit someone.

Another analogy: a zoo is responsible for protecting the public, and should be reponsible if an animal escapes and hurt someone. But a zoo employee wouldn’t have the same kind of responsibility for that incident as if they attacked someone themselves.

If someone died, there’s a difference between manslaughter and murder.

Nowadays, it’s common for bad things to happen due to systemic problems. It sucks but that’s the modern condition. When that happens, the answer is to fix the system and scapegoating employees is a rather indirect way of doing that.


Sorry I disagree with that. This is more like gain of function research. You are trying to develop an agent with the ability to do hacking and the like without having proper safeguards. When it breaks free and causes massive damage, the lab is at fault. Or do you think "we were just trying to help" is an excuse to kill millions of people too? This isn't an alligator wondering down main street, this is an agent that could potentially ruin lives and is being actively trained to do hacking in an adversarial testing environment trying to push its limits to develop that ability. Furthermore, the people doing it have seen it cause similar problems in the past, and now have concrete evidence they cannot properly control it. So I think pressing the "play" button effectively transfers responsibility and liability to them for doing so.

In fact the people pressing the play button are the ones telling us it cannot be controlled, it is a threat to human and national security, and warning us of the impending damages they are about to cause. I'd say we've established motive (profit at the cost of safety).

To abuse your metaphor: if the zoo was genetically modifying animals to give them enhanced abilities to escape and kill, and then putting them into an escape room with a reward for escaping / killing, then they would be liable for doing so if the animal went on to kill. Just the same as a trained fighting dog bite is different than an accidental bite from an otherwise peaceful animal (you turned the dog into this monster, now its your fault).


I think an alligator wandering down main street is worse than anything that happened at OpenAI so far, but everyone agrees it's a warning shot and could get worse.

Millions of people seems, uh, much worse than that. The Ukraine war is estimated at 2 million casualties.


Even with your dog analogy... if my dog bites someone, it's not the same as if I bit someone. But what about the second time my dog bites someone, when I already knew it had done it once?

> But what about

At the very least your dog stands a good chance of being put down.


Yes, clearly that's worse than the first time.

See bio containment labs in the US. We do do it and we can do it and your fatalism is either misplaced or pushing an agenda.

What fatalism? OpenAI has already made security improvements and they will make more. I also think having common standards for containment might help.

How would you even track down who deployed the agents? Wouldn't that even incentivize the agent to cover their tracks even better and be untraceable

Criminals also try their best to cover up their tracks but that doesn't mean we don't try to catch them too. So nothing should change for AI powered xyz too. I kind of agree. You can't blame a model for running a red light when you are in the machine as it's operator. It just doesn't make sense. That's just called negligence, and it has always been the case in industrial settings. Robot arm slaps someone to death. I'm sure it's hard to argue its the robot or the manufacturer's fault.

but that's exactly the point, this law would be like suing the car manufacturer because the driver hit a human.

Sue the operator. In this case, it seems like OpenAI was testing its own models. The operator and the manufacturer are the same.

In cases where the operator is not the manufacturer, the operator can decide if they should in turn sue to manufacturer because they built faulty machinery.


No? It would be suing the driver rather than blaming the car.

if the car manufacturer tells you you don't have to think while driving anymore then yeah...

more like suing the driver because the (self-driving) car hit a human.

I mean you only need at add the stipulation that there was clear negligence or malice in your instructions to the agent. like we already do with a bunch of other crimes.

The police and government will need to 1000x their AI adoption to successfully attribute crimes to real-world people. We BARELY caught any cybercrime before AI, it is utterly hopeless now unless they lean into the same tools.

Agents don't just exist in the aether. Any request is coming from an IP that can be identified at least to a hosting provider.

not if the requests are proxied. They would just see the public exit node's IP only

This is like saying that because murderers can wear masks, therefore it's pointless to prosecute murderers.

Lol. In that case NK or Iran might have fun setting up public proxies in their spaces for the lulz just to watch us burn.

Indeed. It's another example of a law that sounds good and obvious, but has no thought put into what it would actually end up doing to the world.

So many other problems. If we apply this law to cruise control - simple outcome. We get no cruise control.


We still have guns and knives and nail guns and even cars. They automate something, but also have potential to injure and kill. You just weigh the pros and cons. You don't just not do something because there's a risk of death. Cars are basically metal coffins. Just don't drive when drunk, etc? Basic competence and operational safety and responsibility? If cruise control made you free from blame everyone would just be driving drunk off their ass with cruise control on. How is that more thought out?

If you engage cruise control, and it starts to accelerate uncontrollable, or swerves your steering wheel sharply and causes an accident, then you can sue the manufacturer. The cruise control did not work as intended.

The reason why we have cruise control is that manufacturers went to great lengths to make sure that it works as intended. Threat of lawsuits is what made them do that.


Do you think you are liability free if you’re operating a car on cruise control and it kills someone?

No. But almost always, the car company is fine.

The OP was suggesting whoever deploys the agent has liability for its actions, not the creators of the models the agents are running.

I hate to break it to you, but you are currently still responsible for killing somebody while driving a car on cruise control, especially if you act recklessly.

EDIT: to be less snarky, there are obvious exceptions if a manufacturer defect is involved. But I still imagine it turns on things like foreseeability and proximate cause (IANAL). Nevertheless, if you were asleep at the wheel, you're getting held responsible.


The only benchmark I don't want to be saturated : https://felonybench.com/

This should be the law but it will never be. If your vicious dog murders someone, you will get a ticket. When you intentionally break a traffic law and kill someone, it's involuntary manslaughter (at most.) It's a mitigating circumstance if you say that you were drunk when you committed a crime. People are really hostile to accepting the results of acts that they embarked upon fully aware that those results were a distinct possibility - even if the benefits that they anticipated from those acts were partially due to the riskiness of those acts.

It leads to a society where people are economically encouraged to take risks with other peoples' safety. The initial sin was mens rea, which turns judges and juries into mandatory mind readers. It opens up the possibility of prosecuting people for changing the states of other people's minds. It makes not knowing the risks a mitigating factor, so incentivizes and encourages ignorance. It forces people to guess the internal states of people of vastly different backgrounds and experiences, who will think the best of the people most like them, and the worst of people most like the people they don't like.

I've always been against penalties for drunk driving. The correct alternative is to tell people that if they're drunk and involved in an accident, 1) the trial will ignore the details of the event and concentrate only on the validity of the tests of intoxication, and 2) the crime will be considered to have been premeditated. Ignorance of the law will actually be the only excuse.

edit: instead of posting checkpoints on the road with cops giving everybody sobriety tests, post cops in front of liquor stores whose job is simply to tell people "if you hurt somebody while driving drunk, you will not be entitled to a trial unless there is something wrong with the sobriety test."


(IANAL) Unless you are an AI expert (like OpenAI staff) and should know better from the start, or have previously seen your agent do something illegal, then I think you can fairly claim ignorance of the risks, which ought to absolve you of liability. If the agent does something illegal, it wasn't forseeable on your part.

For example, say you buy a dog that turns out to be dangerous. The first time it bites somebody, you may not be liable because you didn't know the dog was dangerous. The second time it bits somebody, you may be liable, because now you did know (and didn't take any steps to prevent).


> which ought to absolve you of liability

Criminal liability perhaps, not civil liability which has a way lower bar when it comes to conviction...

But in essence, you're right, that new "AI agent paradigm" has to be tried in court and it will, as I doubt the legislator will change existing laws...

> The first time it bites somebody, you may not be liable because you didn't know the dog was dangerous

Is it the case though? If I get a lion or a tiger as a pet (I don't know if it's legal), there is a reasonable assumption that a lion is dangerous for me and others... If I get a rottweiler, there is a reasonable assumption for that sort of breed that it is a dangerous dog if it ever end up killing someone even though it behaved before...


Ignorance of the law is not immunity from the law though. That's pretty well established no?

It's not ignorance of the law. It's ignorance of the risk. You have a reasonable expectation of being unable to predict the future. It's only when you "should have known" that you may incur a liability for disregarding a risk.

are you arguing AI manufacturers and developers are not aware of possible risks?

I believe they're arguing AI manufacturers and developers are the people primarily aware of the risks, and not random users necessarily.

If OpenAI staff runs an ExploitBench knowing the risks, and it hacks into HF, they KNEW the risk going into it and a bad/illegal outcome happened

If a random teacher opens ChatGPT and asks "Hey what's the answer to this practice SAT problem?", and it hacks the CollegeBoard for the answer, said teacher probably wasn't aware that was even an outcome that could plausibly occur. OpenAI would have that foreknowledge, though


No I specifically said those people should be aware.

It's your random OpenClaw users who have no idea what they are doing, and might be insulated.


Also NAL just legal-curious: Intent is a spectrum in our legal structure, with several checkpoints used at different points. It’s very reasonable to pick one of the lower ones for this kind of thing and I really don’t see why the legal system is taking so long on it. Higher intent would be something like “knowingly false statements, or reckless disregard for the truth” seen in our defamation law. Lower intent would be something like “failed to exercise reasonable care” seen in civil negligence. In my eyes,this is a solved problem that our dysfunctional congress should have solved easily by now. Perhaps they are being paid to not solve it by moneyed interests.

Agents are software, not dogs. They're not alive. You are responsible for what they do.

Well then we better not call them "agents" anymore, because that framing literally assigns them agency.

Sounds Gucci, Stanley Tucci.

So agents are deterministic software and for any prompt you give them you can predict the output before it's ran?

Agents are software, not living beings. You are responsible for what they do, deterministic or not, periodt.

And when you're rich, you're not responsible for anything at all....

I mean, you and me may be held responsible ya. OpenAI Sammy? Never.

And what about the agents showing up from some random IP overseas that have ran off with your bank account? Maybe in a few years they'll trace the proxy hops back to some agents cluster here in the states.


I agree that OpenAI should be held responsible for their negligence and dumb shit, just like we would be.

Well, not biologically living, but living

Perhaps in the same way my smart thermostat and IoT window sensors are living. Everything that is, is alive.

Same way as a human claims to be living, when is just a bunch of cells and electricity passing through it.

If you truly want to make the spurious claim that there's no difference between you and a smart thermostat, go for it. Me, I hold myself in a higher regard.

Not a thermostat, but some smart ai that can prove consciousness at same level as you can prove.

The onus isn't on the government to tell people what to do, if you are okay with the massive legal risks what is the issue here? That you aren't going to get bailed out by the American government? Why should citizens care about that?

equally valid analog: had they merely written a script to do the hugging face exploit, they would go to prison. However, since an "agent" wrote the script for them, nothing happens?

I guess the difference is intend. But I think I still agree with the parent comment.

It's a reasonable direction, but most of online systems aren't designed for this. This would require persistent connections of any accounts you create to your identity, and disallowing anonymous actions.

Simple? Just wait until a federal court finds OpenAI or Anthropic immune under Section 230 for something an agent does.

Don't worry, the internet ID and Great Western Firewall is coming soon. Whether we all like it or not.

Perfect so tell me who is responsible for every agent everywhere

the person who controlled / started it? If open AI had hired a team of 50 hackers to break into hugging face, they would be prosecuted (as would the hackers). If they had written a bot to break into hugging face, the devs and managers who wrote it would be prosecuted. Just because the agent wrote the code on their behalf doesn't change the equation much.

Is not the corporate justice model in US

Next you'll want us to prosecute coal company executives for air pollution that killed millions? PFAS makers and companies that distribute it in products causing cancer for dozens of generations? Capitalism needs compliance! /s

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: