> say "better than 95% of humans on 95% of intellectual tasks" but if we used that definition we already have AGI and almost no one thinks we have achieved AGI
What matters isn't 95% of humans, it's 95% of actual professionals. Benchmarking an AI accountant against people with zero accounting experience is worse than worthless.
That's what 95% is intended to capture. That some level of expertise in an area is should be captured by 95% sample of the population, you could push it to 99% of 99.9%, but you want it to be quantified by a number to avoid arguments that something is not AGI because obscure field or expert exists that AI can not do.
95% better at 95% of the population is already approaching ASI. One could even argue that AGI is 50% better than 50% of the population.
All that said, this over rotation on benchmarking misses something critical. What is general intelligence? We assume that humans have it and we assume it is captured by benchmarks on "intellectual tasks", but it is probably the case that general intelligence is based displayed by judgement on uncertain outcomes. Benchmarks by their very nature have certain outcomes, they have a wrong and right answer.
Test AIs on questions we don't have the answers to and there is no clear right answer, but there will be at some point in the future. What will the economy do? Which US senators will be be re-elected that polling correctly suggests will not be re-elected. Which US senators, currently not in office, will actually pass bills representing the wishes of their voting base? What published papers will be seen are groundbreaking in 5, 10 15 years? What approach to unifying physics should be taken?
Narrow pre-LLM models that spit out content and ad recommendations have never been capable of also suggesting, let alone implementing, self-improvements.
That point about Elo is irrelevant. Games on Fishtest use bullet-like time controls (60+0.6). Significantly shorter than 1 min per move.
As originally suggested, when running modern Stockfish at 60 seconds per move the games overwhelmingly tend to draw. Even when the older engines make slightly worse moves.
> Nothing to do with some CEO's evil plans, but more to do with the hundreds/thousands of mid-level "not my job" or "doing my best" workers who are actually in charge of handling my data.
The CEO is responsible for what their company does. If a major breach can occur through the oversight or "incompetence" of one worker, the CEO has already failed, whether through negligence or malice.
In China, CEOs go to prison or are executed. Not all the time, but enough. In the US, they are given a golden parachute and make more money at their next posting. There is mostly only failing upwards.
> How much do you think an organization like Google would spend on, for example, AI tokens or compute to detect this internally before it was found and exploited in the wild?
On average, probably not that much. What's the amortized cost of all testing, static analysis, and audit / code review, per "prevented potential bug"?
If you ballpark it as a single 3 hour downtime window per week and iid Poisson, then overlapping downtime probability of 2 providers is approximately the expected occurrence rate per 3 hours, 1/56. Not particularly surprising at all.
The law could make mergers executed provisionally for up to X years, with a binding plan to "unmerge" that must be updated every Y months. The FTC already half-does this with post-merge divestiture requirements.
Yes, that's the point, making something hard means that sometimes it will not happen. Only comparing to the cases where it does happen is missing the entire point.
What matters isn't 95% of humans, it's 95% of actual professionals. Benchmarking an AI accountant against people with zero accounting experience is worse than worthless.
reply