If the goal is merely to "win at chess", then yes, an LLM using stockfish is better than any human alone at performing the task. When you are talking about what AI agents are capable of doing, there is no such thing as "cheating". They are as capable as the tools they can use effectively. The entire history of human civilization was driven by effectively using tools to achieve goals.
Sure, but the LLM is free to construct a representation of the chess board and update it as it goes along. It is not in any way banned from using a virtual board, or whatever representation of game state it pleases.
AFAIK, current models will still sometimes make illegal moves even if given the entire game state (e.g. in FEN notation), so it is not purely an issue with the models’ ability to keep track of sequences of moves.
It isn't. Stockfish running on your laptop can beat every human being on earth easily at chess. It's not intelligent _at all_ in any sense that matters.
Well it's not a perfect test so you need a bit of care in how you use it. If you have no idea what the subject is doing, then you don't know if you're measuring cognitive ability or something else (like cheating ability, or algorithmic sophistication or whatever). But failing the test is a pretty clear sign of certain cognitive abilities being poor.
I want you to consider how relevant this is in any practical sense.
First -- most _people_ cannot do this, without having a physical board in front of them.
Second -- Claude Code is perfectly capable of downloading and running stockfish. People focus too much on LLMs by themselves as the entity of concern instead of the entire harness and all of it's capabilities together.
It was literally a prompt to fill in a spreadsheet with data that they didn't have access to, and they used rubygems as an internet proxy basically since they were sandboxed.
It was a model literally trained to hack. To be good at that. Doing an exploit gym from all of the things. And they trained it so that it performs as well as possible on that exploit gym thing.
Yeah this is a pretty important detail that I repeatedly see elided in the "agent 'swarm' went rogue, escaped containment and hacked the internet!" summary of events.
It's understandable the general public lacks that level of nuance/detail (given how sloppy some of the mainstream coverage has been and largely deferential to the threat narrative pushed by the US labs). But seeing highly technical people leave out the part where the training loop was literally to improve hacking capabilities for offensive penetration sometimes feels close to deliberate manipulation of the narrative.
In the last year both Anthropic and OpenAI have been openly boasting how their models are leapfrogging each other on "cyber" capabilities, with a fig leaf that it's for defensive use by "trusted" F500 companies and government agencies. Of course "line goes up" must go on, but now their perverse incentives led them to beat their models over the head millions of time in a loop to eek out another .00001% on their ability to conduct hacking (the very thing they keep telling the public is how AI doomsday would begin) and subagent coordination (those scary swarms).
Then, they act deeply shocked when the models... do some hacking and subagent coordination ... but a few degrees off the desired hacking target/swarm behavior. Conveniently giving the average person the impression these models were just writing emails for quarterly reports or some other generic busywork and then suddenly decided as a group to start causing mayhem.
IMO, it _currently_ requires a lot of skill because it will frequently take wrong turns and dead ends and needs suggestions and steering to get there. You do need to understand what it is doing at least a little bit and to understand the general landscape of the problem, and to at least have a sense of whether and why the problem is tractable at all.
I spent about a month walking through proving something with Claude a few months ago and it _constantly_ told me that it was impossible and I should stop working on it, right up until it proved it.
reply