> Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.
This makes sense and is very insightful, thank you. But with that solution it seems that only LLM companies will be in the position to create new languages.
"LLMs dont create anything new, if programmers stop reading the code technology will be forever frozen to 2022, no new programming languages, operating systems, concurrency primitives, databases, networking protocols, UI frameworks everything will be based on the training data and future generations will forget about all the primitives we now take for granted.
If someone creates a new programming language/ framework or new better way to do async or whatever, no one will use it because it is not in the training data and it wont take off because everyone is using LLMs. It will be like using the same Lego pieces over and over."
Mmmm idk. I used an LLM to write assembler in my made up API so I don’t think what you say is true. The value in LLMs is precisely that they are not just regurgitating training data, but rather inferring concepts extracted from trained data. If a programming language used concepts completely disconnected from existing paradigms you’re probably right… but that would also be quite challenging for humans to learn to use, since by intrinsic construction it would also be widely separated from human language.
Esoteric languages like BrainFuck are esoteric and difficult precisely because they go out of their way to eschew conceptual links to existing languages or paradigms.
So if you invented a new type of esoteric language with arbitrary syntax and strange operators (not sure how you’d do that, exactly, iirc all fundamental binary operators are known) it might be impossible to use with an LLM even if the user manual was in context… but aside from that, languages and the underlying concepts are extremely generalizable.
Yeah. I have my own file formats for music, pixel art, and levels in a little game I’m building with the kids. LLMs are really good at understanding these proprietary formats that exist nowhere else in the world except on my old laptop.
> I used an LLM to write assembler in my made up API so I don’t think what you say is true.
I think you are underestimating the amount of data/context that is required to use a battle tested general purpose language, for both humans and LLMs. The ecosystem requires official docs, stack overflow answers, blog posts, tutorials, existing source code, subreddits, issue/PR discussions of undocumented features, obscure mailing list threads with rare insights, books, youtube videos, benchmarks, tests suits... and the ecosystem of libraries for the language that also need their own official docs, stack overflow answers...
You also need the collective audit by the community and assurance that this language has been used in production by countless others.
In the age of intelligence on tap, any new language not specifically designed for human-only use will be born into the world with an attendant plague of documentation. And languages are shockingly generalizable. There are few things in any language that cannot be done in any other. Even those are computationally equivalent to some other set of instructions.
What might be the case though, is that new languages will use more context to think about until they are well represented in the training set.
LLMs dont create anything new, if programmers stop reading the code technology will be forever frozen to 2022, no new programming languages, operating systems, concurrency primitives, databases, networking protocols, UI frameworks everything will be based on the training data and future generations will forget about all the primitives we now take for granted.
If someone creates a new programming language/ framework or new better way to do async or whatever, no one will use it because it is not in the training data and it wont take off because everyone is using LLMs. It will be like using the same Lego pieces over and over.
What if programming languages, operating systems, concurrency primitives, databases, networking protocols, UI frameworks are already good enough, and the innovation lies elsewhere?
You can do a lot of cool stuff with the same lego pieces.
What does an advance in music even look like? Shifting tastes for pop music? Or new techniques? New music theory? Or just experimentation?
Considering all music is subjectively influenced by the culture in which it's born (see the difference between Asian traditions of music, European traditions of music, African traditions, and traditions of the Americas) not even all of those have a given structure that is present today like the typical 4/4 and have polyrhythmic and multitonal structures by design. The fact that everything on the radio has converged towards 4/4 165bpm major chord progressions is evidence of that cultural phenomenon.
But what if the fundaments of all these, in the human produced literature, actually contain hidden circularities and holes which make very hard the progress?
IMO for the moment the greatest value from these AI tools is that we can start an audit and hopefully proceed on a saner foundation, after we use the tools and think about it.
This is different than too many AI generated proofs or panic reactions from the academic system with its stupid incentives.
False dichotomy. Chess/Go can still be played between two humans and there is allot of value in that because humans compare each other to other humans, when you see a skillful Grandmaster play you know they are good compared to yourself or the average human, that is why people still watch, play chess/go and train hard to get good. Programming is different because you are creating something not necessarily trying to win a game.
Most programming tasks are exactly like that. Is this agent able to complete this task? Is this agent able to optimize a kernel beyond previous attempts?
Of course some are subjective and that's where progress is harder, like "Is this website pretty?". But for tasks that can be objectively measured, LLMs will go beyond human level, just like with Chess and Go.
My point is that LLMs depend on training data so the code they produce will be stuck in 2022, no new languages, techniques beyond that because new techniques are not in the training data (at least not enough of it for training because most coders are now using LLMs).
Chess/Go continues to progress because it is primarily a human vs human activity, people will always be learning to play chess and chess will continue to develop.
> Chess/Go continues to progress because it is primarily a human vs human activity, people will always be learning to play chess and chess will continue to develop.
AIs are not continuing to get better at chess/go because humans continue to play at levels far below themselves who discover new techniques. They get better because they play against other AIs and discover new techniques that have a higher win rate that way.
I would bet that even if humans stopped playing chess/go and people were still willing to run these AI models against each other they would continue to get better.
I am not talking about the advancement of AI, I am talking about the advancement of chess.
Two things can be true AI drastically contribute to the advancement of chess and humans playing against each other also contribute (even if slowly) to the advancement of chess as it has always been since the invention of the game. The point is that because chess is primarily a human vs human game humans will always have the knowledge of chess, unlike with programmers who are giving it up to prompting, and programming being much more complex than chess (checkmate and win) will be stuck in 2022 because of the training data.
> Pre training data is in large part synthetic these days
How much of that data can lead to innovation? Can you predict all innovation map it out on paper.
> Computer Chess progress has nothing to do with human vs human activity.
The point is that humans will always be learning chess because it primarily a human vs human activity they will be contributing games to the chess database, unlike with programmers who are stopping to code and only prompting, generating code stuck in 2022.
> AlphaGo Zero used no human game data at all.
Sure, but that instance of AlphaGo is still dependent on its training, its intelligence, so it is a question of is that the best and only way to win a game of Go. Just a few weeks ago, a Go Grandmaster found a way to beat one of the strongest Go AIs.
So a specific instance of an LLM might be the smartest based on what we know and need today but that is not the limit of how far we can go, this is why it is important for humans to always have an intimate connection with the code, math, science, chess etc for progress to continue.
> Sure, but that instance of AlphaGo is still dependent on its training, its intelligence, so it is a question of is that the best and only way to win a game of Go.
If this were true then it would be impossible for these models to ever exceed the top human level as there would exist no training data that allows them to exceed the top human level.
However, despite there being no training data on ability to beat the top humans these models have achieved it.
> this is why it is important for humans to always have an intimate connection with the code, math, science, chess etc for progress to continue.
This is just you wanting to remain relevant rather than actually based on evidence.
> If this were true then it would be impossible for these models to ever exceed the top human level as there would exist no training data that allows them to exceed the top human level.
Of course AI exceeds humans at chess, I never denied that. I am saying because chess is primarily a human vs human game, humans will always be learning and playing chess, their games will add to the chess knowledge base, AI also adds to this knowledge base. But programming is not primarily a human vs human activity so there is a risk programmers will forget how to code and all software will be stuck in 2022 because of the training data, this stifles innovation.
> I am saying because chess is primarily a human vs human game, humans will always be learning and playing chess, their games will add to the chess knowledge base
I guess I'm contesting that idea you are putting forward that the data from the games these humans are playing, which are at a vastly lower level that the top AIs are meaningfully important for helping the AIs to improve at the top level.
Would more people learning their times tables be helpful for top mathematicians in their fields to get better at the frontier of maths? Probably not right. Same applies here.
Would AIs advance at the same rate for the top level of chess in a world where humans completely stopped playing chess vs the world we have today. I would say they would as the human level data is of limited value to the frontier which is dominated by AI and AI game data, you are claiming that it does.
> But programming is not primarily a human vs human activity so there is a risk programmers will forget how to code and all software will be stuck in 2022 because of the training data, this stifles innovation.
Does it? Or will AI be able to run its own experiments and find better/more efficient abstractions that propagate because they are better and this will find its way into training data for future AI.
> I guess I'm contesting that idea you are putting forward that the data from the games these humans are playing, which are at a vastly lower level that the top AIs are meaningfully important for helping the AIs to improve at the top level.
I never said human games are meaningfully important for training AI. Human games are still important for the advancement of chess, maybe Magnus Carlson can learn from games between two Super AIs but most humans still learn from games by humans, Grandmasters are continuously developing the opening, middle-game and end-game systems, adding to the chess knowledge base. Every serious chess player still reviews and study games by prominent Grandmasters, every serious chess player documents their own games, writing down every move. All rated games are recorded and added to the chess database that every player can review and study.
>Does it? Or will AI be able to run its own experiments and find better/more efficient abstractions...
> Human games are still important for the advancement of chess
Are they? Why?
For a human vs human game sure but at the very top level? No of course not because it's all done by AI.
> Only if it is in the training data.
This is trivially not true, as how has AI managed to become better than humans if the knowledge of how to do so never existed in the training data.
We are well past AI can't do X unless X is in the training data. If your claim were true then AI could never surpass human expertise in any field because by definition all the available training data will at best be at the current human frontier and not beyond.
> For a human vs human game sure but at the very top level? No of course not because it's all done by AI.
Glad that you finally agree,this is what I was saying the whole time.
> This is trivially not true, as how has AI managed to become better than humans if the knowledge of how to do so never existed in the training data.
AI can do more work, faster, AI it only needs sufficient compute and data. But that does not mean it is more intelligent than humans, it still uses the same code, algorithms, frameworks, protocols etc etc that are in the training data, sourced from human open source code on the web.
Human games databases are completely irrelevant to the strongest chess engines. We are ants in comparison. The Go thing you mention is playing against handicap. Sorry I won't go into more detail explaining why your premises are wrong, I'm tired of this discussion.
Human games databases might not contribute to AI training but they do contribute to the chess database itself like it has always been, AI doesn't change that fact.
LLMs receive new data via input context, not just training data.
Thought experiment: How effective will 2026 LLMs be for humans in 2526?
It's not game over just because 500 years are missing from the training data. The important question is how well can 2526 humans make culture and knowledge navigable to LLMs via tool calls.
Today's LLMs might need for example sub agents to translate to 2526 English, sub agents to read 2526 docs.
It's _really not clear_ whether 2026 LLMs will be useless. To believe that reflects an enormous misunderstanding.
> LLMs receive new data via input context, not just training data.
Be more specific about the "new data". If everyone is using LLMs for work (generating code), especially the juniors who won't get the chance to learn from first principles, LLMs will be training on the data they generated. How will new code enter the system at large enough quantity that it can be used for training?
> It's _really not clear_ whether 2026 LLMs will be useless. To believe that reflects an enormous misunderstanding.
They won't be useless, they will just be frozen knowing only whats in their training data. No new programming languages will emerge, in 2526 they'll still be using Rust and javascript, same exact code from 2022 which dominates the training data.
The "new data" is: person A prompts an LLM to create or modify a tool, person A distributes code person B, person B's LLM uses the tool via docs/help/error. That is a direct path for an LLM to "know more" from a human than what's in its training data.
If we get a new programming language not in the training dataset, we could give an LLM a decent compiler with compile errors, and some sample code and it would be able to write code in the new language without training.
That may be true for popular projects with allot of contributers, but for small projects, being closed source in the age of LLMs is not a bad idea especially for server side code, it can be more secure simply because the bots don't have access to your code and can't do analysis on it.
A trained AI I believe is capable of reverse compiling from machine language back to the source code. Access to the source code is not necessary for them I think.
No sarcasm: play a different game. If you play their game, of course you'll suffer. But if you play a game they're repulsed by, you can win a life they'll never have access to: one filled with people who love and respect you and an environment where your best self and life can emerge.
The parent comments are talking about how the game is negatively impacting society. Your suggestion is to check out and leave society to them. One of the many problems with that suggestion is that it's hard for your best self/life to emerge in a bleak future.
I think it's possible to "pick and choose" to some extent. I've found most of the bleak future stuff goes away if I turn off my phone, but since I need it to contact people in 2026, I try and give things their due, and, to the best of my ability, ignore the more unpleasant aspects of modern life. I think the landline telephone is a good example. It was fun to chit-chat with friends on the phone, but when a telemarketer called, I either hung up immediately, or said "stop calling," then hung up. It can be the same way in modern society. Weight the conversations with warm-bodied people higher, and stop scrolling as much (to put it overly-simplistically).
It’s done wonders for my free time and mood by deleting my Reddit account and restricting HN to toilet time. Instead of reading rage bait American news I do a daily scan of BBC. All those are sources of negativity and stress that bring nothing of value to your life. I invest my newfound time and energy into my family, hobbies, and health.
That's not what I said. I didn't say anything about leaving society, I alluded to leaving their society. That's the thing: people keep giving these guys money and attention. If you take that away, they become more than irrelevant. Simplified: if they're going West, you go East.
The more "small time" people reject this narrative that we need a bunch of elitist white guys shoving their mediocrity down our throats, the faster we rebuild a society they're barred from accessing.
I didn't say anything about leaving society, I alluded to leaving their society. That's the thing: people keep giving these guys money and attention. If you take that away, they become more than irrelevant.
Can you explain how to do that without either drawing attention to them via discussions like this or leaving society altogether? I don't personally fund a16z, but my taxes do. Calls to coordinated action to stop giving them money also count as attention.
Leave Babylon behind and go build local businesses and communities in what they refer to as "flyover country." That's what I'm doing. Easy? Fuck no. But does it take their relevancy away from them? At the scale of everyone who isn't a coastal barnacle: big time.
They view the rest of this country with disdain and indifference. I'm suggesting that others quit looking at these people as lords and sages to be placed up on a pedestal. Most of them are quite incompetent, skill-less, and frankly, not cool enough to want to hang out with (why they're making "jokes" about there only being two jobs left in the world in the future: Anthropic employees and hookers for Anthropic employees [1]). I'm sorry but I'm not taking direction from a bunch of dorks who can't get laid without paying for it. And I'd take a wild guess that given the climate, the lion's share of America is with me.
What does flyover country have to do with the present discussion? Kalshi is just as big a problem (or bigger even) in Nebraska as New York. Taxes going to Anduril happens whether you live in Malibu or Missouri.
You can, of course, distance yourself from their world and work on things they would never see as valuable. However, you can't check out of the most grand game we're playing, the game of life and existence. Those people desperately want to toy with things that could eventually prevent you from being able to live a life where you still have food, water and shelter. You can ignore it while they're still not all-powerful, but as their wealth and influence grows, you will eventually have to confront it.
Then so be it. Hopefully they've considered all of the angles and can defend themselves against people being radicalized to the point of hunting them down.
Turns out their obsession with moats was prescient.
That’s a cop-out. It’s the most powerful “players” who are making the “game” what it is, and they make the rules be whatever benefits them.
It would be pretty reasonable to hate the player who brings a machine gun to a stadium, kills every opponent to start scoring and when people complain says “Nothing in the rules says we can’t kill other players to score. If I didn’t gun everyone down, the opposing team would’ve. Don’t hate the player, hate the game”.
That is like hating a hacker who hacks into your vulnerable server that is outdated and allows ssh password login. If there is vulnerabilities in the system someone will eventually exploit them, sure not you or me but there is always someone out there that will.
That same hacker can choose to responsibly disclose the vulnerability through the proper channels, or take advantage of it to benefit themselves at the expense of others. Which one they choose objectively determines the quality of person they are.
“That is like hating a thief who robs your home which has an old lock and a window you forgot open. If there is an open window in the house someone will eventually climb through it, sure not you or me but there is always someone out there that will.”
It is perfectly reasonable to hate the person who robbed your house or hacked your server. “If that someone hadn’t, another someone would” is in no way a defence. It is not OK to actively harm someone just because they didn’t protect themselves.
Yes. My point is that until you learn to lock your door, thief will enter and steal your property. You can hate them all you want, that wont change the facts or magically lock your door.
Unfortunately the thief has introduced legislation that bans any better locks, and they have a magical veto that allows them to strike down any attempts at changing said laws.
I prefer to live somewhere where I can be relatively sure that forgetting to close my windows doesn't get me robbed. That means punishing those who take advantage of open windows. Turns out deterrence can be quite effective.
Security is always about tradeoffs and in a functioning society you don't need to have perfect security (and pay the extreme cost required for it) because we do indeed put blame on individuals that abuse the system.
It’s not the players we should hate but the ones who make the rules of the game. Specifically not the ones who make the rules for the greater good but for themselves.
“Also, I have no intention of acknowledging that not all people have an equal chance at winning the game, nor doing anything about it for that matter.”
Allot of people claiming it's because of increased load from AI. If that's true why do their systems sometime take hours to come back online, just restart the servers and you're back online, it's not like a nuclear reactor.
That was mostly Bayesian probability filters, followed by SPF and more recently DKIM.. Maybe you could consider the Bayesian filter databases being trained as ML, but I feel it's a far way from our typical AI/LLM surge
Not really, though. Sequences of tokens are predicted as {spam, not spam} via training on pair sets. It's like a token predictor where the output vocabulary is two tokens.
Well, I just said that. Since what is predicated is one of two tokens <is-spam>, <is-not-spam>. Since spam is not made out of those tokens, it doesn't generate spam.
(Even spam could be encoded into those two tokens via binary or whatever, it still wouldn't because it's predicting a classification, and not the next sybmol; i.e. not simply the next bit in a after a string prefix, but a 0 or 1 value indicating whether the bit string is spam, which comes from an asserted external judgment.)
It's all the same sort of thing though: some functions trained to predict a value associated with an input.
You did say that. And then despite that you reached a conclusion that I don't see as compatible.
There is a massive qualitative and quantitative gap between the two. Even if you had a real token predictor for the word spam, you would need a thousand of these to touch LLM capability. But it's not a real predictor. It's not locational. It's something vastly weaker.
> It's all the same sort of thing though: some functions trained to predict a value associated with an input.
Now you made it even worse. If not for the word 'trained' this would be every program. If two programs fall under the umbrella of machine learning in any way, that kind of logic treats them the same. What would you consider a "different kind of AI"? Your definition there fits basically the entire field.
It's true however I do feel that when talking about scams AI can be used offensively way easier. Using it for things like the prevention of phone scams seems quite a bit more difficult to me to get accurate.
I think you are taking that line of reasoning a little too far. The point of a scam is to trick people out of their belongings. There may be certain advantages of targeting and filtering for gullible people, but it's not the point itself.
"proper scam"; if you get too many hooks, you're not going to be able to reel them in. I'm not defining a scam; if you create a really convincing scam, your problem becomes both having too many people and drawing to much attention.
The main fuel is the "need" to have browsers that can execute code, coupled with the ability to transfer money via the internet. JavaScript was a mistake, e-commerce was a mistake, and so is everything else that came after.
Let's ask Vint Cerf whether the web was meant to run logic, or just be a document format. I bet no matter what he answers, the world would be like the GIF pronunciation "I don't care if the original creator said it's pronounced like giraffe, I'm gonna say it like gift because I KnOw BeTtEr"
Internet 3 when (because I know that academia does already have an "Internet 2", although I bet due to Internet 1 security issues, Internet 2 is already compromised too)
Douglas Adams put the source of the rot a bit further back:
> This planet has, or had, a problem, which was this. Most of the people living on it were unhappy for pretty much of the time. Many solutions were suggested for this problem, but most of these were largely concerned with the movements of small, green pieces of paper, which is odd, because on the whole, it wasn't the small, green pieces of paper which were unhappy. And so the problem remained, and lots of the people were mean, and most of them were miserable, even the ones with digital watches. Many were increasingly of the opinion that they'd all made a big mistake coming down from the trees in the first place, and some said that even the trees had been a bad move, and that no-one should ever have left the oceans.
> browsers that can execute code, coupled with the ability to transfer money via the internet
But those are completely unrelated topics. You don't need a hint of JavaScript to do online banking.
And you don't need online banking to send someone your credit card details or banking information. Any kind of communication can turn into a scam. You can steal money over the phone or by mail.
This is not surprising. Would you rather have handmade Italian shoes or factory made shoes sewn by a robot. Humans assign quality relative to themselves.
A chess game between to bots is boring, no matter how strong the bots are. Humans generally compare themselves to other humans and that's part of how they value.
> Would you rather have handmade Italian shoes or factory made shoes sewn by a robot.
The latter, of course, which is why that's what I own. Aside from that, this is an absurd and grossly dishonest analogy -- that's not at all the choice here.
Some of the most fascinating games are between Stockfish and Leela. (e.g., https://www.youtube.com/watch?v=h_84B_VwayU -- this is narrated by National Master "Jerry", one of the most respected chess narrators on YouTube.) People who don't know this know nothing about chess. (And only immature clueless fish respond by challenging someone to a game. Lack of knowledge of chess was already established by their ridiculous claim. BTW, he has radically changed his response , which is now very different from his original claim. It's not my misunderstanding, it's his bad faith.)
Some of the most stupid things I have ever read are from those suffering from radical anti-AI hysteria, and it is becoming a primary means by which I judge people's intelligence. ("radical" is a key word here -- there are numerous legitimate reasons to be critical of numerous aspects of the AI boom.)
I think you misunderstood what I said. Let me rephrase.
Would you rather watch a chess tournament between two of the best human players in world, or between two of the best bots in the world?
I know that chess is deeper than this, its about the positions that are reached and what each game adds to the existing knowledge and in that case it doesn't matter if its a human or bot. I understand very well that there is allot a Grandmaster can learn from games between two strong bots.
But here I am talking about human nature on how we value things, if a human can do something that I can't do than I value that more, I know Usain Bolt is faster than me, I know he is extremly fast compared to myself and other humans.
I dunno which I’d rather have, but I do know the only kind I’ve ever bought is the factory-made kind. So if the aim is appealing to a wide audience, you’re arguing involving robots in the production process is key?
This makes sense and is very insightful, thank you. But with that solution it seems that only LLM companies will be in the position to create new languages.
reply