He's talking about learning that makes you operational. If you can speak with someone, then you've learned language. If you can get a computer to do something, then you've learned some programming.
The equivalent for biology would be to grow a plant or a few plants and animals successfully. That's operational at a certain level, you could also be operational at a lower or higher level.
No, it doesn't necessarily need to be applicable, it just needs to be testable. The size of the Earth for example may not affect you directly in any way, but it's something you can verify to some degree of precision. But if you just read about it, can you really say that you know it? What if all you believe about the size and shape of the Earth is from what you've asked an LLM? Do you still know it?
EDIT: Perhaps not the best example, because the size and shape of the Earth are data that are repeated often enough that an LLM would be unlikely to quote it grossly incorrectly, but I think my point still comes across.
I’ve been burned by this enough that I no longer say that I’ve learned something if my only interaction with it is explanation from books. You can get snippets of knowledge and a framework of understanding, but true learning only comes with deep interaction of the concepts (practice, simulation, experiments, observations) and not merely reading.
Curious, which model do you use for Codex?
I'm very happy with the solutions '5.5 high' finds. It's like it understands exactly what I mean and it also anticipates all sorts of situations.
Before I used '5.5 medium' for some time and it was a bit underwhelming. It may sound funny but it's like it didn't care that much to do a good job.
I use GPT 5.5 High Fast, I often benchmark versus Fable (and previously did versus Opus) and it's night and day.
Claude still (and has always) writes far too much code to fulfill a given spec or plan. It misses edge cases and is generally far too verbose.
Claude also is (and even more so with Fable) super tokenmaxxing, i.e. it seems tuned to use the max amount of tokens per task, whereas Codex will simply get your job done as you specified with the minimum fuss and tokens.
Codex feels way more steerable and just more "professional" as though I'm working with a seasoned engineer, versus someone smart but over excitable, like a super smart associate engineer.
I'm also unemployed. So far the models that I've used the most are Kimi and GLM. I haven't done that much agentic coding though, I've mostly used them for studying math and general conversations and I'm generally happy with their performance.