One of the most decorated minds in AI is striking it alone. Richard Sutton shared the 2024 Turing Award for founding modern reinforcement learning. On Monday he said he will leave Keen Technologies, John Carmack’s startup, to start a new one, Oak Lab.
He announced it clearly in X. He praised Carmack and Keen, and later said that he and collaborator Khurram Javed had “broken off to start our own startup” to follow “a slightly different path toward understanding intelligence.”
His diagnosis of the field is forceful. Current deep learning methods, he wrote, are “weak and inefficient, needing not further tweaking but fundamentally new ideas and extensive reworking.”
Learn from experience, not data sets
The central argument of Oak Lab is about the origin of intelligence. Sutton has long maintained that it is created and maintained from runtime experience, not a clean, human-curated data set.
That distinction matters more than it seems. Current models learn from the data that people have collected, cleaned and filtered. The real experience is more complicated. Some of it is predictable and a lot of it is just noise.
in the laboratory first research positionSutton and Javed put numbers to the problem. The standard optimizer, SGD, cannot differentiate between them. Neither can cousins like Adam. It distributes the blame for each error across all its parameters, so it silently absorbs noise.
His solution updates an old idea of Sutton’s. An algorithm called IDBD, and a new neural version they call NetworkIDBD, learns to selectively allocate credit. Reward only the signs that really predict something. In his tests he learns the real pattern where SGD drowns in garbage.
An agent that runs on 20 watts.
The goal of all this is efficiency. Its methods learn from a flow of experience, step by step, without storing or reproducing data. Oak Lab says that requires orders of magnitude less computing and power than the current approach.
That leads to the lab’s declared holy grail: a trillion parameters Agent that learns and plans in real time with 20 watts. Twenty watts is approximately what the human brain consumes. Current frontier models are trained once, in megawatt-consuming data centers, and then remain frozen. Sutton wants a system that never stops learning, with a pinch of power.
The opposite bet
Sutton has spent his career arguing against the grain. His 2019 essay “The Bitter Lesson” is constantly cited in AI. His textbook with Andrew Barto trained a generation of researchers. But he doubts that scaling up pre-trained linguistic models is the path to real intelligence.
That puts him in interesting company. Yann LeCun has presented a similar case, leaving Meta’s orbit to bet Billion dollars in world models instead of larger chatbots. AlphaGo’s David Silver has placed his go for a different route. They all think that a machine should learn like a child does, from experience, not from a frozen snapshot of the Internet.
The moment suits the mood. The AI race has quietly advanced It stopped being just about the biggest model. Cost and efficiency now matter as much as scale, and researchers are investigating how models really reason instead of just making them bigger.
Whether Oak Lab delivers is another question. It is pursuing a goal that the entire field would love and no one has achieved. But Sutton is betting a historic career on it. He believes the future of AI looks less like a bigger brain in a bigger building and more like a small one that never stops learning.






