I think the bigger problem is that people completely inflate what a human brain is and can do. IF we have an overinflated idea of what a brain are and can do, then any llm/AI will always be an ‘inferior machine’.
There are the ‘free will’ group (ability to do otherwise, in exact same situation), that thinks we have some ability to do anything other than ordinary cause & effect in our brains. Cause and effect is automatic, can’t be avoided, and kicks in from the first bit signal that either enters our brain extrinsicly or intrinsicly. Everything from there just compounds and gets evaluated not much different from a mix of current llm architectures. The fact is that we are biological computers that have evolved many layers of computation that all leads to action states - ‘behavior’ - whether it be movement or thinking.
There’s the ‘consciousness’ group (not well defined concept), that think that being conscious is something cosmic, everlasting and super special. There are a few theories on this fluffy concept, and they mostly - to my knowledge - consist of information theory. A slightly primitive analogue would be images that comes in so fast that they become a sequence - a movie. At some point in evolution, we reach a stage where our brains detects themselves in a balanced feedback loop. This little feedback system has a ‘standing wave’ quality where input signals match our cognitive capacity, and our brains now have an ‘outer loop’ that can recognise its own activity. My guess is that consciousness is a gradient that starts very weak and very early in evolution of all life, but is not ‘special’.
Then there’s the ‘machines can’t be intelligent’ group. A bit difficult to scope what this group thinks, but ‘intelligence’ is also not very well defined in the big human diaspora, and they refuse to see a brain/body as a computing device (our whole body is a part of our ‘brain’ as everything are completely inter-connected). ‘Learning’ - or continual learning/habituation’ for current AI is still crude, but getting better.
There are other groups that hinge their anti-AI stance on the real harm the combination of Capitalism and llms generate: fear of big pscho corp - which is real, fear of loosing jobs - which is real, ‘Slop’, real and so on. These are all true, but they are all artefacts from Capitalism, as already existing problems. Without Capitalism, there is no problems in this group, but instead of attacking ‘C-slop’, ‘C-unemployment’, C-copyrights, C-climate issues and other already existing Capitalist threats, patches and behaviours, they see the real and insane acceleration of this already yucky behaviour through this new ‘AI’, and simply hones in on the new technology as the ‘bad guy’. They have drawn a line in the sand, that they will not cross. AI will never become intelligent or useful. It’s imho not that intelligent to attack the visible artefacts of a Capitalist system that causes it all. Always attack the root-cause, not the effect.
The article here tries to illude to the fact that there is more on top of the first layer of next token prediction, and that is also true. We already know that these algorithms are trained on a ‘projection-layer’ of our cognitive activities - our language and behaviour, but that they lack the evolutionary scaffold that causes that behaviour, so we could argue that AI is just a repeater of that projection only, and would never reach what our brains can compute. But that is not really true from the transformer model and onwards. Earlier ‘AI’ just (sorry for simplifying) learned entities and their relations and spit out weird content, but the transformer introduce ‘position’ between entities in the first layer. ‘position’ in our actions are a primitive temporal ability that constantly adjust output actions on the likelihood according to time, and bam! we had the first reasonable sentences. The next thing is resonance/dissonance, or ‘certainty/uncertainty’ that currently are only detected by scaling up the entity+time layers. sudden uncertainty = surprise, and the ability to react to surprise, stop->adjust cognitive scope->re-engage are the holy grail of intelligent AI.
These are the basics of an intelligent machine, but we need more. We need to run continuously, to navigate the problem-scape we are in, be able to self-adjust the amount of information so it matches our cognitive capacity and be able to return to a ‘self’ state etc (cognition is just navigation of internal states until surprises are resolved or deferred). The absolutely most crude cognitive ‘navigation’ is something like ‘moltbot’ or similar primitive architectures. It basically consist of the llm being yanked back to the “Do this” list every N’th minutes. It will become enormously better very soon. ‘Cognition’ and all other features of our brains, is just the right engineering and it will come to surpass us at some point as it already does in many areas.
I’m sure I forgot something important in this outline, but I think that the longer people stay in one or more of these denial-groups, the harder it will be to get out, the more they will sound hysterical/coping, and the bigger problems they will create for themselves later. People can believe what they want oc, but the less they’ll use AI to fight the things they don’t like in this society, the more AI will allow all the rich pschos to continue their constant sick power seeking behaviour that thrashes our world via their pet ideology - Capitalism, and we’ll all have much higher chance of ending up in a “Blade Runner” dystopian future.
Personally I hope the anti-AI groups learns to swallow their frustration, turn around and start using AI and all new technology to shape society in a better direction, organise, prevent more Trump’s, or Biden’s, prevent the existence of any of their Oligarch ‘Epstein’ fckers - the whole Epstein Class that quietly sit in the back and pulls the real strings of the president puppets. These pscho fckers WILL use AI 24/7 to their exclusive benefit no matter who/what are harmed, and we should all try to prevent that any way we can. We can’t fight back with sticks and harsh words alone…
However, that is just what I believe and want.
Very well put, and people hating on this tech really need to understand that the rational thing to focus on is how to develop it outside corporate control. I wish there were more discussion regarding training models using distributed computing or optimizing local models to be more capable while using fewer resources.
The oligarchs aren’t debating AI ethics while they’re investing billions to own and control this technology, and they absolutely do not care that there’s a fringe left movement whinging about it. Our choice is to cower in nostalgia or fight to have a stake in our future. We should fight to have a say in how this technology is developed and who controls it. People who weren’t happy with proprietary operating systems didn’t just whinge about it incessantly, but rolled up their sleeves and built their own alternatives with things like Linux and BSD. That should be the guiding mindset today as well. And thanks to Chinese companies releasing open weights, we already have a good foundation to work from.
I very much agree with your argument regarding the nature of cognition and consciousness as well. People like to mystify the brain and claim there’s some intangible soul, but that’s just what we’ve always done trying to make ourselves feel special and separate from the rest of nature. The reality is that the more we learn about cognition, the more evident it becomes that it is an emergent phenomenon that exists on a gradient, from the most primitive biological feedback loops to the stunning complexity of the human mind. While LLMs are just a piece of the puzzle, they demonstrate how a lot of the things our brains are capable of can, in fact, be understood and replicated. There’s absolutely no reason to believe that we won’t be able to make systems that are capable of reasoning in the same way as us in the future.
Yeah, my main problems with AI are more about who controls it and what they do with it rather than the technology as a whole, I find techbros obnoxious but if they didn’t have the power they do they wouldn’t bother me nearly as much
Yeah pretty much, I’m very excited for the bubble to finally pop too.
hell yeah i miss buying hardware lol
That’s a lot of words. But well done and well written. I’m on your side of the fence with all this.
Fucking garbage article. Completely inappropriate metaphor between a game with finite states and a win condition on the one hand and literally total open ended language expression on the other
The key point, which is entirely correct, is that next token predictor label is technically true in a narrow sense but it completely misses what LLMs actually do. Pre training is about matching existing sequences, but RLVR means that the model generates its own sequences and learns from the outcomes instead of just imitating the training data.
The chess analogy has nothing to do with chess having finite states. What the article is saying is that a system trained only on a specific set of games would be a next move predictor, but one that chooses moves dynamically based on winning probability uses live context for its heuristic. So the mechanism encoded in that loop is able to discover and reinforce new patterns that were never part of the original training.
but RLVR means that the model generates its own sequences and learns from the outcomes instead of just imitating the training data.
It generates it own sequences by predicting the next token based on the training data
And then it updates it’s weights based on the outcomes of those predictions.
It’s still a probabalistic next token generator.
But RLVR allows humans to define deterministic fitness tests and then continuously run the next token predictor and let the deterministic software use the results to backward propagate adjustments to the weights.
Which all presupposes that the output of an LLM is probabalistic and based statistical weights of model parameters, which returns is back to probabilistic token generation by a statistical model that doesn’t represent anything resemble knowledge or fitness. In theory, you could tweak parameters with RLVR for a millennia and the LLM will just thrash between various states of meaninglessness, because the conjecture of LLMs is that meaning is exclusively encoded in the statistical distribution of tokens.
Yes, it is a stochastic system, that’s the whole point tough. There are two things at play here, one is that the model does prediction based on the context of the current data it’s looking at rather than just whatever data it was trained on. That’s what the article is highlighting. The second part is the feedback loop such as what you see in agentic coding. The model makes a prediction, that prediction is tested against the environment, the model gets feedback, and it iterates. And that’s what grounds it in reality addressing the issue of it producing meaningless outputs. The system as a whole behaves similarly to a genetic algorithm where the solution evolves through the cycle of trial and error.
Incidentally, this is true for humans as well. This is why we need tools like the scientific method and peer review in science. People hallucinate things all the time, and when we loose the connection between the predictions the brain generates and sensory feedback we call that schizophrenia.
one is that the model does prediction based on the context of the current data it’s looking at rather than just whatever data it was trained on
That’s always true. That’s how neural networks work. The model is a statistical transform from input to output. Neural networks work by taking input to produce an output. The “context of the current data” is just the input. “Rather than just whatever data it was trained on” is meaningless.
The second part is the feedback loop
Yes, as I said, there’s a fitness algorithm and it back propagates adjustments to parameter weights. But that’s all it can do, because the model is just weighted parameters. It can’t learn facts, it can only adjust its probabilities.
The model makes a prediction
Or rather it produces a random output based on the input and its parameter weights
that prediction is tested against the environment
By something other than model itself that has knowledge of what “success” is and what “failure” is
the model gets feedback
In the form of adjustments to parameter weights
and it iterates
The training apparatus outside the model does this repeatedly, yes, under the thesis that tweaking parameter weights will result in fewer failures to the deterministic fitness algorithm. That’s a theory.
And that’s what grounds it in reality addressing the issue of it producing meaningless outputs
No. That’s a leap that has no basis. The training data is no less a part of reality as the current prompt context is part of reality. What you’re describing is that the output of the probabilistic transformer gets tested against various forms of curated fitness tests. The problem with is that the only thing one can do with the the results of fitness tests is to change the probabilities of the opaque parameter space. So you can create a fitness test for how many "r"s are in “strawberry” but the results of that test can only be expressed by weight changes. And those fine tuning adjustments are applied to an opaque network of weights that also includes the opaque probabilistic representations of the fitness tests for how many "r"s are in “perrywinkle” and how many "b"s are in “strawberry syrup”.
At no point is the LLM getting closer to learning facts, and the thesis that knowledge or skill is representable as a statistical model is unproven and seems increasingly unlikely.
The system as a whole behaves similarly to a genetic algorithm where the solution evolves through the cycle of trial and error.
Yes, it uses the same concepts as a genetic algorithm but it the representation is still the problem. Genetic algorithms for path finding are great because they have discrete actions and limited scope. Applying the same technique to fine tuning an LLM is a better use of time than manually fine tuning, but that doesn’t make it any less a probabilistic next-token generator that can’t represent stable facts and rules and where every fine tune for one input is always in tension with the fine tubes for all other inputs.
Incidentally, this is true for humans as well.
Yes, but just because algorithms are analogous doesn’t mean they are functionally equivalent. Humans also have an opaque neural network that functionally behaves like a statistical model. But we have more subsystems than LLMs do, we have more dimensions to our encoding, and we have greater self-governing and modification abilities. So while the genetic algorithm approach is useful, it doesn’t make the LLM become closer to reality, it just automates a portion of the fine tuning curation process.
Or rather it produces a random output based on the input and its parameter weights
The bias is precisely what makes it not random, but rather stochastic. There’s a very big difference here.
In the form of adjustments to parameter weights
I’m talking about feedback from the environment it operates in. That’s the actual test that allows the model to keep adjusting outputs towards a specific target rather than them being random. And that’s what makes the whole thing useful in the end.
The training apparatus outside the model does this repeatedly, yes, under the thesis that tweaking parameter weights will result in fewer failures to the deterministic fitness algorithm. That’s a theory.
No, that’s not a theory, that is precisely what we measurably observe in practice with coding harnesses. And having built one myself, I can tell you for a fact that this works exactly the same way a genetic algorithm does, and large part of making an effective harness comes from ensuring that the model gets actionable feedback.
No. That’s a leap that has no basis.
The basis is me having worked on a harness and observed how the model outputs improve based on the feedback. There’s also plenty of research on the subject explaining how and why this works in detail. The parameter space is also not nearly as opaque as you seem to think.
At no point is the LLM getting closer to learning facts, and the thesis that knowledge or skill is representable as a statistical model is unproven and seems increasingly unlikely.
That’s missing the point entirely. The question isn’t about whether LLM is getting closer to learning facts. It’s about whether the biasing from the feedback loop causes the LLM to produce relevant outputs. Also, the thesis that knowledge or skill is representable as a statistical model is very much demonstrated by world models where a temporally consistent simulation of the environment is maintained.
Applying the same technique to fine tuning an LLM is a better use of time than manually fine tuning, but that doesn’t make it any less a probabilistic next-token generator that can’t represent stable facts and rules and where every fine tune for one input is always in tension with the fine tubes for all other inputs.
That’s not how any of this works at all. You’re not trying to get it to represent stable facts, you use things like compilers, test harnesses, formals specs, and so on, to create the selection pressure. Then the model is the stochastic part of the system which finds a path that satisfies the selection criteria. Or, with robotics, you have models interact with the physical world and use the feedback to adjust predictions within the model.
Yes, but just because algorithms are analogous doesn’t mean they are functionally equivalent.
Yet, they are functionally equivalent in accomplishing many tasks now. And of course, biological brains have many more subsystems and are more complex in general. I’m not arguing that part at all. My point was that what grounds our mental models in reality is the same feedback loop we use to ground LLMs, and it’s effective for the exact same reason. I also don’t think LLMs are the pinnacle of AI, they’re just one piece of the puzzle, and as I noted earlier, people are already moving towards world models now.
I find world models to be fundamentally more interesting than plain LLMs because if a model encodes the rules of how the physical world works, that provides a foundation for meaningful communication. Humans can talk to each other easily precisely because we all have a shared context which is the environment we live in. And we see how the rate of misunderstanding quickly goes up when we start talking about abstract topic because they can be interpreted in many different ways. So, if models can share the understanding of the physical world with us, it becomes a lot easier to tell them what you want, to correct them, and to have them genuinely understand requirements in a human sense.
What you’re describing for coding, though, is alignment between prompts+context and known good solutions. Yes, it’s entirely possible to have the LLM produce novel code solutions, just like it can produce novel sentences - stochastically - but that doesn’t mean it’s getting a greater basis in reality. It means that it is mapping the highly variable request and existing code to it’s training corpus and it keeps adding more maps between prompts and valid code solutions via rewards-based training. Which is still a next token generator no matter how you slice it.
What I’m describing is the general feedback loop. Coding is just one application here, and plenty of problems can be encoded in the same way. Again, it doesn’t actually matter if it’s getting a greater basis in reality or not. All it needs to do is to generate plausible outputs within a particular context, and these can be tested, and iterated on to solve a problem. And if you go back and read through the thread, nobody is arguing that it’s not a token generator. What’s being said is that this is a reductive way to look at what’s actually happening. It’s like saying that human is a cell reproduction machine. Technically true, and completely useless for understanding what humans do.
Yeah and he basically argues that you shouldn’t see it as next-token even though he agrees that it is lol
And a mouse is technically a blob of molecules. But does calling a mouse a blob of molecules explain why it moves away from cats?
But it’s different because “it does it’s own exploration of infinite space which is exactly like knowing how to win all chess positions through a brute force solution to the finite game space”



