Yes, it is a stochastic system, that’s the whole point tough. There are two things at play here, one is that the model does prediction based on the context of the current data it’s looking at rather than just whatever data it was trained on. That’s what the article is highlighting. The second part is the feedback loop such as what you see in agentic coding. The model makes a prediction, that prediction is tested against the environment, the model gets feedback, and it iterates. And that’s what grounds it in reality addressing the issue of it producing meaningless outputs. The system as a whole behaves similarly to a genetic algorithm where the solution evolves through the cycle of trial and error.
Incidentally, this is true for humans as well. This is why we need tools like the scientific method and peer review in science. People hallucinate things all the time, and when we loose the connection between the predictions the brain generates and sensory feedback we call that schizophrenia.
one is that the model does prediction based on the context of the current data it’s looking at rather than just whatever data it was trained on
That’s always true. That’s how neural networks work. The model is a statistical transform from input to output. Neural networks work by taking input to produce an output. The “context of the current data” is just the input. “Rather than just whatever data it was trained on” is meaningless.
The second part is the feedback loop
Yes, as I said, there’s a fitness algorithm and it back propagates adjustments to parameter weights. But that’s all it can do, because the model is just weighted parameters. It can’t learn facts, it can only adjust its probabilities.
The model makes a prediction
Or rather it produces a random output based on the input and its parameter weights
that prediction is tested against the environment
By something other than model itself that has knowledge of what “success” is and what “failure” is
the model gets feedback
In the form of adjustments to parameter weights
and it iterates
The training apparatus outside the model does this repeatedly, yes, under the thesis that tweaking parameter weights will result in fewer failures to the deterministic fitness algorithm. That’s a theory.
And that’s what grounds it in reality addressing the issue of it producing meaningless outputs
No. That’s a leap that has no basis. The training data is no less a part of reality as the current prompt context is part of reality. What you’re describing is that the output of the probabilistic transformer gets tested against various forms of curated fitness tests. The problem with is that the only thing one can do with the the results of fitness tests is to change the probabilities of the opaque parameter space. So you can create a fitness test for how many "r"s are in “strawberry” but the results of that test can only be expressed by weight changes. And those fine tuning adjustments are applied to an opaque network of weights that also includes the opaque probabilistic representations of the fitness tests for how many "r"s are in “perrywinkle” and how many "b"s are in “strawberry syrup”.
At no point is the LLM getting closer to learning facts, and the thesis that knowledge or skill is representable as a statistical model is unproven and seems increasingly unlikely.
The system as a whole behaves similarly to a genetic algorithm where the solution evolves through the cycle of trial and error.
Yes, it uses the same concepts as a genetic algorithm but it the representation is still the problem. Genetic algorithms for path finding are great because they have discrete actions and limited scope. Applying the same technique to fine tuning an LLM is a better use of time than manually fine tuning, but that doesn’t make it any less a probabilistic next-token generator that can’t represent stable facts and rules and where every fine tune for one input is always in tension with the fine tubes for all other inputs.
Incidentally, this is true for humans as well.
Yes, but just because algorithms are analogous doesn’t mean they are functionally equivalent. Humans also have an opaque neural network that functionally behaves like a statistical model. But we have more subsystems than LLMs do, we have more dimensions to our encoding, and we have greater self-governing and modification abilities. So while the genetic algorithm approach is useful, it doesn’t make the LLM become closer to reality, it just automates a portion of the fine tuning curation process.
Or rather it produces a random output based on the input and its parameter weights
The bias is precisely what makes it not random, but rather stochastic. There’s a very big difference here.
In the form of adjustments to parameter weights
I’m talking about feedback from the environment it operates in. That’s the actual test that allows the model to keep adjusting outputs towards a specific target rather than them being random. And that’s what makes the whole thing useful in the end.
The training apparatus outside the model does this repeatedly, yes, under the thesis that tweaking parameter weights will result in fewer failures to the deterministic fitness algorithm. That’s a theory.
No, that’s not a theory, that is precisely what we measurably observe in practice with coding harnesses. And having built one myself, I can tell you for a fact that this works exactly the same way a genetic algorithm does, and large part of making an effective harness comes from ensuring that the model gets actionable feedback.
No. That’s a leap that has no basis.
The basis is me having worked on a harness and observed how the model outputs improve based on the feedback. There’s also plenty of research on the subject explaining how and why this works in detail. The parameter space is also not nearly as opaque as you seem to think.
At no point is the LLM getting closer to learning facts, and the thesis that knowledge or skill is representable as a statistical model is unproven and seems increasingly unlikely.
That’s missing the point entirely. The question isn’t about whether LLM is getting closer to learning facts. It’s about whether the biasing from the feedback loop causes the LLM to produce relevant outputs. Also, the thesis that knowledge or skill is representable as a statistical model is very much demonstrated by world models where a temporally consistent simulation of the environment is maintained.
Applying the same technique to fine tuning an LLM is a better use of time than manually fine tuning, but that doesn’t make it any less a probabilistic next-token generator that can’t represent stable facts and rules and where every fine tune for one input is always in tension with the fine tubes for all other inputs.
That’s not how any of this works at all. You’re not trying to get it to represent stable facts, you use things like compilers, test harnesses, formals specs, and so on, to create the selection pressure. Then the model is the stochastic part of the system which finds a path that satisfies the selection criteria. Or, with robotics, you have models interact with the physical world and use the feedback to adjust predictions within the model.
Yes, but just because algorithms are analogous doesn’t mean they are functionally equivalent.
Yet, they are functionally equivalent in accomplishing many tasks now. And of course, biological brains have many more subsystems and are more complex in general. I’m not arguing that part at all. My point was that what grounds our mental models in reality is the same feedback loop we use to ground LLMs, and it’s effective for the exact same reason. I also don’t think LLMs are the pinnacle of AI, they’re just one piece of the puzzle, and as I noted earlier, people are already moving towards world models now.
I find world models to be fundamentally more interesting than plain LLMs because if a model encodes the rules of how the physical world works, that provides a foundation for meaningful communication. Humans can talk to each other easily precisely because we all have a shared context which is the environment we live in. And we see how the rate of misunderstanding quickly goes up when we start talking about abstract topic because they can be interpreted in many different ways. So, if models can share the understanding of the physical world with us, it becomes a lot easier to tell them what you want, to correct them, and to have them genuinely understand requirements in a human sense.
What you’re describing for coding, though, is alignment between prompts+context and known good solutions. Yes, it’s entirely possible to have the LLM produce novel code solutions, just like it can produce novel sentences - stochastically - but that doesn’t mean it’s getting a greater basis in reality. It means that it is mapping the highly variable request and existing code to it’s training corpus and it keeps adding more maps between prompts and valid code solutions via rewards-based training. Which is still a next token generator no matter how you slice it.
What I’m describing is the general feedback loop. Coding is just one application here, and plenty of problems can be encoded in the same way. Again, it doesn’t actually matter if it’s getting a greater basis in reality or not. All it needs to do is to generate plausible outputs within a particular context, and these can be tested, and iterated on to solve a problem. And if you go back and read through the thread, nobody is arguing that it’s not a token generator. What’s being said is that this is a reductive way to look at what’s actually happening. It’s like saying that human is a cell reproduction machine. Technically true, and completely useless for understanding what humans do.
I don’t really think it’s useless as an explanation. It’s a great explanation for why an LLM can produce entire code bases but can’t tell you how many Rs are in Strawberry and why they can’t do math.
Yes, people should understand the ways that they are being scaffolded and why “reasoning” multiplies the cost of inference, and also why it doesn’t solve all problems and why despite the many many advances LLMs will essentially never be able to reliably answer statements of fact even in self-referential systems and why world models are part of the solution and how to reason (as a human) about the limits of world models in addressing the gaps people are seeing.
The whole thing with R’s in strawberry hasn’t been true for a while now. Turns out you can use RL to get the model to do basic calculation. Notably, this is the exact same problem humans have. The way our brains work is also stochastic, and we struggle to do complex math in our heads. But of course, we can reinforce train ourselves to get better at it. And what we typically do is use an external aid like pen and paper to work through problems, which is basically no different from an LLM harness. If you hook up an LLM to REPL in octave, then it can do math quite well all of a sudden.
Understanding the limitations of LLMs and how to use them effectively requires moving past reductive thinking. While token generation is the base operation, focusing on that is like trying to understand the brain by looking at individual neuron firings. What’s actually interesting in both cases are the high level patterns that end up being produced which I’d argue are substrate independent. Meanwhile, a combination of an LLM with a harness can be seen as a type of a neurosymbolic system. The neural network generates novel patterns, while the symbolic engine provides the rails for it to function within.
The whole thing with R’s in strawberry hasn’t been true for a while now.
I don’t know why you think this is true. Transformers literally do not have any concept of algorithms (like the ones humans use to count). You have to add a “skill” - a deterministic implementation of an algorithm - and then you need to use RL to tune the parameters until the transformer outputs the tokens required to pipe the output of the skill to the user instead of producing a normal text output.
Turns out you can use RL to get the model to do basic calculation
Absolutely not. That’s just not how transformers work. It’s a technical impossibility. Maybe you’re talking about non-transformer LLMs, like Mamba, but very few people are using Mamba-based LLMs outside of the research field. When we move beyond transformers, yes, counting is something they can do because they’re effectively FSMs.
Notably, this is the exact same problem humans have.
Not really. It might have similarities, but I would never say it’s the exact same problem.
The way our brains work is also stochastic
Sure.
and we struggle to do complex math in our heads
But for different reasons. Transformers because they are state destroying. Humans because they have limited and volatile working memory.
But of course, we can reinforce train ourselves to get better at it.
We can reinforce train ourselves to get better at math in so many different ways. We can memorize facts. We have metacognition and can often assess when we do or don’t know a fact. We can develop cognitive behavioral algorithms that model physical systems. We can develop cognitive behavioral algorithms that model abstract systems. We can combine all of these together in nested, iterative, and recursive structures. It has only a little bit to do with reinforcing the association between abstract symbolic tokens.
And what we typically do is use an external aid like pen and paper to work through problems, which is basically no different from an LLM harness
It’s certainly different because pen and paper are most often used to enhance working memory, but the fact-based reasoning and the algorithmic state machines are encoded in our brain which is impossible for a transformer LLM. It remains to be seen what Mamba-likes are truly capable of, but they certainly are addressing many of the critiques I’m raising.
If you hook up an LLM to REPL in octave, then it can do math quite well all of a sudden. […] Understanding the limitations of LLMs and how to use them effectively requires moving past reductive thinking
It also requires understanding how things actually work. Hooking up the LLM to a REPL and then iteratively fine tuning it changes the model from outputting a stochastic answer via next-token prediction to outputting a stochastic algorithm via next-token prediction (that will answer the question for you). The LLM did not get better at doing math. It was reweighted to answer questions in the form of “here’s a solution that will answer your question for you since I cannot answer you because I have no ability to do math”, and within that structure you will STILL get hallucinations and can still fuck the model up by posing sufficiently complex or misdirecting word problems. But the corpus for converting word problems to algorithms is massive, and combined with fine-tuning AND ensemble sampling (which increases your real inference cost in multiples) you’re going to produce decent algorithms from math word problems relatively consistently.
Which is the same way it gets better at coding and yet still can’t actually solve complex problems in design space, constantly has to use ensemble sampling, and constantly has to be told to re-roll the dice whenever the test fails. And that behavior is so costly under the hood that it’s eye watering.
While token generation is the base operation, focusing on that is like trying to understand the brain by looking at individual neuron firings.
Not really. Watching individual neuron firings would be equivalent to watching individual parameter weights and the outputs of each step of the transformer. Token generation is literally the entire functioning of transformers.
What’s actually interesting in both cases are the high level patterns that end up being produced which I’d argue are substrate independent
Yeah, patterns are, by definition, substrate independent. But transformers only maintain high level patterns on a per-token basis. High level patterns can and do emerge from weighted parameter space, and in many surprising ways, but they are fundamentally limited in transformers because transformers are, at base, next-token predictors so even though we get emergent high-level patterns that can, for example, sort lists, we STILL get hallucinations specifically because the high-level patterns are ephemeral on a per-token basis. Reinforcement learning and back propagation can only go so far in fitting the parameter space before it results in contention with other outcomes.
Meanwhile, a combination of an LLM with a harness can be seen as a type of a neurosymbolic system. The neural network generates novel patterns, while the symbolic engine provides the rails for it to function within.
Yes. Inference -> fitness check -> iterate. Agentic retry. It’s incredibly expensive precisely because it uses next-token predictors to generate an answer with an already-known fitness algorithm and then just re-runs inference until the answer passes the fitness test. It’s an automated human-in-the-loop system where a human says “No, that’s not right, try again” and we all know how quickly we run of free tokens when we do that, and we’ve also all had the experience of the damn thing never getting anywhere near close to the solution after a dozen attempts at reprompting.
Yes, modern transformer harnesses do a TON of work and actually make these parrots useful instead of novelties. But it doesn’t change the fact that they are fundamentally statistically weighted parameter-space stochastic next-token predictors, no matter how much you add to them.
Instead of arguing against the technical reality, why not focus on the truth about the harnesses - they add a ton of value and make next-token prediction much more useful in some contexts, especially contexts like producing working code.
I mean you can just try it with DeepSeek or any other large model yourself. This is literally a solved problem now.
Absolutely not.
Evidently you need to read up on how reasoning chains work.
Not really. It might have similarities, but I would never say it’s the exact same problem.
It literally is the same problem. Your brains didn’t evolve to do formal logic natively. We emulate it exactly the way the LLM does.
But for different reasons. Transformers because they are state destroying. Humans because they have limited and volatile working memory.
No, for the same fundamental reasons. It’s got nothing to do with state being destroyed either. It has to do with the fact that stochastic systems aren’t a good fit for doing symbolic logic.
It’s certainly different because pen and paper are most often used to enhance working memory, but the fact-based reasoning and the algorithmic state machines are encoded in our brain which is impossible for a transformer LLM.
Fact based reasoning is something our brains are famously terrible at doing actually. That’s why we use tools like computers in the first place. Our brains can be trained to express patterns of formal logic, and an artificial neural network can be trained to do the same thing. That’s why modern LLMs can reliably tell you the number of R’s in strawberry.
It also requires understanding how things actually work.
It doesn’t, that’s the whole beauty of genetic algorithms. All you have to do is specify your selection pressures and your goal criteria, and the system evolves a solution to fit the shape your desire. The LLM doesn’t need to get better at doing math, the stochastic approach means it converges on a solution given the right environmental pressures. And that’s why hallucinations don’t matter, they get weeded out by the attempts being tested against the environment.
Which is the same way it gets better at coding and yet still can’t actually solve complex problems in design space, constantly has to use ensemble sampling, and constantly has to be told to re-roll the dice whenever the test fails. And that behavior is so costly under the hood that it’s eye watering.
I can tell you haven’t actually worked with these tools recently.
Not really. Watching individual neuron firings would be equivalent to watching individual parameter weights and the outputs of each step of the transformer. Token generation is literally the entire functioning of transformers.
As a communist, I expect you to understand the concept of quantity transforming into quality.
Yeah, patterns are, by definition, substrate independent. But transformers only maintain high level patterns on a per-token basis. High level patterns can and do emerge from weighted parameter space, and in many surprising ways, but they are fundamentally limited in transformers because transformers are, at base, next-token predictors so even though we get emergent high-level patterns that can, for example, sort lists, we STILL get hallucinations specifically because the high-level patterns are ephemeral on a per-token basis.
Do explain how this is different from saying that human brains are fundamentally limited in that neurons are just next state predictors.
Yes. Inference -> fitness check -> iterate. Agentic retry. It’s incredibly expensive precisely because it uses next-token predictors to generate an answer with an already-known fitness algorithm and then just re-runs inference until the answer passes the fitness test.
Except it’s not incredibly expensive because the system works on the principle of gradient dissent. It isn’t just producing a random value each turn, it produces a plausible value within the context which is precisely what allows it to quickly converge on a solution.
Yes, modern transformer harnesses do a TON of work and actually make these parrots useful instead of novelties. But it doesn’t change the fact that they are fundamentally statistically weighted parameter-space stochastic next-token predictors, no matter how much you add to them.
Exactly the way the neurons in your brain are stochastic next state predictors.
Instead of arguing against the technical reality, why not focus on the truth about the harnesses - they add a ton of value and make next-token prediction much more useful in some contexts, especially contexts like producing working code.
I would urge you to spend a bit of time actually learning about the technical reality instead of continuing to argue here.
Yes, it is a stochastic system, that’s the whole point tough. There are two things at play here, one is that the model does prediction based on the context of the current data it’s looking at rather than just whatever data it was trained on. That’s what the article is highlighting. The second part is the feedback loop such as what you see in agentic coding. The model makes a prediction, that prediction is tested against the environment, the model gets feedback, and it iterates. And that’s what grounds it in reality addressing the issue of it producing meaningless outputs. The system as a whole behaves similarly to a genetic algorithm where the solution evolves through the cycle of trial and error.
Incidentally, this is true for humans as well. This is why we need tools like the scientific method and peer review in science. People hallucinate things all the time, and when we loose the connection between the predictions the brain generates and sensory feedback we call that schizophrenia.
That’s always true. That’s how neural networks work. The model is a statistical transform from input to output. Neural networks work by taking input to produce an output. The “context of the current data” is just the input. “Rather than just whatever data it was trained on” is meaningless.
Yes, as I said, there’s a fitness algorithm and it back propagates adjustments to parameter weights. But that’s all it can do, because the model is just weighted parameters. It can’t learn facts, it can only adjust its probabilities.
Or rather it produces a random output based on the input and its parameter weights
By something other than model itself that has knowledge of what “success” is and what “failure” is
In the form of adjustments to parameter weights
The training apparatus outside the model does this repeatedly, yes, under the thesis that tweaking parameter weights will result in fewer failures to the deterministic fitness algorithm. That’s a theory.
No. That’s a leap that has no basis. The training data is no less a part of reality as the current prompt context is part of reality. What you’re describing is that the output of the probabilistic transformer gets tested against various forms of curated fitness tests. The problem with is that the only thing one can do with the the results of fitness tests is to change the probabilities of the opaque parameter space. So you can create a fitness test for how many "r"s are in “strawberry” but the results of that test can only be expressed by weight changes. And those fine tuning adjustments are applied to an opaque network of weights that also includes the opaque probabilistic representations of the fitness tests for how many "r"s are in “perrywinkle” and how many "b"s are in “strawberry syrup”.
At no point is the LLM getting closer to learning facts, and the thesis that knowledge or skill is representable as a statistical model is unproven and seems increasingly unlikely.
Yes, it uses the same concepts as a genetic algorithm but it the representation is still the problem. Genetic algorithms for path finding are great because they have discrete actions and limited scope. Applying the same technique to fine tuning an LLM is a better use of time than manually fine tuning, but that doesn’t make it any less a probabilistic next-token generator that can’t represent stable facts and rules and where every fine tune for one input is always in tension with the fine tubes for all other inputs.
Yes, but just because algorithms are analogous doesn’t mean they are functionally equivalent. Humans also have an opaque neural network that functionally behaves like a statistical model. But we have more subsystems than LLMs do, we have more dimensions to our encoding, and we have greater self-governing and modification abilities. So while the genetic algorithm approach is useful, it doesn’t make the LLM become closer to reality, it just automates a portion of the fine tuning curation process.
The bias is precisely what makes it not random, but rather stochastic. There’s a very big difference here.
I’m talking about feedback from the environment it operates in. That’s the actual test that allows the model to keep adjusting outputs towards a specific target rather than them being random. And that’s what makes the whole thing useful in the end.
No, that’s not a theory, that is precisely what we measurably observe in practice with coding harnesses. And having built one myself, I can tell you for a fact that this works exactly the same way a genetic algorithm does, and large part of making an effective harness comes from ensuring that the model gets actionable feedback.
The basis is me having worked on a harness and observed how the model outputs improve based on the feedback. There’s also plenty of research on the subject explaining how and why this works in detail. The parameter space is also not nearly as opaque as you seem to think.
That’s missing the point entirely. The question isn’t about whether LLM is getting closer to learning facts. It’s about whether the biasing from the feedback loop causes the LLM to produce relevant outputs. Also, the thesis that knowledge or skill is representable as a statistical model is very much demonstrated by world models where a temporally consistent simulation of the environment is maintained.
That’s not how any of this works at all. You’re not trying to get it to represent stable facts, you use things like compilers, test harnesses, formals specs, and so on, to create the selection pressure. Then the model is the stochastic part of the system which finds a path that satisfies the selection criteria. Or, with robotics, you have models interact with the physical world and use the feedback to adjust predictions within the model.
Yet, they are functionally equivalent in accomplishing many tasks now. And of course, biological brains have many more subsystems and are more complex in general. I’m not arguing that part at all. My point was that what grounds our mental models in reality is the same feedback loop we use to ground LLMs, and it’s effective for the exact same reason. I also don’t think LLMs are the pinnacle of AI, they’re just one piece of the puzzle, and as I noted earlier, people are already moving towards world models now.
I find world models to be fundamentally more interesting than plain LLMs because if a model encodes the rules of how the physical world works, that provides a foundation for meaningful communication. Humans can talk to each other easily precisely because we all have a shared context which is the environment we live in. And we see how the rate of misunderstanding quickly goes up when we start talking about abstract topic because they can be interpreted in many different ways. So, if models can share the understanding of the physical world with us, it becomes a lot easier to tell them what you want, to correct them, and to have them genuinely understand requirements in a human sense.
What you’re describing for coding, though, is alignment between prompts+context and known good solutions. Yes, it’s entirely possible to have the LLM produce novel code solutions, just like it can produce novel sentences - stochastically - but that doesn’t mean it’s getting a greater basis in reality. It means that it is mapping the highly variable request and existing code to it’s training corpus and it keeps adding more maps between prompts and valid code solutions via rewards-based training. Which is still a next token generator no matter how you slice it.
What I’m describing is the general feedback loop. Coding is just one application here, and plenty of problems can be encoded in the same way. Again, it doesn’t actually matter if it’s getting a greater basis in reality or not. All it needs to do is to generate plausible outputs within a particular context, and these can be tested, and iterated on to solve a problem. And if you go back and read through the thread, nobody is arguing that it’s not a token generator. What’s being said is that this is a reductive way to look at what’s actually happening. It’s like saying that human is a cell reproduction machine. Technically true, and completely useless for understanding what humans do.
I don’t really think it’s useless as an explanation. It’s a great explanation for why an LLM can produce entire code bases but can’t tell you how many Rs are in Strawberry and why they can’t do math.
Yes, people should understand the ways that they are being scaffolded and why “reasoning” multiplies the cost of inference, and also why it doesn’t solve all problems and why despite the many many advances LLMs will essentially never be able to reliably answer statements of fact even in self-referential systems and why world models are part of the solution and how to reason (as a human) about the limits of world models in addressing the gaps people are seeing.
The whole thing with R’s in strawberry hasn’t been true for a while now. Turns out you can use RL to get the model to do basic calculation. Notably, this is the exact same problem humans have. The way our brains work is also stochastic, and we struggle to do complex math in our heads. But of course, we can reinforce train ourselves to get better at it. And what we typically do is use an external aid like pen and paper to work through problems, which is basically no different from an LLM harness. If you hook up an LLM to REPL in octave, then it can do math quite well all of a sudden.
Understanding the limitations of LLMs and how to use them effectively requires moving past reductive thinking. While token generation is the base operation, focusing on that is like trying to understand the brain by looking at individual neuron firings. What’s actually interesting in both cases are the high level patterns that end up being produced which I’d argue are substrate independent. Meanwhile, a combination of an LLM with a harness can be seen as a type of a neurosymbolic system. The neural network generates novel patterns, while the symbolic engine provides the rails for it to function within.
I don’t know why you think this is true. Transformers literally do not have any concept of algorithms (like the ones humans use to count). You have to add a “skill” - a deterministic implementation of an algorithm - and then you need to use RL to tune the parameters until the transformer outputs the tokens required to pipe the output of the skill to the user instead of producing a normal text output.
Absolutely not. That’s just not how transformers work. It’s a technical impossibility. Maybe you’re talking about non-transformer LLMs, like Mamba, but very few people are using Mamba-based LLMs outside of the research field. When we move beyond transformers, yes, counting is something they can do because they’re effectively FSMs.
Not really. It might have similarities, but I would never say it’s the exact same problem.
Sure.
But for different reasons. Transformers because they are state destroying. Humans because they have limited and volatile working memory.
We can reinforce train ourselves to get better at math in so many different ways. We can memorize facts. We have metacognition and can often assess when we do or don’t know a fact. We can develop cognitive behavioral algorithms that model physical systems. We can develop cognitive behavioral algorithms that model abstract systems. We can combine all of these together in nested, iterative, and recursive structures. It has only a little bit to do with reinforcing the association between abstract symbolic tokens.
It’s certainly different because pen and paper are most often used to enhance working memory, but the fact-based reasoning and the algorithmic state machines are encoded in our brain which is impossible for a transformer LLM. It remains to be seen what Mamba-likes are truly capable of, but they certainly are addressing many of the critiques I’m raising.
It also requires understanding how things actually work. Hooking up the LLM to a REPL and then iteratively fine tuning it changes the model from outputting a stochastic answer via next-token prediction to outputting a stochastic algorithm via next-token prediction (that will answer the question for you). The LLM did not get better at doing math. It was reweighted to answer questions in the form of “here’s a solution that will answer your question for you since I cannot answer you because I have no ability to do math”, and within that structure you will STILL get hallucinations and can still fuck the model up by posing sufficiently complex or misdirecting word problems. But the corpus for converting word problems to algorithms is massive, and combined with fine-tuning AND ensemble sampling (which increases your real inference cost in multiples) you’re going to produce decent algorithms from math word problems relatively consistently.
Which is the same way it gets better at coding and yet still can’t actually solve complex problems in design space, constantly has to use ensemble sampling, and constantly has to be told to re-roll the dice whenever the test fails. And that behavior is so costly under the hood that it’s eye watering.
Not really. Watching individual neuron firings would be equivalent to watching individual parameter weights and the outputs of each step of the transformer. Token generation is literally the entire functioning of transformers.
Yeah, patterns are, by definition, substrate independent. But transformers only maintain high level patterns on a per-token basis. High level patterns can and do emerge from weighted parameter space, and in many surprising ways, but they are fundamentally limited in transformers because transformers are, at base, next-token predictors so even though we get emergent high-level patterns that can, for example, sort lists, we STILL get hallucinations specifically because the high-level patterns are ephemeral on a per-token basis. Reinforcement learning and back propagation can only go so far in fitting the parameter space before it results in contention with other outcomes.
Yes. Inference -> fitness check -> iterate. Agentic retry. It’s incredibly expensive precisely because it uses next-token predictors to generate an answer with an already-known fitness algorithm and then just re-runs inference until the answer passes the fitness test. It’s an automated human-in-the-loop system where a human says “No, that’s not right, try again” and we all know how quickly we run of free tokens when we do that, and we’ve also all had the experience of the damn thing never getting anywhere near close to the solution after a dozen attempts at reprompting.
Yes, modern transformer harnesses do a TON of work and actually make these parrots useful instead of novelties. But it doesn’t change the fact that they are fundamentally statistically weighted parameter-space stochastic next-token predictors, no matter how much you add to them.
Instead of arguing against the technical reality, why not focus on the truth about the harnesses - they add a ton of value and make next-token prediction much more useful in some contexts, especially contexts like producing working code.
I mean you can just try it with DeepSeek or any other large model yourself. This is literally a solved problem now.
Evidently you need to read up on how reasoning chains work.
It literally is the same problem. Your brains didn’t evolve to do formal logic natively. We emulate it exactly the way the LLM does.
No, for the same fundamental reasons. It’s got nothing to do with state being destroyed either. It has to do with the fact that stochastic systems aren’t a good fit for doing symbolic logic.
Fact based reasoning is something our brains are famously terrible at doing actually. That’s why we use tools like computers in the first place. Our brains can be trained to express patterns of formal logic, and an artificial neural network can be trained to do the same thing. That’s why modern LLMs can reliably tell you the number of R’s in strawberry.
It doesn’t, that’s the whole beauty of genetic algorithms. All you have to do is specify your selection pressures and your goal criteria, and the system evolves a solution to fit the shape your desire. The LLM doesn’t need to get better at doing math, the stochastic approach means it converges on a solution given the right environmental pressures. And that’s why hallucinations don’t matter, they get weeded out by the attempts being tested against the environment.
I can tell you haven’t actually worked with these tools recently.
As a communist, I expect you to understand the concept of quantity transforming into quality.
Do explain how this is different from saying that human brains are fundamentally limited in that neurons are just next state predictors.
Except it’s not incredibly expensive because the system works on the principle of gradient dissent. It isn’t just producing a random value each turn, it produces a plausible value within the context which is precisely what allows it to quickly converge on a solution.
Exactly the way the neurons in your brain are stochastic next state predictors.
I would urge you to spend a bit of time actually learning about the technical reality instead of continuing to argue here.