• ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
    link
    fedilink
    arrow-up
    2
    ·
    18 hours ago

    What I’m describing is the general feedback loop. Coding is just one application here, and plenty of problems can be encoded in the same way. Again, it doesn’t actually matter if it’s getting a greater basis in reality or not. All it needs to do is to generate plausible outputs within a particular context, and these can be tested, and iterated on to solve a problem. And if you go back and read through the thread, nobody is arguing that it’s not a token generator. What’s being said is that this is a reductive way to look at what’s actually happening. It’s like saying that human is a cell reproduction machine. Technically true, and completely useless for understanding what humans do.

    • freagle@lemmy.ml
      link
      fedilink
      English
      arrow-up
      2
      ·
      17 hours ago

      I don’t really think it’s useless as an explanation. It’s a great explanation for why an LLM can produce entire code bases but can’t tell you how many Rs are in Strawberry and why they can’t do math.

      Yes, people should understand the ways that they are being scaffolded and why “reasoning” multiplies the cost of inference, and also why it doesn’t solve all problems and why despite the many many advances LLMs will essentially never be able to reliably answer statements of fact even in self-referential systems and why world models are part of the solution and how to reason (as a human) about the limits of world models in addressing the gaps people are seeing.

      • ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
        link
        fedilink
        arrow-up
        2
        ·
        17 hours ago

        The whole thing with R’s in strawberry hasn’t been true for a while now. Turns out you can use RL to get the model to do basic calculation. Notably, this is the exact same problem humans have. The way our brains work is also stochastic, and we struggle to do complex math in our heads. But of course, we can reinforce train ourselves to get better at it. And what we typically do is use an external aid like pen and paper to work through problems, which is basically no different from an LLM harness. If you hook up an LLM to REPL in octave, then it can do math quite well all of a sudden.

        Understanding the limitations of LLMs and how to use them effectively requires moving past reductive thinking. While token generation is the base operation, focusing on that is like trying to understand the brain by looking at individual neuron firings. What’s actually interesting in both cases are the high level patterns that end up being produced which I’d argue are substrate independent. Meanwhile, a combination of an LLM with a harness can be seen as a type of a neurosymbolic system. The neural network generates novel patterns, while the symbolic engine provides the rails for it to function within.

        • freagle@lemmy.ml
          link
          fedilink
          English
          arrow-up
          2
          ·
          13 hours ago

          The whole thing with R’s in strawberry hasn’t been true for a while now.

          I don’t know why you think this is true. Transformers literally do not have any concept of algorithms (like the ones humans use to count). You have to add a “skill” - a deterministic implementation of an algorithm - and then you need to use RL to tune the parameters until the transformer outputs the tokens required to pipe the output of the skill to the user instead of producing a normal text output.

          Turns out you can use RL to get the model to do basic calculation

          Absolutely not. That’s just not how transformers work. It’s a technical impossibility. Maybe you’re talking about non-transformer LLMs, like Mamba, but very few people are using Mamba-based LLMs outside of the research field. When we move beyond transformers, yes, counting is something they can do because they’re effectively FSMs.

          Notably, this is the exact same problem humans have.

          Not really. It might have similarities, but I would never say it’s the exact same problem.

          The way our brains work is also stochastic

          Sure.

          and we struggle to do complex math in our heads

          But for different reasons. Transformers because they are state destroying. Humans because they have limited and volatile working memory.

          But of course, we can reinforce train ourselves to get better at it.

          We can reinforce train ourselves to get better at math in so many different ways. We can memorize facts. We have metacognition and can often assess when we do or don’t know a fact. We can develop cognitive behavioral algorithms that model physical systems. We can develop cognitive behavioral algorithms that model abstract systems. We can combine all of these together in nested, iterative, and recursive structures. It has only a little bit to do with reinforcing the association between abstract symbolic tokens.

          And what we typically do is use an external aid like pen and paper to work through problems, which is basically no different from an LLM harness

          It’s certainly different because pen and paper are most often used to enhance working memory, but the fact-based reasoning and the algorithmic state machines are encoded in our brain which is impossible for a transformer LLM. It remains to be seen what Mamba-likes are truly capable of, but they certainly are addressing many of the critiques I’m raising.

          If you hook up an LLM to REPL in octave, then it can do math quite well all of a sudden. […] Understanding the limitations of LLMs and how to use them effectively requires moving past reductive thinking

          It also requires understanding how things actually work. Hooking up the LLM to a REPL and then iteratively fine tuning it changes the model from outputting a stochastic answer via next-token prediction to outputting a stochastic algorithm via next-token prediction (that will answer the question for you). The LLM did not get better at doing math. It was reweighted to answer questions in the form of “here’s a solution that will answer your question for you since I cannot answer you because I have no ability to do math”, and within that structure you will STILL get hallucinations and can still fuck the model up by posing sufficiently complex or misdirecting word problems. But the corpus for converting word problems to algorithms is massive, and combined with fine-tuning AND ensemble sampling (which increases your real inference cost in multiples) you’re going to produce decent algorithms from math word problems relatively consistently.

          Which is the same way it gets better at coding and yet still can’t actually solve complex problems in design space, constantly has to use ensemble sampling, and constantly has to be told to re-roll the dice whenever the test fails. And that behavior is so costly under the hood that it’s eye watering.

          While token generation is the base operation, focusing on that is like trying to understand the brain by looking at individual neuron firings.

          Not really. Watching individual neuron firings would be equivalent to watching individual parameter weights and the outputs of each step of the transformer. Token generation is literally the entire functioning of transformers.

          What’s actually interesting in both cases are the high level patterns that end up being produced which I’d argue are substrate independent

          Yeah, patterns are, by definition, substrate independent. But transformers only maintain high level patterns on a per-token basis. High level patterns can and do emerge from weighted parameter space, and in many surprising ways, but they are fundamentally limited in transformers because transformers are, at base, next-token predictors so even though we get emergent high-level patterns that can, for example, sort lists, we STILL get hallucinations specifically because the high-level patterns are ephemeral on a per-token basis. Reinforcement learning and back propagation can only go so far in fitting the parameter space before it results in contention with other outcomes.

          Meanwhile, a combination of an LLM with a harness can be seen as a type of a neurosymbolic system. The neural network generates novel patterns, while the symbolic engine provides the rails for it to function within.

          Yes. Inference -> fitness check -> iterate. Agentic retry. It’s incredibly expensive precisely because it uses next-token predictors to generate an answer with an already-known fitness algorithm and then just re-runs inference until the answer passes the fitness test. It’s an automated human-in-the-loop system where a human says “No, that’s not right, try again” and we all know how quickly we run of free tokens when we do that, and we’ve also all had the experience of the damn thing never getting anywhere near close to the solution after a dozen attempts at reprompting.

          Yes, modern transformer harnesses do a TON of work and actually make these parrots useful instead of novelties. But it doesn’t change the fact that they are fundamentally statistically weighted parameter-space stochastic next-token predictors, no matter how much you add to them.

          Instead of arguing against the technical reality, why not focus on the truth about the harnesses - they add a ton of value and make next-token prediction much more useful in some contexts, especially contexts like producing working code.

          • ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
            link
            fedilink
            arrow-up
            2
            ·
            11 hours ago

            I don’t know why you think this is true.

            I mean you can just try it with DeepSeek or any other large model yourself. This is literally a solved problem now.

            Absolutely not.

            Evidently you need to read up on how reasoning chains work.

            Not really. It might have similarities, but I would never say it’s the exact same problem.

            It literally is the same problem. Your brains didn’t evolve to do formal logic natively. We emulate it exactly the way the LLM does.

            But for different reasons. Transformers because they are state destroying. Humans because they have limited and volatile working memory.

            No, for the same fundamental reasons. It’s got nothing to do with state being destroyed either. It has to do with the fact that stochastic systems aren’t a good fit for doing symbolic logic.

            It’s certainly different because pen and paper are most often used to enhance working memory, but the fact-based reasoning and the algorithmic state machines are encoded in our brain which is impossible for a transformer LLM.

            Fact based reasoning is something our brains are famously terrible at doing actually. That’s why we use tools like computers in the first place. Our brains can be trained to express patterns of formal logic, and an artificial neural network can be trained to do the same thing. That’s why modern LLMs can reliably tell you the number of R’s in strawberry.

            It also requires understanding how things actually work.

            It doesn’t, that’s the whole beauty of genetic algorithms. All you have to do is specify your selection pressures and your goal criteria, and the system evolves a solution to fit the shape your desire. The LLM doesn’t need to get better at doing math, the stochastic approach means it converges on a solution given the right environmental pressures. And that’s why hallucinations don’t matter, they get weeded out by the attempts being tested against the environment.

            Which is the same way it gets better at coding and yet still can’t actually solve complex problems in design space, constantly has to use ensemble sampling, and constantly has to be told to re-roll the dice whenever the test fails. And that behavior is so costly under the hood that it’s eye watering.

            I can tell you haven’t actually worked with these tools recently.

            Not really. Watching individual neuron firings would be equivalent to watching individual parameter weights and the outputs of each step of the transformer. Token generation is literally the entire functioning of transformers.

            As a communist, I expect you to understand the concept of quantity transforming into quality.

            Yeah, patterns are, by definition, substrate independent. But transformers only maintain high level patterns on a per-token basis. High level patterns can and do emerge from weighted parameter space, and in many surprising ways, but they are fundamentally limited in transformers because transformers are, at base, next-token predictors so even though we get emergent high-level patterns that can, for example, sort lists, we STILL get hallucinations specifically because the high-level patterns are ephemeral on a per-token basis.

            Do explain how this is different from saying that human brains are fundamentally limited in that neurons are just next state predictors.

            Yes. Inference -> fitness check -> iterate. Agentic retry. It’s incredibly expensive precisely because it uses next-token predictors to generate an answer with an already-known fitness algorithm and then just re-runs inference until the answer passes the fitness test.

            Except it’s not incredibly expensive because the system works on the principle of gradient dissent. It isn’t just producing a random value each turn, it produces a plausible value within the context which is precisely what allows it to quickly converge on a solution.

            Yes, modern transformer harnesses do a TON of work and actually make these parrots useful instead of novelties. But it doesn’t change the fact that they are fundamentally statistically weighted parameter-space stochastic next-token predictors, no matter how much you add to them.

            Exactly the way the neurons in your brain are stochastic next state predictors.

            Instead of arguing against the technical reality, why not focus on the truth about the harnesses - they add a ton of value and make next-token prediction much more useful in some contexts, especially contexts like producing working code.

            I would urge you to spend a bit of time actually learning about the technical reality instead of continuing to argue here.

            • freagle@lemmy.ml
              link
              fedilink
              English
              arrow-up
              1
              ·
              40 minutes ago

              Yog, you know I respect you, but we have fundamentally different understandings here and I think yours is wrong.

              An individual neuron is not a next state predictor. An individual neuron is analogous to a NN parameter. It’s not exactly like a NN parameter, it’s more complex, which allows animal neural networks to do a lot more than the NNs we have built, but it’s an analogy. It’s more appropriate to say that a NN parameter is attempting to be analogous to an animal neuron. It’s not there yet, and no amount of training will get it there, because of structural differences.

              The insistence that LLMs are exactly like human brains is reductive. Calling an LLM a next token predictor is not reductive. Adding a harness doesn’t make LLMs not next token predictors. And the difference been an animal brain and an LLM Is not the quality of the harness nor the amount of training.

              You’re right. My understanding of the use of ensemble methods was outdated. Everyone has moved over to mixture of experts routers, a non-emergent property designed and engineered into the transformer. They still generate the next token, but they use a hard coded algorithm to generate a bunch together and then weight the results.

              Yes, reasoning models are doing math. But the didn’t learn to do math better. Reasoning models uses a hidden text buffer where the LLM produces intermediary tokens that are statistically most likely to produce next tokens that will pass the fitness requirements. Through RL, reasoning models began to statistically weight patterns of tokens that resulted in rerunning the “math” multiple times, because it produced more and more context that constrained the probability space. This is not like how human reasoning works in the vast majority of cases, and calling it “exactly like” human reasoning is begging the question.

              As for human analog, the “number of letters” test that most models have gotten better at (but still hallucinate on sometimes), also succeed on nonsense words that are not in their training set. And they work because they are engineered, hard-coded, to break down words into individual characters. That is not an emergent property of the NN like it is in humans.

              I recognize that changes in quantity can induce changes in quality, and the emergent functional structures in energized parameter space of LLMs is exactly that, and that some of those things are analogous to what happens in animal neurology. But it’s not an exact match, as you are saying, and it’s not all the same structures.

              When we force humans into next token predictors, they are surprisingly convergent. 70% of times that someone is asked to pick a number between 1 and 10, they pick 7. LLMs exhibit that behavior too, but at 90%. The reason LLMs do it is because they are trained on data that constantly conjoins text about guessing a random number and the curious human behavior of guessing 7. They produce similar, analogous behavior but for completely different reasons.

              Back to the article, the claim that reinforcement learning makes them “not next token generators” or that somehow calling them “next token generators” is reductive just doesn’t follow. NNs have been producing emergent behavior long behavior LLMs and transformers. The emergent behavior is worth point out to people who think that the output is entirely random (which is reductive). But the fact that transformers are fundamentally next token generators is constantly reinforced by all the ways that frontier model engineers have to design their transformer algorithms - everything becomes about producing tokens in a sequence that constraints the probability space.

              And the point about transformers having ephemeral state on a per token basis, that’s how transformers work, even with the mixture of experts routers. That’s exactly why they give transformers scratch space for “reasoning”, because they are trying to capture the value of the emergent NN structures in the token stream its generating. Contrast this with Mamba models that have a persistent working memory directly in their models and don’t rely exclusively on the session history of token stream for this response to probabilistically re-energize the same emergent structures. This is why I don’t think transformer-based “reasoning chains” are analogous to human reasoning chains. They might be analogous to a very small subset of limited human mental behaviors, but that’s not the same thing.