• jj4211@lemmy.world
    link
    fedilink
    English
    arrow-up
    14
    ·
    2 months ago

    That’s what is wild about it. At any given point in time, the model is wholly consumed only with the very next token. Maybe that token is a running narrative of ‘reasoning’ or directly in the output, either way, the AI does not have anything to model anything beyond the very next token. It doesn’t have a destination in mind and is just finding the words to get there, it’s building it up word by word. The overall ‘meaning’ is an emergent property of just picking the very next token and seeing what happens.

    Honestly, it’s shocking it works as well as it does. More shockingly, there are AI enthusiasts that argue that’s how the human brain works, which I can’t imagine someone going through life with every thought rooted in building it up word by word.

    • krypt@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      2 months ago

      its not that simple. whatever opinion you might on llms have you have to agree this is oversimplifying at best.

      • jj4211@lemmy.world
        link
        fedilink
        English
        arrow-up
        4
        ·
        2 months ago

        I assumed that but everything I have seen as I dug deeper has been that at some level that is what is happening. If it is ‘reasoning’, it’s generating a ‘reasoning chain’ next token by next token and using that to influence the final output tokens. The reasoning chain is discarded and since the actual output is a continuation of the reasoning chain it may conceptually be described as allowing the model to ‘rethink’ things, but even as the generation of a ‘reasoning chain’ has results that more closely resemble reasoning, it is still a scenario where it’s building it one token at a time and we get to see meaning as an emergent property, rather than trying to find words to build to a more abstract concept like humans do. It just gets to throw away the intermediate work and the extra tokens manage to improve the ‘accuracy’ of the preserved final output.

      • khepri@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        edit-2
        2 months ago

        Its just a ball rolling down a hill trying to find the lowest point, but its like a super fancy ball with a ton of rules, and the hills are really detailed with like little walls and bridges and stuff.

    • HrabiaVulpes@europe.pub
      link
      fedilink
      English
      arrow-up
      5
      ·
      2 months ago

      My masters degree in intelligent systems wept in the corner after seeing your explanation.

      But I guess cars are just four wheels and a fancy basket people sit in.

      • jj4211@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 months ago

        Well that description of a car is actually fairly close to the fundamentals, add an engine or motor and a steering wheel, and you’ve got it. Yes, a lot of engineering goes into the best possible realization of those basics, efficiency, suspension, safety, maintenance, and just a ton of more stuff, and it is a very valued execution above and beyond what, say, the Model T delivered. Automotive engineers have done hard and valuable work and complicated work, but no one is surprised that Model T led to faster, more comfortable, safer, more convenient vehicles that move around. It is a bit more surprising that LLM architecture works as well as it does while always focusing on the next token without ability to go further, best case running through and messing up and regenerating until you have operator appropriate output.

        The ‘seahorse emoji’ was a pretty fun example of this at work, as it didn’t have a seahorse emoji, but since it wasn’t trying to generate the emoji, it started by building up the words to confirm and introduce the emoji since obviously there will be one, then putting up a wrong thing, then the words that would go after the wrong thing, but the weight still suggested there should be a correct answer and to start generating words for another try, and so on. “Reasoning” does the job of incurring this hit out of sight a lot of the time. Looking at the reasoning chains you’ll see this behavior a fair amount, that the model suggests words that build toward an answer but fails on the key word and retries until something tests right or it models that it tried enough and it can’t find the key word that would have been expected. It can of course digest it’s own output and summarize the result without showing the operator spinning out, but it at all times is operating on the fundamental principle of model+very cleverly managed context influencing an answer one token at a time and ideally discarding the first run.

    • kreskin@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 months ago

      Well yes and no. it is steered by a buffer of context as well that sorts/ranks/informs what the next word should be. That context differentiates if you are talking about apple the fruit or apple the company or apple the device. Heres a great overview if anyone is interested. And no, its not my video. Its a youtube intro to how AI works. Best watched with duckduckgo browser which trims out youtubes overly frequent ad interruptions. https://www.youtube.com/watch?v=OYvlznJ4IZQ

      • jj4211@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        2 months ago

        But that context is a mix of model output and other sources. The model output portion was generated token by token, and is combined in interesting ways with things like human response, search results, software output. It’s still a backward looking mechanism, rather than having established a concept as a goal and then trying to build the words to reach that concept like we do.

        Size and strategy for managing the context has been critical for improved subjective results, but it still doesn’t exhibit the behavior of the words as a tool to address some concept, everything about the model is about the words themselves. So we end up with something very good at generating what seems right and there’s a super high chance of what seems to be right actually being right. Especially when the software can automatically execute commands and the good or bad results reach into the context window, enabling it to effectively get automatically second guessed. The potential for automatic verification in some scenarios automatically feeding the context window is what makes it particularly appealing for software folks, though not universal.

    • MasterBlaster@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      2 months ago

      Believe it or not, there are actually people who only think in pictures and cannot actually form words in their minds.

      When I learned about that it blew my mind.

      • dil@lemmy.zip
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 months ago

        Ppl who think in abstracts, ppl who can’t visualize, people who have no inner dialouge, etc. fat variety