• partofthevoice@lemmy.zip
      link
      fedilink
      English
      arrow-up
      3
      ·
      2 days ago

      The head of IT at my org is dying on this hill. He says specialized models with proper routing can outperform the big boys. … I really want him to be right, but I don’t believe it. I work with both. The 27B models are like working with ChatGPT on release day.

      • theneverfox@pawb.social
        link
        fedilink
        English
        arrow-up
        2
        ·
        2 days ago

        Well of course they outperform them if they’re specialized to a purpose… Theres a like 3B model that plays Minecraft and generally kicks the ass of all the more general llms

        And in general, the big models are primarily better at remembering instructions and context… An 8B model can hold a basic conversation or perform simple tasks on a similar level to a frontier model

        Really, really depends on what they’re specialized to do though