This is an MoE model with 1.6T-A49B

The weights were up briefly then taken down due to some issues in the repo files apparently, now they’re back up:

GGUFs are out as well:

DeepSeek published benchmarks for reference:

    • SirDimples@programming.devOP
      link
      fedilink
      English
      arrow-up
      7
      ·
      2 days ago

      free to use if you have the hardware. For this model because of its size, the main problem is the hardware availability/cost. But in general there are 3 ways to run an open weights model:

      • pay a provider like DeepSeek/OpenRouter per usage
      • buy hardware that can run it locally: actually not a bad investment for a business
      • rent hardware that can run it from a cloud provider, hardware can be rented dedicated or time-shared (so called serverless).
      • Jo Miran@lemmy.ml
        link
        fedilink
        English
        arrow-up
        1
        ·
        1 day ago

        Is there some sort of calculator to help one determine the best model to run?

          • Jo Miran@lemmy.ml
            link
            fedilink
            English
            arrow-up
            2
            ·
            1 day ago

            I’m actually curious as to what’s the most I can run on an Apple M4 Max system.

                • e0qdk@reddthat.com
                  link
                  fedilink
                  English
                  arrow-up
                  2
                  ·
                  1 day ago

                  If I understand the nature of your hardware correctly, you should be able to run the MoE models like Gemma4 26B-A4B or Qwen3.6 35B-A3B at a high quantization fairly performantly.

                  You could try running some of the dense models (like today’s Qwen 3.8 27B) as well, but I expect they’ll be pretty slow (judging by my own experience with a unified RAM system that has a Strix Halo APU). Might still be useful for tasks that you can leave running on their own for a long time instead of for interactive chat style interaction though.

                  You’ve got enough RAM to load larger models, but there hasn’t been much released in between the “it fits on a 24GB or 32GB GPU that a gamer might own” and the “oh god you need HOW MUCH RAM!?” scales lately…