• notfromhere@lemmy.ml
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    1
    ·
    9 days ago

    the 6B actives should easily fit into an 8GB card

    That’s not how MoE models work. There are many “expert’ models and there is a static router model (dense) which determines which “expert” models to route the tokens through. What you want at a minimum is the dense portion of the model to be on VRAM and all of the weights to be in RAM/VRAM for best performance.