In the space of 1 week, a second open-source Chinese AI model equals the best investors are pouring tens of billions of dollars into.

schizoidman@lemm.ee · 9 months ago

In the space of 1 week, a second open-source Chinese AI model equals the best investors are pouring tens of billions of dollars into.

SmokeyDope@lemmy.world · 9 months ago

It depends on how low you’re willing to go on the quant and what you consider acceptable token speeds. Qwen 32b q3ks can be partially offloaded on my 8gb vram 1070ti and runs at about 2t/s which is just barely what I consider usable for real time conversation.