- cross-posted to:
- hackernews
- cross-posted to:
- hackernews
Crossposted from https://lemmy.ml/post/47429470
You must log in or # to comment.
My oversimplified and possibly wrong understanding: this is like speculative decoding, but instead of a separate draft model (which does its own prompt processing), they use some diffusion thing strapped on top of the main model. The diffusion reuses the high-quality prompt processing result of the main model.
The 7.8x faster claim sounds almost too good to be true. But even if we get like 3x then this is still a huge revolution in localLLMing.
deleted by creator
They said they’re working on Orthus for Qwen 3.5. It’ll be amazing!
deleted by creator


