FYI Lumo 2.0 just dropped and it seems to be comprised of open weight models (GLM 5.2 as the big model, Qwen 3.6 35B as the small, few others to boot)
https://proton.me/support/lumo-privacy
Having played around with it a bit…I dunno. It seems pretty locked down / security theatre heavy / refusals on benign things.
I can’t run GLM locally (who can lol) but all things given I dunno how good value for money it is vs other options.
(Security and privacy promises notwithstanding that is - which perhaps, are the whole point with Lumo).
Has anyone used Lumo for any period of time? What quants do they use?
I have a strong feeling that lite is Qwen 3.5 27B and Max is GLM 5.2, based on the AAII scores on protons announcement post. 27B thinking is closest to 34 AAII and only GLM hits 51. So as fingerprints go…
But the fact that Lumo won’t reveal specifics / inhibits that at system prompt level (despite broad family id being cooked into weights) sort of belies the transparency angle for me. What else is being obfuscated? I can’t even uncover what fair use is - how many messages per hour etc
Logical inference then - opacity is driven by profit margin preservation and competitive insulation. If you can’t calculate the mark up by query (vs OR) then you might be paying 5-20x mark up.
Sits funny amidst all the privacy and transparency cosplay.
At least they’re not building autonomous weapons but treating open weight models as commercial secrets is sorta off putting.
Lumo is like my 3rd tier AI that I use, I use Mistral (paid) but if I don’t like the code it’s producing or how it’s thinking I use Claude free, if that runs out of tokens then I use GLM and if I’m not happy with that then Lumo but I’ve heard good things about 2.0
I might switch up and use Lumo after Claude if I run into issues in future
I use proton for email and cloud drive (non critical files obvs), so I’m playing around with this a bit out of curiosity (and because I can’t run 27B or GLM at speed locally).
For basic, throw away tasks (ala super Google search, fact check what I writing, edit this email etc), it seems to work OK. Worth dicking around with while building out own infra.
One interesting thing (above ZDR) that I’ve noticed is that the actual prompts seem to be obfuscated cryptogenically. Can’t say I’ve seen that before. If it’s truly E2E, that’s something.
Given that GLM is the brains of the operation, I’m willing to poke around with Proton a bit as an alternative to OAI and Anthropic for simple tasks. I’m suspicious of their claims but willing to give them a fair shake and do my own due diligence.
On that topic, some discussions uncovered
https://forum.qubes-os.org/t/lumo-protons-ai-assistant/35373
EDIT: scuttle butt has is that Lumo 2.0 Max (which is GLM 5.2 undoubtedly) has --ctx 128K. They also have compaction ; I hit auto compaction today at 68K, which is somewhat conservative.
Still no idea on quant etc but it does seem that answers a few questions - lite is most likely Qwen 3.5 27B thinking, Max is GLM5.2, context is 128K. I haven’t noticed any of the famous Qwen thinking loops yet; perhaps that explains why compaction is set to 50% total ctx.
Total usage / daily requests are still obfuscated and quants not known. The android app seems to use a very poor version of Vosk for STT (basically unusable). Image gen and OCR are decent.
In a perfect world, preference would be given to running GLM and Qwen 3.5 at home and then accessing via wireguard or head scale. This may be an interesting middle ground for (we) GPU peasants.
Not sure where I sit with this; am documenting what I find for others / not endorsement of proton (I have no relationship to them).
Given that most of us use cloud AI also, this might be worth considerations.
NB: there appears to be a Lumo plug in for VS codium; I have no indication as to if there are separate usage pools for chat and code (like ChatGPT) or all roled in one (Claude Pro). Probably the latter.
Proton is in support of fascism, don’t support Proton


