▲ 5 signals · 2 sources
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
I wanted to see how far I could push a fairly ordinary laptop with a huge MoE model. Turns out, Qwen3.8 Flash Next 176B can run on: RTX 3080 Laptop — 16GB VRAM 32GB system RAM SSD No 128GB/256GB RAM workstation and no multi-GPU setup. I’m running it with TensorSharp , my open-source local LLM inference engine: TensorSharp on GitHub The interesting part for me wasn't simply getting a 176B model to…
github.com ·