An open-source inference engine called Strata lets consumer gaming PCs run Qwen3.8-Flash-Next, a 125 billion parameter model, by splitting the computation across GPU, RAM and SSD, according to the project's GitHub repository.
Installation is a single script, START-HERE.bat on Windows or setup.sh on Linux, that detects the machine's hardware, recommends a model size based on available memory, downloads about 70 gigabytes of weights and launches a local web interface, the repository says. Strata supports several compressed versions of the model, including coder-focused and faster "Swift" variants.
The project is released under the MIT license and has drawn more than 20,000 stars and over 1,500 commits, with a local API compatible with both OpenAI and Anthropic formats, according to the repository.
Running a 125 billion parameter model on a gaming rig instead of a cloud GPU cluster is the kind of hardware trick that quietly moves running a frontier-scale model yourself from a hobbyist stunt to something a real team could build on.