Running a 177B Parameter Model Locally on a Single RTX 3090
This demo shows that
Qwen3.8-Flash-Next, a model with roughly 177 billion parameters and a size of 111 GB, can run entirely locally on a single RTX 3090 with 24 GB of VRAM and 128 GB of system RAM.
How Is This Possible?
The key is the
Mixture of Experts (MoE) architecture. Of the 125 billion parameters in the actual backbone, only about 6 billion are activated per token. In addition, the model includes N-gram embedding tables with around 51 billion parameters. These act as lookup memory, are kept in system RAM and are only accessed when needed.
Running
llama.cpp as a service, the required model weights are distributed between system RAM and GPU memory.
The result shown below was generated in
35 minutes, in a single run, without any manual corrections.
Setup
|
Component
|
Value
|
|
System
|
Linux Mint 22.3
|
|
GPU
|
RTX 3090, 24 GB VRAM
|
|
RAM
|
128 GB DDR4
|
|
Service
|
llama.cpp
|
|
Model
|
Qwen3.8-Flash-Next (MoE)
|
|
Model size
|
111 GB
|
|
Parameters
|
177B
|
|
Quantization
|
Q4_K_M
|
|
Context size
|
128K
|
|
Reasoning
|
On
|
|
Input speed
|
~140 tokens/s
|
|
Output speed
|
~19 tokens/s
|
|
Duration
|
35 min
|
Result
https://www.wartris.com/gfx/solarsystem/
Prompt
Create a 3D model of the solar system with the planets in HTML/JavaScript/CSS. When a planet is selected, the view zooms in and information about the planet is displayed on the right. You can simply have the planets orbit the sun, but don't do any special calculations, just animate them. No timescale calculations, keep it simple, but visually appealing.
Photos
Locally
AI
RTX
3090
Qwen
Views: 17