
Most big AI models live far away from you. They sit in a data center, on servers packed with hundreds of gigabytes of graphics memory, and you rent a small slice of them through an app or an API.
That's fine until you think about the bill. Or privacy. Or the day the service goes down and you're stuck staring at an error.
So when I came across a GitHub project that says you can run a 125-billion-parameter model on an ordinary gaming PC, I didn't believe it at first. I read the whole README, then the docs behind it, to see if it holds up.
The project is called Strata. This post is what I found, in plain words.
Quick note: I read the project's docs. I haven't run Strata on my own machine. The speed numbers below come from the author's own tests, and I'll say so whenever it matters.
Repo: github.com/Niko1221/Strata
So, what is Strata?
Strata is a free program that runs one specific AI model on your own computer. The model is Qwen3.8-Flash-Next, made by the Qwen team.
"125 billion parameters" just means the model has 125 billion numbers inside it. More numbers usually means a smarter model, and also a much bigger one. Models this size normally need a server room.
Once it's running, it can chat, write code, and read pictures if you turn that on. It also plugs into the apps and coding tools you already use. The project says nothing leaves your PC.
Here's the quick summary:
What it is | A local AI engine with a chat app and an API |
What it runs | Qwen3.8-Flash-Next, in compressed sizes |
Works on | Windows 10/11 and Linux |
Needs | NVIDIA or AMD graphics card with 12 GB+ of VRAM |
Price | Free, MIT license |
Popularity | Roughly 9,000 stars and 800+ forks on GitHub |
What your PC needs
This isn't a "runs on any laptop" thing. Here's what the project asks for:
Part | What you need |
|---|---|
Graphics card | NVIDIA RTX 20, 30, 40 or 50 series, or a recent AMD Radeon (RX 7900, 7800 XT, 9070 and similar), with 12 GB of VRAM or more |
RAM | 32 GB or more. This decides which model size fits |
Disk | About 80 GB free, ideally on an SSD |
System | Windows 10/11 or Linux, with an up-to-date graphics driver |
Look at the RAM line twice. Most people only think about the graphics card. With Strata, your RAM does a big part of the job. You'll see why in a minute.
How fast is it?
Speed gets measured in tokens per second. A token is a small piece of a word, about three quarters of one. So 60 tokens per second is faster than most people read.
The author measured two normal gaming PCs:
NVIDIA box: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM
AMD box: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM
"Writes" is how fast the answer appears. "Reads" is how fast it takes in what you send, like a long document.
Model size | RTX 5070 writes | RTX 5070 reads | RX 9070 XT writes | RX 9070 XT reads |
|---|---|---|---|---|
Q2_0 | 94 | 2,650 | 60 | 1,160 |
IQ2_XS | 79 | 2,090 | 52 | 1,110 |
IQ3_XXS | 62 | 1,750 | n/a | n/a |
IQ3_S | 53 | 1,620 | n/a | n/a |
Coder | 55 | 2,180 | 44 | 1,420 |
All numbers are tokens per second, from the project's README.
The empty AMD cells aren't a typo. That PC has 47 GB of RAM, and the project says the bigger sizes don't fit at 48 GB.
The pattern is simple. Smaller and more squeezed means faster. Bigger means a bit smarter and slower. The README also says an RTX 3090 (24 GB) should write somewhere around 100 to 140 tokens per second. That one is an estimate, not a measurement.
Honestly, even the slowest number in that table is quicker than I read.
How does it fit?
This is the part I liked most. A 125-billion-parameter model shouldn't fit on a card with 12 GB of memory. Strata gets around it with a few tricks.
Trick 1: Not everything works on every word
The model isn't one giant block that lights up all at once. It's built from 24,576 small specialists, which the project calls "experts". The README puts it simply: each word needs only 10 of them.
So you don't need all the experts sitting in your fastest memory at the same time. You only need the busy ones close by.
Picture a big workshop with 24,576 tools. Any single job uses about ten.
The workbench is small and very fast. That's your graphics card. It holds the tools you grab all day.
The garage is big and slower. That's your RAM. Every single tool lives there.
A second worker stands in the garage. That's your CPU. If a job needs a tool that isn't on the bench, they use it right where it sits, at the same time as you keep working on the bench.
A box of index cards sits on the shelf. That's your SSD. It holds a big 28.8 GB lookup table, and the model only reads a few small rows from it per word.
Here's the same thing as a picture:
+--------------------------------------------+
| GRAPHICS CARD (12-24 GB) the workbench |
| Fast. Does the work every word needs, |
| plus the most-used experts. |
+--------------------------------------------+
| RAM (32-64 GB+) + CPU the garage |
| Holds ALL 24,576 experts. The CPU works |
| on the ones the GPU doesn't have, at the |
| same time as the GPU. |
+--------------------------------------------+
| SSD index cards |
| A 28.8 GB lookup table. A few rows are |
| read for each word. |
+--------------------------------------------+Two things make this smarter than it first sounds.
First, the graphics card and the CPU work at the same time. Neither one sits around waiting for the other.
Second, the bench isn't fixed. The project says it keeps learning which experts you ask for most, and it adjusts while you chat.
Rule of thumb: According to the docs, more VRAM matters more than a faster GPU. Every extra GB of VRAM holds around 700 more experts, and every expert on the card is one less job for the CPU.
That's also why the README expects an RTX 3090 (24 GB) to write faster than the 12 GB card in the table above, even though the 3090 is an older model.
Trick 2: Guess, then check
Writing text one word at a time is slow, because each word needs a full trip through the model.
Strata speeds this up with a small helper built into the model. The helper quickly guesses the next few words. Then the big model checks all the guesses in one go.
Think of an intern who drafts a line, and an editor who reads the whole line at once instead of word by word. Here's a made-up example:
Helper guesses: "the" "cat" "sat"
Big model checks: OK OK not quite -> writes its own word instead
Result: three words appear in a single stepThe part that matters: the helper only guesses. The big model decides every word. So you get the same answer you'd get without the trick, just 1.6 to 1.8 times sooner.
Trick 3: Reading long text in big gulps
When you paste in a long document, Strata reads it in chunks of up to 8,192 tokens at a time. That's how it gets past 1,000 tokens per second on reading.
After your first message, it keeps the conversation in memory and only reads what's new. So follow-up questions start in seconds.
And the model is squeezed too
There's one more piece. The model files themselves are compressed, which is why you see names like Q2_0 or IQ3_S. The number in the name roughly tells you how many bits each of those 125 billion numbers gets. Fewer bits means a smaller, faster, slightly less sharp model.
A quick sum helped me believe it. 125 billion numbers at around 2.5 bits each comes to about 39 GB. The README says Strata loads 35 to 55 GB into RAM. That's the right neighborhood.
If you want the deep version, the repo has a full explanation and a paper with the measurements.
Which size should you pick?
The installer recommends one for you based on your RAM, so you don't have to guess. Here's the project's own advice:
Your RAM | Pick | Why |
|---|---|---|
32 GB | Coder | It fits, and it's made for code |
48 GB | IQ2_XS (or Q2_0 for max speed) | Bigger sizes don't fit |
64 GB | IQ2_XS, or IQ3_XXS / IQ3_S | Everything fits. IQ3_S is the smartest and slowest |
96 GB+ | IQ3_S, or Unsloth's 4-bit (experimental) | Room for the largest sizes |
A few extras worth knowing:
Coder is a coding version with half of the experts removed. The authors say it keeps 91% of the full model's score on a coding test called SWE-bench Verified. It's weaker outside code, including Chinese and other CJK text.
Swift 1.5 is a fine-tune that thinks for less time before it answers. Same quality, quicker answers.
Unsloth 4-bit is the closest to the full model, but most of it is read from the SSD while it answers. That's about 7 to 8.5 tokens per second on a 64 GB PC. It's marked experimental.
Installing it
There are two ways.
The lazy way. If you use an AI coding assistant like Claude Code, Cursor, Codex or GitHub Copilot, you can paste one line into it:
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.It checks your card, RAM and disk, picks a model that fits, installs it and starts it.
The normal way.
Download the repo as a ZIP and unzip it, or use
git clone.On Windows, double-click
START-HERE.bat. On Linux, run./setup.sh.Answer a few questions: which model, which size, how much context, and whether it should read pictures. Pressing Enter picks the recommended answer every time.
Wait for the download. It's around 70 GB, and if you stop, it continues where it left off.
Your browser opens the Strata app at
http://127.0.0.1:8080.
Heads up: While the model starts, your PC can feel slow or stop responding for 1 to 3 minutes, and longer the very first time. Strata is loading 35 to 55 GB into RAM. The project says this is normal. Wait it out and don't close the window.
Next time, run the same file again and it starts right away. Nothing gets downloaded twice. UPDATE.bat (or ./update.sh) updates it.
Using it with your own tools
The browser app has a Chat tab and a live Monitor that shows the model and your GPU, CPU and RAM use.
For your other apps, Strata pretends to be a normal AI service on your own machine. Add an "OpenAI-compatible" provider with these settings:
Base URL: http://127.0.0.1:8080/v1
API key: anything
Model: anythingApps that use Anthropic's API can use http://127.0.0.1:8080/v1/messages. For Claude Code, the README says to set ANTHROPIC_BASE_URL=http://127.0.0.1:8080.
You can also pick how hard it thinks: off, low, medium or high. Off is fastest. High is best for tricky questions.
Want to reach it from your phone or another PC? The README gives this command, and it says to always use a key:
START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>Don't skip the key. Opening Strata to your network without one means anyone on that network can use your PC to run the model.
What to know before you try it
I'd rather you hear the downsides from me than discover them after a 70 GB download.
One request at a time. It isn't built to serve a crowd.
The first message is slow. The project says it reads about 30,000 tokens per minute on that first message. Follow-ups are fast.
It's big. About 70 GB to download, 80 GB of free disk, and lots of RAM.
It's compressed. The smaller sizes are a bit less sharp than the full model. The project says so itself.
Pictures on AMD cards work on Linux through the processor, but not yet on Windows.
It moves fast. The repo has over 600 commits, so check the README for the latest before you follow any step above.
If something breaks, the README has a short troubleshooting list. The common ones are a frozen PC on first start (wait, or pick a smaller size) and "not enough RAM" slowdowns (close your browser, it eats a lot).
My take
Who is this for? Someone with a decent gaming PC, 32 GB or more of RAM, and a reason to keep their AI at home. Maybe you care about privacy. Maybe you're tired of paying per token. Maybe you just like making hardware do something it wasn't meant to.
It's also a nice fit for developers. A local, OpenAI-style endpoint that your coding tools can talk to is handy, and the setup is much gentler than I expected.
Who should skip it? Anyone on a laptop with weak graphics, a card under 12 GB, or a need to serve many people at once.
What I like most is how the README is written. It talks to normal people. It says "press Enter for the recommended answer." It warns you that your PC will freeze for a bit. It lists what goes wrong and how to fix it. That kind of honesty is rare in AI projects.
The project also gives credit where it's due. The model comes from the Qwen team. The compressed versions come from ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth. And the engine is built with parts of llama.cpp and ggml.
If you try it, I'd love to hear your graphics card, your RAM, and the tokens per second you get. Those real numbers are worth more than any chart.
Comments (1)
Join the discussion by logging into your account.
Igor Ganapolsky
The part that decides whether this stays usable as a local coding endpoint is not the 12 GB VRAM floor. It is expert residency. If each token only activates about 10 of those 24,576 experts, the GPU bench only wins while the hot set stays stable. A session that jumps between a long diff, a test log, and a new file will churn that set, and then the CPU path stops being overlapped work and becomes the decode bottleneck. The 1.6 to 1.8 times speedup from the draft helper also only holds when the guesses land. A miss still spends a full check by the big model. And because it serves one request at a time, two parallel calls do not speed the session up. They serialize behind the first cold read of 35 to 55 GB.