Strata
Run a 125-billion-parameter AI model on your own gaming PC
NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source
A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) · full video (49 s)
Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, smart AI model that usually needs a server. It chats, writes code, reads pictures and works with your apps and coding agents. Nothing leaves your PC.
How fast is it?
We measured it on two ordinary gaming PCs. A token is about ¾ of a word.
- Writes answers: how fast the reply appears in a short chat. 60 tokens per second is faster than you can read.
- Reads your prompt: how fast it takes in what you send (here a 32K-token document, code or chat history).
| NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM | AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM |
|---|---|
| Size | |
| --- | --: |
| Q2_0 | 94 tokens/s |
| IQ2_XS | 79 tokens/s |
| IQ3_XXS | 62 tokens/s |
| IQ3_S | 53 tokens/s |
| Coder | 55 tokens/s |
| Size | |
| --- | --: |
| Q2_0 | 60 tokens/s |
| IQ2_XS | 52 tokens/s |
| Coder | 44 tokens/s |
NVIDIA: Q2_0 with engine 0.1.36, the other rows with 0.1.26 (4K answers, 32K prompts). The full tables are in DETAILS.md. A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and other cards: speed of each model, community results.
Strata is free. If it runs well on your PC, a coffee keeps the work on it going.
What you need
| Graphics card | NVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more. |
| RAM | 32 GB or more. Your RAM decides which model fits. 64 GB runs every size. |
| Disk | About 80 GB free. Use an SSD if you can: the first start is much faster. |
| System | Windows 10 / 11 or Linux, and a current graphics driver from NVIDIA or AMD. |
The installer sets up everything else. Two or three cards can share the model ( multi-GPU).
Experimental, written and tested by community members on their own machines:
- Older graphics cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs.
- Intel Arc, built from source on Linux: Intel Arc.
- Older processors without AVX2: they work, but slowly. Older CPUs.
The full list: docs/INSTALL.md.
Install
Let your AI set it up
Do you use an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, ...)? Paste this into it:
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
It checks your graphics card, RAM and disk and picks the model that fits. Then it installs and starts it and tells you how to connect your apps. AI tools can also install, start and stop Strata through its MCP server.
Or do it yourself
Download Strata and unzip it (or git clone it).
Windows: double-click START-HERE.bat. Linux: run ./setup.sh in the Strata folder.
The steps are the same for NVIDIA and AMD. The installer finds your card and sets up the right engine for it. It asks you a few questions:
- which model and which size,
- how much context (how much text the model keeps in mind),
- whether it should read pictures.
Press Enter each time for the recommended answer. Then it downloads the model (about 70 GB) and starts it. If the
download stops, run it again: it continues where it left off. Your browser opens the Strata app at
http://127.0.0.1:8080.
While the model starts, your PC can be slow or stop responding for 1-3 minutes (longest the first time). Strata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don't close the





