The Coe Lab
← Back to Blog

Strata: When a 125B AI Model Ran on a Gaming PC at 100 Tokens Per Second

October 5, 20266 min read
AIopen-sourcelocal-inferenceQwenprivacy

A new open-source tool called Strata lets you run Qwen 3.8 Flash Next — a 125-billion-parameter model — on an ordinary gaming PC with an RTX 4090. Nothing leaves your machine, and it's faster than you can read.

The Frontier Model in Your Bedroom

For most of 2026, running a 125-billion-parameter AI model meant booking time on a cloud GPU cluster or paying per-token API fees that add up faster than a coffee habit. Strata, a new open-source project that shot to the top of Hacker News this week with over 800 upvotes, changes that math entirely. It runs Qwen 3.8 Flash Next — a model that rivals frontier offerings from OpenAI and Anthropic — on a consumer gaming PC. No cloud. No API. No data leaving your machine.

The numbers are genuinely startling. On an NVIDIA RTX 5070 with 12 GB of VRAM and 64 GB of system RAM, Strata generates text at 53 to 94 tokens per second depending on quantization. On an RTX 3090 with 24 GB of VRAM, community benchmarks push past 100 tokens per second. For context, the average human reads at about 4 to 5 words per second. Strata writes faster than you can keep up.

How It Actually Works

Strata is built around Qwen 3.8 Flash Next, a Mixture-of-Experts (MoE) model from Alibaba's Qwen team. MoE architectures activate only a fraction of their total parameters during inference, which means a 125B model might only use 4 to 8 percent of its parameters for any given token. That is the key trick — it is why a model that sounds impossibly large for a home GPU suddenly becomes feasible.

The tool uses quantization to compress the model further. Strata offers multiple compression levels, from Q2_0 (the most aggressive, fastest, and smallest) to IQ3_S (larger, slightly smarter, still fast). The trade-off is straightforward: smaller quantization means more speed at a small cost to reasoning quality. Even at the most compressed setting, the model handles chat, code generation, image understanding, and tool use competently.

Strata loads 35 to 55 GB of model weights into system RAM and streams the active layers to the GPU for computation. This is not a neat trick or a toy demo — it is a production-quality inference engine with a one-click installer for Windows and Linux.

What You Need to Run It

The hardware requirements are surprisingly modest for what the tool accomplishes:

  • A graphics card with 12 GB+ VRAM (NVIDIA RTX 20-series or newer, or AMD Radeon RX 6800+ series)
  • 32 GB of system RAM minimum (64 GB recommended for all quantization levels)
  • About 80 GB of disk space for model weights (an SSD makes the first load much faster)
  • Windows 10/11 or Linux with current graphics drivers

If you have a recent gaming PC, you probably already qualify. The installer detects your hardware, recommends the best model size for your RAM, and handles everything else automatically. You download, double-click, and wait.

Why This Matters Beyond Speed

Speed is the headline, but the real story is privacy and sovereignty. When you run a model locally, your data never touches a server. Your code stays on your machine. Your conversations are not logged, analyzed, or training some future model. For developers working with proprietary codebases, researchers handling sensitive data, or anyone who has read the terms of service of a cloud AI provider and felt uneasy, Strata eliminates that entire category of concern.

There is also the cost angle. A cloud API charging even a fraction of a cent per token can rack up hundreds of dollars a month for a heavy user. An RTX 3090 picked up used for around $700 runs Strata indefinitely for the cost of electricity. The economics shift dramatically when the inference is free.

The Coding Angle: AI Agents on Your Own Hardware

Strata is not just a chatbot. It supports vision (reading images and screenshots), tool use, and integration with coding agents like Claude Code, Cursor, and GitHub Copilot. There is a Coder variant that strips out half the experts to fit in 32 GB of RAM while still reaching 91 percent of the full model's SWE-bench Verified score.

This means a developer could run a competent AI coding assistant entirely offline. No API keys, no rate limits, no worrying that the provider changes the model behavior overnight. The Coder variant writes at 55 tokens per second on a mid-range RTX 5070 — plenty fast for interactive coding sessions.

Strata also ships with an MCP server, meaning AI tools can programmatically install, start, and stop the model. You can paste a single command into your coding agent and it sets up the entire local inference stack for you.

The Catch: What You Give Up

Strata is not magic. The quantization that makes it fast also costs some quality. At Q2_0 compression, the model is noticeably less sharp on complex reasoning tasks compared to the full-precision version running on a data center GPU. Chinese and other CJK language performance takes a hit with the Coder variant specifically, since half the experts are removed.

There is also the memory ceiling. While 32 GB of RAM works with the Coder variant, the full model with all experts active needs 64 GB to really shine. And the first load is genuinely heavy — Strata locks up 35 to 55 GB of RAM and can make your PC unresponsive for 1 to 3 minutes while it loads. This is not something you run alongside a video edit and a browser with 200 tabs.

But these are the trade-offs of frontier-level AI on consumer hardware. The fact that the trade-offs are this mild — slightly lower reasoning quality, a heavy memory footprint, a brief load time — is itself remarkable. A year ago, running anything close to this scale required an H100.

The Bigger Picture

Strata represents a broader shift in AI: the frontier is moving to the edge faster than anyone predicted. Open-weights models from Qwen, DeepSeek, and Mistral have closed the gap with proprietary models to the point where the difference is often imperceptible for everyday tasks. Tools like Strata, llama.cpp, and Ollama are making those models accessible to anyone with a gaming PC.

The implications are significant. If a $700 used GPU can run a model that competes with GPT-class intelligence for free, the value proposition of cloud AI APIs weakens considerably. Not for everyone — enterprises will still need the reliability, scale, and integration that managed platforms provide. But for developers, hobbyists, privacy-conscious users, and anyone in a region where API access is restricted or expensive, local inference is becoming a genuinely viable primary option.

Strata is free and open source. If you have a gaming PC collecting dust, it might be time to put those GPU cycles to work.

Related Posts

When AI Agents Spend Your Money While You Sleep: Why Hard Budget Caps Are Becoming Non-Negotiable

AWS and Google Cloud finally launched hard spending limits in the same month. It's not a coincidence — it's a response to AI agents that can rack up thousands of dollars before you wake up.

Oct 4, 2026• 6 min

When Utah Banned VPNs: How a Court Stopped a Law That Demanded the Technically Impossible

A federal judge just blocked Utah's unprecedented anti-VPN law, ruling that lawmakers cannot mandate perfect geolocation — a technical impossibility. The case reveals a deeper problem: when legislation outruns engineering.

Oct 3, 2026• 7 min

Cloudflare Clef: When the Edge Network Learned to Make Decisions

Cloudflare's new open-source decision models run at the edge with 40ms latency, beating Jev on accuracy while adding vision support and a 64k context window. Here's why decision models are the missing piece in agentic AI.

Oct 2, 2026• 6 min