OPEN SOURCE Strata Engine by Niko1221 • Run 125B+ MoE Models (Qwen3.8-Flash-Next) on 12GB VRAM GPUs!
⚡
STRATA LLM
Open-Source Local MoE Inference Engine

RUN 125B+ MoE ON CONSUMER GPUS

Powered by Niko1221's dynamic sparse router offloading. Run server-grade frontier models like Qwen3.8-Flash-Next 125B on a single 12GB NVIDIA graphics card and 64GB system RAM—with zero cloud dependencies.

One-Click Quickstart (Windows PowerShell / CMD)
MIT License
git clone https://github.com/Niko1221/Strata.git
cd Strata
.\START-HERE.bat
Min GPU VRAM 12 GB VRAM
Peak Speed 25 - 38 tok/s
API Standard OpenAI & Anthropic
Author Niko1221
Engineering Architecture

How Strata Solves the VRAM Wall

Traditional inference engines crash when a 125B model exceeds GPU memory. Strata utilizes the sparse nature of Mixture-of-Experts architectures to stream only active weights:

01

Predictive Sparse Routing

In MoE models like Qwen 125B, only 2-8 out of 64 experts fire per token. Strata predicts which expert layers will be called 1-2 tokens in advance, queuing them for instant DMA PCIe transfer.

02

Tiered VRAM-RAM Cache

Critical attention heads and frequently triggered experts reside permanently inside your 12GB GPU VRAM. The remaining 90GB of dormant weights stream through high-speed 64GB DDR5 memory.

03

Zero-Copy Pinned Memory

Weights are transferred over asynchronous CUDA streams via PCIe 4.0/5.0 directly into tensor cores, completely bypassing CPU compute cycles and preventing stutter.

Hardware Blueprint

System Requirements Matrix

Tested on Windows 11 & Ubuntu 24.04 LTS
Minimum Required Entry Level
Graphics Card (GPU) NVIDIA RTX 3060 / 4060 / 3070 12 GB VRAM
System RAM 64 GB DDR4/DDR5 (Dual Channel)
Storage PCIe 3.0/4.0 NVMe SSD (50GB+ Free)
Estimated Generation Speed 12 - 18 Tokens/sec (MoE Offloaded)
Recommended Models:
  • Qwen3.8-Flash-Next 125B (4-bit Q_K)
  • DeepSeek-V2-Lite 16B
Sweet Spot Recommended
Graphics Card (GPU) NVIDIA RTX 4070 Ti / 4080 / 3090 16 GB - 24 GB VRAM
System RAM 64 GB - 128 GB DDR5
Storage PCIe 4.0 NVMe SSD (7000 MB/s read)
Estimated Generation Speed 25 - 38 Tokens/sec
Recommended Models:
  • Qwen3.8-Flash-Next 125B (5-bit Q_M)
  • Mixtral 8x22B Instruct
Peak Velocity Enthusiast / Studio
Graphics Card (GPU) NVIDIA RTX 4090 (24 GB) or Dual RTX 3090 24 GB - 48 GB VRAM
System RAM 128 GB DDR5 Quad Channel
Storage PCIe 5.0 / Gen4 NVMe DirectStorage
Estimated Generation Speed 45 - 65 Tokens/sec
Recommended Models:
  • Full-context 125B-236B MoE Models
  • Continuous Batching Serving
Empirical Results

Engine Performance Comparison

Testing Qwen3.8-Flash-Next (125 Billion Parameters) on a standard single RTX 4070 (12GB VRAM) paired with 64GB DDR5:

Inference Engine Hardware Target Tokens / Sec RAM Allocation Status & Stability
⚡ Strata LLM (Niko1221) RTX 4070 (12GB) + 64GB RAM 21.4 tok/s 52 GB Native Support (Zero OOM Crash)
Ollama (Standard) RTX 4070 (12GB) + 64GB RAM Fail / OOM Exceeds limit Out of Memory / Requires Quantization Drop
vLLM (CPU-Offload) RTX 4070 (12GB) + 64GB RAM 3.8 tok/s 58 GB Severe PCIe Bus Bottleneck
Developer Q&A

Frequently Asked Questions

Everything you need to know about setting up Strata and configuring MoE quantization.

What is Strata and who developed it? +
Strata is an open-source local LLM inference engine developed by Niko1221 (available on GitHub under the MIT License). It is specifically engineered to run massive Mixture-of-Experts (MoE) models—such as Qwen3.8-Flash-Next 125B—on everyday consumer hardware by dynamically offloading dormant expert layers between GPU VRAM, system RAM, and high-speed NVMe storage.
What are the minimum hardware requirements to run Strata? +
The bare minimum setup requires an NVIDIA graphics card with at least 12GB of VRAM (RTX 3060 12GB, RTX 4060 Ti 16GB, RTX 3070/4070) and 64GB of system RAM. A fast NVMe SSD (PCIe 3.0 or 4.0) is strongly recommended for seamless weight streaming.
How is Strata different from Ollama, llama.cpp, and vLLM? +
Traditional engines like Ollama and vLLM typically require the entire model or its active layer pipeline to fit squarely into GPU VRAM, causing out-of-memory (OOM) crashes on 125B+ models unless running on multiple expensive A100/H100 enterprise GPUs. Strata uses a predictive MoE sparse router cache that keeps only frequently activated experts in VRAM while streaming dormant experts directly across system RAM in real time, achieving 18-35+ tokens/second on consumer gaming rigs.
Does Strata offer an OpenAI-compatible API? +
Yes! Strata runs a high-performance local web server listening on port 8080 by default. It provides drop-in /v1/chat/completions and /v1/models endpoints fully compatible with OpenAI and Anthropic client SDKs, Continue.dev, Cursor, Open-WebUI, and LangChain.
How do I install and launch Strata on Windows? +
Windows installation is straightforward: Clone the official repository or download the release zip, run `START-HERE.bat`, which automatically prepares the Python virtual environment, configures CUDA 12.x wheels, downloads the quantized Qwen weights, and boots the local API server and Web UI.
Can I use Strata on Linux or WSL2? +
Yes, Strata fully supports Ubuntu 22.04+, Debian, Arch Linux, and Windows Subsystem for Linux (WSL2) with NVIDIA Container Toolkit support. Run `./setup.sh` to install system dependencies and launch the background daemon.