OPEN SOURCE Strata Engine by Niko1221 • Run 125B+ MoE Models (Qwen3.8-Flash-Next) on 12GB VRAM GPUs!
⚡
STRATA LLM
REST API & SDKs

OpenAI & Anthropic Compatible API

Strata includes an embedded high-concurrency web server listening on http://localhost:8080. It functions as a drop-in replacement for OpenAI and Anthropic cloud APIs.

OpenAI Compatible Endpoint http://localhost:8080/v1 Supports /chat/completions, /completions, and /models
Anthropic Compatible Endpoint http://localhost:8080/anthropic/v1 Supports /messages format with streaming

1. Python Client (OpenAI SDK)

pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8080/v1",
    api_key="strata-local-token"  # Any string works locally
)

response = client.chat.completions.create(
    model="qwen3.8-125b-q4",
    messages=[
        {"role": "system", "content": "You are a senior system architect."},
        {"role": "user", "content": "Explain how sparse MoE routing works in 3 bullet points."}
    ],
    temperature=0.7,
    stream=True
)

for chunk in response:
    content = chunk.choices[0].delta.content or ""
    print(content, end="", flush=True)

2. TypeScript / Node.js

npm install openai
import OpenAI from 'openai';

const openai = new OpenAI({
  baseURL: 'http://localhost:8080/v1',
  apiKey: 'strata-local',
});

async function main() {
  const stream = await openai.chat.completions.create({
    model: 'qwen3.8-125b-q4',
    messages: [{ role: 'user', content: 'Write a quicksort in Rust.' }],
    stream: true,
  });

  for await (const chunk of stream) {
    process.stdout.write(chunk.choices[0]?.delta?.content || '');
  }
}

main();

3. Continue.dev (VS Code / JetBrains) Config

Add the following block into your ~/.continue/config.json to use Strata as your free offline local coding assistant:

{
  "models": [
    {
      "title": "Strata Qwen 125B (Local)",
      "provider": "openai",
      "model": "qwen3.8-125b-q4",
      "apiBase": "http://localhost:8080/v1",
      "apiKey": "local"
    }
  ]
}