REST API & SDKs
OpenAI & Anthropic Compatible API
Strata includes an embedded high-concurrency web server listening on http://localhost:8080. It functions as a drop-in replacement for OpenAI and Anthropic cloud APIs.
OpenAI Compatible Endpoint
http://localhost:8080/v1 Supports /chat/completions, /completions, and /models Anthropic Compatible Endpoint
http://localhost:8080/anthropic/v1 Supports /messages format with streaming 1. Python Client (OpenAI SDK)
pip install openaifrom openai import OpenAI
client = OpenAI(
base_url="http://localhost:8080/v1",
api_key="strata-local-token" # Any string works locally
)
response = client.chat.completions.create(
model="qwen3.8-125b-q4",
messages=[
{"role": "system", "content": "You are a senior system architect."},
{"role": "user", "content": "Explain how sparse MoE routing works in 3 bullet points."}
],
temperature=0.7,
stream=True
)
for chunk in response:
content = chunk.choices[0].delta.content or ""
print(content, end="", flush=True) 2. TypeScript / Node.js
npm install openaiimport OpenAI from 'openai';
const openai = new OpenAI({
baseURL: 'http://localhost:8080/v1',
apiKey: 'strata-local',
});
async function main() {
const stream = await openai.chat.completions.create({
model: 'qwen3.8-125b-q4',
messages: [{ role: 'user', content: 'Write a quicksort in Rust.' }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || '');
}
}
main(); 3. Continue.dev (VS Code / JetBrains) Config
Add the following block into your ~/.continue/config.json to use Strata as your free offline local coding assistant:
{
"models": [
{
"title": "Strata Qwen 125B (Local)",
"provider": "openai",
"model": "qwen3.8-125b-q4",
"apiBase": "http://localhost:8080/v1",
"apiKey": "local"
}
]
}