model="auto"Drop-in OpenAI SDK. One base URL. The router handles the rest — models, providers, failover, billing.
# nexoauth — the whole integration from openai import OpenAI import os client = OpenAI( base_url="https://api.nexoai.dev/v1", api_key=os.environ["NEXOAI_KEY"], ) resp = client.chat.completions.create( model="auto", # routed: latency + cost + reliability messages=[{"role": "user", "content": "hello"}], ) print(resp.routed_to, resp.latency_ms) # claude-sonnet-5 412ms
OpenAI SDK, LangChain, LlamaIndex, every library that speaks OpenAI. Point them at NexoAI and the whole fleet follows.
export OPENAI_BASE_URL=https://api.nexoai.dev/v1Use "auto" or pin any of 24 models by id. Pin with a fallback list — the router keeps your preference, your users keep their responses.
model="auto" fallbacks=["claude-sonnet-5", "gpt-5.6"]Mid-stream provider failures switch routes without dropping the stream. The token stops being born in one lab and continues in another.
stream=true · failover < 300ms| model | provider | p50 latency | input / output | status |
|---|---|---|---|---|
| claude-opus-5 | anthropic | 688ms | $5.00 / $25.00 | live |
| claude-sonnet-5 | anthropic | 412ms | $2.00 / $10.00 | live |
| gpt-5.6 | openai | 298ms | $1.25 / $10.00 | live |
| gemini-3-pro | 355ms | $1.25 / $5.00 | live | |
| deepseek-v3 | deepseek | 521ms | $0.27 / $1.10 | live |
| mistral-large-3 | mistral | 391ms | $2.00 / $6.00 | live |
| +18 more | gemini-flash · grok-4.5 · llama-4 · qwen-3 … | |||
2 minutes to first request · usage-based · cancel anytime