Quickstart
Change two lines: the base URL and the key. Any OpenAI-compatible SDK, framework or plain HTTP works with every model in the catalog.
- Sign in and top up (card, or USDC on Solana).
- Create an API key in your dashboard. It is shown once.
- Point your SDK at the base URL below and pick any model id from the catalog.
OPENAI_BASE_URL=https://www.zinf.ai/v1 OPENAI_API_KEY=xk_live_...
Life of a request
- KEYYour app sends the request with its API key. We check the key (we keep only a keyed hash of it) and its limit: 600 requests a minute.
- HOLDAn estimate is set aside from your balance: the prompt's bytes ÷ 4, plus 20%, as input tokens, and max_tokens (4,096 if you leave it out) as output, at list price.
- ENOUGH?Too little available: 402 insufficient_balance, and nothing runs.
- MODELThe model runs. A failed attempt is retried once before the first byte (media never is).
- STREAMThe answer streams back as it is made. A started stream is never switched mid-way.
- METEREDThe tokens actually used are priced at list (Elite accounts: the Elite price).
- SETTLEDThe hold is released and the price charged once, expiring credits first. 0.5% of it buys back your token (up to 2% at Elite).
from openai import OpenAI
client = OpenAI(
base_url="https://www.zinf.ai/v1",
api_key="xk_live_...", # your Xava Inference key
)
r = client.chat.completions.create(
model="openai/gpt-5.4",
messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://www.zinf.ai/v1",
apiKey: process.env.XINF_API_KEY,
});
const stream = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5",
messages: [{ role: "user", content: "Hello!" }],
stream: true,
});
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? "");curl https://www.zinf.ai/v1/chat/completions \
-H "Authorization: Bearer $XINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-5.4","messages":[{"role":"user","content":"Hello!"}]}'Privacy. We never store your prompts: we never store the content of your requests or of model responses. Zero data retention end to end on models marked ZDR supported; uploaded references and generated files auto-delete after 24 hours. On models marked ZDR supported, the model provider keeps no prompts or outputs either, so retention is zero end to end; on other models the model provider may keep request data under its own policy. Generated images, video and audio are kept at a private, unguessable link for 24 hours so you can download them, then deleted. Reference files you upload for a model to read are kept at a private, signed link for 24 hours (or until you delete them), then deleted. We keep billing metadata only.
Generated files. Images, video and audio are served from this origin and deleted 24 hours after creation (expires_at on every response; an expired URL answers 410). Download what you want to keep. Details.