docs
docs

Quickstart

Change two lines: the base URL and the key. Any OpenAI-compatible SDK, framework or plain HTTP works with every model in the catalog.

  1. Sign in and top up (card, or USDC on Solana).
  2. Create an API key in your dashboard. It is shown once.
  3. Point your SDK at the base URL below and pick any model id from the catalog.
environment
OPENAI_BASE_URL=https://zinf.ai/v1
OPENAI_API_KEY=xk_live_...

Life of a request

Life of a requestA request with an API key: the key and its limit of 600 requests a minute are checked; an estimate is held from the balance (the prompt's bytes divided by 4, plus 20%, as input tokens, and max_tokens or 4,096 as output, at list price); with too little available the request is refused with 402 and nothing runs; the model runs and streams its answer; the tokens actually used are priced at list; the hold is released and the price charged once. Example: a 40 KB prompt with no max_tokens at $3 and $15 per million tokens holds $0.0974 (12,000 input and 4,096 output tokens); using 10,210 and 850 it is charged $0.0434 and $0.0541 is released.KEYchecked · 600/minHOLDestimate set asideMODELruns your requestSTREAMtokens as they comeMETEREDactual tokens × listSETTLEDcharged onceenough?no: 402, nothing runsYOURACCOUNTYOUR BALANCEHOLD $0.0974 · estimateCHARGED $0.0434 · $0.0541 releasedexample: 40 KB prompt, no max_tokens, $3 / $15 per 1M tokensLife of a requestA request with an API key: the key and its limit of 600 requests a minute are checked; an estimate is held from the balance (the prompt's bytes divided by 4, plus 20%, as input tokens, and max_tokens or 4,096 as output, at list price); with too little available the request is refused with 402 and nothing runs; the model runs and streams its answer; the tokens actually used are priced at list; the hold is released and the price charged once. Example: a 40 KB prompt with no max_tokens at $3 and $15 per million tokens holds $0.0974 (12,000 input and 4,096 output tokens); using 10,210 and 850 it is charged $0.0434 and $0.0541 is released.KEYchecked · 600/minHOLDestimate set asideMODELruns your requestSTREAMtokens as they comeMETEREDactual tokens × listSETTLEDcharged onceenough?no: 402,nothing runsYOUR BALANCEHOLD $0.0974 · estimateCHARGED $0.0434 · $0.0541 releasedexample: 40 KB prompt, no max_tokens
  • KEYYour app sends the request with its API key. We check the key (we keep only a keyed hash of it) and its limit: 600 requests a minute.
  • HOLDAn estimate is set aside from your balance: the prompt's bytes ÷ 4, plus 20%, as input tokens, and max_tokens (4,096 if you leave it out) as output, at list price.
  • ENOUGH?Too little available: 402 insufficient_balance, and nothing runs.
  • MODELThe model runs. A failed attempt is retried once before the first byte (media never is).
  • STREAMThe answer streams back as it is made. A started stream is never switched mid-way.
  • METEREDThe tokens actually used are priced at list (Elite accounts: the Elite price).
  • SETTLEDThe hold is released and the price charged once, expiring credits first. 0.5% of it buys back your token (up to 2% at Elite).
python
from openai import OpenAI

client = OpenAI(
    base_url="https://zinf.ai/v1",
    api_key="xk_live_...",  # your Xava Inference key
)

r = client.chat.completions.create(
    model="openai/gpt-5.4",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)
typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://zinf.ai/v1",
  apiKey: process.env.XINF_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-5",
  messages: [{ role: "user", content: "Hello!" }],
  stream: true,
});
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
curl
curl https://zinf.ai/v1/chat/completions \
  -H "Authorization: Bearer $XINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-5.4","messages":[{"role":"user","content":"Hello!"}]}'

Privacy. We never store your prompts: we never store the content of your requests or of model responses. Zero data retention end to end on models marked ZDR supported; uploaded references and generated files auto-delete after 24 hours. On models marked ZDR supported, the model provider keeps no prompts or outputs either, so retention is zero end to end; on other models the model provider may keep request data under its own policy. Generated images, video and audio are kept at a private, unguessable link for 24 hours so you can download them, then deleted. Reference files you upload for a model to read are kept at a private, signed link for 24 hours (or until you delete them), then deleted. We keep billing metadata only.

Generated files. Images, video and audio are served from this origin and deleted 24 hours after creation (expires_at on every response; an expired URL answers 410). Download what you want to keep. Details.