docs
docs

OpenAI-compatible API

Chat, images, video, audio and embeddings on one base URL. Model ids are provider/model.

typeendpointexample model
catalogGET /v1/modelsevery model with its list price and Elite price
textPOST /v1/chat/completionsaisingapore/gemma-sea-lion-v4-27b-it
imagePOST /v1/images/generationsalibaba/qwen-image-3.0-pro
videoPOST /v1/media/generationsalibaba/hh1-i2v
audioPOST /v1/media/generationsassemblyai/universal-3-pro
embeddingsPOST /v1/embeddingsbaai/bge-base-en-v1.5

Chat

Standard chat completions, streaming or not. Streamed responses end with a usage chunk. Tool calls, JSON mode and vision work where the model supports them.

python
from openai import OpenAI

client = OpenAI(
    base_url="https://zinf.ai/v1",
    api_key="xk_live_...",  # your Xava Inference key
)

r = client.chat.completions.create(
    model="openai/gpt-5.4",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)
curl
curl https://zinf.ai/v1/chat/completions \
  -H "Authorization: Bearer $XINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-5.4","messages":[{"role":"user","content":"Hello!"}]}'

Images

curl
curl https://zinf.ai/v1/images/generations \
  -H "Authorization: Bearer $XINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"alibaba/qwen-image-3.0-pro","prompt":"a pixel-art rocket","size":"1024x1024"}'

Video

curl
curl https://zinf.ai/v1/media/generations \
  -H "Authorization: Bearer $XINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"alibaba/hh1-i2v","input":{"prompt":"a pixel-art rocket lifting off","duration":4}}'

Audio

curl
curl https://zinf.ai/v1/media/generations \
  -H "Authorization: Bearer $XINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"assemblyai/universal-3-pro","input":{"text":"Hello from one key."}}'

Embeddings

curl
curl https://zinf.ai/v1/embeddings \
  -H "Authorization: Bearer $XINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"baai/bge-base-en-v1.5","input":"The quick brown fox"}'

Native formats

The same key also speaks three providers' own APIs, for SDKs and agents built on them:

formatendpointsmodels
Anthropic MessagesPOST /v1/messages, POST /v1/messages/count_tokens, GET /v1/models with anthropic-versionAnthropic models; native ids such as claude-haiku-4-5-20251001 or claude-sonnet-4-6 map to ours. Auth: x-api-key or Authorization: Bearer
OpenAI ResponsesPOST /v1/responsesOpenAI models (bare ids such as gpt-5.4-mini work)
Gemini APIPOST /v1beta/models/{model}:generateContent, :streamGenerateContent?alt=sse, :countTokens, GET /v1beta/modelsGemini ids such as gemini-2.5-flash, or any chat model by its full id. Auth: x-goog-api-key

Errors use each format's own shape. Token counts are free estimates. Setup for Claude Code, Codex and Gemini CLI: agent docs.

Media endpoints are in preview: request shapes may change before launch. Media is billed per output unit (image, second of video, character or minute of audio), shown on each model's page.

Generated files are deleted after 24 hours. Every media URL (/media/<id> on this origin) comes with expires_at (unix seconds, 24 hours after creation); download or copy the file before then. After that the URL answers 410 Gone.