OpenAI-compatible API
Chat, images, video, audio and embeddings on one base URL. Model ids are provider/model.
| type | endpoint | example model |
|---|---|---|
| catalog | GET /v1/models | every model with its list price and Elite price |
| text | POST /v1/chat/completions | aisingapore/gemma-sea-lion-v4-27b-it |
| image | POST /v1/images/generations | alibaba/qwen-image-3.0-pro |
| video | POST /v1/media/generations | alibaba/hh1-i2v |
| audio | POST /v1/media/generations | assemblyai/universal-3-pro |
| embeddings | POST /v1/embeddings | baai/bge-base-en-v1.5 |
Chat
Standard chat completions, streaming or not. Streamed responses end with a usage chunk. Tool calls, JSON mode and vision work where the model supports them.
from openai import OpenAI
client = OpenAI(
base_url="https://zinf.ai/v1",
api_key="xk_live_...", # your Xava Inference key
)
r = client.chat.completions.create(
model="openai/gpt-5.4",
messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)curl https://zinf.ai/v1/chat/completions \
-H "Authorization: Bearer $XINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-5.4","messages":[{"role":"user","content":"Hello!"}]}'Images
curl https://zinf.ai/v1/images/generations \
-H "Authorization: Bearer $XINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"alibaba/qwen-image-3.0-pro","prompt":"a pixel-art rocket","size":"1024x1024"}'Video
curl https://zinf.ai/v1/media/generations \
-H "Authorization: Bearer $XINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"alibaba/hh1-i2v","input":{"prompt":"a pixel-art rocket lifting off","duration":4}}'Audio
curl https://zinf.ai/v1/media/generations \
-H "Authorization: Bearer $XINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"assemblyai/universal-3-pro","input":{"text":"Hello from one key."}}'Embeddings
curl https://zinf.ai/v1/embeddings \
-H "Authorization: Bearer $XINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"baai/bge-base-en-v1.5","input":"The quick brown fox"}'Native formats
The same key also speaks three providers' own APIs, for SDKs and agents built on them:
| format | endpoints | models |
|---|---|---|
| Anthropic Messages | POST /v1/messages, POST /v1/messages/count_tokens, GET /v1/models with anthropic-version | Anthropic models; native ids such as claude-haiku-4-5-20251001 or claude-sonnet-4-6 map to ours. Auth: x-api-key or Authorization: Bearer |
| OpenAI Responses | POST /v1/responses | OpenAI models (bare ids such as gpt-5.4-mini work) |
| Gemini API | POST /v1beta/models/{model}:generateContent, :streamGenerateContent?alt=sse, :countTokens, GET /v1beta/models | Gemini ids such as gemini-2.5-flash, or any chat model by its full id. Auth: x-goog-api-key |
Errors use each format's own shape. Token counts are free estimates. Setup for Claude Code, Codex and Gemini CLI: agent docs.
Media endpoints are in preview: request shapes may change before launch. Media is billed per output unit (image, second of video, character or minute of audio), shown on each model's page.
Generated files are deleted after 24 hours. Every media URL (/media/<id> on this origin) comes with expires_at (unix seconds, 24 hours after creation); download or copy the file before then. After that the URL answers 410 Gone.