docs
models / Meta / Llama 2 7B Chat HF LoRA

Llama 2 7B Chat HF LoRA

textmeta-llama/llama-2-7b-chat-hf-loraZDR supported

Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 7B fine-tuned model, optimized for dialogue use cases and converted for the Hugging Face Transformers format.

$0.011list priceper 1K compute units
pricing
unitlist → Elite price
per 1K compute units$0.011

No Elite discount on this model right now: everyone pays the list price.

Everyone pays the list price. Reaching Elite (hold or stake 1,337 $XAVA, or hold 1,337,000 $XINF) unlocks the Elite price and a bigger buyback of your token. How it works

openai sdk (python)
from openai import OpenAI

client = OpenAI(
    base_url="https://zinf.ai/v1",
    api_key="xk_live_...",  # your Xava Inference key
)

r = client.chat.completions.create(
    model="meta-llama/llama-2-7b-chat-hf-lora",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)
curl
curl https://zinf.ai/v1/chat/completions \
  -H "Authorization: Bearer $XINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"meta-llama/llama-2-7b-chat-hf-lora","messages":[{"role":"user","content":"Hello!"}]}'
limits
providerMeta
typetext
context window8,192 tokens
max output—
prompt storagenone
provider retentionnone ZDR supported
livesoon
requests 24h—
p50 latency—
uptime 7d—