# Xava Inference > Hold $XINF, earn free AI credits. One OpenAI-compatible API key for text, image, video, audio and embedding models from many providers, at list prices (Elite prices for $XINF and XAVA holders). Every prompt buys back a token you pick: 0.5% of each spend, up to 2% with Elite Reward Status. Prepaid credit backed by USDC in a public reserve (card, or USDC on Solana for 5% less). Agents can also pay per media request with x402 (USDC on Solana), no account needed. Base URL: https://www.zinf.ai/v1 (OpenAI-compatible; send `Authorization: Bearer `). Model ids are `provider/model`, e.g. `openai/gpt-5.4`, `meta/llama-3.1-8b-instruct-fp8`. ## In plain language - What is Xava Inference? One API key for text, image, video, audio and embedding models from many model makers, at the model makers' own list prices. It is OpenAI-compatible: change the base URL and the key. - How is it priced? Everyone pays each model's list price, per unit (per 1M tokens, per image, per second of video). Elite accounts pay the model's Elite price instead: https://www.zinf.ai/models - What is Reward Status? Every account starts as Noob. Hold or stake 1,337 XAVA on Avalanche, or hold 1,337,000 $XINF on Solana, in a wallet linked to your account to reach Elite: Elite prices on every model and a bigger buyback. - How do I pay? Prepaid credit, no subscription: by card at the posted price, or with USDC on Solana and save 5% ($100 of credit costs $105 by card or $100 in USDC). Credit is not redeemable for cash. - What is the buyback token? Every account picks one: $XINF by default, or any eligible Solana token. 0.5% of every spend buys it back (Elite: 2% for $XINF, 1.5% for a custom token), out of our margin, never on top of the price. Buybacks are executed on-chain and published. - What rewards are there? $XINF is paired with USDC. 3% Hold Rewards Tax: every $XINF transfer pays holders in USDC. Our 33% lock earns its share in USDC; an on-chain autobuy share (0-70%, set by a public controller multisig) buys $XINF back into the lock, and the rest becomes AI credits for holders (hold at least 1,337 $XINF on the hourly average). Claim within 30 days; earned and promotional credits expire after 365 days unused, and credits you buy don't expire; burned credits go 50% to XAVA stakers and 50% to $XINF buybacks locked forever. Details: https://www.zinf.ai/rewards - What is the 3% Hold Rewards Tax? 3% of every $XINF transfer, wallet to wallet included, is paid to holders in USDC by the launchpad (StonkFun), automatically, with nothing to claim. - Do you keep my prompts? No. We never store your prompts or replies; we keep billing details only: model, amount and time. Zero data retention end to end on models marked ZDR supported; uploaded references and generated files auto-delete after 24 hours. On a model marked ZDR supported, neither we nor the model provider store your prompts or outputs; on other models the model provider's own retention policy applies. - How long are generated files kept? Images, video and audio are served from this site and deleted 24 hours after creation: every media response carries `expires_at` (unix seconds), and an expired link answers 410 Gone. Download what you want to keep. - Can I give a model my own image, video or audio? Yes: pass an https URL in the model's reference parameter, or upload the file (`POST /v1/uploads`) and pass the upload id. Uploads are charged as "Reference storage" (from $0.001 a file) and deleted 24 hours after upload. - How do coding agents use it? Install the plugin or add the remote MCP server (https://www.zinf.ai/agents), then sign in once in the browser: the agent spends the account balance within the daily cap picked at consent. No wallet or private key is ever needed on the user's machine. - Can agents pay without an account? Only agents that already have a wallet provider with spending policies (e.g. a hosted or custodial agent wallet): image, video and audio can be paid per request in USDC on Solana with x402 (https://www.zinf.ai/docs/x402). ## Pages - [Models and pricing](https://www.zinf.ai/models): every model with its list price and Elite price - [Rewards](https://www.zinf.ai/rewards): AI credits for $XINF holders and USDC for XAVA stakers - [Agents](https://www.zinf.ai/agents): the plugin, MCP server and x402 for AI agents - [Whitepaper](https://www.zinf.ai/whitepaper): the whole design, including risks - [$XINF](https://www.zinf.ai/xinf): the token and the protocol lock ## Docs - [Quickstart](https://www.zinf.ai/docs/quickstart): base URL + key, Python / TypeScript / curl - [API keys](https://www.zinf.ai/docs/api-keys): creating, sending and revoking keys - [OpenAI-compatible API](https://www.zinf.ai/docs/api): chat, images, video, audio, embeddings - [Models & pricing](https://www.zinf.ai/docs/models): ids, units, list price vs Elite price - [Errors & rate limits](https://www.zinf.ai/docs/errors): OpenAI error shape, status codes, retries - [Billing](https://www.zinf.ai/docs/billing): prepaid credit, card and USDC top-ups ## Agents - [MCP server](https://www.zinf.ai/docs/mcp): remote MCP at https://www.zinf.ai/mcp (Streamable HTTP), tools list_models, get_model_pricing, chat, generate_image, generate_video, generate_audio, get_balance, get_usage_summary - [Agent plugins](https://www.zinf.ai/docs/agents): setup for Claude Code, Codex, Cursor, Gemini CLI, OpenCode, Kimi Code, Pi, Cline, Windsurf, Claude Desktop and any OpenAI SDK (each install ends with the client's own browser sign-in) - [x402 pay-per-request](https://www.zinf.ai/docs/x402): for agents with a wallet provider that enforces spending policies, pay image / video / audio requests in USDC on Solana without an account ## Machine-readable - [OpenAPI spec](https://www.zinf.ai/openapi.json): the /v1 API - [Model catalog](https://www.zinf.ai/v1/models): every model with its list price and Elite price (JSON, no key needed) - [Full text for LLMs](https://www.zinf.ai/llms-full.txt): these docs expanded, with the model list ## Optional - [Buyback token & Reward Status](https://www.zinf.ai/docs/buyback): how every spend buys back $XINF or a token you choose, and what reaching Elite gives - [Buyback proofs and credit reserves](https://www.zinf.ai/buybacks) - [Rewards](https://www.zinf.ai/rewards): AI credits for $XINF holders and USDC for stakers, claimed weekly in the dashboard (https://www.zinf.ai/dashboard/rewards) --- # Xava Inference, summarised for LLMs ## The product Xava Inference sells AI inference at list price: one OpenAI-compatible API key reaches every text, image, video, audio and embedding model in the catalog, from many model makers. Requests go to the model you name; nothing is swapped or downgraded. Prompts and responses are never stored; uploaded references and generated files auto-delete after 24 hours. Base URL: https://www.zinf.ai/v1. ## The pricing model - List price: the model maker's own published price, what everyone pays. Elite price: the list price minus the model's Elite discount, for Elite accounts (a linked wallet holding 1,337 XAVA, balance plus credited stake, or 1,337,000 $XINF). An Elite price never goes below what the model costs us. Both are shown per unit on every model page and in `GET /v1/models` (`list_usd`, `elite_usd`). - Credit is prepaid (minimum top-up $5): by card at the posted price ($100 of credit costs $105), or with USDC on Solana, 1:1 (save 5%). One credit is $1 of inference, held as a non-transferable token on Solana and backed 1:1 by USDC in a public reserve (https://www.zinf.ai/buybacks#reserves). Credit cannot be cashed out. - Your buyback token: one account setting, $XINF by default or any eligible Solana token. 0.5% of every spend buys it back (Elite: 2% for $XINF, 1.5% for a custom token), taken from our margin and capped at it. Buybacks are batched, executed on-chain and published with their transactions on https://www.zinf.ai/buybacks. - Where the margin goes: our margin on a request is what it was charged minus what it actually cost us upstream. Your buyback comes out of it first; the rest is split at every burn, 50% goes to XAVA stakers in USDC and 50% buys $XINF into the protocol lock (never sold). There is no team share. - Credit expiry: earned and promotional credits expire after 365 days unused; credits you buy don't expire. Requests spend your expiring credits first, oldest first, so they rarely go unused; credits you buy are used after them. Expired credits are burned through the same split. ## Rewards - $XINF holders: $XINF is paired with USDC. 3% Hold Rewards Tax: every $XINF transfer pays holders in USDC, directly. Our 33% protocol lock (bought at launch, plus every buyback) earns its share in USDC. The credits program splits it on-chain: an autobuy share (0-70%, hard-capped at 70%, set by a public controller multisig) buys $XINF in the main pool straight into the lock, and the rest becomes AI credits for $XINF holders, shared every hour by time-weighted balance. A wallet needs an hourly average of at least 1,337 $XINF to earn in that hour. Credits are claimed weekly in the dashboard (https://www.zinf.ai/dashboard/rewards). - XAVA stakers: 50% of every burn in USDC (the margin of spent credits, and all of the USDC of unclaimed or expired credits), weighted by credited stake in the Avalaunch staking contract (Avalanche C-chain) and a loyalty multiplier that rises to 3x over 180 days. Only staked XAVA counts. - Claim within 30 days: every weekly claim list is published with a fingerprint anyone can recompute. Unclaimed holder credits are burned: 50% to XAVA stakers, 50% to $XINF buybacks locked forever. - Risks: the launchpad controls the 3% rate and whether the lock is paid, and may change either; the whitepaper lists every risk (https://www.zinf.ai/whitepaper#risks). $XINF is a utility token; nothing here is investment advice. --- # Xava Inference API reference (for LLMs) ## Authentication Every /v1 request takes `Authorization: Bearer ` (also accepted: `x-api-key`, for Anthropic-style SDKs). Keys look like `xk_live__` and are created in the dashboard (https://www.zinf.ai/dashboard). The key carries the account's balance, buyback token and Reward Status. Without a key, requests get 401 -- except the media endpoints, which answer 402 with x402 payment requirements when x402 is enabled (see below). ## Endpoints - `POST /v1/chat/completions` -- OpenAI chat completions, `stream: true` supported (SSE, final usage chunk). Tools, JSON mode and vision where the model supports them. Anthropic, Google and other models all work through this one shape. - `POST /v1/messages` -- Anthropic Messages format for Anthropic models (Anthropic SDKs work with base URL https://www.zinf.ai). - `POST /v1/responses` -- OpenAI Responses format for OpenAI models. - `POST /v1/embeddings` -- OpenAI embeddings (`encoding_format`: float or base64). - `POST /v1/images/generations` -- OpenAI images: `{model, prompt, n?, size?, response_format?}`; returns `data[].url` (and `b64_json` if asked). - `POST /v1/media/generations` -- any video, audio or other media model: `{"model": "...", "input": {...model parameters}}`; returns `{id, output, files[]}` with URLs on this origin. - Generated files (images, video, audio) are deleted 24 hours after creation: every media response carries `expires_at` (unix seconds; per file and top level), and an expired URL answers 410 Gone. Download what you want to keep. - `POST /v1/uploads` -- upload a reference file for models that read one (Seedance references, image-to-video, video-to-video, image edit, speech-to-text): send the file as the raw body (`Content-Type` any, `Content-Length` required; up to 95 MB) or as multipart/form-data `file`; returns `{id, url, mime, kind, bytes, created_at, expires_at, status, cost_usd}`. Files over 95 MB (up to 200 MB): POST `{"bytes": N, "mime": "video/mp4"}` first, PUT each `part_urls[i]`, then POST `complete_url`. Types are read from the file's bytes: images JPEG, PNG, WebP, GIF, BMP, TIFF, HEIC/HEIF, AVIF; video MP4, MOV, WebM; audio MP3, WAV, FLAC, OGG/Opus, M4A, AAC. Limits per file: images 30 MB, video 200 MB, audio 100 MB (smaller where a model says so); at most 100 live files / 2 GB per account. Charged once as "Reference storage": $0.02 per GB for the 24 hours, minimum $0.001 per file, from credits first then cash; no rewards. Pass the `id` (`upl_...`) wherever a model takes a file (e.g. `"input": {"image": "upl_..."}`) and we send the model our signed link (or the bytes, for models that take files inline); an https URL works too (plain http links are refused). `GET /v1/uploads[/{id}]`, `DELETE /v1/uploads/{id}`. Uploads are deleted 24 hours after upload; an expired upload answers 410. Uploads need an account (API key or signed-in Playground); x402 callers pass URLs. - `GET /v1/models`, `GET /v1/models/{id}` -- the catalog; each entry has `x_pricing` (per unit: `list_usd` = the list price, `elite_usd` = the Elite price) and `x_elite_discount_bps`. No key needed. Endpoint per model type: text -> /v1/chat/completions, embeddings -> /v1/embeddings, image -> /v1/images/generations, video -> /v1/media/generations, audio -> /v1/media/generations, other -> /v1/media/generations. Every response has `x-request-id`; billed responses also have `x-xm-cost-usd` (what you paid), `x-xm-list-usd` (the list price) and `x-xm-buyback-usd` (what this request bought back of your token). ## Examples ```bash curl https://www.zinf.ai/v1/chat/completions \ -H "Authorization: Bearer $XINF_API_KEY" -H "Content-Type: application/json" \ -d '{"model":"meta/llama-3.1-8b-instruct-fp8","messages":[{"role":"user","content":"Hello"}]}' curl https://www.zinf.ai/v1/images/generations \ -H "Authorization: Bearer $XINF_API_KEY" -H "Content-Type: application/json" \ -d '{"model":"black-forest-labs/flux-1-schnell","prompt":"a lighthouse at dusk, pixel art"}' ``` ```python import os from openai import OpenAI client = OpenAI(base_url="https://www.zinf.ai/v1", api_key=os.environ["XINF_API_KEY"]) r = client.chat.completions.create(model="openai/gpt-5.4", messages=[{"role": "user", "content": "Hello"}]) ``` ## Errors OpenAI's error shape: `{"error": {"message", "type", "code", "param"}, "request_id"}`. 400 invalid_request, 401 missing_api_key / invalid_api_key, 402 insufficient_balance (top up) or x402 payment required (keyless media), 403 account_frozen, 404 model_not_found, 429 rate_limit_exceeded (honour Retry-After), 503 no_capacity (retry shortly), 504 timeout (media; not charged). ## MCP server Remote MCP server: `https://www.zinf.ai/mcp` (Streamable HTTP, stateless, JSON responses). Sign-in (MCP authorization spec, OAuth 2.1 + PKCE): add `https://www.zinf.ai/mcp/account` to any MCP client that supports OAuth. It answers an unauthenticated connection with 401 and `WWW-Authenticate: Bearer resource_metadata="https://www.zinf.ai/.well-known/oauth-protected-resource/mcp/account"`; the client registers itself (dynamic client registration or a Client ID Metadata Document), opens the browser, the user signs in and allows the app with a daily spending cap, and the client gets short-lived access tokens with rotating refresh tokens. Metadata: `https://www.zinf.ai/.well-known/oauth-authorization-server`. On `https://www.zinf.ai/mcp` only the account tools (chat, get_balance, get_usage_summary) answer 401 without a credential; the catalog and x402 media tools stay open. Clients without OAuth: the device login (`POST https://www.zinf.ai/oauth/device`, confirm at `https://www.zinf.ai/login/device`) yields a connection key for `Authorization: Bearer`. An API key (`Authorization: Bearer `) works everywhere too. Tools: - `list_models` {type?, provider?, search?, limit?} -- catalog search with list price per unit (no key needed) - `get_model_pricing` {model} -- list price and Elite price for one model (no key needed) - `chat` {model, prompt | messages, system?, max_tokens?, temperature?} -- non-streaming chat (key) - `generate_image` {model, prompt, size?, n?, params?, buyback_token?} -- returns file URLs (key or x402; `buyback_token` applies to x402 calls) - `generate_video` {model, prompt, duration?, resolution?, params?, buyback_token?} (key or x402) - `generate_audio` {model, text? | prompt?, params?, buyback_token?} (key or x402) - `get_balance` {} -- available and held credit (key) - `get_usage_summary` {days?} -- spend, requests and buybacks, total and per model (key) Claude Code, the plugin in one line (the last command opens the browser sign-in): `claude plugin marketplace add xavadao/xinf-plugin && claude plugin install xinf@xinf --config base_url=https://www.zinf.ai && claude mcp login plugin:xinf:xinf`; or the MCP server alone: `claude mcp add --transport http xinf https://www.zinf.ai/mcp/account && claude mcp login xinf`. Codex: `codex mcp add xinf --url https://www.zinf.ai/mcp/account` (opens the browser sign-in by itself). ## x402 pay-per-request Not enabled on this deployment. ## Models (229) ### text - `aisingapore/gemma-sea-lion-v4-27b-it` -- AI Singapore Gemma SEA Lion v4 27B IT: input_tokens $0.351 per 1M input tokens; output_tokens $0.555 per 1M output tokens - `ai4bharat/indictrans2-en-indic-1b` -- AI4Bharat IndicTrans2 EN Indic 1B: input_tokens $0.342 per 1M input tokens; output_tokens $0.342 per 1M output tokens - `alibaba/qwen3-max` -- Alibaba Qwen 3 Max: input_tokens $1.20 per 1M input tokens; output_tokens $6.00 per 1M output tokens - `alibaba/qwen3.5-397b-a17b` -- Alibaba Qwen 3.5 397B A17B: input_tokens $0.60 per 1M input tokens; output_tokens $3.60 per 1M output tokens - `alibaba/qwen3.7-max` -- Alibaba Qwen 3.7 Max: cached_input_tokens $0.25 per 1M cached input tokens; input_tokens $1.25 per 1M input tokens; output_tokens $3.75 per 1M output tokens - `alibaba/qwen3.7-plus` -- Alibaba Qwen 3.7 Plus: cached_input_tokens $0.064 per 1M cached input tokens; input_tokens $0.32 per 1M input tokens; output_tokens $1.28 per 1M output tokens - `alibaba/qwen3.8-max` -- Alibaba Qwen 3.8 Max: cached_input_tokens $0.25 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $6.00 per 1M output tokens - `anthropic/claude-fable-5` -- Anthropic Claude Fable 5: cache_write_tokens $12.50 per 1M cache write tokens; cached_input_tokens $1.00 per 1M cached input tokens; input_tokens $10.00 per 1M input tokens; output_tokens $50.00 per 1M output tokens - `anthropic/claude-fable-5.1` -- Anthropic Claude Fable 5.1: cache_write_tokens $12.50 per 1M cache write tokens; cached_input_tokens $0.25 per 1M cached input tokens; input_tokens $10.00 per 1M input tokens; output_tokens $50.00 per 1M output tokens - `anthropic/claude-haiku-4.5` -- Anthropic Claude Haiku 4.5: cache_write_tokens $1.25 per 1M cache write tokens; cached_input_tokens $0.10 per 1M cached input tokens; input_tokens $1.00 per 1M input tokens; output_tokens $5.00 per 1M output tokens - `anthropic/claude-opus-4.5` -- Anthropic Claude Opus 4.5: cache_write_tokens $6.25 per 1M cache write tokens; cached_input_tokens $0.50 per 1M cached input tokens; input_tokens $5.00 per 1M input tokens; output_tokens $25.00 per 1M output tokens - `anthropic/claude-opus-4.6` -- Anthropic Claude Opus 4.6: cache_write_tokens $6.25 per 1M cache write tokens; cached_input_tokens $0.50 per 1M cached input tokens; input_tokens $5.00 per 1M input tokens; output_tokens $25.00 per 1M output tokens - `anthropic/claude-opus-4.7` -- Anthropic Claude Opus 4.7: cache_write_tokens $6.25 per 1M cache write tokens; cached_input_tokens $0.50 per 1M cached input tokens; input_tokens $5.00 per 1M input tokens; output_tokens $25.00 per 1M output tokens - `anthropic/claude-opus-4.8` -- Anthropic Claude Opus 4.8: cache_write_tokens $6.25 per 1M cache write tokens; cached_input_tokens $0.50 per 1M cached input tokens; input_tokens $5.00 per 1M input tokens; output_tokens $25.00 per 1M output tokens - `anthropic/claude-opus-5` -- Anthropic Claude Opus 5: cache_write_tokens $6.25 per 1M cache write tokens; cached_input_tokens $0.50 per 1M cached input tokens; input_tokens $5.00 per 1M input tokens; output_tokens $25.00 per 1M output tokens - `anthropic/claude-opus-5.5` -- Anthropic Claude Opus 5.5: cache_write_tokens $5.00 per 1M cache write tokens; cached_input_tokens $0.20 per 1M cached input tokens; input_tokens $4.00 per 1M input tokens; output_tokens $20.00 per 1M output tokens - `anthropic/claude-sonnet-4.5` -- Anthropic Claude Sonnet 4.5: cache_write_tokens $3.75 per 1M cache write tokens; cached_input_tokens $0.30 per 1M cached input tokens; input_tokens $3.00 per 1M input tokens; output_tokens $15.00 per 1M output tokens - `anthropic/claude-sonnet-4.6` -- Anthropic Claude Sonnet 4.6: cache_write_tokens $3.75 per 1M cache write tokens; cached_input_tokens $0.30 per 1M cached input tokens; input_tokens $3.00 per 1M input tokens; output_tokens $15.00 per 1M output tokens - `anthropic/claude-sonnet-5` -- Anthropic Claude Sonnet 5: cache_write_tokens $2.50 per 1M cache write tokens; cached_input_tokens $0.20 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $10.00 per 1M output tokens - `deepseek-ai/deepseek-r1-distill-qwen-32b` -- DeepSeek DeepSeek R1 Distill Qwen 32B: input_tokens $0.497 per 1M input tokens; output_tokens $4.881 per 1M output tokens - `deepseek/deepseek-v4-pro` -- DeepSeek DeepSeek V4 Pro: cached_input_tokens $0.145 per 1M cached input tokens; input_tokens $1.74 per 1M input tokens; output_tokens $3.48 per 1M output tokens - `deepseek-ai/deepseek-v4-flash-0731` -- DeepSeek DeepSeek v4 Flash 0731: cached_input_tokens $0.014 per 1M cached input tokens; input_tokens $0.44 per 1M input tokens; output_tokens $1.32 per 1M output tokens - `deepseek-ai/deepseek-v4-pro-0813` -- DeepSeek DeepSeek v4 Pro 0813: cached_input_tokens $0.044 per 1M cached input tokens; input_tokens $1.32 per 1M input tokens; output_tokens $3.96 per 1M output tokens - `google/gemini-2.5-flash` -- Google Gemini 2.5 Flash: cached_input_tokens $0.03 per 1M cached input tokens; input_tokens $0.30 per 1M input tokens; output_tokens $2.50 per 1M output tokens - `google/gemini-2.5-flash-lite` -- Google Gemini 2.5 Flash Lite: cached_input_tokens $0.01 per 1M cached input tokens; input_tokens $0.10 per 1M input tokens; output_tokens $0.40 per 1M output tokens - `google/gemini-2.5-pro` -- Google Gemini 2.5 Pro: cached_input_tokens $0.125 per 1M cached input tokens; input_tokens $1.25 per 1M input tokens; output_tokens $10.00 per 1M output tokens - `google/gemini-3-flash` -- Google Gemini 3 Flash: cached_input_tokens:over_200k $0.05 per 1M cached input tokens (prompts over 200K tokens); cached_input_tokens:under_200k $0.05 per 1M cached input tokens (prompts up to 200K tokens); input_tokens:over_200k $0.50 per 1M input tokens (prompts over 200K tokens); input_tokens:under_200k $0.50 per 1M input tokens (prompts up to 200K tokens); output_tokens:over_200k $3.00 per 1M output tokens (prompts over 200K tokens); output_tokens:under_200k $3.00 per 1M output tokens (prompts up to 200K tokens) - `google/gemini-3.1-flash-lite` -- Google Gemini 3.1 Flash Lite: cached_input_tokens:over_200k $0.03 per 1M cached input tokens (prompts over 200K tokens); cached_input_tokens:under_200k $0.03 per 1M cached input tokens (prompts up to 200K tokens); input_tokens:over_200k $0.25 per 1M input tokens (prompts over 200K tokens); input_tokens:under_200k $0.25 per 1M input tokens (prompts up to 200K tokens); output_tokens:over_200k $1.50 per 1M output tokens (prompts over 200K tokens); output_tokens:under_200k $1.50 per 1M output tokens (prompts up to 200K tokens) - `google/gemini-3.1-pro` -- Google Gemini 3.1 Pro: cached_input_tokens:over_200k $0.40 per 1M cached input tokens (prompts over 200K tokens); cached_input_tokens:under_200k $0.20 per 1M cached input tokens (prompts up to 200K tokens); input_tokens:over_200k $4.00 per 1M input tokens (prompts over 200K tokens); input_tokens:under_200k $2.00 per 1M input tokens (prompts up to 200K tokens); output_tokens:over_200k $18.00 per 1M output tokens (prompts over 200K tokens); output_tokens:under_200k $12.00 per 1M output tokens (prompts up to 200K tokens) - `google/gemini-3.5-flash` -- Google Gemini 3.5 Flash: cached_input_tokens:over_200k $0.15 per 1M cached input tokens (prompts over 200K tokens); cached_input_tokens:under_200k $0.15 per 1M cached input tokens (prompts up to 200K tokens); input_tokens:over_200k $1.50 per 1M input tokens (prompts over 200K tokens); input_tokens:under_200k $1.50 per 1M input tokens (prompts up to 200K tokens); output_tokens:over_200k $9.00 per 1M output tokens (prompts over 200K tokens); output_tokens:under_200k $9.00 per 1M output tokens (prompts up to 200K tokens) - `google/gemini-3.5-flash-lite` -- Google Gemini 3.5 Flash-Lite: cache_write_tokens $0.03 per 1M cache write tokens; cached_input_tokens $0.03 per 1M cached input tokens; input_tokens $0.30 per 1M input tokens; output_tokens $2.50 per 1M output tokens - `google/gemini-3.6-flash` -- Google Gemini 3.6 Flash: cache_write_tokens $1.50 per 1M cache write tokens; cached_input_tokens $0.15 per 1M cached input tokens; input_tokens $1.50 per 1M input tokens; output_tokens $7.50 per 1M output tokens - `google/gemini-3.7-flash` -- Google Gemini 3.7 Flash: cached_input_tokens $0.075 per 1M cached input tokens; input_tokens $0.75 per 1M input tokens; output_tokens $3.75 per 1M output tokens - `google/gemini-3.8-flash` -- Google Gemini 3.8 Flash: cached_input_tokens $0.075 per 1M cached input tokens; input_tokens $0.75 per 1M input tokens; output_tokens $3.75 per 1M output tokens - `google/gemma-2b-it-lora` -- Google Gemma 2B IT LoRA: compute_units $0.011 per 1K compute units - `google/gemma-4-26b-a4b-it` -- Google Gemma 4 26B A4B IT: input_tokens $0.10 per 1M input tokens; output_tokens $0.30 per 1M output tokens - `google/gemma-7b-it-lora` -- Google Gemma 7B IT LoRA: compute_units $0.011 per 1K compute units - `ibm-granite/granite-4.0-h-micro` -- IBM Granite 4.0 H Micro: input_tokens $0.017 per 1M input tokens; output_tokens $0.112 per 1M output tokens - `llava-hf/llava-1.5-7b-hf` -- LLaVA LLaVA 1.5 7B HF: compute_units $0.011 per 1K compute units - `meta-llama/llama-2-7b-chat-hf-lora` -- Meta Llama 2 7B Chat HF LoRA: compute_units $0.011 per 1K compute units - `meta/llama-3.1-8b-instruct-fp8` -- Meta Llama 3.1 8B Instruct FP8: input_tokens $0.152 per 1M input tokens; output_tokens $0.287 per 1M output tokens - `meta/llama-3.2-11b-vision-instruct` -- Meta Llama 3.2 11B Vision Instruct: input_tokens $0.0485 per 1M input tokens; output_tokens $0.676 per 1M output tokens - `meta/llama-3.2-1b-instruct` -- Meta Llama 3.2 1B Instruct: input_tokens $0.027 per 1M input tokens; output_tokens $0.201 per 1M output tokens - `meta/llama-3.2-3b-instruct` -- Meta Llama 3.2 3B Instruct: input_tokens $0.0509 per 1M input tokens; output_tokens $0.335 per 1M output tokens - `meta/llama-3.3-70b-instruct-fp8-fast` -- Meta Llama 3.3 70B Instruct FP8 Fast: input_tokens $0.293 per 1M input tokens; output_tokens $2.253 per 1M output tokens - `meta/llama-4-scout-17b-16e-instruct` -- Meta Llama 4 Scout 17B 16E Instruct: input_tokens $0.27 per 1M input tokens; output_tokens $0.85 per 1M output tokens - `meta/llama-guard-3-8b` -- Meta Llama Guard 3 8B: input_tokens $0.484 per 1M input tokens; output_tokens $0.03 per 1M output tokens - `meta/m2m100-1.2b` -- Meta M2M100 1.2B: input_tokens $0.342 per 1M input tokens; output_tokens $0.342 per 1M output tokens - `minimax/m2.7` -- MiniMax MiniMax M2.7: cache_write_tokens $0.375 per 1M cache write tokens; cached_input_tokens $0.06 per 1M cached input tokens; input_tokens $0.30 per 1M input tokens; output_tokens $1.20 per 1M output tokens - `minimax/m3` -- MiniMax MiniMax M3: cached_input_tokens:over_512k $0.24 per 1M cached input tokens (prompts over 512K tokens); cached_input_tokens:under_512k $0.06 per 1M cached input tokens (prompts up to 512K tokens); input_tokens:over_512k $1.20 per 1M input tokens (prompts over 512K tokens); input_tokens:under_512k $0.30 per 1M input tokens (prompts up to 512K tokens); output_tokens:over_512k $4.80 per 1M output tokens (prompts over 512K tokens); output_tokens:under_512k $1.20 per 1M output tokens (prompts up to 512K tokens) - `mistral/mistral-7b-instruct-v0.2-lora` -- Mistral AI Mistral 7B Instruct v0.2 LoRA: compute_units $0.011 per 1K compute units - `mistralai/mistral-small-3.1-24b-instruct` -- Mistral AI Mistral Small 3.1 24B Instruct: input_tokens $0.351 per 1M input tokens; output_tokens $0.555 per 1M output tokens - `moondream/moondream3.1-9b-a2b` -- Moondream Moondream3.1 9B A2B: input_tokens $0.30 per 1M input tokens; output_tokens $1.00 per 1M output tokens - `moonshotai/kimi-k2.6` -- Moonshot AI Kimi K2.6: cached_input_tokens $0.16 per 1M cached input tokens; input_tokens $0.95 per 1M input tokens; output_tokens $4.00 per 1M output tokens - `moonshotai/kimi-k2.7-code` -- Moonshot AI Kimi K2.7 Code: cached_input_tokens $0.19 per 1M cached input tokens; input_tokens $0.95 per 1M input tokens; output_tokens $4.00 per 1M output tokens - `moonshotai/kimi-k3` -- Moonshot AI Kimi K3: cached_input_tokens $0.30 per 1M cached input tokens; input_tokens $3.00 per 1M input tokens; output_tokens $15.00 per 1M output tokens - `nvidia/nemotron-3-120b-a12b` -- NVIDIA Nemotron 3 120B A12B: input_tokens $0.50 per 1M input tokens; output_tokens $1.50 per 1M output tokens - `openai/gpt-oss-120b` -- OpenAI GPT OSS 120B: input_tokens $0.35 per 1M input tokens; output_tokens $0.75 per 1M output tokens - `openai/gpt-oss-20b` -- OpenAI GPT OSS 20B: input_tokens $0.20 per 1M input tokens; output_tokens $0.30 per 1M output tokens - `openai/gpt-4.1` -- OpenAI GPT-4.1: cached_input_tokens $0.50 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $8.00 per 1M output tokens - `openai/gpt-4.1-mini` -- OpenAI GPT-4.1 mini: cached_input_tokens $0.10 per 1M cached input tokens; input_tokens $0.40 per 1M input tokens; output_tokens $1.60 per 1M output tokens - `openai/gpt-4.1-nano` -- OpenAI GPT-4.1 nano: cached_input_tokens $0.025 per 1M cached input tokens; input_tokens $0.10 per 1M input tokens; output_tokens $0.40 per 1M output tokens - `openai/gpt-4o` -- OpenAI GPT-4o: cached_input_tokens $0.625 per 1M cached input tokens; input_tokens $1.25 per 1M input tokens; output_tokens $5.00 per 1M output tokens - `openai/gpt-4o-mini` -- OpenAI GPT-4o mini: cached_input_tokens $0.0375 per 1M cached input tokens; input_tokens $0.075 per 1M input tokens; output_tokens $0.30 per 1M output tokens - `openai/gpt-5` -- OpenAI GPT-5: cached_input_tokens $0.125 per 1M cached input tokens; input_tokens $1.25 per 1M input tokens; output_tokens $10.00 per 1M output tokens - `openai/gpt-5-mini` -- OpenAI GPT-5 mini: cached_input_tokens $0.025 per 1M cached input tokens; input_tokens $0.25 per 1M input tokens; output_tokens $2.00 per 1M output tokens - `openai/gpt-5-nano` -- OpenAI GPT-5 nano: cached_input_tokens $0.005 per 1M cached input tokens; input_tokens $0.05 per 1M input tokens; output_tokens $0.40 per 1M output tokens - `openai/gpt-5.1` -- OpenAI GPT-5.1: cached_input_tokens $0.125 per 1M cached input tokens; input_tokens $1.25 per 1M input tokens; output_tokens $10.00 per 1M output tokens - `openai/gpt-5.4` -- OpenAI GPT-5.4: cached_input_tokens $0.25 per 1M cached input tokens; input_tokens $2.50 per 1M input tokens; output_tokens $15.00 per 1M output tokens - `openai/gpt-5.4-mini` -- OpenAI GPT-5.4 mini: cached_input_tokens $0.075 per 1M cached input tokens; input_tokens $0.75 per 1M input tokens; output_tokens $4.50 per 1M output tokens - `openai/gpt-5.4-nano` -- OpenAI GPT-5.4 nano: cached_input_tokens $0.02 per 1M cached input tokens; input_tokens $0.20 per 1M input tokens; output_tokens $1.25 per 1M output tokens - `openai/gpt-5.4-pro` -- OpenAI GPT-5.4 pro: input_tokens $30.00 per 1M input tokens; output_tokens $180.00 per 1M output tokens - `openai/gpt-5.5` -- OpenAI GPT-5.5: input_tokens $5.00 per 1M input tokens; output_tokens $30.00 per 1M output tokens - `openai/gpt-5.5-pro` -- OpenAI GPT-5.5 pro: input_tokens $30.00 per 1M input tokens; output_tokens $180.00 per 1M output tokens - `openai/gpt-5.6-luna` -- OpenAI GPT-5.6 Luna: cache_write_tokens $0.25 per 1M cache write tokens; cached_input_tokens $0.02 per 1M cached input tokens; input_tokens $0.20 per 1M input tokens; output_tokens $1.20 per 1M output tokens - `openai/gpt-5.6-sol` -- OpenAI GPT-5.6 Sol: cache_write_tokens $3.125 per 1M cache write tokens; cached_input_tokens $0.25 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $10.00 per 1M output tokens - `openai/gpt-5.6-terra` -- OpenAI GPT-5.6 Terra: cache_write_tokens $2.50 per 1M cache write tokens; cached_input_tokens $0.20 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $12.00 per 1M output tokens - `openai/gpt-6-astra` -- OpenAI GPT-6 Astra: cache_write_tokens:long_context $25.00 per 1M cache write tokens (long context); cache_write_tokens:short_context $12.00 per 1M cache write tokens (short context); cached_input_tokens:long_context $2.00 per 1M cached input tokens (long context); cached_input_tokens:short_context $1.00 per 1M cached input tokens (short context); input_tokens:long_context $20.00 per 1M input tokens (long context); input_tokens:short_context $10.00 per 1M input tokens (short context); output_tokens:long_context $75.00 per 1M output tokens (long context); output_tokens:short_context $50.00 per 1M output tokens (short context) - `openai/gpt-6-luna` -- OpenAI GPT-6 Luna: cache_write_tokens:long_context $0.25 per 1M cache write tokens (long context); cache_write_tokens:short_context $0.125 per 1M cache write tokens (short context); cached_input_tokens:long_context $0.02 per 1M cached input tokens (long context); cached_input_tokens:short_context $0.01 per 1M cached input tokens (short context); input_tokens:long_context $0.20 per 1M input tokens (long context); input_tokens:short_context $0.10 per 1M input tokens (short context); output_tokens:long_context $0.75 per 1M output tokens (long context); output_tokens:short_context $0.50 per 1M output tokens (short context) - `openai/gpt-6-sol` -- OpenAI GPT-6 Sol: cache_write_tokens:long_context $5.00 per 1M cache write tokens (long context); cache_write_tokens:short_context $2.50 per 1M cache write tokens (short context); cached_input_tokens:long_context $0.40 per 1M cached input tokens (long context); cached_input_tokens:short_context $0.20 per 1M cached input tokens (short context); input_tokens:long_context $4.00 per 1M input tokens (long context); input_tokens:short_context $2.00 per 1M input tokens (short context); output_tokens:long_context $15.00 per 1M output tokens (long context); output_tokens:short_context $10.00 per 1M output tokens (short context) - `openai/o3` -- OpenAI o3: cached_input_tokens $0.50 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $8.00 per 1M output tokens - `openai/o3-mini` -- OpenAI o3-mini: cached_input_tokens $0.55 per 1M cached input tokens; input_tokens $1.10 per 1M input tokens; output_tokens $4.40 per 1M output tokens - `openai/o4-mini` -- OpenAI o4-mini: cached_input_tokens $0.275 per 1M cached input tokens; input_tokens $1.10 per 1M input tokens; output_tokens $4.40 per 1M output tokens - `qwen/qwq-32b` -- Qwen QwQ 32B: input_tokens $0.66 per 1M input tokens; output_tokens $1.00 per 1M output tokens - `qwen/qwen2.5-coder-32b-instruct` -- Qwen Qwen2.5 Coder 32B Instruct: input_tokens $0.66 per 1M input tokens; output_tokens $1.00 per 1M output tokens - `qwen/qwen3-30b-a3b-fp8` -- Qwen Qwen3 30B A3B FP8: input_tokens $0.0509 per 1M input tokens; output_tokens $0.335 per 1M output tokens - `qwen/qwen3.8-27b` -- Qwen Qwen3.8 27B: cached_input_tokens $0.05 per 1M cached input tokens; input_tokens $0.45 per 1M input tokens; output_tokens $3.20 per 1M output tokens - `thinkingmachines/inkling` -- Thinking Machines Inkling: cache_write_tokens $1.87 per 1M cache write tokens; cached_input_tokens $0.374 per 1M cached input tokens; input_tokens $1.87 per 1M input tokens; output_tokens $4.68 per 1M output tokens - `thinkingmachines/inkling-256k` -- Thinking Machines Inkling 256K: cache_write_tokens $3.74 per 1M cache write tokens; cached_input_tokens $0.748 per 1M cached input tokens; input_tokens $3.74 per 1M input tokens; output_tokens $9.36 per 1M output tokens - `typesafe/jev` -- TypeSafe Jev: cached_input_tokens $0.00 per 1M cached input tokens; input_tokens $0.042 per 1M input tokens; output_tokens $0.00 per 1M output tokens - `unbiased/pareto` -- Unbiased Pareto: cached_input_tokens $0.25 per 1M cached input tokens; input_tokens $2.50 per 1M input tokens; output_tokens $7.50 per 1M output tokens - `zai-org/glm-4.7-flash` -- Z.ai GLM 4.7 Flash: input_tokens $0.0605 per 1M input tokens; output_tokens $0.40 per 1M output tokens - `zai-org/glm-5.2` -- Z.ai GLM 5.2: cached_input_tokens $0.26 per 1M cached input tokens; input_tokens $1.40 per 1M input tokens; output_tokens $4.40 per 1M output tokens - `zai-org/glm-5.3` -- Z.ai GLM 5.3: cached_input_tokens $0.26 per 1M cached input tokens; input_tokens $1.40 per 1M input tokens; output_tokens $4.40 per 1M output tokens - `zai-org/glm-5.3-flash` -- Z.ai GLM 5.3 Flash: cached_input_tokens $0.03 per 1M cached input tokens; input_tokens $0.15 per 1M input tokens; output_tokens $0.50 per 1M output tokens - `xai/grok-4.20-multi-agent-0309` -- xAI Grok 4.20 Multi-Agent: cached_input_tokens $0.20 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $6.00 per 1M output tokens - `xai/grok-4.20-0309-non-reasoning` -- xAI Grok 4.20 Non-Reasoning: cached_input_tokens $0.20 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $6.00 per 1M output tokens - `xai/grok-4.20-0309-reasoning` -- xAI Grok 4.20 Reasoning: cached_input_tokens $0.20 per 1M cached input tokens; input_tokens $2.00 per 1M input tokens; output_tokens $6.00 per 1M output tokens - `xai/grok-4.3` -- xAI Grok 4.3: cached_input_tokens $0.20 per 1M cached input tokens; input_tokens $1.25 per 1M input tokens; output_tokens $2.50 per 1M output tokens - `xai/grok-4.5` -- xAI Grok 4.5: cached_input_tokens:over_200k $0.60 per 1M cached input tokens (prompts over 200K tokens); cached_input_tokens:under_200k $0.30 per 1M cached input tokens (prompts up to 200K tokens); input_tokens:over_200k $4.00 per 1M input tokens (prompts over 200K tokens); input_tokens:under_200k $2.00 per 1M input tokens (prompts up to 200K tokens); output_tokens:over_200k $12.00 per 1M output tokens (prompts over 200K tokens); output_tokens:under_200k $6.00 per 1M output tokens (prompts up to 200K tokens) - `xai/grok-4.6` -- xAI Grok 4.6: cached_input_tokens:over_200k $1.00 per 1M cached input tokens (prompts over 200K tokens); cached_input_tokens:under_200k $0.50 per 1M cached input tokens (prompts up to 200K tokens); input_tokens:over_200k $4.00 per 1M input tokens (prompts over 200K tokens); input_tokens:under_200k $2.00 per 1M input tokens (prompts up to 200K tokens); output_tokens:over_200k $12.00 per 1M output tokens (prompts over 200K tokens); output_tokens:under_200k $6.00 per 1M output tokens (prompts up to 200K tokens) - `xai/grok-4.7` -- xAI Grok 4.7: cached_input_tokens:over_200k $1.00 per 1M cached input tokens (prompts over 200K tokens); cached_input_tokens:under_200k $0.50 per 1M cached input tokens (prompts up to 200K tokens); input_tokens:over_200k $4.00 per 1M input tokens (prompts over 200K tokens); input_tokens:under_200k $2.00 per 1M input tokens (prompts up to 200K tokens); output_tokens:over_200k $12.00 per 1M output tokens (prompts over 200K tokens); output_tokens:under_200k $6.00 per 1M output tokens (prompts up to 200K tokens) ### image - `alibaba/qwen-image-3.0-pro` -- Alibaba Qwen Image 3.0 Pro: output_images $0.04 per image; output_images:1k $0.04 per image (1K); output_images:2k $0.07 per image (2K) - `alibaba/wan-2.6-image` -- Alibaba Wan 2.6 Image: output_images $0.03 per image - `black-forest-labs/flux-1-schnell` -- Black Forest Labs FLUX 1 Schnell: compute_units $0.011 per 1K compute units - `black-forest-labs/flux-2-dev` -- Black Forest Labs FLUX 2 Dev: compute_units $0.011 per 1K compute units - `black-forest-labs/flux-2-klein-4b` -- Black Forest Labs FLUX 2 Klein 4B: compute_units $0.011 per 1K compute units - `black-forest-labs/flux-2-klein-9b` -- Black Forest Labs FLUX 2 Klein 9B: compute_units $0.011 per 1K compute units - `black-forest-labs/flux-1-kontext-max` -- Black Forest Labs FLUX.1 Kontext [max]: output_images $0.08 per image - `black-forest-labs/flux-1-kontext-pro` -- Black Forest Labs FLUX.1 Kontext [pro]: output_images $0.04 per image - `black-forest-labs/flux-2-flex` -- Black Forest Labs FLUX.2 [flex]: input_megapixels $0.05 per input megapixel; output_megapixels_additional $0.05 per additional output megapixel; output_megapixels_first $0.05 per first output megapixel - `black-forest-labs/flux-2-max` -- Black Forest Labs FLUX.2 [max]: input_megapixels $0.03 per input megapixel; output_megapixels_additional $0.03 per additional output megapixel; output_megapixels_first $0.07 per first output megapixel - `black-forest-labs/flux-2-pro-preview` -- Black Forest Labs FLUX.2 [pro]: input_megapixels $0.015 per input megapixel; output_megapixels_additional $0.015 per additional output megapixel; output_megapixels_first $0.03 per first output megapixel - `bytedance/seedream-4.0` -- ByteDance Seedream 4.0: output_images $0.03 per image - `bytedance/seedream-4.5` -- ByteDance Seedream 4.5: output_images $0.04 per image - `bytedance/seedream-5-lite` -- ByteDance Seedream 5 Lite: output_images $0.035 per image - `bytedance/seedream-5-pro` -- ByteDance Seedream 5 Pro: input_images $0.03 per input image; output_images $0.045 per image - `bytedance/stable-diffusion-xl-lightning` -- ByteDance Stable Diffusion XL Lightning: compute_units $0.011 per 1K compute units - `google/nano-banana` -- Google Nano Banana: cache_write_tokens $0.083333 per 1M cache write tokens; cached_input_tokens $0.03 per 1M cached input tokens; input_tokens $0.30 per 1M input tokens; output_tokens $30.00 per 1M output tokens - `google/nano-banana-2` -- Google Nano Banana 2: input_tokens $0.50 per 1M input tokens; output_tokens $60.00 per 1M output tokens - `google/nano-banana-2-lite` -- Google Nano Banana 2 Lite: input_tokens $0.25 per 1M input tokens; output_tokens $30.00 per 1M output tokens - `google/nano-banana-pro` -- Google Nano Banana Pro: input_tokens $2.00 per 1M input tokens; output_tokens $120.00 per 1M output tokens - `krea/krea-2-large` -- Krea Krea 2 Large: output_images $0.06 per image - `krea/krea-2-medium` -- Krea Krea 2 Medium: output_images $0.03 per image - `krea/krea-2-medium-turbo` -- Krea Krea 2 Medium Turbo: output_images $0.015 per image - `leonardo/lucid-origin` -- Leonardo.Ai Lucid Origin: steps $0.000132 per diffusion step; tiles_512 $0.007 per 512×512 tile - `leonardo/phoenix-1.0` -- Leonardo.Ai Phoenix 1.0: steps $0.00011 per diffusion step; tiles_512 $0.00583 per 512×512 tile - `lykon/dreamshaper-8-lcm` -- Lykon DreamShaper 8 LCM: compute_units $0.011 per 1K compute units - `openai/gpt-image-1.5` -- OpenAI GPT Image 1.5: cached_input_tokens $1.25 per 1M cached input tokens; input_tokens $5.00 per 1M input tokens; output_tokens $10.00 per 1M output tokens - `openai/gpt-image-2` -- OpenAI GPT Image 2: cached_input_image_tokens $2.00 per 1M cached input image tokens; cached_input_tokens $1.25 per 1M cached input tokens; input_image_tokens $8.00 per 1M input image tokens; input_tokens $5.00 per 1M input tokens; output_image_tokens $30.00 per 1M output image tokens; output_tokens $10.00 per 1M output tokens - `openai/gpt-image-2.5-flare` -- OpenAI GPT Image 2.5 Flare: cached_input_image_tokens $3.00 per 1M cached input image tokens; cached_input_tokens $1.25 per 1M cached input tokens; input_image_tokens $8.00 per 1M input image tokens; input_tokens $5.00 per 1M input tokens; output_image_tokens $30.00 per 1M output image tokens - `openai/gpt-image-2.5-sunburst` -- OpenAI GPT Image 2.5 Sunburst: cached_input_image_tokens $3.00 per 1M cached input image tokens; cached_input_tokens $1.25 per 1M cached input tokens; input_image_tokens $8.00 per 1M input image tokens; input_tokens $5.00 per 1M input tokens; output_image_tokens $30.00 per 1M output image tokens - `pruna/p-image` -- Pruna P-Image: output_images $0.005 per image - `pruna/p-image-try-on` -- Pruna P-Image Try-On: input_images $0.008 per input image; output_images $0.015 per image - `pruna/p-image-edit` -- Pruna P-Image-Edit: output_images $0.01 per image - `pruna/p-image-upscale` -- Pruna P-Image-Upscale: output_images:mp_17_32 $0.04 per image (17–32 MP); output_images:mp_1_4 $0.005 per image (1–4 MP); output_images:mp_33_64 $0.06 per image (33–64 MP); output_images:mp_5_8 $0.01 per image (5–8 MP); output_images:mp_65_128 $0.12 per image (65–128 MP); output_images:mp_9_16 $0.02 per image (9–16 MP) - `recraft/recraftv3` -- Recraft Recraft V3: output_images $0.04 per image - `recraft/recraftv4` -- Recraft Recraft V4: output_images $0.04 per image - `recraft/recraftv4-pro` -- Recraft Recraft V4 Pro: output_images $0.25 per image - `recraft/recraftv4-pro-vector` -- Recraft Recraft V4 Pro SVG: output_images $0.30 per image - `recraft/recraftv4-vector` -- Recraft Recraft V4 SVG: output_images $0.08 per image - `recraft/recraftv4-1` -- Recraft Recraft V4.1: output_images $0.04 per image - `recraft/recraftv4-1-pro` -- Recraft Recraft V4.1 Pro: output_images $0.25 per image - `recraft/recraftv4-1-pro-vector` -- Recraft Recraft V4.1 Pro SVG: output_images $0.30 per image - `recraft/recraftv4-1-vector` -- Recraft Recraft V4.1 SVG: output_images $0.08 per image - `recraft/recraftv4-1-utility` -- Recraft Recraft V4.1 Utility: output_images $0.04 per image - `recraft/recraftv4-1-utility-pro` -- Recraft Recraft V4.1 Utility Pro: output_images $0.25 per image - `recraft/recraftv4-1-utility-pro-vector` -- Recraft Recraft V4.1 Utility Pro SVG: output_images $0.30 per image - `recraft/recraftv4-1-utility-vector` -- Recraft Recraft V4.1 Utility SVG: output_images $0.08 per image - `runwayml/stable-diffusion-v1-5-inpainting` -- Runway Stable Diffusion v1 5 Inpainting: compute_units $0.011 per 1K compute units - `stabilityai/stable-diffusion-xl-base-1.0` -- Stability AI Stable Diffusion XL Base 1.0: compute_units $0.011 per 1K compute units - `xai/grok-imagine-image` -- xAI Grok Imagine Image: input_images $0.002 per input image; output_images $0.02 per image - `xai/grok-imagine-image-2.0` -- xAI Grok Imagine Image 2.0: input_images $0.01 per input image; output_images $0.04 per image - `xai/grok-imagine-image-quality` -- xAI Grok Imagine Image Quality: input_images $0.01 per input image; output_images $0.05 per image ### video - `alibaba/hh1-i2v` -- Alibaba HappyHorse 1.0 I2V: output_video_seconds $0.28 per second of video; output_video_seconds:1080p $0.28 per second of 1080p video; output_video_seconds:720p $0.14 per second of 720p video - `alibaba/hh1-t2v` -- Alibaba HappyHorse 1.0 T2V: output_video_seconds $0.28 per second of video; output_video_seconds:1080p $0.28 per second of 1080p video; output_video_seconds:720p $0.14 per second of 720p video - `alibaba/hh1.1-i2v` -- Alibaba HappyHorse 1.1 I2V: output_video_seconds $0.18 per second of video; output_video_seconds:1080p $0.18 per second of 1080p video; output_video_seconds:720p $0.14 per second of 720p video - `alibaba/hh1.1-r2v` -- Alibaba HappyHorse 1.1 R2V: output_video_seconds $0.18 per second of video; output_video_seconds:1080p $0.18 per second of 1080p video; output_video_seconds:720p $0.14 per second of 720p video - `alibaba/hh1.1-t2v` -- Alibaba HappyHorse 1.1 T2V: output_video_seconds $0.18 per second of video; output_video_seconds:1080p $0.18 per second of 1080p video; output_video_seconds:720p $0.14 per second of 720p video - `alibaba/wan-2.7-i2v` -- Alibaba Wan 2.7 I2V: output_video_seconds $0.10 per second of video; output_video_seconds:1080p $0.15 per second of 1080p video; output_video_seconds:720p $0.10 per second of 720p video - `alibaba/wan-3.0-prime` -- Alibaba Wan 3.0 Prime Video: output_video_seconds $0.068 per second of video; output_video_seconds:1080p $0.28 per second of 1080p video; output_video_seconds:480p $0.068 per second of 480p video; output_video_seconds:720p $0.14 per second of 720p video - `alibaba/wan-3.0` -- Alibaba Wan 3.0 Video: output_video_seconds $0.05 per second of video; output_video_seconds:1080p $0.20 per second of 1080p video; output_video_seconds:480p $0.05 per second of 480p video; output_video_seconds:720p $0.10 per second of 720p video - `black-forest-labs/flux-3-video` -- Black Forest Labs FLUX 3 Video: output_video_seconds:fhd $0.29 per second of Full HD video; output_video_seconds:hd $0.17 per second of HD video; output_video_seconds:hd_draft $0.06 per second of HD video (draft); output_video_seconds:v2v_fhd $0.53 per second of video-to-video Full HD video; output_video_seconds:v2v_hd $0.41 per second of video-to-video HD video; output_video_seconds:v2v_hd_draft $0.12 per second of video-to-video HD video (draft) - `black-forest-labs/flux-video-upscale` -- Black Forest Labs FLUX Video Upscale: video_megapixel_seconds:creative $0.10 per megapixel-second of video (creative); video_megapixel_seconds:precise $0.07 per megapixel-second of video (precise) - `bytedance/seedance-2.0` -- ByteDance Seedance 2.0: output_video_seconds $0.15 per second of video; output_video_seconds:1080p_non_video_in $0.37 per second of 1080p video (without video input); output_video_seconds:1080p_video_in $0.914 per second of 1080p video (with video input); output_video_seconds:480p_non_video_in $0.07 per second of 480p video (without video input); output_video_seconds:480p_video_in $0.172 per second of 480p video (with video input); output_video_seconds:4k_non_video_in $0.78 per second of 4K video (without video input); output_video_seconds:4k_video_in $1.866 per second of 4K video (with video input); output_video_seconds:720p_non_video_in $0.15 per second of 720p video (without video input); output_video_seconds:720p_video_in $0.372 per second of 720p video (with video input) - `bytedance/seedance-2.0-fast` -- ByteDance Seedance 2.0 Fast: output_video_seconds $0.12 per second of video; output_video_seconds:480p_non_video_in $0.06 per second of 480p video (without video input); output_video_seconds:480p_video_in $0.132 per second of 480p video (with video input); output_video_seconds:720p_non_video_in $0.12 per second of 720p video (without video input); output_video_seconds:720p_video_in $0.286 per second of 720p video (with video input) - `bytedance/seedance-2.0-mini` -- ByteDance Seedance 2.0 Mini: output_video_seconds $0.09 per second of video; output_video_seconds:480p_non_video_in $0.04 per second of 480p video (without video input); output_video_seconds:480p_video_in $0.084 per second of 480p video (with video input); output_video_seconds:720p_non_video_in $0.09 per second of 720p video (without video input); output_video_seconds:720p_video_in $0.182 per second of 720p video (with video input) - `bytedance/seedance-2.5` -- ByteDance Seedance 2.5: output_video_seconds $0.2312 per second of video; output_video_seconds:480p_non_video_in $0.1028 per second of 480p video (without video input); output_video_seconds:480p_video_in $0.4304 per second of 480p video (with video input); output_video_seconds:720p_non_video_in $0.2312 per second of 720p video (without video input); output_video_seconds:720p_video_in $0.9676 per second of 720p video (with video input) - `google/gemini-omni-flash` -- Google Gemini Omni Flash: input_audio_tokens $1.50 per 1M input audio tokens; input_image_tokens $1.50 per 1M input image tokens; input_tokens $1.50 per 1M input tokens; input_video_tokens $1.50 per 1M input video tokens; output_tokens $9.00 per 1M output tokens; output_video_tokens $17.50 per 1M output video tokens; reasoning_tokens $9.00 per 1M reasoning tokens - `google/gemini-omni-1.1-flash` -- Google Gemini Omni Flash 1.1: input_audio_tokens $1.50 per 1M input audio tokens; input_image_tokens $1.50 per 1M input image tokens; input_tokens $1.50 per 1M input tokens; input_video_tokens $1.50 per 1M input video tokens; output_tokens $9.00 per 1M output tokens; output_video_tokens $17.50 per 1M output video tokens; reasoning_tokens $9.00 per 1M reasoning tokens - `google/veo-3.1` -- Google Veo 3.1: output_video_seconds $0.40 per second of video; output_video_seconds:1080p $0.20 per second of 1080p video; output_video_seconds:1080p_audio $0.40 per second of 1080p video with audio; output_video_seconds:4k $0.40 per second of 4K video; output_video_seconds:4k_audio $0.60 per second of 4K video with audio; output_video_seconds:720p $0.20 per second of 720p video; output_video_seconds:720p_audio $0.40 per second of 720p video with audio - `google/veo-3.1-fast` -- Google Veo 3.1 Fast: output_video_seconds $0.08 per second of video; output_video_seconds:1080p $0.10 per second of 1080p video; output_video_seconds:1080p_audio $0.12 per second of 1080p video with audio; output_video_seconds:4k $0.25 per second of 4K video; output_video_seconds:4k_audio $0.30 per second of 4K video with audio; output_video_seconds:720p $0.08 per second of 720p video; output_video_seconds:720p_audio $0.10 per second of 720p video with audio - `lightricks/ltx-2-5-fast` -- Lightricks LTX-2.5 Fast: output_video_seconds $0.09 per second of video; output_video_seconds:1080p $0.15 per second of 1080p video; output_video_seconds:2k $0.19 per second of 2K video; output_video_seconds:4k $0.37 per second of 4K video; output_video_seconds:720p $0.09 per second of 720p video - `minimax/h3` -- MiniMax MiniMax H3: output_video_seconds $0.08 per second of video; output_video_seconds:2k $0.13 per second of 2K video; output_video_seconds:768p $0.08 per second of 768p video - `minimax/h3-max` -- MiniMax MiniMax H3 Max: output_video_seconds $0.05 per second of video; output_video_seconds:480p $0.05 per second of 480p video; output_video_seconds:768p $0.08 per second of 768p video - `minimax/hailuo-2.3` -- MiniMax MiniMax Hailuo 2.3: output_video_clip:1080p_6s $0.49 per 6 s 1080p clip; output_video_clip:768p_10s $0.56 per 10 s 768p clip; output_video_clip:768p_6s $0.28 per 6 s 768p clip; output_video_seconds $0.047 per second of video - `minimax/hailuo-2.3-fast` -- MiniMax MiniMax Hailuo 2.3 Fast: output_video_clip:1080p_6s $0.33 per 6 s 1080p clip; output_video_clip:768p_10s $0.32 per 10 s 768p clip; output_video_clip:768p_6s $0.19 per 6 s 768p clip; output_video_seconds $0.032 per second of video - `pixverse/v5.6` -- PixVerse Pixverse v5.6: output_video_clip:1080p_5s $0.375 per 5 s 1080p clip; output_video_clip:1080p_5s_audio $0.75 per 5 s 1080p clip with audio; output_video_clip:1080p_8s $0.75 per 8 s 1080p clip; output_video_clip:1080p_8s_audio $0.975 per 8 s 1080p clip with audio; output_video_clip:360p_10s $0.385 per 10 s 360p clip; output_video_clip:360p_10s_audio $0.61 per 10 s 360p clip with audio; output_video_clip:360p_5s $0.175 per 5 s 360p clip; output_video_clip:360p_5s_audio $0.40 per 5 s 360p clip with audio; output_video_clip:360p_8s $0.35 per 8 s 360p clip; output_video_clip:360p_8s_audio $0.575 per 8 s 360p clip with audio; output_video_clip:540p_10s $0.385 per 10 s 540p clip; output_video_clip:540p_10s_audio $0.61 per 10 s 540p clip with audio; output_video_clip:540p_5s $0.175 per 5 s 540p clip; output_video_clip:540p_5s_audio $0.45 per 5 s 540p clip with audio; output_video_clip:540p_8s $0.35 per 8 s 540p clip; output_video_clip:540p_8s_audio $0.575 per 8 s 540p clip with audio; output_video_clip:720p_10s $0.495 per 10 s 720p clip; output_video_clip:720p_10s_audio $0.72 per 10 s 720p clip with audio; output_video_clip:720p_5s $0.225 per 5 s 720p clip; output_video_clip:720p_5s_audio $0.40 per 5 s 720p clip with audio; output_video_clip:720p_8s $0.45 per 8 s 720p clip; output_video_clip:720p_8s_audio $0.675 per 8 s 720p clip with audio; output_video_seconds $0.08 per second of video - `pixverse/v6` -- PixVerse Pixverse v6: output_video_seconds $0.06 per second of video; output_video_seconds:1080p $0.09 per second of 1080p video; output_video_seconds:1080p_audio $0.115 per second of 1080p video with audio; output_video_seconds:360p $0.025 per second of 360p video; output_video_seconds:360p_audio $0.035 per second of 360p video with audio; output_video_seconds:540p $0.035 per second of 540p video; output_video_seconds:540p_audio $0.045 per second of 540p video with audio; output_video_seconds:720p $0.045 per second of 720p video; output_video_seconds:720p_audio $0.06 per second of 720p video with audio - `pruna/p-video` -- Pruna P-Video: output_video_seconds $0.02 per second of video; output_video_seconds:1080p $0.04 per second of 1080p video; output_video_seconds:1080p_draft $0.01 per second of 1080p video (draft); output_video_seconds:720p $0.02 per second of 720p video; output_video_seconds:720p_draft $0.005 per second of 720p video (draft) - `pruna/p-video-animate` -- Pruna P-Video-Animate: output_video_seconds $0.03 per second of video; output_video_seconds:1080p $0.06 per second of 1080p video; output_video_seconds:720p $0.03 per second of 720p video - `pruna/p-video-avatar` -- Pruna P-Video-Avatar: output_video_seconds $0.025 per second of video; output_video_seconds:1080p $0.045 per second of 1080p video; output_video_seconds:720p $0.025 per second of 720p video - `pruna/p-video-replace` -- Pruna P-Video-Replace: output_video_seconds $0.03 per second of video; output_video_seconds:1080p $0.06 per second of 1080p video; output_video_seconds:720p $0.03 per second of 720p video - `runwayml/aleph-2` -- Runway RunwayML Aleph 2: output_video_seconds $0.336 per second of video - `runwayml/gen-4.5` -- Runway RunwayML Gen-4.5: output_video_seconds $0.12 per second of video - `vidu/q3-pro` -- Vidu Vidu Q3 Pro: output_video_seconds $0.125 per second of video; output_video_seconds:1080p $0.15 per second of 1080p video; output_video_seconds:540p $0.05 per second of 540p video; output_video_seconds:720p $0.125 per second of 720p video - `vidu/q3-turbo` -- Vidu Vidu Q3 Turbo: output_video_seconds $0.06 per second of video; output_video_seconds:1080p $0.07 per second of 1080p video; output_video_seconds:540p $0.04 per second of 540p video; output_video_seconds:720p $0.06 per second of 720p video - `xai/grok-imagine-video` -- xAI Grok Imagine Video: output_video_seconds $0.05 per second of video; output_video_seconds:480p $0.05 per second of 480p video; output_video_seconds:720p $0.07 per second of 720p video - `xai/grok-imagine-video-1.5-preview` -- xAI Grok Imagine Video 1.5 Preview: output_video_seconds $0.08 per second of video; output_video_seconds:480p $0.08 per second of 480p video; output_video_seconds:720p $0.14 per second of 720p video ### audio - `assemblyai/universal-3-pro` -- AssemblyAI AssemblyAI Universal-3 Pro: input_audio_minutes $0.0035 per audio minute - `assemblyai/universal-3.5-pro` -- AssemblyAI AssemblyAI Universal-3.5 Pro: input_audio_minutes $0.0035 per audio minute; input_audio_minutes:keyterms $0.00083 per audio minute (key terms); input_audio_minutes:medical $0.0025 per audio minute (medical); input_audio_minutes:speaker_diarization $0.0003 per audio minute (speaker diarization) - `deepgram/aura-1` -- Deepgram Aura 1: input_characters $0.015 per 1K characters - `deepgram/aura-2-en` -- Deepgram Aura 2 EN: input_characters $0.03 per 1K characters - `deepgram/aura-2-es` -- Deepgram Aura 2 ES: input_characters $0.03 per 1K characters - `deepgram/flux` -- Deepgram FLUX: input_audio_minutes:websocket $0.0077 per audio minute (streaming) - `deepgram/nova-3` -- Deepgram Nova 3: input_audio_minutes $0.0052 per audio minute; input_audio_minutes:websocket $0.0092 per audio minute (streaming) - `elevenlabs/eleven-flash-v2-5` -- ElevenLabs Eleven Flash v2.5: input_characters $0.05 per 1K characters - `elevenlabs/eleven-multilingual-v2` -- ElevenLabs Eleven Multilingual v2: input_characters $0.10 per 1K characters - `elevenlabs/eleven-turbo-v2-5` -- ElevenLabs Eleven Turbo v2.5: input_characters $0.05 per 1K characters - `elevenlabs/eleven-v3` -- ElevenLabs Eleven v3: input_characters $0.10 per 1K characters - `elevenlabs/music-v2` -- ElevenLabs ElevenLabs Music v2: output_audio_seconds $0.0025 per second of generated audio - `google/gemini-3.1-flash-tts` -- Google Gemini 3.1 Flash TTS: input_audio_tokens $3.00 per 1M input audio tokens; input_tokens $0.75 per 1M input tokens; output_audio_tokens $12.00 per 1M output audio tokens; output_tokens $4.50 per 1M output tokens - `inworld/tts-1.5-max` -- Inworld Inworld TTS 1.5 Max: input_characters $0.035 per 1K characters - `inworld/tts-1.5-mini` -- Inworld Inworld TTS 1.5 Mini: input_characters $0.015 per 1K characters - `inworld/tts-2` -- Inworld Inworld TTS 2: input_characters $0.025 per 1K characters - `minimax/music-2.6` -- MiniMax MiniMax Music 2.6: output_lyrics $0.01 per lyrics generation; output_tracks $0.15 per track - `minimax/speech-2.8-hd` -- MiniMax MiniMax Speech 2.8 HD: input_characters $0.10 per 1K characters - `minimax/speech-2.8-turbo` -- MiniMax MiniMax Speech 2.8 Turbo: input_characters $0.06 per 1K characters - `myshell-ai/melotts` -- MyShell MeloTTS: output_audio_seconds $0.000205 per minute of generated audio - `openai/gpt-4o-transcribe` -- OpenAI GPT-4o Transcribe: input_audio_minutes $0.006 per audio minute - `openai/tts-1` -- OpenAI TTS-1: input_characters $0.015 per 1K characters - `openai/tts-1-hd` -- OpenAI TTS-1 HD: input_characters $0.03 per 1K characters - `openai/whisper` -- OpenAI Whisper: input_audio_minutes $0.000453 per audio minute - `openai/whisper-large-v3-turbo` -- OpenAI Whisper Large v3 Turbo: input_audio_minutes $0.000513 per audio minute - `openai/whisper-tiny-en` -- OpenAI Whisper Tiny EN: compute_units $0.011 per 1K compute units - `pipecat-ai/smart-turn-v2` -- Pipecat Smart Turn v2: input_audio_minutes $0.000338 per audio minute - `xai/grok-stt` -- xAI Grok STT: input_audio_minutes $0.001667 per audio minute; input_audio_minutes:streaming $0.003334 per audio minute (streaming) - `xai/grok-tts` -- xAI Grok TTS: input_characters $0.015 per 1K characters - `xai/grok-voice` -- xAI Grok Voice: input_messages $0.004 per text message; total_audio_minutes $0.05 per minute of voice session ### embeddings - `baai/bge-base-en-v1.5` -- BAAI BGE Base EN v1.5: input_tokens $0.0666 per 1M input tokens - `baai/bge-large-en-v1.5` -- BAAI BGE Large EN v1.5: input_tokens $0.204 per 1M input tokens - `baai/bge-m3` -- BAAI BGE M3: input_tokens $0.0118 per 1M input tokens - `baai/bge-small-en-v1.5` -- BAAI BGE Small EN v1.5: input_tokens $0.0202 per 1M input tokens - `google/embeddinggemma-300m` -- Google EmbeddingGemma 300M: compute_units $0.011 per 1K compute units - `pfnet/plamo-embedding-1b` -- Preferred Networks PLaMo Embedding 1B: input_tokens $0.0186 per 1M input tokens - `qwen/qwen3-embedding-0.6b` -- Qwen Qwen3 Embedding 0.6B: input_tokens $0.0118 per 1M input tokens ### other - `baai/bge-reranker-base` -- BAAI BGE Reranker Base: input_tokens $0.00311 per 1M input tokens - `huggingface/distilbert-sst-2-int8` -- Hugging Face DistilBERT SST 2 INT8: input_tokens $0.0263 per 1M input tokens - `microsoft/resnet-50` -- Microsoft ResNet 50: requests $0.00000251 per request