Inkling
thinkingmachines/inklingInkling is Thinking Machines' open-weights hybrid reasoning model, built on a mixture-of-experts architecture. It reasons by default, exposing its chain-of-thought as leading thinking content blocks, with reasoning effort tunable via a Tinker-specific output_config.effort parameter. Available through Tinker's beta Anthropic Messages-compatible endpoint alongside tool use, streaming, and multi-turn conversations. Currently intended for low-traffic testing and internal use rather than high-throughput production deployments; prompt caching, citations, and audio input are not supported through this endpoint.
| unit | list → Elite price |
|---|---|
| per 1M cache write tokens | $1.87 |
| per 1M cached input tokens | $0.374 |
| per 1M input tokens | $1.87 |
| per 1M output tokens | $4.68 |
No Elite discount on this model right now: everyone pays the list price.
Everyone pays the list price. Reaching Elite (hold or stake 1,337 $XAVA, or hold 1,337,000 $XINF) unlocks the Elite price and a bigger buyback of your token. How it works
from openai import OpenAI
client = OpenAI(
base_url="https://www.zinf.ai/v1",
api_key="xk_live_...", # your Xava Inference key
)
r = client.chat.completions.create(
model="thinkingmachines/inkling",
messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)curl https://www.zinf.ai/v1/chat/completions \
-H "Authorization: Bearer $XINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"thinkingmachines/inkling","messages":[{"role":"user","content":"Hello!"}]}'