Fast responses. Lower-cost cached context. One API key for the OpenAI and Anthropic SDKs.
Nemotron 3.5 Lightning 30B, 1 of 4
Lower cached input price (USD / 1M tokens)
Session-aware routing keeps repeat calls close to cached context.
One session, one replica.
Repeat calls prefer the replica holding their context.
No hint needed.
Provide a session ID, or let automatic affinity use the conversation when available.
Reported, billed as cached.
Cache hits appear in usage and are billed at the cached rate. Hits stay best effort.
Per 1M input tokens, at each model's measured cache hit rate.
Price per 1M input tokens, log scale
cheaper, 98.8% cache hit
cheaper, 90.7% cache hit
Streams report cached tokens when usage is requested. Your usage page shows your own billed figures.
4 models
Model capabilities: Tool calling / JSON mode / Streaming / Image input
Effective input $0.012 per 1M at that hit rate Effective input rate $0.012 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours.
Model capabilities: Tool calling / JSON mode / Streaming
Model capabilities: Tool calling / JSON mode / Streaming / Image input
Effective input $0.019 per 1M at that hit rate Effective input rate $0.019 per 1M tokens, from the measured cache share on billed agent traffic in the last 24 hours.
Model capabilities: Tool calling / JSON mode / Streaming / Image input
Loading agent setup.
One command points the agent you already run at RunInfra. One takes it back.
Every flag, and what preview still needs. Read the guide
Automatic setup, verified live against RunInfra.
npx @runinfra/connect claudeNode 20 or newer.
Writes~/.claude/settings.json
The check request uses credits.
Undo with runinfra-connect claude off.