Skip to main content

AI for your own applications

Beyond the console panel, an organization owner or admin can bind AI to a deployed application — like binding a database: bind, and env vars appear in the app.

How to bind: open the application's environment in the console, use Create Service → AI Model Access, pick the application and an SDK preset, and confirm. The binding is created for that app (you never enter an application id), the env vars are injected, and the app redeploys. Manage it afterwards from its card under the environment's services — Rotate key, adjust budget/rate limits, or Unbind.

Create Service → AI Model Access: pick the application and the SDK whose env names to inject, then Bind AI

An environment page: the app’s stack, and the AI binding under Attached services as “AI — AI model access”, active, with Configure

An AI binding: the env names it injects (ANTHROPIC_BASE_URL, ANTHROPIC_API_KEY), its gateway key, budget and rate limits, the models the app can call, and Rotate key / Unbind

The binding writes two environment variables into the application and redeploys it. You pick a preset to match whatever SDK your app uses — the gateway speaks both major dialects:

Your app usesBind with
Anthropic SDKANTHROPIC_BASE_URL + ANTHROPIC_API_KEY
OpenAI SDKOPENAI_BASE_URL + OPENAI_API_KEY
Something elseWhatever env names your code reads for base URL + key

Don't let vendor-flavored names mislead you: the values are not a vendor account. The URL points at the platform's model gateway, and the key is a per-application gateway key with its own monthly budget and rate limits. Your app's SDK talks to the platform gateway with zero code changes.

Who provides what, and who pays:

LayerProvided byPaid by
The two env valuesThe platform, at bind time—
Free-tier modelsThe platform's built-in inferenceThe operator (compute)
Premium modelsThe operator's vendor account behind the gatewayThe operator pays the vendor; you pay via your plan
Your app's usage—Bounded by the app's budget/rate limits and your plan's cap

If your app already sets its own vendor key in those env names (the do-it-yourself alternative — your key, your vendor bill, the platform uninvolved), binding will refuse rather than overwrite it: remove the manual key first, or bind under different env names. Rotating a binding's key invalidates the old one immediately and redeploys the app — use it if a key leaks, not to pick up a plan change (those apply on their own); unbinding revokes the key and removes the env vars.

Your app never holds a vendor key: it talks to the gateway, and the gateway enforces the binding's budget and your plan's cap.

Choosing which model your app calls​

The binding injects only the base URL and the key — it does not set which model your app requests. That's your app's own config, because the platform can't know the env var your code reads for it (one app calls it GENERATION_MODEL, another MODEL, another AI_MODEL) or which model you want. So after binding, set your model-selection env var yourself on the application's Environment tab and redeploy.

Two rules for the value:

  • Use a gateway model name, not a raw vendor name. The key is scoped to the platform's model groups, so request e.g. free/llama3.2 or paid/claude-opus — not a vendor id like claude-opus-5 (the gateway rejects unknown names).
  • The names you're allowed depend on your plan's AI tier, and on which providers your plan includes. The free tier always offers free/*; paid/* models are available only if your plan is on the paid AI tier (and the operator has configured that provider). A plan can also be limited to specific providers — for example Groq-only, or Claude-only — in which case only those paid/* names appear for your keys. A request for a name outside your plan is refused even though the model exists — see your plan, or ask your operator. When your plan's tier changes, the models your key (and every app binding) can call update automatically within a few minutes — you don't rotate, re-bind, or redeploy to get them. Re-query /v1/models to see the new set.

For a quick smoke test, free/llama3.2 works on any plan with the AI assistant enabled.

Discovering what you can call, in code. You don't have to hardcode the list: your app can query the gateway's standard models endpoint with the injected key and get exactly the models this binding is allowed (it's scoped to your plan's tier). For example GET $ANTHROPIC_BASE_URL/v1/models (or the SDK's models.list()) returns them — pick one for GENERATION_MODEL at runtime, or just to verify your tier.

Which model for which job​

Model names map to a size/speed class. Pick the smallest one that does the job — smaller is faster and cheaper, and for narrow tasks it's often just as good. (Exact names depend on what your operator has enabled; query /v1/models for your list.)

ModelClassGood at (concretely)Example
free/llama3.2Tiny, free, localSmoke tests, simple/short drafting, basic classification. Free tier runs on CPU — expect slower replies and weaker facts."Generate a placeholder product blurb for this SKU."
paid/claude-haikuSmall, fast, cheapHigh-volume extraction, classification, intent-reading, tidy structured output, first-draft copy. The workhorse for app features that run a lot."Extract {name, email, plan_interest} from this inbound message." · "Is this review positive, neutral, or negative?"
paid/claude-sonnetMid, balancedMost app features that need real reasoning and good writing — summaries, support replies, content generation from context."Write a support reply from this ticket + these two KB snippets."
paid/claude-opusFrontier, slower, priciestHard multi-step reasoning, complex agents, high-stakes or ambiguous work where quality matters more than cost."Given this 40-page contract, list the termination clauses and their risks."
paid/gptMid (OpenAI)General-purpose alternative when you specifically want a GPT model."Rewrite this paragraph in a friendlier tone."
paid/groq-llama-70bFast, open-weightsLatency-critical or cost-sensitive tasks that still need a capable model; good for tool-use loops."Route this request to the right handler with a one-word label."

Rule of thumb: haiku for volume, sonnet for most features, opus for the hard 5%. Start at haiku and move up only if quality falls short — not the other way around.

Prompt caching (paid Claude models) — an optimisation, safe to skip on a first read

Prompt caching​

If your app sends the same large prefix on many requests — a big system prompt, a long instruction block, tool definitions, or a growing conversation — mark it with Anthropic prompt caching and the platform gateway passes it straight through to the model. A cache read is billed at roughly 10% of normal input, so a stable prefix reused across calls cuts cost sharply.

A stable prefix is the part of your request that's identical on every call — the system prompt, tool definitions, a long standing instruction or reference document. Mark the end of it with a cache breakpoint; put the part that changes (the user's message) after the breakpoint.

  • Your system message is cached for you. On every paid/claude-* model the gateway automatically marks the request's system message as a cache breakpoint. Anthropic caches everything up to a breakpoint, and tool definitions come before the system block, so a stable system prompt and your tool definitions get cache reads with no code change. The same applies to the Org Agent, whose system prompt and tool definitions repeat on every turn of a tool-call loop.
  • Anything in the conversation itself (a long reference document in a user turn, a growing multi-turn history) is a per-request choice you make in your own code: put "cache_control": {"type": "ephemeral"} on the last content block you want cached. (In the Anthropic SDK: a cache_control field on the block; the OpenAI dialect supports the same on message content.) Anthropic allows at most four breakpoints per request; the automatic system one counts as one of them.
  • One caveat of the automatic breakpoint: a cache write costs ~1.25× normal input, so a long system prompt that is sent once and never reused within five minutes pays a small premium instead of saving. Reuse it at least once and it pays for itself (see the example below).
  • It only helps paid Claude models with a large-enough stable prefix — roughly 1K+ tokens for Sonnet/Opus, larger for Haiku (a few-thousand-token prefix is a safe floor). Below the minimum, nothing is cached and you're billed normally. free/llama3.2 ignores it.

How you're billed — each request splits its input into three separately-priced buckets:

BucketWhat it isPrice vs normal input
Cache writethe prefix, the first time it's seen~1.25× (one-time premium)
Cache readthe same prefix, every later call~0.10×
Normal inputeverything after the breakpoint (the user's message)1×

So a big prefix costs a little extra once, then a tenth of its price on every repeat — while your variable input is always billed normally.

Worked example — a 16,800-token system prefix on Sonnet (input at $3 / million tokens):

Without cachingWith caching
1st call$0.050$0.063 (write)
Each later call$0.050$0.005 (read)
10 calls total$0.50$0.11

≈ 78% cheaper across 10 calls, and it pays for itself by the 2nd. The more you reuse the prefix, the closer the savings get to 90%.

Metering stays accurate: the gateway records cache-read and cache-write tokens at these rates, so your plan's spend reflects the savings automatically — no action needed.