← Back to all stories

Check prompt cache savings before switching inference providers

A practical preflight for measuring cached input, protecting tenant boundaries, and comparing the full cost of serverless inference.

Disclosure: Tylor.nz has a DigitalOcean affiliate relationship. This article uses direct documentation links and contains no commission-bearing signup link. It is a documentation-based guide, not a hands-on benchmark.

If your application repeatedly sends a long policy document, tool catalogue, or instruction set to a model, check what you actually pay for repeated input before changing providers. A lower cached-input rate can matter. So can output length, unsuccessful answers, and the engineering work needed to move.

On October 8, 2026, DigitalOcean announced general availability of automatic prompt caching across its hosted models. Cache writes have no additional charge, and matching prefixes receive automatic cache-aware routing. This is a specific hosted-model change; it is not a new launch of every caching feature available through DigitalOcean. Official announcement

The useful next step is a small, budgeted comparison against your current setup. Keep the model and task fixed first. Decide whether to migrate only after measuring useful answers, latency, and total cost.

Check fit and access first

DigitalOcean’s serverless route avoids provisioning an inference server, while dedicated inference trades token billing for GPU-hour billing and more environment control. The company positions dedicated capacity for sustained high-throughput workloads. For a spiky application, start the comparison with serverless; for continuously busy workloads, include a dedicated-capacity quote and operating overhead. Serverless overview

Check the current model catalogue for the exact model ID and supported capabilities. The catalogue explicitly says EU data residency is unavailable. A nearby Droplet region does not establish that inference meets your residency requirement. If location is a hard constraint, resolve it before uploading representative data. Model availability

You need an account, access to the selected model, and sufficient prepaid funds or applicable credits. Serverless access stops when eligible credits and balances are exhausted. Obtain current input, cached-input and output rates from the pricing page; this article deliberately avoids freezing a price into a purchasing recommendation. Current pricing and prepayment

Use a narrowly scoped model access key in a server-side secret store. DigitalOcean supports model scoping and optional VPC restrictions; keys must not appear in frontend code. Check the documented VPC DNS prerequisite if you restrict access that way. Model access keys

Understand the cache boundary

For hosted models, prefix matching is exact. Keep shared instructions first and changing content later. The optional prompt_cache_key separates reuse within an account; use stable opaque tenant or workload identifiers, at most 64 characters. Different keys do not share cached prefixes. It does not replace authorization.

The cache contains internal KV state rather than prompt text. Entries have no configured expiry and can be evicted under load. Hosted models do not support prompt_cache_retention or a caching opt-out. Minimum reusable blocks vary by model, up to 256 tokens. Busy routing can still produce misses. Changing a namespace does not prove deletion of old state.

Responses expose cached input through cache_read_input_tokens and prompt_tokens_details.cached_tokens. Treat these as alternative reports of the same cached tokens, not additive charges. Anthropic and OpenAI caching have separate rules; do not copy their retention controls into this hosted-model workflow. Caching behavior and limitations

DigitalOcean says it does not train models on customer inputs or send hosted-model data back to the model creator. Read that privacy statement alongside the cache documentation. Avoid turning “no training” into a promise of no transient processing state or a deletion deadline. If your approval depends on exact retention semantics, ask the provider before sending sensitive material. Data privacy

Design a useful comparison

Start with synthetic or approved non-sensitive examples. Choose a task with an objective review rule, such as extracting a support category and citing the relevant policy paragraph. Write the acceptance criteria before looking at cost.

Prepare four small groups:

  • Repeated context: several different questions about the same reference document.
  • Changed context: the same question set after a deliberate policy revision.
  • Separate tenants: equivalent tasks under two distinct cache namespaces.
  • Ordinary traffic: a sample resembling the mix and spacing your application really expects.

The first group explores reuse. The second checks whether your prompt assembly supplies the updated material. The third exercises your own namespace construction and access-control tests; a cache counter cannot prove tenant security. The fourth prevents a rapid repeated-request demonstration from becoming an unrealistic monthly forecast.

Preserve your current provider as a control. Use matching documents, questions, output limits and evaluation criteria. Record which settings differ instead of hiding them. Keep failure and retry costs in the comparison.

DigitalOcean’s Chat Completions guide requires a model ID and messages containing the necessary context. It recommends max_completion_tokens for controlling response length and marks max_tokens deprecated. Follow the current endpoint guide when implementing the test. Do not assume a cached prefix means you can omit required context from later requests. Chat Completions setup

Calculate the whole request cost

For the hosted-model workflow, use this worksheet with current rates in the same currency per million tokens:

Estimated token cost = ((input tokens − cached input tokens) × standard input rate + cached input tokens × cached input rate + output tokens × output rate) ÷ 1,000,000.

This is a planning calculation, excluding taxes, credits, discounts and any separately billed features. Use the billing record for reconciliation. Do not apply it unchanged to a provider that separately charges cache writes.

Here is a hypothetical example, not a DigitalOcean price quote or measured result. Suppose a request uses 10,000 input tokens, 8,000 cached input tokens and 1,000 output tokens. Let input cost 1 unit per million, cached input 0.1 units and output 2 units. The estimate is 0.0048 units, versus 0.012 without reuse: 60% lower total token cost. The cached-input rate is 90% lower, but the whole request is not 90% cheaper.

Next divide the total spend for each test group by its number of accepted answers. If no answers meet your criteria, record the group as a failure instead of reporting an attractive cost per answer. Add migration time and recurring maintenance separately when deciding whether the difference is worthwhile.

Reconcile results before buying more capacity

In the control panel, open Inference Engine, Serverless Inference, then Analyze. Use the platform metrics alongside your application log. Metrics navigation

Track time to first token separately from complete-response latency. DigitalOcean defines the former around queueing and prefill, while the latter also includes generation. A faster initial response can still finish slowly, so choose the metric that matches what your users experience. Include errors and retries in the same review. Metric definitions

Billing’s Cost Analysis tab now includes Serverless Inference Live Usage. It refreshes every ten minutes, covers thirty days and includes token-based inference only. Inspect the access key, model, token breakdown and estimated cost for the test period. Amounts below a cent appear as less than $0.01; that is not zero usage. Estimates exclude credits, promotional discounts and taxes, and the invoice remains the final billing reference. Billing reconciliation

Make the provider decision

Proceed with a limited rollout when your accepted-answer cost improves enough to justify switching, your latency targets hold under representative traffic, and your data requirements are met. Set a rollback trigger and retain the comparison notes.

Stay with the current provider when quality regresses or the savings disappear after retries and migration work. For a low-volume application, simplifying its prompt or reducing unnecessary calls may be more valuable than changing infrastructure. For steady high utilisation, compare dedicated inference using a realistic capacity and staffing estimate.

Caching gives you another measurable variable. Buy on the result your application needs: acceptable answers, within its response-time and data-handling requirements, at a sustainable total cost.

Research checked October 9, 2026. No account resources were created and no paid inference requests were run for this article.