03 / SEMANTIC CACHING

Fewer repeat calls.
Carefully reused answers.

When a workload repeats, a suitable cached response can avoid another model call. We engineer semantic caching around answer quality, context boundaries and freshness, then measure whether it earns its place.

Discuss your engineering priorities ARCHITECTURE THROUGH TO OPERATION

Similarity is a signal. Reuse is a policy.

A similar question can require a different answer. A cache hit must be suitable for the caller, the current context and the task.

Scoped lookup / eligibility / response

  1. 01

    Establish scope

    Identify tenant, permissions and the context and policy versions.

  2. 02

    Find candidates

    Compare within the allowed namespace using an appropriate similarity metric.

  3. 03

    Check eligibility

    Apply thresholds, expiry and workload-specific exclusions.

  4. 04

    Serve or infer

    Reuse a qualifying answer; otherwise call the model and apply storage policy.

Policy-controlled outcomes

  • REUSE

    Eligible hit

    Reuse current, non-personalised content within the permitted scope.

  • INFERENCE

    Miss or bypass

    Generate a new response for live, personalised or otherwise excluded work.

Measure eligible hit rate, false matches, response quality, end-to-end latency and total operating cost.

Illustrative cache flow. Cache eligibility is enforced alongside permissions, context versions and the application’s data policy.

Start with workloads
where reuse makes sense.

Stable public information and repeated, low-risk explanatory answers can be good candidates. Live account balances, user-specific decisions and action-taking workflows often need bypass rules or stricter exact-match controls.

Semantic response caching, exact response caching and a provider’s prompt-prefix caching are different mechanisms. We establish which problem you are solving, what is stored and where the benefit is expected.

Engineer the conditions
for a useful cache.

Matching and evaluation

Select an embedding approach and similarity threshold. Evaluate false matches as well as cache misses on representative requests.

Isolation

Separate tenants and permission scopes. Include relevant context, model, prompt and policy versions so answers are not reused across incompatible boundaries.

Freshness and invalidation

Define retention, expiry, deletion and invalidation when source information, access policy or generation behaviour changes.

Economics and operations

Compare avoided inference against embedding, lookup, storage and operational overhead. Monitor quality alongside latency and hit rate.

A cache that fits
the wider architecture.

Kong’s AI Semantic Cache provides a gateway-level option. Application-managed caching may offer more control over business context and eligibility. We assess supported stores, deployment boundaries and the full request path before selecting an approach.

Decisions worth getting right.

How much will semantic caching save?

That depends on repetition, eligibility, hit quality and the relative cost of inference and caching. We measure a baseline and evaluate the total cost; we do not assume a fixed saving.

Is a similar prompt enough to return the same answer?

No. The caller’s permissions, live context and the meaning of the task can differ. Similarity is considered within an eligibility policy, with bypass and invalidation rules where reuse is unsuitable.

Can a cache be shared across customers?

Only where the content and access policy explicitly permit it. Tenant-specific or sensitive content requires isolation and appropriate storage controls. We make that boundary part of the design and test it.

Make response reuse
a measured decision.

Bring us a workload, a technical constraint or an architecture that needs a second look. We will help define a practical next step.

Discuss your engineering priorities

Architecture advice, focused implementation and support for your engineering team.