Jangada AIJangada AI

Response cache

jangada can cache LLM responses to save tokens and latency — plugged via LLM(..., cache=...). The client checks the cache before calling the provider and populates it after a successful response. Two modes: exact and semantic.

Exact cache

Hits only when the request is identical (same method, provider/model/params scope and messages). LRU with optional max_size and ttl, no embedding cost.

from jangada_ai import LLM, ExactCache

llm = LLM("openai", "gpt-4o-mini", cache=ExactCache(max_size=512, ttl=3600))
llm.complete("Summarize the theory of relativity.")   # calls the provider
llm.complete("Summarize the theory of relativity.")   # served from cache (identical)

Semantic cache

Hits when the question is similar enough to a previous one in the same scope (cosine similarity ≥ threshold). Reuses LLM.embed + a RAG vector_store. After a semantic hit it applies an exact scope filter (provider/model/params/method) for high precision.

from jangada_ai import LLM, SemanticCache

embedder = LLM("openai", "text-embedding-3-small")
cache = SemanticCache(embedder, threshold=0.85)

llm = LLM("openai", "gpt-4o-mini", cache=cache)
llm.complete("What is the capital of France?")
llm.complete("Tell me the French capital.")   # paraphrase → cache hit

Tune threshold per embedding model

The cosine scale varies a lot between models — there is no magic number:

Embedding modelParaphrase (≈)Distinct question (≈)
text-embedding-3-small (OpenAI)0.620.11
gemini-embedding-001 (Gemini)0.900.49

The default is 0.85 (a middle ground). Measure paraphrases vs. distinct questions on your model and pick a cutoff between the two distributions. Too high = dead cache; too low = wrong answers (false positives).

What is not cached

  • Calls with tools or MCP (tools=/mcp_servers=).
  • Streaming (stream/astream).
  • Responses coming from fallback (only the primary candidate populates).
  • Guardrail refusals.

complete/parse and their async variants share the same key. The cache is process-local (the Completion is kept in memory; the vector store is used only for similarity).

What changed in 1.9.0

  • The key includes the parse schema: parse(p, A) followed by parse(p, B) no longer returns the A-typed object.
  • SemanticCache is scoped by context: system, previous history and non-text parts (images) are part of the scope. "yes" or "continue" in different conversations no longer hit another conversation's answer.
  • A hit returns a copy with cost=0.0 and cached=True — summing costs in a Flow/Agent doesn't count an unpaid call again, and changing the returned object doesn't change what is cached.
  • Eviction and expiry clean the SemanticCache vector store (dead entries used to stay there and lower the hit rate).
  • A cache error doesn't break the call: a failure in get/set (e.g. the semantic cache embedder got rate limited) becomes a miss, logged at DEBUG.
  • ExactCache and SemanticCache are thread-safe.

On this page