Response cache
jangada can cache LLM responses to save tokens and latency — plugged via
LLM(..., cache=...). The client checks the cache before calling the provider
and populates it after a successful response. Two modes: exact and
semantic.
Exact cache
Hits only when the request is identical (same method, provider/model/params
scope and messages). LRU with optional max_size and ttl, no embedding cost.
from jangada_ai import LLM, ExactCache
llm = LLM("openai", "gpt-4o-mini", cache=ExactCache(max_size=512, ttl=3600))
llm.complete("Summarize the theory of relativity.") # calls the provider
llm.complete("Summarize the theory of relativity.") # served from cache (identical)Semantic cache
Hits when the question is similar enough to a previous one in the same scope
(cosine similarity ≥ threshold). Reuses LLM.embed + a RAG vector_store.
After a semantic hit it applies an exact scope filter
(provider/model/params/method) for high precision.
from jangada_ai import LLM, SemanticCache
embedder = LLM("openai", "text-embedding-3-small")
cache = SemanticCache(embedder, threshold=0.85)
llm = LLM("openai", "gpt-4o-mini", cache=cache)
llm.complete("What is the capital of France?")
llm.complete("Tell me the French capital.") # paraphrase → cache hitTune threshold per embedding model
The cosine scale varies a lot between models — there is no magic number:
| Embedding model | Paraphrase (≈) | Distinct question (≈) |
|---|---|---|
text-embedding-3-small (OpenAI) | 0.62 | 0.11 |
gemini-embedding-001 (Gemini) | 0.90 | 0.49 |
The default is 0.85 (a middle ground). Measure paraphrases vs. distinct
questions on your model and pick a cutoff between the two distributions. Too
high = dead cache; too low = wrong answers (false positives).
What is not cached
- Calls with tools or MCP (
tools=/mcp_servers=). - Streaming (
stream/astream). - Responses coming from fallback (only the primary candidate populates).
- Guardrail refusals.
complete/parse and their async variants share the same key. The cache is
process-local (the Completion is kept in memory; the vector store is used
only for similarity).
What changed in 1.9.0
- The key includes the
parseschema:parse(p, A)followed byparse(p, B)no longer returns theA-typed object. SemanticCacheis scoped by context: system, previous history and non-text parts (images) are part of the scope. "yes" or "continue" in different conversations no longer hit another conversation's answer.- A hit returns a copy with
cost=0.0andcached=True— summing costs in aFlow/Agentdoesn't count an unpaid call again, and changing the returned object doesn't change what is cached. - Eviction and expiry clean the
SemanticCachevector store (dead entries used to stay there and lower the hit rate). - A cache error doesn't break the call: a failure in
get/set(e.g. the semantic cache embedder got rate limited) becomes a miss, logged atDEBUG. ExactCacheandSemanticCacheare thread-safe.