Cost and tokens
Every successful response comes back with usage (tokens) and cost (estimated USD).
comp = llm.complete("...")
print(comp.usage) # {"input_tokens": ..., "output_tokens": ...}
print(comp.cost) # e.g.: 0.000123 (USD, approximate)The client calls pricing.compute_cost() after each success and sets
Completion.cost. FlowResult and GraphResult aggregate usage/cost
along the chain.
Pricing table
Prices come from a JSON catalog (approximate, per 1M tokens). You can register/adjust them at runtime:
from jangada_ai import register_price, price_for
register_price("my-model", 0.5, 1.5) # USD per 1M tokens (input, output)
print(price_for("gpt-4o-mini"))⚠️ The values are approximate and serve for estimation/observability — do not treat them as a billing source.
Always-fresh prices (without updating the lib)
Prices live in a JSON catalog (no longer hardcoded). The copy bundled in the
package is just the offline fallback; jangada publishes the catalog at
jangada.dev.br/prices.json and applies it on its own, with no code from you —
and without running pip install -U.
Automatic (default). The first time a cost is computed (right below each provider call), the lib fires a background refresh of the catalog, cached for 1 day. You just use the LLM normally:
comp = llm.complete("...")
print(comp.cost) # already tends to use today's prices (background refresh)- Non-blocking: runs on a daemon thread;
importnever touches the network. - At most once/day: cached at
~/.cache/jangada/prices.json(and once per process). - Resilient: network failed? stays on cache/embedded — never raises.
- Disable: env
JANGADA_NO_PRICE_REFRESH=1.
Manual (refresh_prices). For explicit/synchronous control (boot, force now,
custom URL):
jangada_ai.refresh_prices() # synchronous, respects the cache
jangada_ai.refresh_prices(ttl=3600, force=True) # revalidate / force nowURL defaults to https://jangada.dev.br/prices.json (override with url= or the
JANGADA_PRICES_URL env var; no .env required). The manual override
(register_price) takes priority over everything.
Multimodal cost
-
Image (vision): there's no separate price — providers already count the image's tokens inside
input_tokens. Image cost already comes out of the normal table. -
Audio (transcription): billed per minute, not per token. The cost only appears when
usagecarries the duration (audio_seconds): passresponse_format="verbose_json"on transcription or set the duration onAudio.from_bytes(data, mime, duration=...). Register/adjust the per-minute price withregister_audio_price:from jangada_ai import register_audio_price register_audio_price("whisper-1", 0.006) # USD per minute -
Detection:
detect_objects/adetect_objectsreturn onlylist[Detection](no cost). For the cost, usedetect_objects_full/adetect_objects_full, which return aDetectionResultwith.detectionsand.completion/.cost/.usage:from jangada_ai import detect_objects_full res = detect_objects_full(llm, "photo.jpg") print(res.detections, res.cost)
Where this shows up
Completion.cost/Completion.usageon every call.- Aggregated totals in Flows and Graph.
- In Step-by-step debug, the cost of each step is shown in the trace.
What changed in 1.9.0
Usage contract
Every adapter returns usage in the same shape:
| Key | Meaning |
|---|---|
input_tokens | total input, cache tokens included |
output_tokens | total output (on Gemini it includes thinking tokens) |
cache_read_tokens | part of the input read from cache (optional) |
cache_write_tokens | part of the input written to cache (optional) |
reasoning_tokens | part of the output spent reasoning (informative only) |
server_tool_requests | native tool calls, e.g. {"web_search": 2} |
compute_cost charges uncached input at full price, cache reads/writes at the
model's cache price, output once (reasoning is not charged twice) and adds
the per-call fee of native tools. Previously cache and thinking were ignored —
Gemini with thinking and Claude with prompt caching came out far cheaper than the
real bill.
Cache prices, above-200k tier and tool fees
- Each rule may have
cache_read/cache_write(USD per 1M tokens). Without them the library uses the family's documented multiplier (Claude 0.1×/1.25×; Gemini 0.1×; gpt-5 0.1×; gpt-4.1/o3/o4-mini 0.25×; gpt-4o/o1/o3-mini 0.5×; unknown 1×). above_200k=(in, out)applies another price when the prompt exceeds 200k tokens (e.g. Gemini 2.5 Pro, 3.1 Pro).- Per-call fees for native tools (
tool_feescatalog): Anthropic and OpenAI web search US$10/1k, Google Search on Gemini 3.x US$14/1k queries, OpenAI file search US$2.50/1k, Mistral connectors, etc. They go intoCompletion.cost. Free tiers are not deducted. - Model matching gained boundaries:
claude-opus-4is 15/75 again (it no longer inherits the Opus 4.5+ price),gpt-5-pro/o3-minidon't take the base model rule, and a model without a rule returnsNoneinstead of a wrong price. - A hit on the library's own cache returns
cost=0.0andcached=True.
from jangada_ai import register_price
from jangada_ai.pricing import register_tool_fee, compute_cost
register_price(r"my-model", 1.0, 4.0, cache_read=0.1, above_200k=(2.0, 8.0))
register_tool_fee("web_search", 12.0, provider="my-provider") # USD per 1000 calls
compute_cost("claude-sonnet-5", {"input_tokens": 10_000, "cache_read_tokens": 8_000,
"output_tokens": 500,
"server_tool_requests": {"web_search": 1}},
provider="anthropic")The remote catalog refresh now only accepts https, validates types and sizes,
rejects dangerous regexes and replaces the remote batch on each refresh (it
used to accumulate).
Example
examples/retry_cost_example.py — runnable script.
examples/pricing_refresh_example.py — dynamic prices with refresh_prices().
Retry and fallback
jangada combines two defenses against API failures: retry with backoff on the same candidate and fallback to another model/provider.
Response cache
Cache LLM responses to save tokens and latency — exact (ExactCache) and semantic (SemanticCache, by embedding similarity). Plugged via LLM(..., cache=...).