Jangada AIJangada AI

Cost and tokens

Every successful response comes back with usage (tokens) and cost (estimated USD).

comp = llm.complete("...")
print(comp.usage)   # {"input_tokens": ..., "output_tokens": ...}
print(comp.cost)    # e.g.: 0.000123  (USD, approximate)

The client calls pricing.compute_cost() after each success and sets Completion.cost. FlowResult and GraphResult aggregate usage/cost along the chain.

Pricing table

Prices come from a JSON catalog (approximate, per 1M tokens). You can register/adjust them at runtime:

from jangada_ai import register_price, price_for

register_price("my-model", 0.5, 1.5)   # USD per 1M tokens (input, output)
print(price_for("gpt-4o-mini"))

⚠️ The values are approximate and serve for estimation/observability — do not treat them as a billing source.

Always-fresh prices (without updating the lib)

Prices live in a JSON catalog (no longer hardcoded). The copy bundled in the package is just the offline fallback; jangada publishes the catalog at jangada.dev.br/prices.json and applies it on its own, with no code from you — and without running pip install -U.

Automatic (default). The first time a cost is computed (right below each provider call), the lib fires a background refresh of the catalog, cached for 1 day. You just use the LLM normally:

comp = llm.complete("...")
print(comp.cost)   # already tends to use today's prices (background refresh)
  • Non-blocking: runs on a daemon thread; import never touches the network.
  • At most once/day: cached at ~/.cache/jangada/prices.json (and once per process).
  • Resilient: network failed? stays on cache/embedded — never raises.
  • Disable: env JANGADA_NO_PRICE_REFRESH=1.

Manual (refresh_prices). For explicit/synchronous control (boot, force now, custom URL):

jangada_ai.refresh_prices()                       # synchronous, respects the cache
jangada_ai.refresh_prices(ttl=3600, force=True)   # revalidate / force now

URL defaults to https://jangada.dev.br/prices.json (override with url= or the JANGADA_PRICES_URL env var; no .env required). The manual override (register_price) takes priority over everything.

Multimodal cost

  • Image (vision): there's no separate price — providers already count the image's tokens inside input_tokens. Image cost already comes out of the normal table.

  • Audio (transcription): billed per minute, not per token. The cost only appears when usage carries the duration (audio_seconds): pass response_format="verbose_json" on transcription or set the duration on Audio.from_bytes(data, mime, duration=...). Register/adjust the per-minute price with register_audio_price:

    from jangada_ai import register_audio_price
    register_audio_price("whisper-1", 0.006)   # USD per minute
  • Detection: detect_objects/adetect_objects return only list[Detection] (no cost). For the cost, use detect_objects_full/adetect_objects_full, which return a DetectionResult with .detections and .completion/.cost/.usage:

    from jangada_ai import detect_objects_full
    res = detect_objects_full(llm, "photo.jpg")
    print(res.detections, res.cost)

Where this shows up

What changed in 1.9.0

Usage contract

Every adapter returns usage in the same shape:

KeyMeaning
input_tokenstotal input, cache tokens included
output_tokenstotal output (on Gemini it includes thinking tokens)
cache_read_tokenspart of the input read from cache (optional)
cache_write_tokenspart of the input written to cache (optional)
reasoning_tokenspart of the output spent reasoning (informative only)
server_tool_requestsnative tool calls, e.g. {"web_search": 2}

compute_cost charges uncached input at full price, cache reads/writes at the model's cache price, output once (reasoning is not charged twice) and adds the per-call fee of native tools. Previously cache and thinking were ignored — Gemini with thinking and Claude with prompt caching came out far cheaper than the real bill.

Cache prices, above-200k tier and tool fees

  • Each rule may have cache_read/cache_write (USD per 1M tokens). Without them the library uses the family's documented multiplier (Claude 0.1×/1.25×; Gemini 0.1×; gpt-5 0.1×; gpt-4.1/o3/o4-mini 0.25×; gpt-4o/o1/o3-mini 0.5×; unknown 1×).
  • above_200k=(in, out) applies another price when the prompt exceeds 200k tokens (e.g. Gemini 2.5 Pro, 3.1 Pro).
  • Per-call fees for native tools (tool_fees catalog): Anthropic and OpenAI web search US$10/1k, Google Search on Gemini 3.x US$14/1k queries, OpenAI file search US$2.50/1k, Mistral connectors, etc. They go into Completion.cost. Free tiers are not deducted.
  • Model matching gained boundaries: claude-opus-4 is 15/75 again (it no longer inherits the Opus 4.5+ price), gpt-5-pro/o3-mini don't take the base model rule, and a model without a rule returns None instead of a wrong price.
  • A hit on the library's own cache returns cost=0.0 and cached=True.
from jangada_ai import register_price
from jangada_ai.pricing import register_tool_fee, compute_cost

register_price(r"my-model", 1.0, 4.0, cache_read=0.1, above_200k=(2.0, 8.0))
register_tool_fee("web_search", 12.0, provider="my-provider")   # USD per 1000 calls

compute_cost("claude-sonnet-5", {"input_tokens": 10_000, "cache_read_tokens": 8_000,
                                  "output_tokens": 500,
                                  "server_tool_requests": {"web_search": 1}},
             provider="anthropic")

The remote catalog refresh now only accepts https, validates types and sizes, rejects dangerous regexes and replaces the remote batch on each refresh (it used to accumulate).

Example

examples/retry_cost_example.py — runnable script.

examples/pricing_refresh_example.py — dynamic prices with refresh_prices().

On this page