Jangada AIJangada AI

Observability (automatic)

Jangada sends your LLM calls to the observability platform automatically (zero-config): just set up your .env. Each call becomes an observation with provider, model, tokens, cost, latency, tool calls and the capabilities of AI used — sent in the background, with no code instrumentation.

Enable (zero-config)

# .env
JANGADA_OBSERVABILITY=true
JANGADA_OBSERVABILITY_API_KEY=lobs_xxx        # project key (dashboard)
# optional (defaults to the official platform):
# JANGADA_OBSERVABILITY_ENDPOINT=https://api.jangada.dev.br
from jangada_ai import LLM

llm = LLM("openai", "gpt-4o-mini")
resp = llm.complete("Summarize: ...")   # already sent to the platform, on its own

Embeddings are also instrumented automatically since v1.4.1:

embedder = LLM("openai", "text-embedding-3-small")

with observability_session(name="rag.documents.ingest"):
    vectors = embedder.embed(["first chunk", "second chunk"])

The observation records latency, input tokens, estimated cost, vector count and dimensions, and the embeddings capability.

For Gemini, counting follows usage_metadata → count_tokens() with the same model/batch → local estimation as the last fallback. Pricing comes from jangada.dev.br/prices.json, never from a value hardcoded in the adapter. The usageSource and usageEstimated fields distinguish real counts from estimates.

The flag must be "truthy" (1/true/yes/on) and the token present; if either is missing, the mode stays off and nothing is sent (zero cost). Network failures never bring your app down — sending is best-effort on a daemon thread.

Shutdown (don't lose the last trace)

Since sending runs on a daemon thread, a script that ends right after its last call would risk losing that last trace: the interpreter kills daemon threads on exit and the POST would die mid-flight. To prevent this, jangada automatically registers an atexit handler that waits for pending sends to finish before shutting down (with a total safety deadline — it won't hang the process if the network is slow). You don't have to do anything.

The only exception is an abrupt shutdown (os._exit(), a signal that bypasses atexit, or a sys.exit() in a context that ignores handlers): then call the flush manually before exiting.

from jangada_ai import LLM, flush_observability

llm = LLM("openai", "gpt-4o-mini")
resp = llm.complete("last question")

flush_observability()   # ensures the trace above was sent
# ... abrupt shutdown ...

Grouping into a batch

By default each call becomes its own trace, named after the entry script that produced it (e.g. python examples/02_multi.py → 02_multi) — so different scripts stay distinguishable in the dashboard instead of a wall of identical traces. Outside a nameable script (REPL, -c, -m, runners) the name falls back to the method (complete, parse, stream, embed…); it never shows up as "(no name)". To group several calls of one request into the same batch (even when sent one at a time) and give it its own name, open a scope with observability_session(name=...): a batch id is generated, every call inside shares it and the scope name takes precedence — the backend groups them into the same trace.

from jangada_ai import LLM, observability_session

llm = LLM("openai", "gpt-4o-mini")

with observability_session(name="summary+translation", user_id="customer-123"):
    r1 = llm.complete("Summarize: ...")     # observation in the same trace
    r2 = llm.complete("Translate: ...")     # same — grouped by the batch id

observability_session accepts id (reuse an external id), name, user_id, session_id and metadata, and returns the batch id. It works in sync and async code (it uses contextvars).

Production feedback

Capture the end user's reaction (👍/👎 or a score) and attach it to the trace with feedback(), using the batch id returned by observability_session. It's best-effort (never crashes your app): True if sent, False if the key is missing or the network failed.

from jangada_ai import LLM, observability_session, feedback

llm = LLM("openai", "gpt-4o-mini")

with observability_session(name="support") as trace_id:
    resp = llm.complete("How do I issue an invoice?")

# later, when the user rates it:
feedback(trace_id, 1, comment="solved my problem")   # 👍
# feedback(trace_id, -1, comment="wrong answer")      # 👎

It becomes a Score on the trace (source api), shows up in the dashboard next to human 👍/👎, and closes the loop: a 👎 can be promoted to a dataset example and become a regression case in evals.

What is captured

From each call: provider, model, promptTokens/completionTokens (from usage), costUsd (from cost), latency, the input (messages or embedding texts; very long content is truncated), the output (response text, or vector count/dimensions), and the tool calls the model requested (tools: id/name/args).

The input keeps the tool history auditable: each tool_call records its name and arguments ([tool_call consultar_estoque {"produto": "cabo HDMI"}]) and each tool_result records the returned content ([tool_result] {"disponivel": 0, "previsao_dias": 12}), flagged as an error when applicable. That lets you check where every number the model stated came from — not just an empty marker. Individually large args and results are truncated (per-part cap), on top of the global input cap.

Each observation has a status: OK, or INCOMPLETE when the response was cut off by the token limit (finish_reason == "length"), along with the normalized stop reason in finishReason. In the dashboard it becomes a badge (amber) and is part of the status filter.

Capabilities

Each observation records which AI capabilities were used — tools, mcp, a2a, vision, audio, documents, rag, structured_output, guardrails, cache, agents, embeddings. In the dashboard they show up as badges, a filter, and the Analytics → Usage by capability breakdown.

Detection is automatic from the call arguments: images= → vision, files= → documents, tools= → tools, mcp_servers= → mcp, parse() → structured_output, guardrails → guardrails. tools is also derived when the model requests tool calls.

embed()/aembed() register embeddings directly, including sessions that contain only RAG ingestion.

In the dashboard

At app.jangada.dev.br you track everything:

  • Traces and detail — each batch and its observations (provider, model, tokens, cost, latency, tool calls and capabilities), as a table or waterfall.
  • Analytics — cost, calls, tokens, error rate and latency (p50/p95/p99), broken down by model, provider and capability, plus a daily time series.
  • Filters and export — filter by model, provider, errors, dates, userId/sessionId, minimum cost and capability; export to CSV/JSON.
  • Live tail — traces in real time, with pause/resume.
  • Anomalies — automatic warnings when cost/latency/errors drift from the 7-day baseline.
  • Alerts — daily-cost or error-rate rules.
  • Scores — per-trace evaluations (human feedback or LLM-as-judge).
  • Budget — monthly cost cap per project, with tracking and projection.

Details

  • The api_key is the project key, generated in the dashboard and set in .env.
  • Reusing the same batch id (via observability_session(id=...)) appends observations to the same trace, idempotently on the backend.
  • The cost/token fields come from Cost and tokens.

What changed in 1.9.0

  • Failures become traces too. When a call exhausts retries and fallbacks, the library sends an observation with status="ERROR" and the error (type + message, truncated) — previously only successes showed up in the dashboard. To report manually: jangada_ai.observability.auto_report_error(error, provider=..., model=...).
  • Real call start. startedAt now marks when the call started (not when it finished); auto_report(..., started_at=...) accepts an epoch or a datetime.
  • Streaming and transcription reported. stream/astream send the accumulated text at the end (capability streaming); transcribe also reports (audio). Cache hits get the cache capability.
  • Bounded queue. Instead of one thread per call there is a queue (1000 events) with a few workers; if the endpoint is slow and the queue fills up, extra events are dropped — dropped_count() tells how many. flush() and the atexit hook wait for the queue to drain, with a deadline.
  • HTTPS only. The endpoint (and the feedback one) must be https:// (http only on localhost), so the key and prompts don't travel in clear text.
  • Output truncated like the input, and embed no longer sends every vector (only dimensions and counts above a threshold).
  • Send failures go to the DEBUG log of the jangada_ai logger (they never break the call).

On this page