Observability (automatic)
Jangada sends your LLM calls to the observability platform automatically
(zero-config): just set up your .env. Each call becomes an observation with
provider, model, tokens, cost, latency, tool calls and the capabilities of AI
used — sent in the background, with no code instrumentation.
Enable (zero-config)
# .env
JANGADA_OBSERVABILITY=true
JANGADA_OBSERVABILITY_API_KEY=lobs_xxx # project key (dashboard)
# optional (defaults to the official platform):
# JANGADA_OBSERVABILITY_ENDPOINT=https://api.jangada.dev.brfrom jangada_ai import LLM
llm = LLM("openai", "gpt-4o-mini")
resp = llm.complete("Summarize: ...") # already sent to the platform, on its ownEmbeddings are also instrumented automatically since v1.4.1:
embedder = LLM("openai", "text-embedding-3-small")
with observability_session(name="rag.documents.ingest"):
vectors = embedder.embed(["first chunk", "second chunk"])The observation records latency, input tokens, estimated cost, vector count and
dimensions, and the embeddings capability.
For Gemini, counting follows usage_metadata → count_tokens() with the same
model/batch → local estimation as the last fallback. Pricing comes from
jangada.dev.br/prices.json, never from a value hardcoded in the adapter. The
usageSource and usageEstimated fields distinguish real counts from estimates.
The flag must be "truthy" (1/true/yes/on) and the token present; if
either is missing, the mode stays off and nothing is sent (zero cost). Network
failures never bring your app down — sending is best-effort on a daemon thread.
Shutdown (don't lose the last trace)
Since sending runs on a daemon thread, a script that ends right after its last
call would risk losing that last trace: the interpreter kills daemon threads on
exit and the POST would die mid-flight. To prevent this, jangada automatically
registers an atexit handler that waits for pending sends to finish before
shutting down (with a total safety deadline — it won't hang the process if the
network is slow). You don't have to do anything.
The only exception is an abrupt shutdown (os._exit(), a signal that
bypasses atexit, or a sys.exit() in a context that ignores handlers): then
call the flush manually before exiting.
from jangada_ai import LLM, flush_observability
llm = LLM("openai", "gpt-4o-mini")
resp = llm.complete("last question")
flush_observability() # ensures the trace above was sent
# ... abrupt shutdown ...Grouping into a batch
By default each call becomes its own trace, named after the entry script that
produced it (e.g. python examples/02_multi.py → 02_multi) — so different
scripts stay distinguishable in the dashboard instead of a wall of identical
traces. Outside a nameable script (REPL, -c, -m, runners) the name falls back
to the method (complete, parse, stream, embed…); it never shows up as
"(no name)". To group several calls of one request into the same batch (even
when sent one at a time) and give it its own name, open a scope with
observability_session(name=...): a batch id is generated, every call inside
shares it and the scope name takes precedence — the backend groups them into the
same trace.
from jangada_ai import LLM, observability_session
llm = LLM("openai", "gpt-4o-mini")
with observability_session(name="summary+translation", user_id="customer-123"):
r1 = llm.complete("Summarize: ...") # observation in the same trace
r2 = llm.complete("Translate: ...") # same — grouped by the batch idobservability_session accepts id (reuse an external id), name, user_id,
session_id and metadata, and returns the batch id. It works in sync and async
code (it uses contextvars).
Production feedback
Capture the end user's reaction (👍/👎 or a score) and attach it to the trace
with feedback(), using the batch id returned by observability_session.
It's best-effort (never crashes your app): True if sent, False if the key
is missing or the network failed.
from jangada_ai import LLM, observability_session, feedback
llm = LLM("openai", "gpt-4o-mini")
with observability_session(name="support") as trace_id:
resp = llm.complete("How do I issue an invoice?")
# later, when the user rates it:
feedback(trace_id, 1, comment="solved my problem") # 👍
# feedback(trace_id, -1, comment="wrong answer") # 👎It becomes a Score on the trace (source api), shows up in the dashboard next
to human 👍/👎, and closes the loop: a 👎 can be promoted to a dataset example
and become a regression case in evals.
What is captured
From each call: provider, model, promptTokens/completionTokens (from
usage), costUsd (from cost), latency, the input (messages or embedding
texts; very long content is truncated), the output (response text, or vector
count/dimensions), and the tool calls the model requested (tools: id/name/args).
The input keeps the tool history auditable: each tool_call records its
name and arguments ([tool_call consultar_estoque {"produto": "cabo HDMI"}])
and each tool_result records the returned content ([tool_result] {"disponivel": 0, "previsao_dias": 12}), flagged as an error when applicable.
That lets you check where every number the model stated came from — not just an
empty marker. Individually large args and results are truncated (per-part cap), on
top of the global input cap.
Each observation has a status: OK, or INCOMPLETE when the response was
cut off by the token limit (finish_reason == "length"), along with the
normalized stop reason in finishReason. In the dashboard it becomes a badge
(amber) and is part of the status filter.
Capabilities
Each observation records which AI capabilities were used — tools, mcp,
a2a, vision, audio, documents, rag, structured_output, guardrails,
cache, agents, embeddings. In the dashboard they show up as badges, a
filter, and the Analytics → Usage by capability breakdown.
Detection is automatic from the call arguments: images= → vision, files= →
documents, tools= → tools, mcp_servers= → mcp, parse() →
structured_output, guardrails → guardrails. tools is also derived when the
model requests tool calls.
embed()/aembed() register embeddings directly, including sessions that
contain only RAG ingestion.
In the dashboard
At app.jangada.dev.br you track everything:
- Traces and detail — each batch and its observations (provider, model, tokens, cost, latency, tool calls and capabilities), as a table or waterfall.
- Analytics — cost, calls, tokens, error rate and latency (p50/p95/p99), broken down by model, provider and capability, plus a daily time series.
- Filters and export — filter by model, provider, errors, dates, userId/sessionId, minimum cost and capability; export to CSV/JSON.
- Live tail — traces in real time, with pause/resume.
- Anomalies — automatic warnings when cost/latency/errors drift from the 7-day baseline.
- Alerts — daily-cost or error-rate rules.
- Scores — per-trace evaluations (human feedback or LLM-as-judge).
- Budget — monthly cost cap per project, with tracking and projection.
Details
- The
api_keyis the project key, generated in the dashboard and set in.env. - Reusing the same batch id (via
observability_session(id=...)) appends observations to the same trace, idempotently on the backend. - The cost/token fields come from Cost and tokens.
What changed in 1.9.0
- Failures become traces too. When a call exhausts retries and fallbacks, the
library sends an observation with
status="ERROR"and the error (type + message, truncated) — previously only successes showed up in the dashboard. To report manually:jangada_ai.observability.auto_report_error(error, provider=..., model=...). - Real call start.
startedAtnow marks when the call started (not when it finished);auto_report(..., started_at=...)accepts an epoch or adatetime. - Streaming and transcription reported.
stream/astreamsend the accumulated text at the end (capabilitystreaming);transcribealso reports (audio). Cache hits get thecachecapability. - Bounded queue. Instead of one thread per call there is a queue (1000 events)
with a few workers; if the endpoint is slow and the queue fills up, extra events
are dropped —
dropped_count()tells how many.flush()and theatexithook wait for the queue to drain, with a deadline. - HTTPS only. The endpoint (and the
feedbackone) must behttps://(httponly on localhost), so the key and prompts don't travel in clear text. - Output truncated like the input, and
embedno longer sends every vector (only dimensions and counts above a threshold). - Send failures go to the
DEBUGlog of thejangada_ailogger (they never break the call).
Scope guardrails
ScopeGuard keeps the LLM within a domain and blocks unwanted utterances — blocklist + judge LLM classifier. Refusal via Completion, on input and/or output, sync and async. Plain Python, no new dependency.
Evaluation (evals)
Datasets, Evaluators and Experiments: measure quality and compare models by score × cost × latency. Heuristic (Evaluator.fn) + LLM judge (Evaluator.judge); runs offline and, with push=True, shows up in the dashboard (Experiments/Datasets).