Jangada AIJangada AI

Agents and teams

Jangada ships a lightweight multi-agent orchestration layer — Agent and Squad — built on top of what already exists (tool calling, MCP, RAG). No new dependency: it is plain Python composing the library itself.

Agent — an agent with a role and tools

An Agent is an LLM with a role/goal, optionally with tools (functions it runs) and memory. It runs the tool-calling loop on its own until the final answer.

from jangada_ai import LLM, Agent

def weather(city: str) -> str:
    "Returns the weather for a city."
    return f"sunny in {city}, 28°C"

meteo = Agent(
    LLM("openai", "gpt-4o-mini"),
    role="Meteorologist",
    goal="report the weather clearly",
    tools=[weather],
)

res = meteo.run("What's the weather like in Recife?")
print(res.text)          # the model called weather("Recife") and replied
print(res.cost, res.usage, res.iterations)
  • tools= are callables — the function runs locally when the model calls it, and the result goes back to the model. They can be sync (def) or async (async def): async def tools are awaited inside arun. With run (sync) use sync tools only.
  • For an MCP server, pass mcp_client=MCPClient(...) and use arun (async): the agent lists the server's tools and uses them alongside its own. mcp_allowed_tools=[...] restricts which MCP tools are visible. mcp_tools_cache=[...] skips list_tools() on every arun call — list once with await mcp_tools(mcp_client) and pass it here; mcp_allowed_tools still filters on top of the already-built list.
  • AgentResult.stopped_by_limit: True when the loop stopped by hitting max_iterations with tool_calls still pending — in that case text/ messages are NOT the model's final answer (a UserWarning is also emitted). False when the model stopped requesting tools on its own.
  • on_tool_call/on_tool_result: callbacks (sync in run; sync or async in arun) that run on every tool call (function or MCP). on_tool_call returning False vetoes the call. AgentResult.tool_trace carries {"call", "result", "is_error"} for every call in the turn.
async with MCPClient("https://your-mcp/mcp/") as mcp:
    agent = Agent(llm, role="Operator", mcp_client=mcp, mcp_allowed_tools=["list_products"],
                  on_tool_call=lambda c: print("calling", c.name))
    res = await agent.arun("List the products")
    print(res.text, res.tool_trace)

Multi-turn conversation (history=)

run/arun accept a history of previous turns (list[Message]), inserted before the new task, giving faithful dialogue continuity. AgentResult.messages returns the turn's history — you persist it (e.g. per conversation_id) and reinject on the next turn:

from jangada_ai.message import Message

history = [
    Message("user", "How much did I spend in May?"),
    Message("assistant", "R$ 3,200 in May."),
]
res = await agent.arun("And the month before?", history=history)

Streaming the response (astream)

astream emits the final response token-by-token. Tool-calls are resolved internally first (the stream protocol does not expose tool_calls); with no tools, it streams directly.

async for token in agent.astream("Summarize my spending this month"):
    print(token, end="", flush=True)

Agent Card and A2A server

card() returns discoverable metadata in the Agent Card vocabulary of the A2A protocol (name, description, version, url, capabilities, skills). The jangada[a2a] extra goes further and exposes the agent as an A2A server (JSON-RPC): discovery, message/send, message/stream (SSE), tasks/get/tasks/cancel, with continuity by contextId.

from jangada_ai.a2a import A2AHandler, build_a2a_app

app = build_a2a_app(A2AHandler(sofia, url="https://app.example/a2a"))
# ASGI Starlette — serve with:  uvicorn module:app
RouteWhat
GET /.well-known/agent-card.json (and the legacy /.well-known/agent.json)Agent Card of the primary agent
GET /agentscatalog (list of Agent Cards)
POST /message/send (JSON), message/stream (SSE), tasks/get, tasks/cancel
POST /agents/{name}same, for a catalog agent

Multi-tenant: agent resolved per request

Instead of a fixed agent, pass a resolver that receives the request context (headers/auth) and returns the right Agent — handy when the agent is built per tenant (e.g. tools closing over the tenant_id from the JWT). With tenant_key, history is isolated per tenant.

def resolver(ctx):                       # ctx = {"headers": {...}, "auth": "Bearer ..."}
    tenant = (ctx.get("auth") or "").removeprefix("Bearer ")
    return build_sofia(tenant_id=tenant)

handler = A2AHandler(resolver=resolver, name="Sofia",
                     tenant_key=lambda c: c.get("auth", ""))
app = build_a2a_app(handler)

Everything is optional (the fixed mode A2AHandler(agent) stays the same): resolver (excludes agent), name (required with resolver), description, tenant_key (isolates history) and context_factory= on build_a2a_app (how to build the context from the Request — swap it to validate/decode the JWT). The lib does not decode JWTs.

Long-term memory (RAG)

RAGMemory gives the agent persistent memory over a RAG: before answering it recalls what is relevant; afterwards, it remembers what happened.

from jangada_ai import LLM, Agent, RAGMemory
from jangada_ai.rag import RAG, InMemoryVectorStore

rag = RAG(LLM("openai", "text-embedding-3-small"), InMemoryVectorStore())
agent = Agent(llm, role="Support", memory=RAGMemory(rag, k=3))

Squad — several agents collaborating

Squad orchestrates a team of agents. Two processes:

Sequential (handoff)

By default each agent runs in order and receives only the previous agent's output as context (not the accumulated transcript):

from jangada_ai import LLM, Agent, Squad

llm = LLM("openai", "gpt-4o-mini")
researcher = Agent(llm, role="Researcher", goal="gather facts")
writer     = Agent(llm, role="Writer", goal="write clear copy")

squad = Squad([researcher, writer])
res = squad.run("Write a paragraph about northeastern Brazilian rafts.")
print(res.text)            # the last agent's output
print(res.outputs)         # {"Researcher": "...", "Writer": "..."}

Context semantics (context=):

  • "last" (default): each agent gets only the immediately previous output. The per-hop input is ~constant — chain cost grows O(N), not O(N²).
  • "full": each agent gets the accumulated transcript (every prior output, labeled by role). More context, O(N²) cost on long chains.
Squad([researcher, writer], context="full")   # whole transcript at each hop

Per-agent observability: res.steps has one entry per agent with (role, usage, cost, cost_complete, dt) — so you can watch input/cost/latency grow (or not) at each hop without unpacking the Squad by hand.

Hierarchical (delegation)

A manager agent receives delegate_to_<role> tools generated automatically from the members and decides whom to delegate each subtask to:

manager = Agent(llm, role="Manager", goal="coordinate the team")
squad = Squad([researcher, writer], manager=manager)
res = squad.run("Produce a summary about topic X.")

Both run and arun aggregate the whole team's usage/cost.

Planning

plan() breaks a goal down into an ordered list of tasks (structured output):

from jangada_ai import plan

for task in plan(llm, "Launch an AI newsletter", max_tasks=5):
    print("-", task)

What changed in 1.9.0

  • Hierarchical Squad is truly async. In Squad.arun the delegation tool is async and calls await member.arun() — the event loop is not blocked and the member keeps its mcp_client and async tools. The manager is a copy of the original agent, so on_tool_call (veto), on_tool_result, MCP and memory also apply in hierarchical mode.
  • Members' cost and trace. SquadResult.usage/cost add up the manager and the delegated members (cost_complete is True only if all have a price); outputs holds each member's output and steps one entry per delegation. Repeated roles get unique keys (Reviewer, Reviewer#2).
  • Delegation tool names are transliterated (accents removed), de-duplicated with a suffix (_2, _3) and truncated to 64 characters.
  • Agent.astream(prompt, chunk_size=24): without tools it streams from the provider; with tools it resolves the loop with acomplete and emits the final text in chunks (no second generation). When done, agent.last_stream_result holds the AgentResult (usage, cost, tool_trace, stopped_by_limit).
  • Mixed native tools. Agent(llm, tools=[web_search(), my_function]) works: the native tool runs at the provider and shows up in tool_trace with "server": True (see Native tools).
  • MCP: allowlist enforced at execution. An MCP tool that was not offered to the model (outside mcp_allowed_tools) is not executed — it comes back as an error tool_result, even if the model makes the name up.
  • BaseModel tool parameters receive the validated model instance (jangada_ai.coerce_args).
  • plan() raises a clear ValueError if the response has no parsed; RAGMemory logs failures (logger jangada_ai) instead of swallowing them and does not store turns that stopped at the iteration limit.

A2A (card and server)

Agent.card() follows the A2A ≥ 0.3 spec: it includes protocolVersion (default "0.3.0") and preferredTransport="JSONRPC", accepts security_schemes=/ security=, and tool parameters become tags (the off-spec skills[].parameters field is gone). A2AHandler serves the card at /.well-known/agent-card.json (and keeps /.well-known/agent.json), bounds memory with max_tasks/task_ttl and max_contexts/context_ttl (default 1000 items / 1 h), isolates tasks/get|cancel per tenant, returns -32002 when cancelling a finished task, -32700/-32600/-32602 for invalid JSON/invalid body/empty question, and does not leak the exception message in -32603 (the detail goes to the log).

from jangada_ai import Agent, LLM, Squad, web_search

researcher = Agent(LLM("anthropic", "claude-sonnet-5"), role="Researcher",
                   tools=[web_search(max_uses=3)])
writer = Agent(LLM("openai", "gpt-5-mini"), role="Writer")
manager = Agent(LLM("openai", "gpt-5"), role="Manager")

res = await Squad([researcher, writer], manager=manager).arun("Summarize what's new in Python 3.14")
print(res.text, res.cost, res.outputs.keys(), len(res.steps))

How it relates to the rest

There is no magic and no new infra: Agent is the tool-calling loop (like run_agent); memory is the RAG; the hierarchical Squad uses tool-based delegation (Tools). You can swap any agent's provider without changing anything else — jangada's thesis holds here too.

On this page