Agents and teams
Jangada ships a lightweight multi-agent orchestration layer — Agent and
Squad — built on top of what already exists (tool calling, MCP, RAG). No new
dependency: it is plain Python composing the library itself.
Agent — an agent with a role and tools
An Agent is an LLM with a role/goal, optionally with tools (functions
it runs) and memory. It runs the tool-calling loop on its own until the
final answer.
from jangada_ai import LLM, Agent
def weather(city: str) -> str:
"Returns the weather for a city."
return f"sunny in {city}, 28°C"
meteo = Agent(
LLM("openai", "gpt-4o-mini"),
role="Meteorologist",
goal="report the weather clearly",
tools=[weather],
)
res = meteo.run("What's the weather like in Recife?")
print(res.text) # the model called weather("Recife") and replied
print(res.cost, res.usage, res.iterations)tools=are callables — the function runs locally when the model calls it, and the result goes back to the model. They can be sync (def) or async (async def):async deftools are awaited insidearun. Withrun(sync) use sync tools only.- For an MCP server, pass
mcp_client=MCPClient(...)and usearun(async): the agent lists the server's tools and uses them alongside its own.mcp_allowed_tools=[...]restricts which MCP tools are visible.mcp_tools_cache=[...]skipslist_tools()on everyaruncall — list once withawait mcp_tools(mcp_client)and pass it here;mcp_allowed_toolsstill filters on top of the already-built list. AgentResult.stopped_by_limit:Truewhen the loop stopped by hittingmax_iterationswithtool_callsstill pending — in that casetext/messagesare NOT the model's final answer (aUserWarningis also emitted).Falsewhen the model stopped requesting tools on its own.on_tool_call/on_tool_result: callbacks (sync inrun; sync or async inarun) that run on every tool call (function or MCP).on_tool_callreturningFalsevetoes the call.AgentResult.tool_tracecarries{"call", "result", "is_error"}for every call in the turn.
async with MCPClient("https://your-mcp/mcp/") as mcp:
agent = Agent(llm, role="Operator", mcp_client=mcp, mcp_allowed_tools=["list_products"],
on_tool_call=lambda c: print("calling", c.name))
res = await agent.arun("List the products")
print(res.text, res.tool_trace)Multi-turn conversation (history=)
run/arun accept a history of previous turns (list[Message]), inserted
before the new task, giving faithful dialogue continuity. AgentResult.messages
returns the turn's history — you persist it (e.g. per conversation_id) and
reinject on the next turn:
from jangada_ai.message import Message
history = [
Message("user", "How much did I spend in May?"),
Message("assistant", "R$ 3,200 in May."),
]
res = await agent.arun("And the month before?", history=history)Streaming the response (astream)
astream emits the final response token-by-token. Tool-calls are resolved
internally first (the stream protocol does not expose tool_calls); with no
tools, it streams directly.
async for token in agent.astream("Summarize my spending this month"):
print(token, end="", flush=True)Agent Card and A2A server
card() returns discoverable metadata in the Agent Card vocabulary of the
A2A protocol (name, description, version,
url, capabilities, skills). The jangada[a2a] extra goes further and
exposes the agent as an A2A server (JSON-RPC): discovery, message/send,
message/stream (SSE), tasks/get/tasks/cancel, with continuity by
contextId.
from jangada_ai.a2a import A2AHandler, build_a2a_app
app = build_a2a_app(A2AHandler(sofia, url="https://app.example/a2a"))
# ASGI Starlette — serve with: uvicorn module:app| Route | What |
|---|---|
GET /.well-known/agent-card.json (and the legacy /.well-known/agent.json) | Agent Card of the primary agent |
GET /agents | catalog (list of Agent Cards) |
POST / | message/send (JSON), message/stream (SSE), tasks/get, tasks/cancel |
POST /agents/{name} | same, for a catalog agent |
Multi-tenant: agent resolved per request
Instead of a fixed agent, pass a resolver that receives the request
context (headers/auth) and returns the right Agent — handy when the agent is
built per tenant (e.g. tools closing over the tenant_id from the JWT). With
tenant_key, history is isolated per tenant.
def resolver(ctx): # ctx = {"headers": {...}, "auth": "Bearer ..."}
tenant = (ctx.get("auth") or "").removeprefix("Bearer ")
return build_sofia(tenant_id=tenant)
handler = A2AHandler(resolver=resolver, name="Sofia",
tenant_key=lambda c: c.get("auth", ""))
app = build_a2a_app(handler)Everything is optional (the fixed mode A2AHandler(agent) stays the same):
resolver (excludes agent), name (required with resolver), description,
tenant_key (isolates history) and context_factory= on build_a2a_app (how to
build the context from the Request — swap it to validate/decode the JWT). The
lib does not decode JWTs.
Long-term memory (RAG)
RAGMemory gives the agent persistent memory over a RAG: before answering it
recalls what is relevant; afterwards, it remembers what happened.
from jangada_ai import LLM, Agent, RAGMemory
from jangada_ai.rag import RAG, InMemoryVectorStore
rag = RAG(LLM("openai", "text-embedding-3-small"), InMemoryVectorStore())
agent = Agent(llm, role="Support", memory=RAGMemory(rag, k=3))Squad — several agents collaborating
Squad orchestrates a team of agents. Two processes:
Sequential (handoff)
By default each agent runs in order and receives only the previous agent's output as context (not the accumulated transcript):
from jangada_ai import LLM, Agent, Squad
llm = LLM("openai", "gpt-4o-mini")
researcher = Agent(llm, role="Researcher", goal="gather facts")
writer = Agent(llm, role="Writer", goal="write clear copy")
squad = Squad([researcher, writer])
res = squad.run("Write a paragraph about northeastern Brazilian rafts.")
print(res.text) # the last agent's output
print(res.outputs) # {"Researcher": "...", "Writer": "..."}Context semantics (context=):
"last"(default): each agent gets only the immediately previous output. The per-hop input is ~constant — chain cost grows O(N), not O(N²)."full": each agent gets the accumulated transcript (every prior output, labeled by role). More context, O(N²) cost on long chains.
Squad([researcher, writer], context="full") # whole transcript at each hopPer-agent observability: res.steps has one entry per agent with
(role, usage, cost, cost_complete, dt) — so you can watch input/cost/latency grow
(or not) at each hop without unpacking the Squad by hand.
Hierarchical (delegation)
A manager agent receives delegate_to_<role> tools generated automatically
from the members and decides whom to delegate each subtask to:
manager = Agent(llm, role="Manager", goal="coordinate the team")
squad = Squad([researcher, writer], manager=manager)
res = squad.run("Produce a summary about topic X.")Both run and arun aggregate the whole team's usage/cost.
Planning
plan() breaks a goal down into an ordered list of tasks (structured output):
from jangada_ai import plan
for task in plan(llm, "Launch an AI newsletter", max_tasks=5):
print("-", task)What changed in 1.9.0
- Hierarchical Squad is truly async. In
Squad.arunthe delegation tool isasyncand callsawait member.arun()— the event loop is not blocked and the member keeps itsmcp_clientand async tools. The manager is a copy of the original agent, soon_tool_call(veto),on_tool_result, MCP and memory also apply in hierarchical mode. - Members' cost and trace.
SquadResult.usage/costadd up the manager and the delegated members (cost_completeisTrueonly if all have a price);outputsholds each member's output andstepsone entry per delegation. Repeated roles get unique keys (Reviewer,Reviewer#2). - Delegation tool names are transliterated (accents removed), de-duplicated
with a suffix (
_2,_3) and truncated to 64 characters. Agent.astream(prompt, chunk_size=24): without tools it streams from the provider; with tools it resolves the loop withacompleteand emits the final text in chunks (no second generation). When done,agent.last_stream_resultholds theAgentResult(usage, cost,tool_trace,stopped_by_limit).- Mixed native tools.
Agent(llm, tools=[web_search(), my_function])works: the native tool runs at the provider and shows up intool_tracewith"server": True(see Native tools). - MCP: allowlist enforced at execution. An MCP tool that was not offered to
the model (outside
mcp_allowed_tools) is not executed — it comes back as an errortool_result, even if the model makes the name up. BaseModeltool parameters receive the validated model instance (jangada_ai.coerce_args).plan()raises a clearValueErrorif the response has noparsed;RAGMemorylogs failures (loggerjangada_ai) instead of swallowing them and does not store turns that stopped at the iteration limit.
A2A (card and server)
Agent.card() follows the A2A ≥ 0.3 spec: it includes protocolVersion (default
"0.3.0") and preferredTransport="JSONRPC", accepts security_schemes=/
security=, and tool parameters become tags (the off-spec
skills[].parameters field is gone). A2AHandler serves the card at
/.well-known/agent-card.json (and keeps /.well-known/agent.json), bounds
memory with max_tasks/task_ttl and max_contexts/context_ttl (default 1000
items / 1 h), isolates tasks/get|cancel per tenant, returns -32002 when
cancelling a finished task, -32700/-32600/-32602 for invalid JSON/invalid
body/empty question, and does not leak the exception message in -32603 (the
detail goes to the log).
from jangada_ai import Agent, LLM, Squad, web_search
researcher = Agent(LLM("anthropic", "claude-sonnet-5"), role="Researcher",
tools=[web_search(max_uses=3)])
writer = Agent(LLM("openai", "gpt-5-mini"), role="Writer")
manager = Agent(LLM("openai", "gpt-5"), role="Manager")
res = await Squad([researcher, writer], manager=manager).arun("Summarize what's new in Python 3.14")
print(res.text, res.cost, res.outputs.keys(), len(res.steps))How it relates to the rest
There is no magic and no new infra: Agent is the tool-calling loop (like
run_agent); memory is the RAG; the hierarchical
Squad uses tool-based delegation (Tools). You can swap any
agent's provider without changing anything else — jangada's thesis holds here
too.
MCP (Model Context Protocol)
Connect MCP servers to your calls via mcp_servers=[...]. All providers support MCP, but in different ways — and jangada uses each SDK's native approach.
Gemini Interactions and Deep Research
GeminiInteractions: Gemini's Interactions API (stateful with previous_interaction_id), streaming, a function-calling loop and managed agents such as Deep Research in the background with polling. Preview; requires google-genai>=2.3.