MCP (Model Context Protocol)
Connect MCP servers to your calls via mcp_servers=[...]. All providers
support MCP, but in different ways — and jangada uses each SDK's native
approach.
Two models (important)
- Remote (URL) —
MCPServer(url=...): the provider connects to the MCP server and runs the tools (server-side). You run nothing. → Anthropic, OpenAI, Groq. - Client-side (session) — you pass a
ClientSessionfrom themcppackage: the SDK calls the tools locally (automatic function calling). → Gemini.
| Provider | Model | How | Note |
|---|---|---|---|
| Anthropic | remote (URL) | Messages API (mcp_servers + MCPToolset, beta header) | beta (mcp-client-2025-11-20) |
| OpenAI | remote (URL) | Responses API (tools=[{type:"mcp"}]) | uses the Responses API, not chat.completions |
| Groq | remote (URL) | Responses API (OpenAI-compatible) | beta — under the hood jangada uses the OpenAI client on Groq's base_url (requires the openai package installed) |
| Gemini | client-side (session) | tools=[session] (automatic function calling) | async only (acomplete) — the session is asynchronous |
Remote (Anthropic / OpenAI / Groq)
from jangada_ai import LLM, MCPServer
llm = LLM("anthropic", "claude-opus-4-8") # or ("openai", "gpt-4o"), ("groq", ...)
comp = llm.complete(
"List the repository's open issues.",
mcp_servers=[MCPServer(
url="https://mcp.example.com/sse",
name="github",
authorization_token="TOKEN", # optional (OAuth/Bearer)
allowed_tools=["list_issues"], # optional (restricts the tools)
)],
)
print(comp.text) # the provider already executed the MCP toolsClient-side (Gemini, async)
In Gemini, MCPServer(url=...) also works — but only in acomplete: jangada
opens its own MCPClient under the hood and hands it to the SDK as a
client-side session (it's the only way Gemini supports it). Requires the
[mcp] extra.
from jangada_ai import LLM, MCPServer
llm = LLM("gemini", "gemini-2.5-flash")
comp = await llm.acomplete(
"List the open issues.",
mcp_servers=[MCPServer(url="https://mcp.example.com/mcp/", name="github",
authorization_token="TOKEN")],
)In
complete()(sync) this still raisesUnsupportedError— Gemini has no way to open an async session outside ofacomplete.MCPServer'sallowed_tools/require_approvalalso raiseUnsupportedErroron this path (instead of being silently ignored) — Gemini's SDK lists and runs the tools on its own, with no filter/approval hook; to restrict tools on Gemini, use jangada's ownMCPClient/run_agent(below), which hasallowed_tools=.
The MCP session is asynchronous (built manually), so use acomplete:
from mcp import ClientSession
from mcp.client.stdio import stdio_client, StdioServerParameters
from jangada_ai import LLM
llm = LLM("gemini", "gemini-2.5-flash")
params = StdioServerParameters(command="npx", args=["-y", "@example/mcp"])
async with stdio_client(params) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
comp = await llm.acomplete("Use tool X", mcp_servers=[session])
print(comp.text)In Gemini,
complete()(sync) with a session raisesUnsupportedErrorasking foracomplete(). And passing a session in Anthropic/OpenAI/Groq raises — they are remote by URL.
Built-in MCP client + agent (portable, any provider)
jangada ships its own MCP client (MCPClient) and an agent loop
(run_agent) that connects to the server, lists the tools, and runs the cycle
(model requests → execute → resend) on its own — on any provider (it uses
tools=, supported across all 4), independent of each SDK's native MCP.
pip install "jangada-ai[mcp]" # MCP client (official `mcp` package)from jangada_ai import LLM
from jangada_ai.mcp import MCPClient, run_agent
llm = LLM("openai", "gpt-4o-mini") # or anthropic/groq/gemini
async with MCPClient("https://my-mcp/mcp/") as mcp: # or command=/args= (stdio)
ans = await run_agent(llm, "Roll some dice", client=mcp)
print(ans.text)Under the hood: await mcp.list_tools() becomes tools=[...], and each tool_call from
the model is executed with await mcp.call_tool(...) and resent via
Message.tool_results(...) — the same tool calling as always, on
autopilot. Want full control? Use MCPClient + tools= by hand.
isErrorfrom the tool result becomesis_error=Truein thetool_result— the model knows the tool failed (unlike a network/protocol exception, which already became an error before).list_tools/list_resources/list_promptsfollownextCursoron their own until pages run out — servers with many tools don't come back incomplete. It stops if a buggy server repeats the same cursor, and has a defensive limit of 10,000 pages for cursors that never repeat but also never run out — hitting that limit without exhausting the cursor emits aUserWarning(the list may be incomplete; it never truncates silently).- Iteration limit: if
run_agenthitsmax_iterationswithtool_callsstill pending, it emits aUserWarning(the returnedCompletionis not necessarily the final answer). InAgent.run/arun(below), the same case also setsAgentResult.stopped_by_limit = True. - Connection error (
MCPClient.__aenter__, e.g. invalid token, server down) becomes the library's ownAPIConnectionErrorwith the real cause (e.g.HTTP 403/HTTP 400), instead of the raw transport error (ExceptionGroup/CancelledErrorfromanyio) — even when the real cause only surfaces when closing the connection, not on opening it.
Transport, authentication and timeout
async with MCPClient(
"https://my-mcp/sse", # URL ending in /sse -> auto-detects SSE
# transport="sse", # or force it explicitly ("sse" | "streamable-http")
auth_token="TOKEN", # becomes an Authorization: Bearer TOKEN header
timeout=30, # seconds, forwarded to the HTTP transport
) as mcp:
...Long-lived connection (connect/aclose/reconnect, keep_alive, ping)
Outside of async with — useful for opening at app startup and closing on
shutdown:
mcp = MCPClient("https://my-mcp/mcp/", auth_token="TOKEN")
await mcp.connect() # idempotent: calling again while connected is a no-op
...
await mcp.ping() # health check (session/ping)
...
await mcp.reconnect() # close and reopen (e.g. noticed the connection dropped)
...
await mcp.aclose() # on shutdownWith keep_alive=True, MCPClient connects on its own on the first call
(no need for connect()/async with) and reconnects once, on its own, if
a call fails with something that looks like a dropped connection — it does not
retry on a tool's LOGIC error (e.g. an invalid argument):
mcp = MCPClient("https://my-mcp/mcp/", auth_token="TOKEN", keep_alive=True)
tools = await mcp.list_tools() # connects on its own hereSafe for concurrent use: if several calls notice the connection dropped at the same time, only the first one actually reconnects — the others wait and reuse the fresh session, instead of racing each other into separate reconnects. If the reconnect itself fails (server down), the calls that were waiting reuse that SAME error for a short cooldown, instead of each one waiting out its own connection timeout from scratch.
Full MCP primitives (in MCPClient)
Beyond tools, MCPClient covers the rest of the protocol:
async with MCPClient("https://my-mcp/mcp/") as mcp:
# Resources — data/context the server exposes
resources = await mcp.list_resources()
text = await mcp.resource_text("file:///guide.md")
# Prompts — reusable server templates -> become list[Message]
msgs = await mcp.prompt_messages("review", {"text": "..."})
resp = await llm.acomplete(None, history=msgs)And the client features (the server calls the client back), configured in the constructor:
from jangada_ai import LLM
async with MCPClient(
command="python", args=["server.py"],
roots=["./workspace"], # filesystem scope (file://)
sampling_llm=LLM("openai", "gpt-4o-mini"), # the server requests generation from YOUR LLM
elicitation_callback=my_handler, # the server requests input from the user
logging_callback=my_logger, # server logs
) as mcp:
...- Roots: the server asks which directories it may use; the client responds with the list.
- Sampling: the server requests an LLM generation (
sampling/createMessage) and jangada runs it with yourLLM— the server stays model-independent and you control cost/permissions. - Elicitation / Logging: you pass a callback (
async) that the SDK calls.
| Primitive | Methods / config |
|---|---|
| Tools | list_tools / call_tool |
| Resources | list_resources / read_resource / resource_text |
| Prompts | list_prompts / get_prompt / prompt_messages |
| Roots | roots=[...] |
| Sampling | sampling_llm=LLM(...) |
| Elicitation | elicitation_callback=... |
| Logging | logging_callback=... / set_logging_level(...) |
run_agent: history, callbacks, allowlist and enriched return
from jangada_ai.message import Message
def veto_delete(call):
return call.name != "delete_file" # False = vetoes the call
async with MCPClient("https://my-mcp/mcp/") as mcp:
ans = await run_agent(
llm, "List and then delete the temp files", client=mcp,
history=[Message("user", "hi"), Message("assistant", "hello!")], # previous turns
allowed_tools=["list_files", "delete_file"], # restricts visible tools
on_tool_call=veto_delete, # (sync or async) False = vetoes the call
on_tool_result=lambda c, r: print(c.name, r.is_error),
)
print(ans.text, ans.iterations, ans.stopped_by_limit)
print(ans.tool_trace) # [{"call", "result", "is_error"}, ...] for ALL rounds
print(ans.usage_total, ans.cost_total, ans.cost_complete) # aggregated over the whole loophistory=injects previous turns before the task;prompt=Nonecontinues from the history alone (no empty user message added).allowed_tools=filters by name before the model sees the tools (also available directly asmcp_tools(client, allowed_tools=[...])).tools=skipslist_tools()(and its round-trip) when you already listed them before — list once and reuse across calls;allowed_tools=still applies on top oftools=(filters the already-built list).on_tool_call(call)/on_tool_result(call, result)(sync or async) run on every tool call;on_tool_callreturningFalsevetoes the call.- The returned
Completiongains extra attributes (usage/costremain only the LAST call to the LLM, as always):
| Extra attribute | What |
|---|---|
tool_trace | list of {"call", "result", "is_error"} for ALL tool calls in the loop |
iterations | how many rounds the loop took |
stopped_by_limit | True if it stopped by hitting max_iterations |
usage_total/cost_total/cost_complete | aggregated over ALL calls to the LLM in the loop |
In Agent/Squad, the same shows up as
mcp_allowed_tools=, on_tool_call=/on_tool_result= in the constructor, and
AgentResult.tool_trace.
Be an MCP server (expose your tools/Agent)
Jangada can also be an MCP server — expose your functions/Agent to clients
(Claude Desktop, Cursor, another agent). Built on the low-level Server (the
protocol, not FastMCP); the schema comes from tools.py.
from jangada_ai import serve_mcp
def add(a: float, b: float) -> float:
"Add two numbers."
return a + b
serve_mcp("my-calc", tools=[add]) # stdio (Claude Desktop/Cursor)Expose an Agent (becomes the ask tool):
serve_mcp("support", agent=Agent(LLM("openai", "gpt-4o-mini"), role="support"))Over HTTP (streamable-http) or mounted in your ASGI app:
serve_mcp("my-calc", tools=[add], transport="streamable-http", port=8000)
# or: app = build_mcp_app("my-calc", tools=[add], path="/mcp") # StarletteExtra jangada-ai[mcp] (HTTP also needs starlette + uvicorn).
serve_mcp/build_mcp_app = server; MCPClient = client.
Ready-made example: jangada-docs-mcp
serves all of jangada's docs to your editor (Claude Code/Desktop/Cursor) — no
clone: uvx --from git+https://github.com/nerigleston/jangada-docs-mcp jangada-docs-mcp.
What changed in 1.9.0
- Compatible with
mcp1.x and 2.x. The[mcp]extra now acceptsmcp>=1.0,<3; client, server, SSE, streamable-HTTP and stdio were validated end to end on both versions. - Retries only where it is safe. With
keep_alive=True, the automatic reconnection only retries side-effect-free operations (list_*,read_resource,get_prompt,ping).call_toolis retried only when the request provably did not go out (connection failure); on a read timeout it reconnects for the next calls but does not retry the current one (avoids a duplicated order). To retry anyway:MCPClient(..., retry_tools=True). 401/403 (and other 4xx except 408/429) do not trigger reconnection. timeout=applies to the whole session, including stdio — a stuck tool on the server no longer hangs the client forever.- Allowlist enforced at execution.
run_agent(..., allowed_tools=[...])(andAgentwithmcp_allowed_tools) only executes tools actually offered to the model; any other name comes back as an errortool_result, not executed. - SSE detected by the URL path (
/sse?token=...works). - Sampling with approval.
MCPClient(..., sampling_llm=llm, on_sampling_request=fn):fn(request)(sync or async) returningFalserefuses the server's request.stopReasonreflects the real reason (maxTokens,endTurn,toolUse) andprompt_messageskeeps images. - Secrets out of
repr. Token and headers no longer show up inrepr(MCPClient)orrepr(MCPServer).
Server: authentication and DNS-rebinding protection
from jangada_ai import serve_mcp
serve_mcp(
"my-tools", tools=[get_order],
transport="streamable-http", host="0.0.0.0", port=8000,
auth_token="secret", # requires Authorization: Bearer secret (401 otherwise)
allowed_hosts=["mcp.mycompany.com"], # DNS-rebinding protection
allowed_origins=["https://app.mycompany.com"],
)build_mcp_app(...) takes the same auth_token/allowed_hosts/
allowed_origins/security_settings. Synchronous tools (and a synchronous
agent) run in a thread, without blocking other sessions. If the exposed agent=
has memory, the library warns you: that memory is shared by every client.
Example
examples/mcp_example.py— MCP client.examples/mcp_server_example.py— MCP server (stdio).
RAG (embeddings + vector/hybrid search)
jangada covers the "LLM parts" of RAG (embeddings + building the context) and ships an optional jangada_ai.rag module with chunking, vector store (pgvector/Mongo), and hybrid search.
Agents and teams
Jangada ships a lightweight multi-agent orchestration layer — Agent and Squad — built on top of what already exists (tool calling, MCP, RAG). No new dependency: it is plain Python composing the library itself.