Gemini Interactions and Deep Research
GeminiInteractions is a thin, Gemini-specific layer over the google-genai
SDK's Interactions API.
It differs from LLM:
- Stateful on the server: each call returns an
id, and the next one continues the conversation withprevious_interaction_id, without resending the history. - It's the only way to reach Google's managed agents, such as Deep Research, which plans, searches the web and writes a report while running in the background for several minutes.
Preview. The Interactions API and the agents are in preview at Google: agent names and fields change often. Requires
google-genai>=2.3. With an older SDK, jangada raisesUnsupportedErrorasking you to upgrade.
pip install "jangada-ai[gemini]" "google-genai>=2.3"Set GEMINI_API_KEY in the environment (or pass api_key=).
from jangada_ai import GeminiInteractions
gi = GeminiInteractions(model="gemini-3.8-flash")
r = gi.create("Who won the 2002 World Cup?", tools=[{"type": "google_search"}])
print(r.text)
for c in r.citations:
print("-", c["title"], c["url"])
# continues the conversation on the server, without resending history
r2 = gi.create("And the top scorer?", previous_interaction_id=r.id)When to use GeminiInteractions and when to use LLM
LLM("gemini", ...): the default path. Switch providers without changing code, retry, fallback, cache, guardrails. Gemini's native tools also work there (see Native tools).GeminiInteractions: when you want server-side state (previous_interaction_id), managed agents (Deep Research) or long background tasks with polling. It has no retry or fallback: it's a thin layer.
API
GeminiInteractions(api_key=None, *, model=None, vertexai=False, **client_kwargs)| Method (sync / async) | What it does |
|---|---|
create / acreate(input, *, model, agent, agent_config, tools, system, previous_interaction_id, store, background, response_format, tool_choice, max_tokens, seed, stop, thinking_level, **extra) | Creates an interaction (or starts an agent) |
stream / astream(input, ...) | Same as create, yielding InteractionEvent |
resume_stream / aresume_stream(id, last_event_id=) | Resumes an interrupted stream |
get / aget(id) | Fetches the current state |
cancel / acancel(id) | Cancels |
delete / adelete(id) | Deletes |
wait / await_(target, *, poll_interval=10, timeout=None, on_update=None) | Polls until it leaves queued/in_progress |
run / arun(input, *, tools=[...], max_iterations=10, on_tool_call=, on_tool_result=) | Function-calling loop with local execution |
deep_research / adeep_research(prompt, *, agent=DEEP_RESEARCH_AGENT, wait=True, ...) | Starts Deep Research in the background |
close()/aclose() close the client.
InteractionResult
id,status(queued,in_progress,requires_action,completed,failed,cancelled,incomplete,budget_exceeded),done,requires_action.text: final text.steps: normalized steps (dicts:model_output,thought,function_call,google_search_call...).citations: list of dicts{url, title, start, end}.function_calls: pending calls (InteractionFunctionCall(id, name, args), with.result(output, is_error=False)to build the reply).usage,cost,parsed(withresponse_format),errors,raw.- Filled in by
run:tool_trace,iterations,stopped_by_limit,usage_total,cost_total,cost_complete. raise_for_status(): raisesProviderErrorif it ended asfailed,cancelledorbudget_exceeded.
Tools
tools= accepts:
- your functions (callables, Pydantic models,
Tool), which become{"type": "function", ...}; - the API's native dicts:
{"type": "google_search"},url_context,code_execution,file_search,google_maps,mcp_server...; - jangada's canonical
NativeTools (web_search(),url_context(),code_execution(),file_search(...),google_maps(...)) andnative_tool("gemini", spec).image_generationand tools of another provider raiseUnsupportedError.
Function calling with local execution (run)
run creates the interaction, runs locally the functions the model asks for and
replies with previous_interaction_id, until the final answer or
max_iterations. A tool that fails, doesn't exist or is vetoed by
on_tool_call (returning False) becomes a result with is_error, and the
model can react.
def exchange_rate(currency: str) -> float:
"""Exchange rate of the currency in reais."""
return {"USD": 5.4, "EUR": 5.9}.get(currency.upper(), 0.0)
r = gi.run("How much are 100 dollars in reais?", tools=[exchange_rate])
print(r.text, [t["name"] for t in r.tool_trace], r.cost_total)To drive the loop by hand, use create and answer r.function_calls:
r = gi.create("How much are 100 dollars?", tools=[exchange_rate])
if r.requires_action:
replies = [c.result(exchange_rate(**c.args)) for c in r.function_calls]
r = gi.create(replies, previous_interaction_id=r.id, tools=[exchange_rate])Structured output
from pydantic import BaseModel
class Summary(BaseModel):
title: str
points: list[str]
r = gi.create("Summarize the history of frevo.", response_format=Summary)
print(r.parsed) # Summary(...); if it doesn't validate, OutputValidationErrorStreaming
stream yields InteractionEvent(type, text, status, step, usage, result, ...),
with type in created, status, step_start, text, thought, delta,
step_stop, completed (with the final .result) and error (raises
ProviderError).
for ev in gi.stream("Explain RAG in one sentence."):
if ev.type == "text":
print(ev.text, end="", flush=True)Deep Research (background)
report = gi.deep_research(
"Overview of the open-source LLM market in 2026, in 5 topics.",
on_update=lambda x: print("status:", x.status),
timeout=1800,
)
print(report.text)- Runs with
background=Trueandstore=True(the API requiresstorewith background) and polls every 10 s. wait=Falsereturns theInteractionResultright away inqueued/in_progress; then usegi.wait(r)orgi.get(r.id).timeoutexceeded →APITimeoutError. The interaction keeps running on the server; you can pick it up again withwait/get.- The default agent is
DEEP_RESEARCH_AGENT("deep-research-preview-04-2026"). Other agents go inagent=(e.g."deep-research-max-preview-04-2026","antigravity-preview-05-2026"), withagent_config=passed through as is. - It takes minutes and consumes a lot of tokens (100 thousand to millions per task): use it with cost in mind.
Usage and cost
Usage follows the library's contract: input_tokens (total input + tool-use
tokens), output_tokens (output + thoughts), reasoning_tokens,
cache_read_tokens, server_tool_requests and total_tokens. Cost is computed
when there's a model=. With agent= it stays None, because there's no price
table per agent.
Limitations
temperature,top_pandtop_kdon't exist in the Interactions API: they're ignored (with a debug-level log).max_tokens,seed,stopandthinking_levelwork.- No retry, fallback, cache or guardrails (use
LLMfor that). streamdoesn't run the tool loop (run).- Vertex AI (
vertexai=True): the SDK routes it, but Google's docs don't confirm availability yet. - Retention at Google: 55 days (paid) or 1 day (free).
Related: Gemini, Native tools, Agents and teams.
Agents and teams
Jangada ships a lightweight multi-agent orchestration layer — Agent and Squad — built on top of what already exists (tool calling, MCP, RAG). No new dependency: it is plain Python composing the library itself.
Scope guardrails
ScopeGuard keeps the LLM within a domain and blocks unwanted utterances — blocklist + judge LLM classifier. Refusal via Completion, on input and/or output, sync and async. Plain Python, no new dependency.