Gemini
Provider gemini. Adapter over the google-genai SDK. It has a single Client; the
async one lives in client.aio.
pip install "jangada-ai[gemini]"provider=:"gemini"- Environment variable:
GEMINI_API_KEY(orGOOGLE_API_KEY)
from jangada_ai import LLM
llm = LLM("gemini", "gemini-2.5-flash")What it does
- Text and streaming (
generate_content/generate_content_stream). - Structured output (
parse):config.response_schema=Model+response_mime_type="application/json"→resp.parsed. - Vision (
images=): images becometypes.Part.from_bytes. - Documents (
files=): local text extraction. - Object detection (
detect_objects): the most accurate — the 0–1000 bounding box format is native to Gemini's training. - Audio transcription (
transcribe): multimodal — audio enters asPart.from_bytesalongside an instruction; it is not a dedicated endpoint.
Structure and quirks
- Messages: the
systemrole becomessystem_instructionin theGenerateContentConfig(not a regular message);assistantbecomesmodel. - Canonical parameters → config:
max_tokens→max_output_tokens,stop→stop_sequences;temperature/top_p/top_k/seedpass through. It is the only provider withtop_k. - Model profile (
profiles.py):gemini-3.xdrops sampling (temperature/top_p/top_k). See Parameters and profiles. - Multi-turn function calling in 3.x — thought signatures (handled by the
library): Gemini 3.x attaches an opaque
thought_signatureto every function call and requires it back, unchanged, when you resend the history. jangada preserves this automatically — the value travels in the opaquemetadata: dictfield ofToolCall/ToolCallPart, and the adapter re-attaches it when rebuilding the history. Sotools=,AgentandMCPClientwork across turns without you touching anything.gemini-2.5does not use this field. (fixed in 1.3.1) - Response (
Completion):usagecomes fromusage_metadata(prompt_token_count/candidates_token_count).
Thinking — you pass only thinking_budget
Gemini has two incompatible conventions: 2.5 only accepts thinking_budget
(in tokens) and 3.x only accepts thinking_level (LOW/HIGH) — mixing them
returns HTTP 400. jangada handles this for you: pass thinking_budget (or
thinking_level) and the library adapts to the model, wrapping it in the native
thinking_config under the hood.
# Works the same on both versions — you change nothing:
LLM("gemini", "gemini-2.5-flash").complete("...", params={"thinking_budget": 1024})
LLM("gemini", "gemini-3-pro").complete("...", params={"thinking_budget": 1024})
# on 2.5 it becomes thinking_config(thinking_budget=1024);
# on 3.x the profile converts it to thinking_config(thinking_level="HIGH").thinking_budget:0disables (when the model allows),-1is automatic.- On 3.x the budget is approximated to a valid level (
LOWif ≤0, elseHIGH). On 2.5 athinking_levelis converted to a budget. You can also pass a readythinking_config, which is respected as-is. - Accepted levels:
thinking_levelrecognizesMINIMAL/LOW/MEDIUM/HIGH. As input on 2.5 all four are converted to athinking_budget(0/1024/8192/24576). Directly on 3.x, Gemini only acceptsLOW/HIGHuniversally (MINIMALis Flash/Lite only,MEDIUMonly on 3.0 Flash; using them as a level on other 3.x returns HTTP 400) — that's why the budget→level conversion uses onlyLOW/HIGH.
Why it's the "Swiss Army knife" here
It is the only one that covers vision, detection, and audio without separate
endpoints — all via generateContent. See Capability matrix.
What changed in 1.9.0
ToolCall.idis theidreturned by the API (or"name#i"when absent) — no longer the function name. Two parallel calls to the same function no longer collide. The real name is still used in thefunction_response, and old histories (id equal to the name) keep working.- A tool parameter called
titleno longer disappears from the schema (the schema cleaner was also removing thetitlekey insideproperties). - A safety-blocked prompt raises
BadRequestErrorwith the reason (prompt_feedback.block_reason), without retry or fallback. - Usage:
output_tokensincludes thinking tokens (reasoning_tokensshows how many) andinput_tokensincludes native-tool result tokens — both are billed; the cost used to be underestimated. - Native tools: Google Search, URL context, code execution, file search, Google Maps and computer use (see Native tools). For the Interactions API and agents (Deep Research), see Interactions.
transcribehonourslanguage=;parsewith an empty response raises instead of returningparsed=None; a system message given as parts has its text extracted.
Groq
Provider groq. Adapter over the groq SDK, which speaks the same chat.completions dialect as OpenAI — that's why it inherits from OpenAICompatible.
Mistral
Provider mistral. Mistral models via the official mistralai SDK through jangada's normalized API: text, vision, streaming, structured output, tools, embeddings, OCR/Document AI and Voxtral transcription.