Generation parameters and per-model profiles
jangada accepts canonical parameter names and each adapter translates them to the SDK's native name, discarding the unsupported ones.
Canonical parameters
| Canonical | OpenAI/Groq | Anthropic | Gemini |
|---|---|---|---|
temperature | temperature | temperature | temperature |
max_tokens | max_tokens | max_tokens | max_output_tokens |
top_p | top_p | top_p | top_p |
top_k | (discarded) | top_k | top_k |
stop | stop | stop_sequences | stop_sequences |
seed | seed | (discarded) | seed |
llm = LLM("anthropic", "claude-opus-4-8", temperature=0.2) # max_tokens=8192
# per-call override
llm.complete("...", params={"temperature": 0.9, "max_tokens": 16384})
# clone with new defaults
creative = llm.with_params(temperature=1.0)Since v1.4.5, max_tokens has a high default of 8192 across all
providers. For models with a different limit or even larger extractions, set
it explicitly:
llm = LLM("openai", "gpt-5", max_tokens=16384)
# or just for this call:
llm.parse("Extract every item.", Items, files=[...], params={"max_tokens": 16384})SDK-specific parameters that have no canonical name go via extra=.
Automatic profiles (per-model quirks)
Models from the same provider sometimes have different contracts. jangada normalizes
this in profiles.py, per model, without you needing to know:
gpt-5and later (gpt-5.x,gpt-6-*…, except the*-chat-*ones) rejectstemperature(HTTP 400) and requiresmax_completion_tokens.gemini-3.xdiscardstemperature/top_p/top_k.- Gemini thinking: you pass only
thinking_budget(orthinking_level) and the library adapts — on 2.5 it becomesthinking_budget, on 3.x it becomesthinking_level—, wrapping it in the nativethinking_config. See Gemini.
thinking_budget/thinking_level are not constructor kwargs: a loose kwarg
would fall into client_kwargs (which goes to genai.Client, the wrong place).
Pass them via extra= on the constructor or via params= on the call:
# on the constructor (applies to every call of this LLM)
llm = LLM("gemini", "gemini-3.5", extra={"thinking_level": "LOW"})
# per call (overrides just this one)
llm.complete("...", params={"thinking_budget": 1024})Don't mix the two conventions in the same value: thinking_level=500 is wrong —
thinking_level is an enum (LOW/HIGH); 500 would be a thinking_budget.
Pick one and the library translates it to what the model accepts.
Order applied in the adapter: _translate() (canonical → native) →
apply_profile() (model quirks) → wrap thinking_config. When supporting a new
model with a different contract, add a rule in profiles.py instead of scattering
ifs around.
What changed in 1.9.0
profile_model=: whenmodelis not the real model name (e.g. the deployment name on Azure), tell the library the base model so it applies the right rules:LLM("azure", "my-deploy", profile_model="gpt-5")dropstemperatureand usesmax_completion_tokens, as for gpt-5. It is not sent to the SDK.- Profile aliases:
vertexinherits thegeminirules andazuretheopenaiones. For your own provider:jangada_ai.profiles.register_alias("mine", "openai"). - o-series (o1/o3/o4):
max_tokensbecomesmax_completion_tokensand sampling is removed, as for gpt-5. thinking_budgeton Gemini 3.x is graduated: up to 1024 →LOW, up to 8192 →MEDIUM(Flash models only), above →HIGH. Ongemini-2.5-pro,thinking_level="MINIMAL"becomes 128 (the minimum accepted), not 0.- Removing a parameter in a derived
LLM:llm.with_params(temperature=REMOVE)(from jangada_ai import REMOVE);Noneis still ignored. - A negative
max_retriesraisesValueErrorwhen creating theLLM.
Example
examples/model_profiles_example.py — runnable script.
Getting started with jangada
jangada is a thin layer over the official LLM SDKs (Anthropic, OpenAI, Groq, Gemini). The goal is to swap provider / model / api_key without changing the rest of your code.
Providers and API keys
jangada supports several providers (Anthropic, OpenAI, Groq, Gemini, Mistral, OpenRouter and cloud gateways), each isolated in an adapter that translates the normalized types (Message/Completion) to the native SDK.