Jangada AIJangada AI

Generation parameters and per-model profiles

jangada accepts canonical parameter names and each adapter translates them to the SDK's native name, discarding the unsupported ones.

Canonical parameters

CanonicalOpenAI/GroqAnthropicGemini
temperaturetemperaturetemperaturetemperature
max_tokensmax_tokensmax_tokensmax_output_tokens
top_ptop_ptop_ptop_p
top_k(discarded)top_ktop_k
stopstopstop_sequencesstop_sequences
seedseed(discarded)seed
llm = LLM("anthropic", "claude-opus-4-8", temperature=0.2)  # max_tokens=8192

# per-call override
llm.complete("...", params={"temperature": 0.9, "max_tokens": 16384})

# clone with new defaults
creative = llm.with_params(temperature=1.0)

Since v1.4.5, max_tokens has a high default of 8192 across all providers. For models with a different limit or even larger extractions, set it explicitly:

llm = LLM("openai", "gpt-5", max_tokens=16384)
# or just for this call:
llm.parse("Extract every item.", Items, files=[...], params={"max_tokens": 16384})

SDK-specific parameters that have no canonical name go via extra=.

Automatic profiles (per-model quirks)

Models from the same provider sometimes have different contracts. jangada normalizes this in profiles.py, per model, without you needing to know:

  • gpt-5 and later (gpt-5.x, gpt-6-*…, except the *-chat-* ones) rejects temperature (HTTP 400) and requires max_completion_tokens.
  • gemini-3.x discards temperature/top_p/top_k.
  • Gemini thinking: you pass only thinking_budget (or thinking_level) and the library adapts — on 2.5 it becomes thinking_budget, on 3.x it becomes thinking_level —, wrapping it in the native thinking_config. See Gemini.

thinking_budget/thinking_level are not constructor kwargs: a loose kwarg would fall into client_kwargs (which goes to genai.Client, the wrong place). Pass them via extra= on the constructor or via params= on the call:

# on the constructor (applies to every call of this LLM)
llm = LLM("gemini", "gemini-3.5", extra={"thinking_level": "LOW"})

# per call (overrides just this one)
llm.complete("...", params={"thinking_budget": 1024})

Don't mix the two conventions in the same value: thinking_level=500 is wrong — thinking_level is an enum (LOW/HIGH); 500 would be a thinking_budget. Pick one and the library translates it to what the model accepts.

Order applied in the adapter: _translate() (canonical → native) → apply_profile() (model quirks) → wrap thinking_config. When supporting a new model with a different contract, add a rule in profiles.py instead of scattering ifs around.

See Providers and Extending.

What changed in 1.9.0

  • profile_model=: when model is not the real model name (e.g. the deployment name on Azure), tell the library the base model so it applies the right rules: LLM("azure", "my-deploy", profile_model="gpt-5") drops temperature and uses max_completion_tokens, as for gpt-5. It is not sent to the SDK.
  • Profile aliases: vertex inherits the gemini rules and azure the openai ones. For your own provider: jangada_ai.profiles.register_alias("mine", "openai").
  • o-series (o1/o3/o4): max_tokens becomes max_completion_tokens and sampling is removed, as for gpt-5.
  • thinking_budget on Gemini 3.x is graduated: up to 1024 → LOW, up to 8192 → MEDIUM (Flash models only), above → HIGH. On gemini-2.5-pro, thinking_level="MINIMAL" becomes 128 (the minimum accepted), not 0.
  • Removing a parameter in a derived LLM: llm.with_params(temperature=REMOVE) (from jangada_ai import REMOVE); None is still ignored.
  • A negative max_retries raises ValueError when creating the LLM.

Example

examples/model_profiles_example.py — runnable script.

On this page