OpenAI
Provider openai. Adapter over the openai SDK (chat.completions dialect).
pip install "jangada-ai[openai]"provider=:"openai"- Environment variable:
OPENAI_API_KEY - Adapter base:
_OpenAICompatible(shared with Groq)
from jangada_ai import LLM
llm = LLM("openai", "gpt-4o-mini")What it does
- Text (
complete/acomplete) and streaming (stream/astream). - Structured output (
parse): uses the native helperchat.completions.parse(response_format=Model)→.message.parsed. - Vision (
images=): images becomeimage_urlwith a data URI. - Documents (
files=): local text extraction (common to all). - Object detection (
detect_objects): via vision + structured. - Audio transcription (
transcribe): dedicated endpointaudio.transcriptions.create. Models:gpt-4o-transcribe,gpt-4o-mini-transcribe,whisper-1.
Structure and quirks
- Parameters: accepts
temperature,max_tokens,top_p,stop,seed. It has notop_k(it is dropped). - Model profile (
profiles.py): thegpt-5family rejectstemperatureand usesmax_completion_tokensinstead ofmax_tokens— jangada normalizes this automatically. See Parameters and profiles. - Response (
Completion):text,usage(input_tokens/output_tokensderived fromprompt_tokens/completion_tokens),rawwith the native object. - Errors: translated by
errors.classify()into the normalized hierarchy.
What changed in 1.9.0
- Native tools → Responses API. With
web_search(),file_search(),code_execution()(code interpreter),image_generation()orcomputer_use()intools=, the call goes through the Responses API (see Native tools). Without native tools it stays on chat.completions. - Model that refuses function tools on chat.completions → Responses API. Some
new models (e.g.
gpt-5.6-luna) refusetools=on chat.completions and ask for the Responses API. Jangada switches route on the first refusal and remembers the model — later tool calls go straight there. Nothing changes on your side. - Strict structured output: the schema now lists every property in
required(a strict-mode requirement), and the fallback to JSON mode covers more "schema not supported" messages. - Truncated output in
parsebecomesTruncatedError; arefusalbecomesOutputValidationError. - Usage:
cache_read_tokens(prompt caching) andreasoning_tokens(reasoning models) inusage; the cost applies the cache discount. - Batched embeddings: large lists are split into several requests (
llm.embed(texts, batch_size=...)). - o-series (o1/o3/o4) has a profile rule:
max_completion_tokensand no sampling.
Structured example
from pydantic import BaseModel
class Person(BaseModel):
name: str; age: int
llm.parse("Extract: João, 30 years old.", Person).parsed # Person(name='João', age=30)Related: Capability matrix, Groq (same dialect), Audio transcription.