Jangada AIJangada AI

OpenAI

Provider openai. Adapter over the openai SDK (chat.completions dialect).

pip install "jangada-ai[openai]"
  • provider=: "openai"
  • Environment variable: OPENAI_API_KEY
  • Adapter base: _OpenAICompatible (shared with Groq)
from jangada_ai import LLM
llm = LLM("openai", "gpt-4o-mini")

What it does

  • Text (complete/acomplete) and streaming (stream/astream).
  • Structured output (parse): uses the native helper chat.completions.parse(response_format=Model) → .message.parsed.
  • Vision (images=): images become image_url with a data URI.
  • Documents (files=): local text extraction (common to all).
  • Object detection (detect_objects): via vision + structured.
  • Audio transcription (transcribe): dedicated endpoint audio.transcriptions.create. Models: gpt-4o-transcribe, gpt-4o-mini-transcribe, whisper-1.

Structure and quirks

  • Parameters: accepts temperature, max_tokens, top_p, stop, seed. It has no top_k (it is dropped).
  • Model profile (profiles.py): the gpt-5 family rejects temperature and uses max_completion_tokens instead of max_tokens — jangada normalizes this automatically. See Parameters and profiles.
  • Response (Completion): text, usage (input_tokens/output_tokens derived from prompt_tokens/completion_tokens), raw with the native object.
  • Errors: translated by errors.classify() into the normalized hierarchy.

What changed in 1.9.0

  • Native tools → Responses API. With web_search(), file_search(), code_execution() (code interpreter), image_generation() or computer_use() in tools=, the call goes through the Responses API (see Native tools). Without native tools it stays on chat.completions.
  • Model that refuses function tools on chat.completions → Responses API. Some new models (e.g. gpt-5.6-luna) refuse tools= on chat.completions and ask for the Responses API. Jangada switches route on the first refusal and remembers the model — later tool calls go straight there. Nothing changes on your side.
  • Strict structured output: the schema now lists every property in required (a strict-mode requirement), and the fallback to JSON mode covers more "schema not supported" messages.
  • Truncated output in parse becomes TruncatedError; a refusal becomes OutputValidationError.
  • Usage: cache_read_tokens (prompt caching) and reasoning_tokens (reasoning models) in usage; the cost applies the cache discount.
  • Batched embeddings: large lists are split into several requests (llm.embed(texts, batch_size=...)).
  • o-series (o1/o3/o4) has a profile rule: max_completion_tokens and no sampling.

Structured example

from pydantic import BaseModel
class Person(BaseModel):
    name: str; age: int

llm.parse("Extract: João, 30 years old.", Person).parsed   # Person(name='João', age=30)

Related: Capability matrix, Groq (same dialect), Audio transcription.

On this page