Jangada AIJangada AI

Streaming

Receive incremental tokens with stream() (sync) or astream() (async).

for token in llm.stream("Tell me about {{x}}", x="João Pessoa"):
    print(token, end="")
async for token in llm.astream("..."):   # e.g.: FastAPI StreamingResponse
    ...

Retry and fallback in streaming

Retry and fallback happen before the first token: if opening the stream fails with a transient error, jangada retries (backoff) and, if needed, falls back to the next candidate — all before you receive any content. Once the first token is out, the stream runs to the end.

See Retry and fallback for the full policy.

Notes

  • stream()/astream() accept the same system=, history=, images=, files=, and params= as the other calls.
  • For cost and tokens use the non-stream calls (complete/parse), which return usage/cost in the response — see Cost and tokens.

What changed in 1.9.0

  • Chunks without content (e.g. Azure's first content-filter chunk) no longer break the stream, and the connection is closed if you stop consuming the iterator halfway.
  • After the stream, the provider keeps usage and finish_reason in llm.provider.last_stream (OpenAI, Azure, DeepSeek, Mistral, Ollama). It is per instance: with several threads sharing the same LLM, read it right after the stream.
  • The stream is reported to observability at the end.
  • With native tools, streaming emits only the text (Anthropic/Gemini/OpenAI); on Mistral and Ollama, streaming + native tools raises UnsupportedError.

Example

examples/async_example.py — executable script.

On this page