Streaming
Receive incremental tokens with stream() (sync) or astream() (async).
for token in llm.stream("Tell me about {{x}}", x="João Pessoa"):
print(token, end="")async for token in llm.astream("..."): # e.g.: FastAPI StreamingResponse
...Retry and fallback in streaming
Retry and fallback happen before the first token: if opening the stream fails with a transient error, jangada retries (backoff) and, if needed, falls back to the next candidate — all before you receive any content. Once the first token is out, the stream runs to the end.
See Retry and fallback for the full policy.
Notes
stream()/astream()accept the samesystem=,history=,images=,files=, andparams=as the other calls.- For cost and tokens use the non-stream calls (
complete/parse), which returnusage/costin the response — see Cost and tokens.
What changed in 1.9.0
- Chunks without content (e.g. Azure's first content-filter chunk) no longer break the stream, and the connection is closed if you stop consuming the iterator halfway.
- After the stream, the provider keeps
usageandfinish_reasoninllm.provider.last_stream(OpenAI, Azure, DeepSeek, Mistral, Ollama). It is per instance: with several threads sharing the sameLLM, read it right after the stream. - The stream is reported to observability at the end.
- With native tools, streaming emits only the text (Anthropic/Gemini/OpenAI); on
Mistral and Ollama, streaming + native tools raises
UnsupportedError.
Example
examples/async_example.py — executable script.
Step-back prompting
step_back() turns a specific question into a conceptually broader one, to retrieve background context in RAG. Works on any provider.
RAG (embeddings + vector/hybrid search)
jangada covers the "LLM parts" of RAG (embeddings + building the context) and ships an optional jangada_ai.rag module with chunking, vector store (pgvector/Mongo), and hybrid search.