Documents (docx, pdf, csv, xlsx)
Attach files to any call with files=. By default jangada extracts the
text from the file locally instead of using vision — it's cheaper and works
on any model, including ones without vision.
pip install "jangada-ai[files]" # pypdf, python-docx, openpyxlfrom jangada_ai import LLM, Document
llm = LLM("openai", "gpt-4o-mini")
# paths: type detected by extension
llm.complete("Summarize:", files=["report.pdf", "contract.docx"])
# xlsx: ALL sheets are included, each one labeled (## Sheet: ...)
llm.complete("Highest total?", files=[Document("sales.xlsx", max_rows=200)])
# in-memory bytes (upload/queue) — provide the name to detect the type
llm.parse("Any duplicates?", Report, files=[Document(blob, name="x.csv")])
# force vision (scanned PDF / when layout matters)
llm.complete("Transcribe:", files=[Document("scan.pdf", mode="vision")])The mode rule
mode | Behavior |
|---|---|
"auto" | (default) extracts text from csv/xlsx/docx/text-PDF; image → vision |
"text" | forces text extraction (error if the format has no text) |
"vision" | forces the image path |
Why not use vision for everything
- Cheaper: text costs far less than image tokens.
- Works on any model, including those without vision.
- Preserves tables as markdown.
A PDF without a text layer (scanned) raises DocumentError suggesting
mode="vision" — it never silently returns an empty block.
Details
files=exists oncomplete/parse/stream(sync and async) and coexists withimages=.- Conversion happens at the client boundary (
files.py→to_part()): each file becomes aTextPartorImagePart, so the adapters never see a document format. - Formats:
.csv,.tsv,.xlsx,.xlsm,.docx,.pdf, plus plain text (.txt,.md,.json, ...).
Related: Vision, Structured output.
What changed in 1.9.0
- Scanned PDFs with
mode="vision"really work: each page is rendered to PNG (up toDocument(..., max_pages=20)) and sent as an image. It usespypdfium2(now in the[files]extra) orpymupdfif installed; without either, the error says what to install. The library never sends a PDF as if it were an image (that caused a 400 at the provider). To build the parts yourself:jangada_ai.files.to_parts(doc)returns one part per page. - CSV encoding: tries UTF-8 (with and without BOM), then cp1252 and latin-1 —
CSVs exported by Excel in Latin locales no longer turn into
�. - Large CSV/TSV are streamed up to
max_rows;.tsvuses its own extractor. - Protected PDF: tries an empty password; otherwise raises a clear
DocumentError(corrupted PDFs too). - xlsx: a formula cell without a cached value shows up as
[fórmula sem valor calculado: =...]instead of empty;|and line breaks in cells are escaped in the markdown table. - Nameless bytes: the format is detected from the content (PDF, PNG, JPEG, GIF, WEBP, BMP, docx, xlsx).
from jangada_ai import LLM, Document
llm = LLM("gemini", "gemini-3.5-flash")
comp = llm.complete("Transcribe the invoice.",
files=[Document("scanned_invoice.pdf", mode="vision", max_pages=3)])Example
examples/files_example.py — runnable script.
Audio transcription (speech-to-text)
LLM.transcribe() converts audio into text. Not supported by every provider — it depends on whether the provider's API accepts audio:
Object detection
detect_objects() detects objects in an image and returns the bounding boxes in absolute pixels. It's vision + structured output, so it works on any provider with vision...