Example: Fiscal Vision
An invoice/receipt reader from photos: it extracts the data as structured JSON via vision, detects regions in the image, and checks the item sum against the printed total using tools. It combines vision + structured output + function calling in a single pipeline.
Folder: pocs/fiscal-vision · Suggested port: 8002
jangada features
Image.from_bytes()— loads the uploaded imagellm.aparse(prompt, Schema, images=[...])— structured output via visionadetect_objects()— region detection (bounding boxes)- Tools with
tool_choice="required"— forces the sum check Document+complete(..., files=[doc])— PDF/XLSX/CSV reports
Core of the example
The invoice schema (app/schemas/nota.py):
class NotaFiscal(BaseModel):
estabelecimento: str | None = Field(default=None)
cnpj: str | None = Field(default=None)
data: str | None = Field(default=None)
itens: list[ItemNota] = Field(default_factory=list)
total_informado: float | None = Field(default=None, description="Total printed on the invoice")Vision extraction (app/routers/fiscal.py):
from jangada_ai import Image
@router.post("/extrair", response_model=NotaFiscal)
async def extract(file: UploadFile = File(...)) -> NotaFiscal:
"""Extracts the invoice data via VISION + structured output."""
llm = vision_llm(provider, model)
img = Image.from_bytes(await file.read(), file.content_type or "image/jpeg")
with observability_session(name="extract-invoice", metadata={"file": file.filename}):
comp = await llm.aparse(PROMPT_NOTA, NotaFiscal, images=[img])
return coerce(comp, NotaFiscal) # coercion fallback over comp.parsed/comp.textThings to watch
parse()returns aCompletion— the Pydantic object lives incomp.parsed. The POC'scoerce()handles theparsed=Nonecase by falling back toModel.model_validate_json(comp.text).- Set a generous
max_tokensfor invoices with many items, or the JSON gets truncated (see Parameters). default=Nonefields keep the model from inventing missing data.
How to run
cd pocs/fiscal-vision
pip install -r requirements.txt
uvicorn app.main:app --reload --port 8002 # http://localhost:8002/docsSee Vision, Structured output, Detection and Documents.