Jangada AIJangada AI

Tutorial: RAG from scratch

Goal: answer questions based on your texts. Jangada does embeddings + hybrid search (lexical BM25 + vector, fused via RRF) + answer generation. Start in memory (no database) and swap in a real vector store later — without changing the rest.

Install the extra:

pip install "jangada-ai[openai,rag]"
export OPENAI_API_KEY=sk-...

1. Assemble the RAG

You need an embedder (generates vectors), a store (keeps them), and a chat (answers):

from jangada_ai import LLM
from jangada_ai.rag import RAG, InMemoryVectorStore

embedder = LLM("openai", "text-embedding-3-small")
chat     = LLM("openai", "gpt-4o-mini")

rag = RAG(embedder, InMemoryVectorStore(), chat=chat, k=3, alpha=0.5)
  • k — how many chunks to retrieve.
  • alpha — search balance: 0 = BM25 only (words), 1 = vector only (semantics), 0.5 = balanced.

2. Index content

rag.add_texts([
    "Jangada switches providers by changing only LLM('provider', 'model').",
    "The fallback is triggered on rate limit, 5xx, or timeout.",
])

# or from files (docx/pdf/csv/xlsx/txt):
rag.add_document("manual.pdf")

3. Ask

resp = rag.ask("How do I switch providers?")
print(resp.text)                 # answer based on the retrieved context
print(len(resp.sources), "chunks used")

ask() retrieves the most relevant chunks, builds the context, and generates the answer. resp.sources brings the chunks (with scores) that backed it.

4. Going to production: a real vector store

Swap only the store — pgvector or Mongo, detected from the connection string:

from jangada_ai.rag import vector_store

store = vector_store("postgresql://user:pass@host/db")   # or "mongodb+srv://..."
rag = RAG(embedder, store, chat=chat)

The rest of the code stays the same. Fine-tuning (min_score, filter, weights, max_context_chars) is in RAG.

Async version (1.9.0+)

The whole flow has an async counterpart — handy in APIs (FastAPI) and to index many documents without blocking the event loop:

await rag.aadd_document("policy.pdf")
resp = await rag.aask("Can I return it after 30 days?")

In hybrid mode (the default), min_score is compared with the vector search cosine similarity — values like 0.3–0.5 make sense.

Next steps

  • Complete recipe: examples/cookbook/03_chatbot_rag.py.
  • Detailed reference: RAG.

On this page