RAG (embeddings + vector/hybrid search)
jangada covers the "LLM parts" of RAG (embeddings + building the context) and ships
an optional jangada_ai.rag module with chunking, vector store (pgvector/Mongo),
and hybrid search.
pip install "jangada-ai[rag]" # psycopg (pgvector) + pymongo (Mongo)Embeddings (embed)
An optional capability, like audio: OpenAI and Gemini support it; Anthropic and
Groq raise UnsupportedError.
from jangada_ai import LLM
emb = LLM("openai", "text-embedding-3-small") # or ("gemini", "gemini-embedding-001")
emb.embed("a sentence") # -> vector (list[float])
emb.embed(["a", "b"]) # -> list of vectors
emb.embed(texts, task="document") # task: "document" when indexing, "query" when searching
emb.embed(texts, batch_size=50) # Gemini: texts per request (1..100)task becomes task_type in Gemini (RETRIEVAL_DOCUMENT/RETRIEVAL_QUERY);
OpenAI ignores it. dimensions (OpenAI) / output_dimensionality (Gemini) go via **opts.
Since v1.4.0, batch_size (1..100, default 100) controls how many texts
Gemini sends per request. The API accepts at most 100 items per
BatchEmbedContents; larger inputs are split automatically while preserving
order and count. OpenAI, Azure, and OpenRouter ignore this parameter.
| Provider | Embeddings? | Typical model |
|---|---|---|
| OpenAI | ✅ | text-embedding-3-small / -large |
| Gemini | ✅ | gemini-embedding-001 / gemini-embedding-2 |
| Anthropic | ❌ | — (use Voyage/Cohere externally) |
| Groq | ❌ | — |
Since v1.4.1, embed()/aembed() also generate automatic observations with
latency, tokens, estimated cost, and the embeddings capability. Inside
observability_session(name="rag.documents.ingest"), they appear in the
ingestion trace.
Since v1.4.4, the observation's output field contains the complete vector
array, always using the batch shape ([[...], [...]]).
Since v1.4.2, when Gemini does not return token usage for embeddings, Jangada
calls count_tokens() with the same model and batch. Cost combines this count
with the price from jangada.dev.br/prices.json; the local approximation of
about 4 characters per token is only used if provider counting also fails.
Full pipeline (RAG)
from jangada_ai import LLM
from jangada_ai.rag import RAG, vector_store
emb = LLM("openai", "text-embedding-3-small")
chat = LLM("openai", "gpt-4o-mini")
# the store is chosen by the CONNECTION STRING (just pass your DATABASE_URL_VECTOR)
store = vector_store("postgresql://user:password@host:5432/db") # or "mongodb+srv://..."
rag = RAG(emb, store, chat=chat, k=5) # k = number of chunks in the context (tunable)
rag.add_document("manual.pdf", metadata={"source": "manual"}) # extract -> chunk -> embed -> store
answer = rag.ask("How do I make a backup?", mode="hybrid") # uses the RAG's k
more = rag.ask("How do I make a backup?", k=10, mode="hybrid") # per-call override
print(answer.text)
for s in answer.sources:
print(s.score, s.chunk.content[:80])Vector store by connection string
vector_store(url) detects the adapter by scheme:
| URL | Adapter | Search |
|---|---|---|
postgresql:// / postgres:// | pgvector (Postgres) | cosine (<=>) + full-text (tsvector) |
mongodb:// / mongodb+srv:// | MongoDB | Atlas $vectorSearch + $text (client-side cosine fallback) |
memory / None | in-memory | cosine + keyword (no deps) |
Tables/collections and indexes are created automatically on first use (setup).
Portuguese (and other) full-text search (pgvector)
PgVectorStore uses the Postgres 'simple' configuration by default (no stemming
or stopwords). For Portuguese content, pass text_config="portuguese" — it
improves the lexical side of hybrid search:
from jangada_ai.rag import vector_store
store = vector_store("postgresql://...", text_config="portuguese")The config is fixed at table creation (the tsv column is GENERATED), so set
it before the first setup. Any Postgres regconfig works (english,
spanish, ...).
Search modes
mode="vector" | "text" | "hybrid":
- vector — embedding similarity only.
- text — lexical: BM25 in the in-memory store (with
rank_bm25; falls back to term counting without it) and native full-text in pgvector (tsvector) / Mongo ($text). - hybrid — combines both via Reciprocal Rank Fusion (RRF).
The vector × lexical balance comes from weights=(vector, text) or the shortcut
alpha (0 = BM25/text only, 1 = vector only; alpha becomes weights=(alpha, 1-alpha)):
RAG(emb, store, chat=chat, alpha=0.5) # balanced; 0.0 = BM25 only, 1.0 = vector onlyrag.search("incremental backup", k=5, mode="vector") # vector only
rag.search("incremental backup", k=5, mode="hybrid") # vector + text (RRF)Reranking (biggest quality jump)
The retriever brings candidates (good recall), but the order isn't always best.
A reranker reorders candidates by relevance and keeps the best ones — the
biggest RAG quality gain per effort. With reranker=, RAG fetches more
candidates (fetch_k, default k*4) and returns the top k reordered.
from jangada_ai.rag import RAG, Reranker, vector_store
rag = RAG(emb, vector_store("memory"), chat=chat, reranker=Reranker.cohere())
rag.ask("How do I run an incremental backup?")Constructors: Reranker.cohere(model="rerank-v3.5"), Reranker.voyage(model="rerank-2.5")
(extra jangada-ai[rerank] + COHERE_API_KEY/VOYAGE_API_KEY) or
Reranker.fn(lambda query, docs: [scores]) for a custom scorer. Per call,
rag.search(q, rerank=False) disables it.
Tunable parameters
Defined in RAG(...) (default) and/or per call:
rag = RAG(
emb, store, chat=chat,
k=5, # number of chunks in the context
min_score=0.25, # discards chunks below this similarity
max_context_chars=6000, # context budget (truncates the excess)
chunker=my_chunker, # function(text)->list[str] (replaces the default chunking)
rrf_k=60, # RRF constant (hybrid)
weights=(1.0, 0.5), # weights (vector, text) in hybrid
chunk_size=1000, overlap=200,
)
# metadata filter (scope by document/source/tenant) + per-call override
rag.ask("backup?", k=8, filter={"source": "manual"}, min_score=0.3, mode="hybrid")
rag.search("backup?", filter={"tenant": "acme"}, mode="vector")filterbecomesmetadata @> ...in pgvector and$matchonmetadata.<key>in Mongo.min_scoreuses the real similarity (cosine in vector, rank in text).max_context_charscuts the chunks that don't fit (keeps at least one).weights=(v, t)weighs the vector and text rankings in the RRF fusion.
Advanced retrieval strategies (opt-in)
Pass strategy= to search/ask (default: plain search, unchanged):
- multi-query — the LLM generates query variations, searches all and fuses by
RRF.
rag.ask(q, strategy="multi_query"). - parent-document — index small children keeping the big parent chunk; search
returns the parent:
rag.add_document("m.pdf", parent_chunk_size=4000)+rag.ask(q, strategy="parent_document"). - contextual compression — the LLM extracts only the relevant part of each
chunk:
rag.ask(q, compress=True).
They combine with each other and with reranker=.
Incremental indexing
Since v1.4.3, sync_document/sync_texts work both in memory and with
PostgreSQL/pgvector. They preserve unchanged chunks without new embeddings,
insert only new hashes, update metadata/position, and remove missing hashes.
result = rag.sync_document(
"manual.pdf",
name="manual.pdf",
document_id="manual",
metadata={"source": "manual.pdf"},
)
# {"added": 3, "removed": 1, "unchanged": 42}For pgvector, migration of the hash column and partial unique index is
idempotent. Each version runs in a transaction with an advisory lock per
document_id, preventing duplicates and mixed versions under concurrency.
Extraction, embedding, or database failures preserve the previous version.
MongoDB does not yet implement incremental synchronization.
sync_document/sync_texts reindex only what changed (dedup by content
hash): embed new chunks and drop removed ones. Requires document_id.
rag.sync_document("manual.pdf", name="manual")
# {'added': 3, 'removed': 1, 'unchanged': 42}Which feature to use? (cheat sheet)
All are opt-in — the lib works without any. Combine as needed:
| I want… | Use |
|---|---|
| Much better ordering of chunks | reranker=Reranker.cohere() |
| Catch synonyms / vague questions | strategy="multi_query" |
| Better context without losing precision | parent_chunk_size= + strategy="parent_document" |
| Fewer context tokens | compress=True |
| Chunks that don't cut ideas | chunker=semantic_chunker(emb) |
| Cheap reindex (only what changed) | sync_document(...) |
Recommended recipe: semantic chunking (index) + reranker (search); add multi-query if questions are vague.
from jangada_ai import LLM
from jangada_ai.rag import RAG, Reranker, semantic_chunker, vector_store
emb = LLM("openai", "text-embedding-3-small")
rag = RAG(emb, vector_store("postgresql://..."), chat=LLM("openai", "gpt-4o-mini"),
chunker=semantic_chunker(emb), reranker=Reranker.cohere())
rag.sync_document("manual.pdf", name="manual")
ans = rag.ask("How do I run a backup?", strategy="multi_query", compress=True)Chunking
from jangada_ai.rag import chunk_text
chunk_text(text, size=1000, overlap=200) # splits without cutting wordsRelated: Documents (text extraction reused in RAG) and Observability.
What changed in 1.9.0
- Full async API:
aadd_texts,aadd_document,async_texts,async_document,asearchandaask(usingaembed/acomplete; the store runs in a thread).Rerankergainedarerank, and there areaexpand_queries/acompress_results. - Hybrid search fixed on pgvector/Mongo. Fusion (RRF) used object identity and, on those databases, never merged the same chunk coming from both searches — it returned duplicates. It now uses a stable key (id → document+hash → content).
min_scorein hybrid mode filters by vector similarity (before fusion), not by the RRF score (which stays near 0.03 and emptied the results).- The in-memory index is no longer mutated by searches:
parent_documentandcompress=Truework on copies. - The default prompt wraps the retrieved context in
<contexto>tags and tells the model to treat it as data, not instructions (mitigates prompt injection coming from documents). Custom prompts with literal braces (e.g. a JSON example) no longer break. - Embeddings in batches of 128 texts (the OpenAI/Azure/OpenRouter and Mistral adapters also split), so large PDFs don't exceed the provider limit.
- pgvector:
- models above 2000 dimensions (e.g.
gemini-embedding-001,text-embedding-3-large, 3072) are indexed viahalfvec(requires pgvector ≥ 0.7), up to 4000; above that the table has no HNSW index, with a warning; add()/add_textsreturn how many chunks were actually inserted (duplicates by document+hash are dropped);- embeddings are computed outside the transaction and the advisory lock, the
connection is lock-protected across threads, and there is
close()+with PgVectorStore(...); - the table name is validated (accepts
schema.table).
- models above 2000 dimensions (e.g.
- Mongo Atlas: stores the chunk
hash, acceptsfilter_fields=[...](metadata fields declared asfilterin the index — required forfilter=in searches), falls back to client-side search only on Atlas errors (with a warning) and hasclose()+ a context manager.
store = vector_store("postgresql://...", table="docs", text_config="english")
with store:
rag = RAG(chat=LLM("openai", "gpt-5-mini"), embedder=LLM("openai", "text-embedding-3-small"), store=store)
await rag.aadd_document("manual.pdf")
resp = await rag.aask("What is the warranty period?", min_score=0.3) # min_score = cosineExample
examples/rag_example.py — executable script.