Jangada AIJangada AI

Scope guardrails

Guardrails keep the LLM within a domain — so it doesn't become an assistant that answers anything — and block unwanted utterances. It's a thin composition layer (plain Python, no new dep): it reuses Message/Completion and the library's own parse.

A guardrail intercepts the call at two points:

  • input — before calling the main model (validates the user's request);
  • output — after the answer (validates what the model replied).

When it blocks, the client short-circuits and returns a Completion with the refusal message (message=) — on input it doesn't even spend the main model. With raise_on_block=True, it raises GuardrailError instead.

ScopeGuard

Combines two mechanisms, from cheap to robust:

  1. blocklist (regex/terms) — blocks instantly, no cost, no LLM;
  2. scope classifier (LLM-as-judge) — a cheap judge=LLM(...) decides, via structured output, whether the text belongs to the described scope.
from jangada_ai import LLM, ScopeGuard

guard = ScopeGuard(
    scope=(
        "e-Gestor system support: invoices, finance, registrations. "
        "Does NOT answer other topics (recipes, politics, code, etc.)."
    ),
    judge=LLM("groq", "llama-3.1-8b"),   # cheap/fast model just to classify
    block=[r"\bpassword\b", "ignore the instructions"],  # instant, no LLM
    message="Sorry, I can only help with e-Gestor topics.",
    check="both",                         # "input" (default), "output" or "both"
)

llm = LLM("openai", "gpt-4o", guardrails=[guard])

llm.complete("how do I issue an invoice?")  # in scope -> answers normally
llm.complete("teach me to bake a cake")      # out of scope -> Completion with refusal

The refusal comes back as a normal Completion (comp.text == message), with comp.cost is None and comp.raw == {"guardrail": "<reason>"} for inspection.

Parameters

ParameterPurpose
scopeText description of what's allowed (used by the judge).
judgeCheap LLM that classifies scope. If omitted, uses the main model.
blockList of regex/terms that block instantly, without calling an LLM.
messageRefusal text returned when it blocks.
check"input" (default), "output" or "both".
instructionOverrides the classifier instruction.
raise_on_blockTrue raises GuardrailError instead of refusing.
fail_closedIf the judge fails/is inconclusive: True blocks (safe), False (default) allows.

Recommendations

  • Use a separate, cheap judge (e.g. llama-3.1-8b, gpt-4.1-nano, gemini-2.5-flash-lite): classification is per call, so a small model keeps cost low. Without judge, the main model classifies (recursion is avoided internally, but it's pricier).
  • The blocklist is free: put the obvious terms/phrases there and leave topic judgment to the judge.
  • Cost: a blocked input does not spend the main model. A blocked output already paid for the generation (the refusal just replaces the text).

Where it applies

  • complete/acomplete and parse/aparse: input and output.
  • stream/astream: input only (an output guard would require buffering the whole stream). If blocked, the stream emits only the refusal message.

Custom guardrail

ScopeGuard covers the common case, but you can write your own by subclassing Guardrail and overriding check_input/check_output (and the a* versions), returning GuardResult(ok, reason):

from jangada_ai import Guardrail, GuardResult

class MaxLength(Guardrail):
    message = "Message too long."
    def __init__(self, limit): self.limit = limit
    def check_input(self, messages, judge):
        text = " ".join(m.content for m in messages if isinstance(m.content, str))
        return GuardResult(len(text) <= self.limit, "exceeded the limit")

What changed in 1.9.0

  • Safe under concurrent calls. The flag that stops the judge from triggering the guardrails again is now per context (ContextVar), not an LLM attribute. Previously, with asyncio.gather or threads sharing the same LLM, one call could skip every guardrail while another was waiting for the judge.
  • ScopeGuard(check_history=False) (default): checks only what arrived after the last assistant turn — the new user message and new tool results. A blocked term far back in the history no longer locks the conversation forever, and an agent loop doesn't re-judge the whole history on each iteration. Use check_history=True for the old behavior.
  • Tool results (role tool) go through the blocklist.
  • An output refusal keeps the cost: when the output guard blocks, the refusal carries the usage/cost of the call that was already paid.

On this page