---
title: "&quot;Guardrail architecture for LLM applications&quot; | Dhruv Doshi"
description: "Guardrails are the controls that keep LLM applications within acceptable bounds: blocking disallowed content, enforcing policies, validating outputs, and constraining agent actions.…"
canonical: https://doshidhruv.com/notes/guardrail-architecture-for-llm-applications/
author: Dhruv Doshi
---

[Skip to content](#main-content)

[Dhruv Doshi](/)

[Work](/projects)[Experience](/resume)[Notes](/notes)[Guides](/guides)[Research](/research)[Investing](/investing)[About](/about)[Contact](/contact)[Search /](/search)dark mode

[All notes](/notes)

[AI governance](/topics/ai-governance)

# "Guardrail architecture for LLM applications"

2026-10-01 · 10 minute read

Guardrails are the controls that keep LLM applications within acceptable bounds: blocking disallowed content, enforcing policies, validating outputs, and constraining agent actions. Every production LLM system needs them. Most teams build them as an afterthought — a content filter bolted onto the output, a few regexes, a hope. This note is about designing guardrails as architecture: layered, testable, and matched to actual risk.

The key insight: guardrails are not one thing. They are a set of control points at different stages of the request lifecycle, each catching different failure classes. A system with only output filtering is a system with one lock on a house with six doors.

## The control points

Map guardrails to the request lifecycle. Each stage gets its own controls.

Input validation happens before the model sees anything. Check that inputs match expected shapes: lengths, formats, character sets for structured fields. Detect and flag suspected injection attempts — not as a blocking control (detection is unreliable), but as a signal that raises scrutiny downstream. For high-risk flows, classify input intent and route accordingly: a request that looks like it is probing boundaries gets the strictest handling.

Retrieval controls govern what knowledge enters the prompt. Enforce access control at retrieval time — the retriever must only return documents the current user is authorized to see. This is the most important guardrail in enterprise RAG and the one most often implemented sloppily. Filter retrieved content for known-bad patterns (injected instructions in documents), and cap the amount and type of context that reaches the model. Retrieval is where tenant isolation lives or dies.

Prompt construction controls ensure the assembled prompt is what you intended. Validate that system instructions are intact, that content boundaries are properly delimited, that no component injected unexpected content into the prompt. This sounds paranoid until the first time a template bug concatenates user input into the system instruction block.

Output validation happens after generation. Check outputs against policy: disallowed content categories, PII patterns, format compliance, citation requirements. This is where most teams put their only guardrail, and it is necessary — but it is the last line, not the first.

Action controls govern what agents do, not just what they say. Every tool call gets checked against policy: is this tool allowed in this context, are the parameters within bounds, does this action need human approval? Action controls are where guardrails meet real consequences — a blocked output is an embarrassment avoided, a blocked tool call is an incident avoided.

## Deterministic vs. model-based guardrails

Guardrails come in two implementation flavors, and you need both.

Deterministic guardrails are rules: regexes, blocklists, schema validators, threshold checks, policy engines. They are fast, cheap, explainable, and testable. They are also brittle — they catch what you anticipated and miss what you did not. Use them for everything that can be specified precisely: output format validation, PII pattern detection, parameter bounds checking, keyword-based policy flags, rate limits. Deterministic guardrails are your reliable foundation.

Model-based guardrails use a second model to evaluate inputs or outputs: toxicity classifiers, intent classifiers, policy-violation detectors, hallucination checkers. They handle the fuzzy cases that rules cannot — nuanced policy judgments, context-dependent appropriateness, sophisticated injection attempts. But they are probabilistic: they have false positives and false negatives, they cost inference on every request, they add latency, and they can be attacked themselves.

The architecture principle: deterministic first, model-based for what rules cannot cover. Every policy that can be expressed as a rule should be a rule — it is cheaper, faster, and auditable. Reserve model-based guardrails for genuinely ambiguous judgments. And never stack model-based guardrails as your only defense on a critical control: a policy that matters (no PII in outputs, no unauthorized actions) needs a deterministic backstop even if a model does the nuanced work.

Calibrate model-based guardrails like any classifier: measure precision and recall on labeled data, tune thresholds to your risk tolerance (high recall for safety-critical, balanced for user experience), and monitor for drift. An uncalibrated classifier guardrail is a random number generator with opinions.

## The layered defense model

No single guardrail catches everything. Design in layers, where each layer catches different failure classes and the layers overlap on the critical ones.

A typical layering for a customer-facing agent: input validation and injection signaling at the edge; retrieval access control and content filtering in the pipeline; prompt construction validation in the orchestration layer; output policy checks (deterministic rules plus a classifier for nuanced cases) after generation; action approval workflows for consequential tool calls; and behavioral monitoring across all of it for anomaly detection.

The layers should be independent where possible. If your output filter and your input classifier are the same model with the same prompt, they share failure modes — an attack that fools one fools both. Diversity in guardrail implementation (different models, different approaches, deterministic plus probabilistic) is what makes layers actually layered rather than merely multiple.

Fail closed on critical controls, fail open (or degrade gracefully) on experience controls. A PII output filter that is uncertain should block or redact — the cost of a false positive is a degraded response; the cost of a false negative is a data leak. A tone classifier that is uncertain should probably let the output through — the cost of a false positive is a worse user experience on every borderline case. Match the failure mode to the stakes.

## Guardrails for agentic systems

Agents need everything above plus controls on their ability to act.

Tool-level guardrails: each tool declares its risk level, and the guardrail layer enforces policy per tool. Read-only tools get light checking. Tools with side effects get parameter validation, and the high-risk ones get human approval. Tools that touch other users' data or external systems get the strictest handling. This policy should live in configuration, not in prompts — it needs to be auditable, version-controlled, and changeable without retraining or re-prompting.

Parameter validation deserves emphasis because it is where agent attacks land. An agent tricked into calling a legitimate tool with malicious parameters — sending data to the wrong recipient, querying outside its scope, issuing an unintended transaction — is the characteristic agentic attack. Validate parameters against the user's stated intent and against policy bounds: amounts within limits, recipients within expected sets, data scopes matching the session's authorization.

Approval workflows for consequential actions need real design. The approver — human or automated policy — needs to see the action, the parameters, the justification, and the relevant context. "Agent requests approval" with no detail is not a control; it is a click-through ritual. For human approvers, design the interface for quick, informed decisions: what is being done, why, what could go wrong, and one-click approve/reject with the ability to drill into detail. For automated approval, the policy rules must be explicit and tested.

Budget and rate guardrails prevent the runaway agent: per-session limits on tool calls, on inference cost, on actions taken. These are the circuit breakers. An agent in a retry loop or a confused planning spiral hits the budget and stops, rather than burning resources until someone notices.

## Testing guardrails

Guardrails you have not tested are decorations. Test them like security controls, because that is what they are.

Adversarial testing: run your attack library against the guardrailed system and measure bypass rates per guardrail and per layer. A guardrail with a high bypass rate against known attacks is not a guardrail. Test each layer independently (does the output filter catch this?) and the system as a whole (does any layer catch this?).

False positive testing: run legitimate traffic through the guardrails and measure how much gets blocked or degraded. A guardrail that blocks 5% of legitimate requests is a user experience problem that will get disabled — and then you have no guardrail. Tune thresholds with both metrics in view: bypass rate and false positive rate, traded against each other explicitly.

Regression testing in CI: guardrail behavior should be pinned by tests. Every prompt change, model upgrade, and policy change re-runs the guardrail test suite. Model upgrades are the silent killer — a new model version can change output distributions enough to shift a calibrated classifier's behavior, and you will not notice until the bypass rate moves.

Red-team the guardrails specifically, not just the system. Dedicated guardrail testing — trying to evade each control — finds the weaknesses that system-level testing misses. The output filter that handles direct requests but fails on encoded content, the PII detector that catches formatted SSNs but misses them in prose, the approval workflow that can be bypassed by splitting an action into smaller steps.

## Operational concerns

Guardrails are production infrastructure. Operate them accordingly.

Latency budgeting: guardrails add latency to every request. Budget it explicitly — input checks, classifier calls, output validation — and design for the p99, not the mean. A classifier that usually responds in 200ms and occasionally takes 3 seconds will blow your latency SLA on exactly the requests you can least afford it. Set timeouts on guardrail calls with defined fallback behavior (fail closed for critical, degrade for non-critical).

Monitoring: track guardrail trigger rates, bypass attempts (from adversarial testing and from production signals), false positive reports, and latency contribution. A sudden change in trigger rates means something changed — an attack campaign, a model behavior shift, a broken rule. Alert on it.

Versioning and change management: guardrail configurations, classifier thresholds, blocklists, and policy rules are production config. Version them, review changes, test before deploying, and be able to roll back. The story of "someone updated the blocklist and broke the checkout flow" is a change-management failure, not a guardrail failure.

Auditability: log guardrail decisions — what was checked, what triggered, what was blocked or allowed and why. For regulated industries, this log is evidence of control operation. For everyone, it is the forensic record when something gets through. A guardrail that blocks without logging is a control you cannot prove existed.

## Guardrail UX: failing gracefully

Guardrails that trigger need a user experience, not just a block. What the user sees when a guardrail fires determines whether they trust the system or route around it.

For input blocks: explain what happened in plain language and offer a path forward. "I can't help with that" with no alternative is a dead end; "I can't process requests containing account numbers in this chat — you can continue in the secure support portal" preserves the user's goal. Never reveal guardrail internals in the refusal — describing exactly which rule triggered teaches attackers how to evade it.

For output interventions: prefer graceful degradation over hard refusal where the policy allows. A response with PII redacted and a note that sensitive details were removed is more useful than a refusal, when redaction satisfies the policy. Partial compliance beats total refusal for user trust, as long as the redaction is reliable.

For approval workflows: make the wait transparent. Tell the user what is being reviewed, why, and how long it typically takes. Approval queues that go silent create support tickets and workarounds — and workarounds bypass the guardrail you built. Monitor approval queue depth and time-to-decision as operational metrics; a guardrail that creates a permanent backlog will be pressured into rubber-stamping, which converts your control into theater.

Log every user-facing guardrail event with the same rigor as the control decision itself. When users complain about over-blocking — and they will — the logs are how you distinguish false positives worth fixing from correct blocks worth explaining.

## Anti-patterns

The single output filter: one content check at the end and nothing elsewhere. It catches the obvious, misses the structural, and creates false confidence.

Prompt-only guardrails: "the system prompt says not to" as the entire control strategy. It is the weakest layer pretending to be the whole building.

Uncalibrated classifiers: model-based guardrails deployed without precision/recall measurement, tuned by vibes.

Guardrails nobody monitors: trigger rates, bypass rates, and latency contributions invisible until an incident.

The untestable policy: "use good judgment" as a guardrail specification. If you cannot write a test for it, it is not a control.

Approval theater: human approval steps with no context, no detail, and no real ability to evaluate — producing clicks, not oversight.

## Stated plainly

Design guardrails as layered architecture across the request lifecycle: input, retrieval, prompt construction, output, and action controls. Deterministic rules for what can be specified; calibrated models for what cannot. Fail closed where the stakes are high. Test adversarially, measure false positives, run regression in CI, and operate guardrails as production infrastructure with monitoring, versioning, and audit logs. A guardrail is a control — and controls that are not tested, monitored, and maintained are just documentation.

## Continue reading

- ["AI red teaming: a practical playbook for enterprise teams"](/notes/ai-red-teaming-a-practical-playbook-for-enterprise-teams) · Note
- ["Third-party AI vendor risk management"](/notes/third-party-ai-vendor-risk-management) · Note
- [Design APIs for agent consumers](/notes/design-apis-for-agent-consumers) · Note
- ["Audit logging for AI systems: what to record and why"](/notes/audit-logging-for-ai-systems-what-to-record-and-why) · Note

**Dhruv Doshi** · Toronto, Canada · [work@doshidhruv.com](mailto:work@doshidhruv.com)

[Resume](/resume)[Notes](/notes)[Guides](/guides)[Topics](/topics)[Search](/search)[Research](/research)[Investing](/investing)[LinkedIn](https://www.linkedin.com/in/dhruvdoshi25071999)[GitHub](https://github.com/DhruvDoshi)

[Sitemap](/sitemap.xml)[RSS](/feed.xml)[LLMs](/llms.txt)

© 2026 Dhruv Doshi
