Skip to content

AI Guardrail Rule Tester

Write keyword and regex guardrail rules that block, flag or redact text, test them against sample inputs and outputs, and export them as JSON.

AI Guardrail Rule Tester

Rule packs:
Rules (4)
3 matches

Every enabled rule runs against this text. A rule's Input/Output scope is metadata saved with exported rules; it is not applied to the in-tool test.

Highlighted matches
Hi, this is Jane Doe from Acme. Email me at jane.doe@example.com or call (555) 010-4477. My card on file is 4111 1111 1111 1111.
Match details
flagEmailjane.doe@example.com
redactPhone (US)(555) 010-4477
blockCredit Card4111 1111 1111 1111

Actions

blockReject the entire message
flagAllow but flag for review
redactReplace matched text with [REDACTED]

Updated . Provided as is. Check the output before you rely on it in production.

How to use AI Guardrail Rule Tester

  1. 1

    Build a rule to block harmful outputs

    Load a rule pack (PII Detection, Prompt Injection or Jailbreak Patterns) or add keyword, regex or pattern rules, each with a block, flag or redact action.

  2. 2

    Test guardrail rules against real LLM outputs

    Generate sample outputs from your LLM, paste into tester. Run each guardrail rule. See which rules trigger, why, and what to adjust. Catch false positives before production.

  3. 3

    Build cascading guardrails with rule priority

    Apply each rule to input, output or both. Match details show which rules fired on your sample text and the resulting action.

  4. 4

    Use regex patterns for flexible matching

    Instead of exact strings, use regex patterns like 'password|api.?key|secret' to catch common sensitive data. Tester shows which patterns matched in your test output.

  5. 5

    Export rules as JSON config for deployment

    Once rules are working in tester, export as JSON. Deploy to your LLM server/middleware to enforce guardrails in production. Rules are deterministic, no additional LLM calls needed.

Questions and answers

What are AI guardrails?
Guardrails are rules that filter, flag, or redact content in AI inputs and outputs. They prevent PII leaks, prompt injection attacks, harmful content generation, and policy violations in production AI applications.
What preset rule packs are available?
Three packs: PII (SSN, email, US phone and credit card patterns), prompt injection (ignore-instructions phrases, system prompt leak requests, role overrides, delimiter injection) and jailbreak (DAN, hypothetical bypass, no-rules roleplay, encoded payloads). Any pack loads into the editor as a starting point.
Can I export rules for NeMo Guardrails?
Yes, export your tested rules as JSON that can be adapted for NVIDIA NeMo Guardrails, LLM Guard, or any custom guardrail framework.
Does this tool detect prompt injection?
It tests YOUR rules against sample text. The preset injection patterns catch common attacks like 'ignore previous instructions', 'you are now', and role confusion attempts. You can customize and extend these patterns.
Is my test data private?
Privacy-first by design. All rule testing runs in your browser. No data is sent to any server. Your test content and rules stay on your machine.
For AI agents: how to call this tool

Machine-readable contract, endpoints and examples. Humans can ignore this section.

Best Path For Builders

Browser workflow

Runs instantly in the browser with private local processing and copy/export-ready output.

Browser Workflow

This tool is optimized for instant in-browser execution with local data handling. Run it here and copy/export the output directly.

/guardrail-rule-tester/

For automation planning, fetch the canonical contract at /api/tool/guardrail-rule-tester.json.