Creating a policy
- Select Safety and Evaluations > Responsible AI in the sidebar.
- Select Create New Policy and give it a name.
- Enable and configure the checks you need across the available categories.
- Select Save in the top right.
- Select Start Testing in the right panel to validate the policy against sample interactions.
Compliance frameworks
Compliance is a category in the Responsible AI policy editor that evaluates agent interactions against named regulatory and industry standards, in addition to the content checks described elsewhere on this page. Open a policy in Safety and Evaluations > Responsible AI, enable Compliance, and select the frameworks to enforce. Then assign the policy to an agent through the Responsible AI feature card in the Agent Builder.
Available frameworks
Select any combination of the following frameworks:| Framework | Standard |
|---|---|
| GDPR | EU General Data Protection Regulation. |
| MAS | Monetary Authority of Singapore regulatory requirements. |
| EU AI Act | EU Artificial Intelligence Act, risk-tiered obligations for AI systems. |
| HIPAA | US Health Insurance Portability and Accountability Act. |
| SOC 2 | AICPA SOC 2 Trust Service Criteria. |
| FedRAMP | US Federal Risk and Authorization Management Program. |
| ISO 27001 | ISO/IEC 27001 Information Security Management. |
| CCPA | California Consumer Privacy Act, including CPRA. |
| PCI DSS | Payment Card Industry Data Security Standard. |
Custom rules
Beyond the built-in frameworks, you can define your own compliance rules. Select Add rule to create a user-defined rule that is enforced in addition to the frameworks you selected.Input and output checks
A compliance framework runs at two points in an interaction, and you can enable either or both.- Input-level check evaluates what the user submits, for example a data retention policy or a storage description, against the selected framework and flags any violations before the agent acts on the input.
- Output-level check evaluates the agent’s response and blocks any non-compliant output, returning an explanation of which rule the response violated.
When to use each check
| Mode | What it delivers |
|---|---|
| Enterprise control | Output-level checks so end users never receive non-compliant responses in production. |
| Self-review | Input-level or output-level checks so a builder can verify an agent’s compliance manually during development. |
Custom guardrails
Custom guardrails let you bring your own HTTP guardrail servers into a Responsible AI policy. Enable Custom Guardrails in the policy editor, then add one or more guardrails that Lyzr calls during an interaction. Use this when your safety or compliance logic lives in your own service rather than in the built-in checks.
verdict field in the response body decides whether the interaction is allowed or denied. Each guardrail has the following settings:
| Setting | What it does |
|---|---|
| Name | Labels the guardrail so you can identify it in the policy. |
| Endpoint URL | The HTTPS endpoint Lyzr sends the POST request to, for example https://guardrail.mycompany.com/check. |
| Operation | Sets how the guardrail acts on the text. Validate (allow / deny only) returns a pass or block decision. Mutate (may rewrite the text) can return modified text. |
| Auth Data | Sets how Lyzr authenticates to your endpoint: No auth, Bearer token (sent as an Authorization: Bearer header), or Basic auth. |
| Mode | Sets how a deny or a server error is handled: Enforce (block on deny), Audit (log only), Fail Open (allow on server error), or Fail Closed (block on server error). |
| Run on | Sets when the guardrail runs: on LLM input or on LLM output. |
| Timeout (seconds) | Sets the maximum time Lyzr waits for your endpoint to respond. |
AWS Bedrock Guardrails
Lyzr supports connecting AWS Bedrock Guardrails as an external content governance layer. This option requires your own AWS credentials and is not enabled by default.
| Filter | What it blocks |
|---|---|
| Sexual Content | Sexually explicit material |
| Violence | Violent content |
| Hate Speech | Hate speech and discrimination |
| Insults | Insulting language |
| Misconduct | Content promoting illegal activities |
| Prompt Attack | Prompt injection attempts |

Toxicity detection
Lyzr validates every LLM output for toxicity before it reaches the user. The system scores responses between 0 and 1. Responses above the configured threshold are blocked, and the LLM is asked to regenerate until a safe response is produced. Default threshold: 0.4. Values closer to 1 allow more content through; lower values are stricter.
Prompt injection protection
Lyzr checks every incoming user message for prompt injection attempts before it is sent to the LLM. The system assigns a risk score from 0 to 1. Messages above the threshold are blocked before they reach the model. Default threshold: 0.3. Lower values are stricter.
Secrets detection
Lyzr automatically detects and redacts sensitive credentials from both inputs and outputs. Detected values are masked before being stored, displayed, or transmitted. Covered: API keys, authentication tokens, JWTs, private keys, and certificate data.Allowed topics
Restrict the agent to responding only to queries within explicitly approved topic domains. Configure by providing comma-separated values:Banned topics
Prevent the agent from discussing specific prohibited topics. Configure by providing comma-separated values:NSFW detection
Detects and blocks not-safe-for-work or inappropriate content before it is processed or returned.
- Sentence-by-sentence: scans each sentence individually for higher precision.
- Full text: evaluates the entire response as a whole for contextual detection.
Keyword management
Block or redact specific words and phrases from both inputs and outputs.

- Literal: exact substring match.
- Regex: regular expression for format-based patterns.
- Cucumber: parameter extraction for advanced logical matching.
- Blocked: the interaction stops if the keyword is detected.
- Redacted: the keyword is masked and the conversation continues.
Personally Identifiable Information (PII)
Configure how the agent handles each category of personal data. Each type can be independently set to Disabled, Blocked, or Redacted.| Data type | Description |
|---|---|
| Credit card numbers | 13 to 16 digit card number patterns |
| Email addresses | Standard email format |
| Phone numbers | International and local formats |
| Names (person) | Common personal name patterns |
| Locations | City, state, country, address |
| IP addresses | IPv4 and IPv6 |
| Social Security Numbers | U.S. SSN format XXX-XX-XXXX |
| URLs | Standard web address patterns |
| Dates and times | Temporal references and specific dates |
Use case reference
| Use case | Checks to enable |
|---|---|
| Customer support chatbot | Toxicity, Secrets, PII (email, phone) |
| Internal HR agent | Allowed Topics (HR/policy), Keywords (names/projects), PII (SSN, names) |
| Public-facing financial assistant | Prompt Injection, Banned Topics (politics), URL redaction, Credit card blocking |
| Legal document Q&A | Secrets, Credit card blocking, Topic control |