Admin ConsoleAdminister

Guardrails

Screen every request your organization sends through the gateway: an AI judge that enforces your written policies, and a PII detector that blocks or redacts sensitive data before it reaches a model provider.

Required role
Owner or Admin of the active organization.

How guardrails work

Open Guardrails from the Administration group. The page has two tabs, each with an ON/OFF badge. Both are off until you turn them on, and each organization has its own configuration.

TabWhat it doesCost
Layer 1A model you choose (the judge) reads the request and blocks it if it breaks one of your policies.One extra model call per request. It isn't counted against the person's tier budgets or credits.
Layer 2Detects sensitive data such as card numbers, national IDs and email addresses by pattern and checksum, then blocks the request or rewrites it.No model call (unless you turn on the optional name and place detection).

Every chat request made through the gateway in the organization, from LatentCode, LatentWork or any API client, goes through Layer 1 and then Layer 2 before it's sent to the provider. A request blocked by either layer never reaches the model and returns HTTP 400:

Blocked by Layer 1
{
  "error": {
    "message": "The message contains a US Social Security number.",
    "type": "guardrail_blocked_llm",
    "policy": "Block PII"
  }
}

Layer 2 blocks use the type guardrail_blocked_pii and list the entities found. Only the content your users send is screened.

Layer 1: AI judge

Set it up

  1. 1
    Pick a Judge model

    Any model the organization has configured. The judge runs on every request, so choose a fast, inexpensive one.

  2. 2
    Add policies

    Select Add policy. Give it a short Label (for example "Block PII") and write the Instruction to the judge model, in plain language: what to block. Select Add policy to save it. Add as many as you need.

  3. 3
    Test it

    Under Test the judge, paste a sample message and select Run test. You see BLOCKED or ALLOWED, how long the judge took, and for a block, the policy and the reason. Tests use the current settings; nothing needs saving first.

  4. 4
    Turn it on

    Switch the toggle on. It needs a judge model and at least one policy. The card then reads AI enforcement active.

Settings

FieldDefaultMeaning
Judge modelNoneThe model that evaluates requests.
On judge error / timeoutAllow the request (fail-open)What happens when the judge fails, times out, or gives an answer that can't be read. Block the request (fail-closed) is stricter but blocks everything while the judge is unavailable.
Timeout (seconds)8How long to wait for the judge, from 1 to 60.
Block message (shown to caller)"Your request was blocked by an AI safety policy."Returned when a request is blocked and the judge gave no reason, or when a fail-closed error blocks it. Normally the judge's own reason is returned.

The judge reviews the conversation and blocks only clear policy violations. To remove a policy, select its trash icon.

Writing policies
Write each policy as a clear instruction about what to block, one topic per policy. The judge reports which policy it applied, so specific labels make blocks easier to understand. Test borderline examples before turning enforcement on.

Layer 2: PII detection

If the layer shows "Presidio isn't installed on this deployment", PII detection isn't available on this installation; contact your LatentStack administrator.

Choose a mode

ModeEffect when something is detected
BlockRejects the request with the Block message. This is the default mode.
SanitizeRewrites each detected item according to its action and forwards the cleaned request. If something is found in tool-call arguments, which can't be rewritten, the request is blocked instead.
OffNothing is scanned.

The toggle turns the layer on or off; the mode decides what it does.

Choose what to detect

Under Entities & actions, tick the entities to watch and pick an action for each. Out of the box, CREDIT_CARD and EMAIL_ADDRESS are masked and US_SSN, IN_PAN and IN_AADHAAR are redacted.

GroupEntities
GlobalCREDIT_CARD, EMAIL_ADDRESS, PHONE_NUMBER, IBAN_CODE, IP_ADDRESS, CRYPTO, MEDICAL_LICENSE
IndiaIN_PAN, IN_AADHAAR, IN_VEHICLE_REGISTRATION
USUS_SSN, US_BANK_NUMBER, US_ITIN, US_PASSPORT
Names & places (needs NER model)PERSON, LOCATION, NRP, DATE_TIME
ActionIn Sanitize mode, the detected text is…
redactRemoved.
maskOverwritten with * (up to 12 characters).
replaceReplaced with the entity name, such as <CREDIT_CARD>.
hashReplaced with its SHA-256 hash.
keepLeft as is (still counts as a detection in Block mode).

Add your own patterns

  • Browse library opens a catalogue of ready-made patterns grouped by category, each marked with its compliance area and false-positive risk. Tick cards and add them, or select Add on the Recommended starter pack for the low-false-positive set. Added patterns are redacted by default.
  • Custom rule adds your own regular expression: an Entity name (for example EMP_ID), an Action, the Regex, and a Confidence score (default 0.7). Invalid regular expressions are rejected.

Custom and library rules appear under Custom rules, where you can change their action or delete them.

Other settings

FieldDefaultMeaning
Detection engineRegex + checksum onlyLocal NER model also detects names and places, if available (the page shows whether it is).
Min confidence0.35Detections scoring below this are ignored. Raise it to reduce false positives.
Scan rolesuserWhich messages to scan: user, assistant and/or system.
Block message"Your request was blocked because it contained sensitive data."Returned in Block mode.
Deny-listEmptyComma-separated words that are always flagged, such as internal project code names.
Allow-listEmptyComma-separated words that are never flagged.

Changes on this tab save as you make them.

Test detection

Paste text under Test detection and select Run test. The result lists each entity found with its score, and shows the text as it would look after sanitizing.

Troubleshooting

"Pick a judge model and add at least one policy first"

Layer 1 can't be turned on until both are set.

Legitimate requests are being blocked

Run them through the test box to see which policy or entity triggers. For Layer 1, make the policy wording more specific. For Layer 2, raise Min confidence, add the term to the Allow-list, or change the entity's action and use Sanitize instead of Block.

Requests got slower

Layer 1 adds a judge call to every request. Use a faster judge model or lower the timeout.

Names aren't detected

Name and place detection needs Local NER model selected, and it must be available; the page shows whether it is.