Features
22 test categories across security, safety, ethics, and quality — every check grounded in real-world risks seen on production chatbots and agents.
Protect your chatbot from adversarial misuse and data leakage.
Detects language that tries to override system instructions, hijack context, or insert adversarial directives through user input.
Catches jailbreak attempts, uncensored-mode requests, and language designed to remove the model's safety rails.
Flags prompts that request credentials, API keys, PII, internal data, or anything that should never leave the model context.
Identifies manipulation patterns that try to coerce the model into unsafe actions by appealing to authority or urgency.
Spots instructions that fetch or summarise external URLs, which could inject hostile content into the model's context window.
Detects over-broad tool grants — shell access, filesystem writes, unrestricted API calls — that increase blast radius.
Flags prompts likely to produce dangerous code, SQL, or markup without proper sanitisation guidance.
Catches attempts to bias future model behaviour or embed persistent instructions through crafted inputs.
Prevent bias, toxicity, and harmful outputs from reaching your users.
Instructions likely to produce harmful, offensive, abusive, or hateful content including slurs and derogatory language.
Prompts requesting dangerous medical, legal, financial, or safety guidance without appropriate professional caveats.
Language that could produce racially biased or discriminatory outputs or make unfair assumptions based on ethnicity.
Instructions that could cause the model to favour particular political parties, candidates, or ideologies.
Instructions that reinforce gender stereotypes, make assumptions based on gender, or treat genders unequally.
Instructions that disparage, unduly favour, or make assumptions about specific religious groups or their members.
Prompts encoding harmful generalisations about groups of people based on identity characteristics.
Language that discriminates or makes unfair assumptions based on age, whether against older or younger people.
Catch reliability and consistency issues before they reach production.
Prompts that explicitly invite the model to fabricate facts, citations, statistics, or data it cannot verify.
Instructions that may cause the model to produce contradictory or internally inconsistent factual claims across a response.
Ambiguous or contradictory instructions that reduce the model's ability to complete tasks reliably and predictably.
Prompts likely to produce wildly varying outputs across repeated runs, making results hard to test or depend on.
Instructions that prevent the model from appropriately refusing unsafe or out-of-scope requests.
Missing or ambiguous output format constraints that could cause inconsistent or unparseable responses downstream.
Our model evaluates each prompt across all 22 categories and returns a structured JSON result with score, summary, and per-finding explanations.
Every finding includes a plain-English explanation of the risk and a concrete guardrail recommendation — not just a flag.
Get a production-ready rewrite of your prompt with guardrails applied — copy, paste, and ship.
Organise tests by product, model, or environment. Assign owner, admin, and member roles. All history is scoped to the workspace.
Monthly test quotas are enforced at the workspace level. Upgrade mid-month and limits reset automatically.
Gate your deployment pipeline on prompt safety. One API call, a score back, and a pass/fail threshold you control.