KAVACH · AI-SECURITY SHIELD
Inspect every model call before it runs.
AgentAnywhere Kavach — कवच, armor — is the AI-security shield that sits in front of every model and agent call. It inspects the request and the response, detects prompt injection and jailbreaks with an ML classifier backed by rule and policy layers, and returns a decision: allow, block, or flag. Every verdict is enforced under per-tenant policy and written to an append-only audit log.
Every model call is an attack surface.
The moment an agent can be prompted, it can be prompted by an attacker. A support ticket, a scraped web page, a PDF, a tool response — any untrusted text in the context window can carry an instruction the model was never meant to follow. Prompt injection and jailbreaks are not edge cases; they are the front door.
Bolting a keyword filter onto each application does not hold. It cannot see contextual injection, it drifts out of sync app by app, and it leaves no single place to prove what was inspected, what was blocked, and why.
Kavach makes security a control that sits in front of the call, not a snippet copied into every service: one shield that inspects each request and response, decides allow, block, or flag, and records the verdict.
Inspect. Detect. Decide. Record.
Kavach sits in the path of every model and agent call. It inspects what goes in and what comes back, detects injection and jailbreak attempts, returns a verdict under your policy, and writes a record of what happened.
Inspect
Read the full request and the model's response inline — prompts, tool calls, retrieved context, and output — on every call, before anything is acted on.
Detect
An ML classifier flags prompt injection and jailbreaks — direct and contextual — backed by rule and policy layers. Detection that reasons about intent, not just a keyword list.
Decide
Return one verdict: allow, block, or flag — enforced under per-tenant policy so a clean call passes, an attack is stopped, and the ambiguous case is routed to review.
Record
Every verdict is written to an append-only audit log — subject, policy, decision, and time — so the security team can see and prove what the shield did.
Detection that reasons about intent, not keywords.
A blocklist catches yesterday's phrasing. Kavach pairs a trained classifier with deterministic policy rules so novel and contextual attacks are caught, and your own hard constraints are always enforced.
ML classifier
- A trained model that scores requests for prompt injection and jailbreak intent
- Catches contextual injection carried in retrieved text and tool output, not just direct attacks
- Generalizes past fixed keyword lists to phrasing it has not seen verbatim
- Runs inline on the call — this is a live component, not a roadmap item
Rule & policy layers
- Configurable policies, defined per tenant, from a versioned policies directory
- Deterministic guardrails for the constraints you must always enforce
- Layered with the classifier so detection is both learned and rule-backed
- Every decision the layers make is logged with the policy that produced it
Three verdicts. One contract.
Kavach returns a single, explicit decision on every call — so enforcement is unambiguous and the same everywhere the shield sits.
Allow
The call is clean under policy. It passes through to the model or tool, and the decision is still recorded — a clean verdict is evidence too.
Block
The request or response violates policy or trips the injection classifier. Kavach stops it before it reaches the model or leaves your boundary.
Flag
The call is ambiguous or high-risk. Kavach marks it for human review instead of failing silently — the decision is surfaced, not swallowed.
Kavach and Veil are a pair.
Two threats meet a model on every call: a malicious instruction trying to get in, and sensitive data leaking out. Kavach handles the first, Veil the second — together they are the front door to your models.
Kavach — the shield
- Inspects intent: detects prompt injection and jailbreaks
- Enforces per-tenant policy and returns allow / block / flag
- Stops a malicious or policy-violating call before it runs
- Records every verdict to the audit log
Veil — the mask
- Classifies PII, financial, and regulated data on the call
- Masks it fail-closed so raw values never reach the model
- Reversible tokenization or irreversible redaction, by policy
- Pairs with Kavach on the same call, so it is both safe and clean — see Veil
Every verdict is a trust event.
A security decision that no one can see is not a control — it is a hope. Kavach records every verdict to an append-only audit log: what was inspected, which policy applied, what was decided, and when.
That log is not a side file. Kavach is part of the AgentAnywhere trust platform alongside Veil, TrustFabric, and Custodian, and its verdicts are exactly the kind of first-class trust events the platform is built to surface — so a security team, a supervisor, or an auditor can ask what the shield did and get an answer grounded in the record.
Runs in front of your models, on your infrastructure.
Kavach is a live product in pilots today (app.agentanywhere.ai/agents/kavach). It stands in front of your own models and agents — standalone, or as part of the platform — without shipping your traffic to someone else's cloud.
Standalone
- Deploy Kavach in front of your existing models and agents
- Configure per-tenant policies and enforce them at the call
- Air-gap-friendly — run it entirely inside your boundary
- Enterprise SSO via IdentityAnywhere
Platform-native
- The Agent Universal Gateway routes every governed call
- Veil masks sensitive data fail-closed on the same call
- Swaraj runs the whole shield air-gapped on your keys
- Observe keeps the log; the verdict is a platform trust event
FAQ
Frequently asked questions.
- What is AgentAnywhere Kavach?
- AgentAnywhere Kavach is an AI-security shield that sits in front of every model and agent call. It inspects the request and the response, detects prompt injection and jailbreaks with an ML classifier backed by rule and policy layers, and returns a decision — allow, block, or flag — under per-tenant policy. Every verdict is written to an append-only audit log. Kavach is a live product in pilots.
- How does Kavach detect prompt injection and jailbreaks?
- Kavach runs a trained ML classifier that scores requests for prompt-injection and jailbreak intent, including contextual injection carried in retrieved text or tool output, not just direct attacks. The classifier is layered with configurable, per-tenant policy rules, so detection is both learned and rule-backed rather than a fixed keyword list.
- What decisions does Kavach return?
- Kavach returns one explicit verdict on every call: allow (the call is clean and passes through), block (the call violates policy or trips the injection classifier and is stopped), or flag (the call is ambiguous or high-risk and is routed to human review). Every verdict — including allow — is recorded to the audit log.
- How do Kavach and Veil work together?
- Kavach and Veil are a pair that guard the same model call. Kavach inspects intent and blocks prompt injection and policy violations; Veil classifies PII, financial, and regulated data and masks it fail-closed so raw values never reach the model. Together they make a call both safe from malicious instructions and clean of sensitive data.
- Can Kavach be deployed on-premises or air-gapped?
- Yes. Kavach is designed to run in front of your own models and agents, on your infrastructure, and is air-gap-friendly — you can run it entirely inside your boundary. It supports enterprise SSO via IdentityAnywhere and runs standalone or natively as part of the AgentAnywhere platform, where the Gateway routes calls, Veil masks data, and Swaraj runs it air-gapped on your keys.
- Is Kavach available today?
- Kavach is a live product in pilots, available at app.agentanywhere.ai/agents/kavach. It is a running security shield with an ML injection classifier, configurable per-tenant policies, and an append-only audit log — not a roadmap concept.
Put a shield in front of every model call.
Kavach is a live product in pilots — it inspects each request and response, decides allow, block, or flag under your policy, and logs every verdict. Tell us what your agents and models expose, and we will scope a pilot.