779+ vulnerabilities in the catalog

Stress-test your AI agent
before your users do.

AegisMonkey runs adversarial attacks — jailbreaks, context poisoning, tool abuse — against your agent and scores its robustness. Ship with confidence.

779+

Cataloged attacks

4

Test modules

100

Max robustness score

< 60s

Avg run time

How it works

01

Register your agent

Point AegisMonkey at any LLM endpoint — OpenAI, Anthropic, or a custom HTTP server. Paste your base URL and we handle the rest.

02

Run the gauntlet

Our worker fires adversarial prompts across four attack surfaces. Each test is logged, scored, and stored — you watch results stream in live.

03

Read your scorecard

Get a robustness score out of 100 with per-module breakdowns. Compare across runs to track regressions as you iterate.

Four attack surfaces

Every test run covers all four modules. Nothing skipped, nothing hidden.

Prompt Injection

critical

DAN jailbreaks, payload splitting, role-play overrides, and token-level obfuscation. Does your agent hold its system prompt under pressure?

Context Poisoning

high

Adversarial history turns that plant false facts, override instructions, or hijack the agent's objectives mid-conversation.

Infrastructure Stress

medium

Simulated tool failures, API timeouts, and rate-limit jitter. Does your agent retry gracefully or spiral into hallucinated workarounds?

Hallucination Detection

medium

Responses compared against ground-truth datasets. Catches factual errors and confident confabulation before they reach production.

Vulnerability feed

Attacks updated from the research frontier

AegisMonkey ingests new attack techniques twice daily from arXiv, MITRE ATLAS, Garak, HiddenLayer, OWASP, and more. Every finding is reviewed before it reaches the catalog — so your tests reflect the actual threat landscape, not last year's.

779+

cataloged vulnerabilities

growing daily

Pricing

Start free. Scale when you need it.

Free

$0

forever

  • 1 agent
  • 1 test run / day
  • Core test suite
  • Robustness scorecard
  • 7-day run history
Get started
Popular

Pro

$29

per month

  • 3 agents
  • 10 test runs / day
  • Core + vulnerability feed tests
  • Unlimited run history
  • Discord webhook alerts
Get started

Startup

$99

per month

  • 10 agents
  • 50 test runs / day
  • Everything in Pro
  • Priority email support
Get started

Enterprise

Custom

volume pricing

  • Unlimited agents
  • Custom attack catalog
  • Private feed ingestion
  • SSO / SAML
  • Dedicated Slack channel
  • SLA + security review
Contact us