Open source prompt injection protection for Agents calling tools (via MCP, CLI or direct function calling). Detect and defend against prompt injection attacks. 22MB, CPU-only, < 10ms latency.
-
Updated
Oct 2, 2026 - TypeScript
Open source prompt injection protection for Agents calling tools (via MCP, CLI or direct function calling). Detect and defend against prompt injection attacks. 22MB, CPU-only, < 10ms latency.
A collection of public writeups on attacks against AI agents: prompt injection, tool abuse, memory poisoning, and data exfiltration.
System-level security for LLM agents: fine-grained policy enforcement on tool calls to defend against indirect prompt injection
AgentForensics is an open-source security framework that monitors complete LLM agent sessions in real time, detecting prompt injection attacks across tool outputs, web pages, documents, and API responses. It uses heuristic rules, a DistilBERT ML classifier, instruction boundary detection, semantic drift, and sliding-window multi-turn detection.
Zero Trust for AI Agents
Multi-hop cross-prompt injection benchmark for multi-agent AI systems. 250 attack cases, 8 taxonomy categories, 4 defenses evaluated. Watch: https://www.youtube.com/watch?v=fGOlMij4HPQ
A prompt-injection firewall for AI agents, with no AI inside: your agent reads anything and takes orders only from you. For Claude Code, Gemini CLI, MCP hosts and LangChain. | Файрвол от промпт-инъекций для ИИ-агентов без ИИ внутри: агент читает что угодно и слушается только вас.
Prompt-injection defenses for Claude Code. A PreToolUse Bash hook blocks compositional credential-exfiltration shapes (secret read plus network, env dump to network, remote script to shell, reverse shells). A sanitizing MCP server wraps untrusted URLs and files in sentinels, strips invisible unicode, flags jailbreaks.
Signed provenance labels and taint-tracking policy for LLM agent security. The core library behind AgentMesh.
Reproducible security benchmarking for the Deconvolute SDK and AI system integrity against adversarial attacks.
Penetration testing for AI agents — find, prove, and measure cross-server confused-deputy / prompt-injection exfiltration chains in an MCP tool mesh. Local-first, bring-your-own-model.
Generate YARA rules automatically from positive and negative examples. For PII detection, secret scanning, and prompt injection.
AgenticAnomaly is an indirect prompt injection CTF for testing how agentic security operations center (SOC) workflows can be exploited through indirect prompt injection.
A verification layer for AI agents that checks proposed actions before execution, blocks or holds unrequested and high-risk behavior, and produces auditable decisions. Includes incident-derived security evaluations, adversarial holdouts, benign controls, known limitations, and public results.
A reproducible prompt-injection benchmark that measures which defenses actually work: each payload is replayed against every defense, the report separates whether the agent COMPLIED from whether the damage was CONTAINED, failure rates carry bootstrap confidence intervals over a measured noise floor, and the defenses that do not work get published.
Picket: a local signature detector for prompt injection, planted and typed - no model, no network, no dependencies
Semantic-layer prompt injection defence that separates untrusted instructions from authority while preserving the legitimate task.
Empirical study of indirect prompt injection against a production LLM voicemail-to-action pipeline: reproducible attack harness, attack-success measurement, and a defense-cost analysis of layered mitigations.
A plain-English explanation of prompt injection for people who use AI assistants at work and have never written a line of code.
OWASP LLM Top 10 #1 — Live simulation of indirect prompt injection hijacking a finance LLM agent. Two-layer detection, HITL remediation, blast radius risk model. EU AI Act Art.9 · DORA · GDPR.
To associate your repository with the indirect-prompt-injection topic, visit your repo's landing page and select "manage topics."