Skip to content
#

indirect-prompt-injection

Here are 47 public repositories matching this topic...

AgentForensics is an open-source security framework that monitors complete LLM agent sessions in real time, detecting prompt injection attacks across tool outputs, web pages, documents, and API responses. It uses heuristic rules, a DistilBERT ML classifier, instruction boundary detection, semantic drift, and sliding-window multi-turn detection.

  • Updated May 9, 2026
  • Python
cordon

A prompt-injection firewall for AI agents, with no AI inside: your agent reads anything and takes orders only from you. For Claude Code, Gemini CLI, MCP hosts and LangChain. | Файрвол от промпт-инъекций для ИИ-агентов без ИИ внутри: агент читает что угодно и слушается только вас.

  • Updated Oct 1, 2026
  • TypeScript

Prompt-injection defenses for Claude Code. A PreToolUse Bash hook blocks compositional credential-exfiltration shapes (secret read plus network, env dump to network, remote script to shell, reverse shells). A sanitizing MCP server wraps untrusted URLs and files in sentinels, strips invisible unicode, flags jailbreaks.

  • Updated Aug 26, 2026
  • Python

AgenticAnomaly is an indirect prompt injection CTF for testing how agentic security operations center (SOC) workflows can be exploited through indirect prompt injection.

  • Updated Jun 23, 2026
  • Python

A verification layer for AI agents that checks proposed actions before execution, blocks or holds unrequested and high-risk behavior, and produces auditable decisions. Includes incident-derived security evaluations, adversarial holdouts, benign controls, known limitations, and public results.

  • Updated Sep 15, 2026

A reproducible prompt-injection benchmark that measures which defenses actually work: each payload is replayed against every defense, the report separates whether the agent COMPLIED from whether the damage was CONTAINED, failure rates carry bootstrap confidence intervals over a measured noise floor, and the defenses that do not work get published.

  • Updated Sep 29, 2026
  • Python

Empirical study of indirect prompt injection against a production LLM voicemail-to-action pipeline: reproducible attack harness, attack-success measurement, and a defense-cost analysis of layered mitigations.

  • Updated Jul 2, 2026
  • Python

Add this topic to your repo

To associate your repository with the indirect-prompt-injection topic, visit your repo's landing page and select "manage topics."

Learn more