AI Security

Prompt Injection Defense & LLM Security

Defending LLMs and AI Agents against indirect prompt injection, jailbreaks, and unauthorized tool execution.

Architecture Overview

Prompt injection occurs when untrusted user or document input manipulates the LLM into ignoring system instructions or executing unauthorized backend function calls.

Key Architecture Concepts

  • ❖Direct Injection (Jailbreaking): User explicitly instructs LLM to bypass safety guardrails.
  • ❖Indirect Injection: Untrusted web data, email, or database record fetched by RAG contains hidden override instructions.
  • ❖Privilege Separation: Separating system prompts, tool call permissions, and data retrieval layers.
  • ❖Input Sanitization & Output Guardrails: Filtering system tags and inspecting tool execution payloads.

Practical Engineering Takeaway

Never execute arbitrary shell commands or database write operations based on unvalidated LLM output.