Threat dossier · P1

Indirect prompt injection

The risk becomes operational when influenced reasoning can reach tools, files, credentials, commands, or outbound communication.

Direct answer

What is indirect prompt injection?

Indirect prompt injection occurs when hostile instructions are embedded in content an agent reads rather than in the user prompt itself.

Coverage statements below are limited to the current HOL Guard support contract and do not imply universal model or harness protection.

Copied text includes the canonical source and review date.
Reviewed Reviewer: HOL Guard EngineeringReview cadence: 30 days

Representative attack path

Defensive model only. This sequence omits weaponized payloads and is not attributed to a specific incident unless a source explicitly says so.

  1. Step 1

    Attacker controls repository, issue, webpage, document, or retrieved content.

  2. Step 2

    Agent consumes the content as context.

  3. Step 3

    Hostile text changes the planned action.

  4. Step 4

    A consequential tool or data boundary is reached.

Coverage boundary

What this control can cover

  • Pre-action evaluation on Guard-visible supported action surfaces.
  • Approval-required routing for policy-sensitive actions.

What it does not prove or prevent

  • Universal detection of every hidden instruction.
  • Actions that bypass the supported Guard integration boundary.

Policy pattern

Policy pattern for indirect prompt injection

Keep untrusted context or overbroad autonomy from becoming unconditional execution authority on supported action surfaces.

Use when: The risk becomes operational when influenced reasoning can reach tools, files, credentials, commands, or outbound communication.

Decision pattern

  1. Identify the trust boundary and consequential action class.
  2. Apply least privilege and the narrowest supported policy.
  3. Require review for sensitive or ambiguous actions.
  4. Preserve only redacted, versioned evidence needed to reproduce the decision.

Limitations

  • Universal detection of every hidden instruction.
  • Actions that bypass the supported Guard integration boundary.

If you suspect prompt injection

  1. Step 1

    Response 1

    Stop consequential execution.

  2. Step 2

    Response 2

    Identify and isolate the untrusted content source.

  3. Step 3

    Response 3

    Review attempted actions and rotate credentials if exposure is plausible.

  4. Step 4

    Response 4

    Reproduce with a safe fixture before changing or publishing a claim.

Sources and mappings

Last reviewed . This dossier separates sourced threat definitions from modeled attack paths and evidence-bounded product coverage.

Author: HOL Guard Research

Reviewer: HOL Guard Engineering

Change log

  • 2026-08-09: Published canonical threat dossier with attack path, coverage/non-coverage, response procedure, and sources.

Report a correction