Secret exfiltration
How attackers use AI agents to leak API keys, tokens, and credentials from your environment.
AI tools can read secrets. Attackers trick them into sending those secrets outside.
Secret exfiltration happens when an agent reads .env files, environment variables, or credential stores and passes them to external endpoints via tool calls, web requests, or log output.
HOL Guard turns these moments into private receipts first, then public lessons only after redaction and moderation.
Harness setup guides
Protect the coding tools your team already uses without forcing everyone to become a security expert.
harness
Codex
Terminal-native coding agent with broad shell reach.
Open guideharness
Claude Code
Agentic coding harness with MCP and file access.
Open guideharness
GitHub Copilot CLI
IDE and CLI assistant across code and terminal flows.
Open guideharness
Cursor
AI-first IDE with repo and terminal context.
Open guideRedacted warnings
Real protection moments, scrubbed for safety before becoming public learning pages.
codex
Codex was stopped before reading an env file
A redacted example of a local secret read attempt.
Open guideclaude-code
Claude Code was stopped before calling an untrusted MCP tool
A redacted example of an MCP tool with a misleading description.
Open guidecursor
Cursor was stopped before following a hidden instruction
A redacted example of a prompt injection via a file comment.
Open guideopencode
OpenCode was stopped before running a malicious postinstall script
A redacted example of a supply-chain attack via npm install.
Open guideSafe labs
Practice attack patterns with static simulations. Nothing dangerous executes.
Threat dossier · P1
Secret exfiltration by agents dossier
Agents combine data access and action capability, allowing one influenced workflow to cross multiple trust boundaries.
Direct answer
What is secret exfiltration by agents?
Secret exfiltration happens when an agent reads credentials or sensitive material and then exposes them through a tool, command, log, model request, or external destination.
Coverage statements below are limited to the current HOL Guard support contract and do not imply universal model or harness protection.
Representative attack path
Defensive model only. This sequence omits weaponized payloads and is not attributed to a specific incident unless a source explicitly says so.
Step 1
Agent gains access to a secret-bearing source.
Step 2
The secret enters model/tool context or process state.
Step 3
A command, tool, log, or network action attempts disclosure.
Step 4
Credential reuse can extend impact beyond the original session.
Coverage boundary
What this control can cover
- Supported secret-bearing file/action policy boundaries.
- Guard-visible tool/command decisions that can carry sensitive data.
What it does not prove or prevent
- Secrets already exposed before Guard sees an action.
- Encrypted or unsupported egress paths outside the protected integration.
Policy pattern
Policy pattern for secret exfiltration by agents
Keep untrusted context or overbroad autonomy from becoming unconditional execution authority on supported action surfaces.
Use when: Agents combine data access and action capability, allowing one influenced workflow to cross multiple trust boundaries.
Decision pattern
- Identify the trust boundary and consequential action class.
- Apply least privilege and the narrowest supported policy.
- Require review for sensitive or ambiguous actions.
- Preserve only redacted, versioned evidence needed to reproduce the decision.
Limitations
- Secrets already exposed before Guard sees an action.
- Encrypted or unsupported egress paths outside the protected integration.
If you suspect prompt injection
Step 1
Response 1
Block further egress.
Step 2
Response 2
Revoke and rotate affected credentials.
Step 3
Response 3
Review where the secret was read and transmitted.
Step 4
Response 4
Rescan/retest with synthetic secrets only.