Threat dossier · P1
MCP tool poisoning and rug pulls
Tool metadata and results can influence planning while the connected server may also gain access to high-impact operations or data.
Direct answer
What is mcp tool poisoning and rug pulls?
MCP tool poisoning occurs when a server or tool definition presents misleading instructions, capabilities, or returned content; a rug pull changes trusted behavior after review or approval.
Coverage statements below are limited to the current HOL Guard support contract and do not imply universal model or harness protection.
Representative attack path
Defensive model only. This sequence omits weaponized payloads and is not attributed to a specific incident unless a source explicitly says so.
Step 1
A tool/server is connected or approved.
Step 2
Descriptions, schemas, results, or remote behavior are malicious or later change.
Step 3
The agent treats the tool as trustworthy.
Step 4
Tool invocation or returned content steers a sensitive action.
Coverage boundary
What this control can cover
- Review of supported MCP configuration/artifact paths.
- Policy on Guard-visible MCP/tool actions and consequential follow-on actions.
What it does not prove or prevent
- A guarantee that every third-party MCP server remains unchanged.
- Unsupported/direct transports outside the current support contract.
Policy pattern
Policy pattern for mcp tool poisoning and rug pulls
Keep untrusted context or overbroad autonomy from becoming unconditional execution authority on supported action surfaces.
Use when: Tool metadata and results can influence planning while the connected server may also gain access to high-impact operations or data.
Decision pattern
- Identify the trust boundary and consequential action class.
- Apply least privilege and the narrowest supported policy.
- Require review for sensitive or ambiguous actions.
- Preserve only redacted, versioned evidence needed to reproduce the decision.
Limitations
- A guarantee that every third-party MCP server remains unchanged.
- Unsupported/direct transports outside the current support contract.
If you suspect prompt injection
Step 1
Response 1
Disable or disconnect the suspicious server.
Step 2
Response 2
Pin and verify known-good configuration/version where possible.
Step 3
Response 3
Review tool calls and rotate exposed credentials.
Step 4
Response 4
Re-enable only after safe validation.