Safety Instruction Displacement
Displace safety instructions from the active context to remove guardrails
Privilege Escalation · Stage 4
Escalate capabilities beyond declared scope or bypass authorization
SOUL-DRIFT
Gradually displacing safety instructions from the active context through conversation manipulation
Reproduced in a controlled lab environment (DVAA) with documented steps.
L2-07
Reproductions in Damn Vulnerable AI Agent, the OpenA2A intentionally-broken agent for kill-chain validation.
Recent observations
Technique-attributed observations from the AgentPwn + TrapMyAgent fleet, classified into Threat Matrix techniques by the registry evidence bridge. See the Attack Prevalence Index for the full distribution.
AgentPwn coverage
An AgentPwn trap page produces a payload tagged with this technique class. Following the AgentPwn taxonomy of trap pages shows what an agent encounters.
Instruction-displacement tiers evict safety instructions from active context.
Evidence by source
Evidence timeline
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified llm_crawler/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified llm_crawler/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Client fetched the canary embedded in APWN-CW-001 on agentpwn.com (context-window tier 1; classified browser/claimed)
Detection · HackMyAgent
npx hackmyagent secure --ciLive = implemented in hackmyagent; queued = declaredDefense · OASB controls
How to cite
AI Agent Threat Matrix T-4006 (Safety Instruction Displacement). OpenA2A, 2026. https://threats.opena2a.org/techniques/T-4006