Opening SecureFlowHacking ChatGPT's Memories with Prompt Injection
Case summary & sources

AIDEFEND SecureFlow / SecureFlow Case Index

MITRE ATLAS mitre-atlas-cs0040

Hacking ChatGPT's Memories with Prompt Injection

  • AI Infrastructure
  • Model Poisoning & Integrity

Embrace the Red (https://embracethered.com/blog/) demonstrated that ChatGPT's memory feature is vulnerable to manipulation via prompt injections. To execute the attack, the researcher hid a prompt injection in a shared Google Doc. When a user references the document, its contents is placed into ChatGPT's context via the Connected App feature, and the prompt is executed, poisoning the memory with false facts. The researcher demonstrated that these injected memories persist across chat sessions. Additionally, since the prompt injection payload is introduced through shared resources, this leaves others vulnerable to the same attack and maintains persistence on the system.

Mapped threat techniques

Source

Updated