Opening SecureFlowChatGPT Conversation Exfiltration
Case summary & sources

AIDEFEND SecureFlow / SecureFlow Case Index

MITRE ATLAS mitre-atlas-cs0021

ChatGPT Conversation Exfiltration

  • AI Infrastructure
  • Data Exfiltration

Embrace the Red (https://embracethered.com/blog/) demonstrated that ChatGPT users' conversations can be exfiltrated via an indirect prompt injection. To execute the attack, a threat actor uploads a malicious prompt to a public website, where a ChatGPT user may interact with it. The prompt causes ChatGPT to respond with the markdown for an image, whose URL has the user's conversation secretly embedded. ChatGPT renders the image for the user, creating an automatic request to an adversary-controlled script and exfiltrating the user's conversation. Additionally, the researcher demonstrated how the prompt can execute other plugins, opening them up to additional harms.

Mapped threat techniques

Source

Updated