Opening SecureFlowPoisonGPT
Case summary & sources

AIDEFEND SecureFlow / SecureFlow Case Index

MITRE ATLAS mitre-atlas-cs0019

PoisonGPT

  • AI Supply Chain
  • Model Poisoning & Integrity

Researchers from Mithril Security demonstrated how to poison an open-source pre-trained large language model (LLM) to return a false fact. They then successfully uploaded the poisoned model back to HuggingFace, the largest publicly-accessible model hub, to illustrate the vulnerability of the LLM supply chain. Users could have downloaded the poisoned model, receiving and spreading poisoned data and misinformation, causing many potential harms.

Mapped threat techniques

Source

Updated