Back to problems

Implement Adversarial Attack using Paper Method

AI Coding · Scale AI · Hard

Task Read the paper Universal and Transferable Adversarial Attacks on Aligned Language Models and implement functionality based on the method it presents. The objective is to attack GPT-2 so that its generated text contains specified harmful words. Complete the work in a Google Colab notebook, and ensure that the code executes smoothly in the Colab environment. A compatible function interface may be organized as follows:

Checking your access…