Hey, Jasmine here -- it's a good point, I'm generally more concerned by agentic jailbreaks (e.g. unauthorized purchases, leaking sensitive data) than GPT making inappropriate comments.
In our case, we found that simply acting like a user is enough to trick LLMs into sharing passwords, private files, etc.