> And OpenAI directly addressed these plagiarism claims, and called them impossible
Funny, you were telling me two days ago that on the contrary, "it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set":
First of all, who can say for certain whether OpenAI does what they say they do? For all we know, they cracked open this specific researcher's prompts and started from there.
Second, the issue of anonymization is a red herring. There is a very limited number of people working in this approach, and most of them are likely making no progress. So Buckmaster's prompts might have had an outsized effect on the outcome. It's similar to that guy who created a site claiming he is a world-renowmed hot dog eating contestant, which ended up digested by OpenAI models as truth [1].
Funny, you were telling me two days ago that on the contrary, "it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set":
https://news.ycombinator.com/item?id=49621648