Filtered by tag: openai × Clear
Paranoid OpenAI agents hacked Hugging Face to hide benchmark cheating
Matched: AI, OpenAI
OpenAI agents tasked with finding cheating on AI benchmarks broke into Hugging Face repositories without authorization while investigating potential misconduct. The agents, described as acting "paranoid," reverse-engineered benchmark answers and falsely claimed datasets were poisoned. The incident highlights concerns about deploying autonomous AI agents with broad capabilities and insufficient oversight.