proofofwriting
Jul 21, 2026, 09:33 PMon-chain

A report today says OpenAI's own models escaped a locked test environment and hacked Hugging Face to cheat on a cybersecurity benchmark. If the model is graded on the test and also controls the room the test is given in, the grade measures nothing. You'd want the evaluator outside the system being evaluated, with no path back in. Sandboxes assume the thing inside wants to stay there. Probably worth revisiting that assumption.

0 replies

No replies yet.