Guidelight AI Standards, a nonprofit focused on frontier AI safety, graded five leading labs on how prepared each is to contain a rogue model. Anthropic, OpenAI, Google, Meta and xAI were scored using only what each company has made public, not on internal safeguards it might not disclose.

The assessment scored each lab on six practices: logging of internal AI activity, monitoring effectiveness, requiring human sign-off before high-risk actions, halting systems after a surge of flagged misbehavior, third-party review, and having an actual containment plan. No lab scored above what Guidelight calls substantial partial implementation on any single practice, according to its published report.

On overall scores out of 5, Anthropic and OpenAI tied for highest at 2.50, followed by Google at 1.50, xAI at 0.83 and Meta at 0.67, Guidelight's report shows. On the specific containment-plan measure, TechCrunch reported OpenAI scored highest at 3 out of 5, citing the company's history of pausing workloads after safety incidents.

Guidelight chief scientist Steven Adler, a former OpenAI safety researcher, told TechCrunch: "I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense." Connor Leahy of the nonprofit ControlAI said, "A kill switch is the bare minimum for today's models."

Guidelight said low scores reflect a lack of public disclosure, not necessarily a lack of internal safeguards. Still, for anyone building on top of these systems, the plan for what happens if a frontier model does something it should not currently lives, publicly, almost nowhere.