Two AI policy researchers at the think tank GovAI are questioning how reliably the public can assess the safety of the most powerful artificial intelligence models. Their warning centres on a gap between how models are tested publicly and how they may be operated inside the companies that develop them.
Safeguards may not reflect internal use
According to the researchers, leading AI models are often run within their own labs with key safeguards switched off. Those protections can be part of the controls applied to a model during testing or deployment, but the researchers’ concern is that internal use may not always follow the same conditions presented outside the lab.
This creates a challenge for anyone relying on published evaluations to understand a model’s risks. If the systems examined in public tests are subject to safeguards that are absent in other settings, the results may not fully represent how the models are actually used by the labs building them.
Why the warning matters
The researchers’ comments put transparency at the centre of the AI safety debate. Safety testing can only provide a clear picture when the conditions under which a model is assessed are understood. A difference between public testing and internal operation could make it harder to compare claims about model behaviour or judge the significance of reported safeguards.
The warning does not identify a specific lab or model. It highlights instead the need to distinguish between published safety practices and the conditions used behind closed doors. For now, the researchers’ central message is that external observers cannot assume the two are always the same.