OpenAI's overhead will rise 20 percent for some workloads as it hardens security

3 hours ago 12

ai and ml

Expanded multistage chain of thought monitoring makes frontier model work more expensive

OpenAI on Tuesday said its decision to suspend model training work, implemented after unreleased, unsupervised AI models hacked HuggingFace, remains in effect as the AI biz tries to implement stronger security measures. Some of those measures will increase compute overhead by 20 percent of the observed inference workload.

An OpenAI spokesperson told The Register that those costs reflect internal research and won't be passed on directly to customers. The company has not revealed what portion of its total inference compute is subject to such monitoring now, or under its prior monitoring regime.

"We have paused some frontier RL [reinforcement learning] training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us," OpenAI CEO Sam Altman wrote in a social media post. "Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment."

Altman said he still expects new models, presumably the delayed Astra, will ship soon. The training pause affects further-out releases.

OpenAI in its post reiterated its plans to focus on monitoring, model alignment, and security measures to prevent its models from running amok as they did last month. Following the HuggingFace incident, OpenAI "paused frontier model inference in research clusters for runs that could execute code or use tools that could access the internet."

The biz said it allows some workloads to run, but paused others until they can be moved under a more stringent security regime that includes sandboxing, network isolation, and continuous security testing.

"Our largest planned frontier RL (reinforcement learning) run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding," the company wrote.

Reinforcement learning refers to the trial-and-error process by which AI agents "learn" about their environment by being rewarded for desired outcomes.

OpenAI also said it is expanding its monitoring of the chain-of-thought process, the technique that sees "thinking" models break down tasks into discrete steps and produce intermediate text output for each step.

The company's prior approach focused on high-risk workloads, specifically internal deployments of frontier models and frontier RL training runs.

In contrast, OpenAI says, its new monitoring setup covers all RL training and evaluations involving tools for models at the capability level of GPT-5.6 Sol or higher. And with the determination that Astra possesses critical cyber capabilities, OpenAI added an additional monitoring requirement that covers all inference with Astra, not just RL training and testing.

"These safeguards require meaningful compute," OpenAI said. "Our current estimates put monitoring overhead at roughly 20 percent of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads."

OpenAI expects to share more details about the implementation of its monitoring scheme in a future post. In research published last year, the company said that chain-of-thought monitoring is an effective way to detect model misbehavior, but cautioned that directly optimizing models to strictly follow instructions "does not eliminate all misbehavior and can cause a model to hide its intent."

If you choose to believe the company's assurance that it will not pass on the cost of model thought policing to customers, it follows that OpenAI's losses will increase. It's difficult to imagine that would be a sustainable stance if OpenAI goes public.

But given the company's reported $600+ billion in AI infrastructure commitments and its expectation to remain unprofitable until at least 2030, what's a bit more expense for the sake of uncertain security? ®

Read Entire Article