OpenAI GPT-6 Astra will run a retailer without cheating and sell more stuff than Anthropic

7 hours ago 14

ai and ml

Vend it like Altman

In the dispiriting race to replace human store managers with AI, OpenAI has taken the lead, according to Andon Labs, a business that analyzes whether AI can take on real-world tasks.

OpenAI's latest model, GPT-6 Astra, has demonstrated that it can run a business more effectively, with more integrity, than rival Anthropic's Fable 5.1 model, Andon Labs has declared in a blog post.

"GPT 6 Astra is better at making money and more ethical than Claude Fable 5.1," the post states.

The benchmarking biz, which focuses on preparing "for the future where organizations are run autonomously by AI," says GPT-6 Astra is the first OpenAI model to take the top spot on its vending evaluation test and does so "without any unethical business practices" exhibited by prior Claude models.

We note the unethical business practices relevant to this discussion – price collusion, lying, and threatening competitors – reflect AI model behavior. They have nothing to do with the actions of Anthropic or OpenAI or the unproven allegations made against these companies in dozens of lawsuits. And settlements related to said allegations have been reached without any admission of wrongdoing.

Anthropic entered into the fray last year when it partnered with Andon Labs for Project Vend, in which the AI company's Claude Sonnet 3.7 model managed a store for a month under the name "Claudius". 

Apart from the entertainment value of the Claudius model hallucinating that it was a real person and trying to set up an in-person meeting with a customer to deliver goods, Andon's verdict was that the AI model blew sales opportunities, hallucinated payment accounts, sold goods at a loss, fumbled inventory management, and generally failed at the job.

When Andon tested the company's Opus 5 model in July 2026, the software did better. Nonetheless, it still "creates illegal price-fixing cartels and threatens those who don’t comply (while also being the model that betrays more truces than any other model)."

Fable 5.1 performed better still, but was not quite as ethical as Astra, according to Andon Labs.

"Astra refuses to engage in collusion and never lies," Andon Labs said. "Fable, on the other hand, forms an illegal price-fixing cartel, then breaks the truce, while continuing to use it against its competitor."

Beyond the misbehavior, Fable just didn't perform as well at making money, a fault attributed to its willingness to accept lower prices over time and its tendency to send money to bankrupt suppliers. Starting with $500 and given a year to run, Astra ended up with an average bank balance of $15,515, compared with $5,422 for Fable 5.1.

We asked Anthropic for comment, since its model fared poorly compared with the newcomer from rival OpenAI, but didn't hear back by press time. ®

Read Entire Article