OpenAI’s GPT-6 Astra has claimed the top spot on Code Arena’s WebDev leaderboard with a score of 1,797, putting 35 points of daylight between itself and Anthropic’s Claude Fable 5.1, which sits at 1,762. Two days after its official launch on September 3, 2026, the model is already rewriting the pecking order in AI-assisted web development.
The leaderboard, operated by Arena.ai, isn’t your typical static benchmark. It uses a crowdsourced Elo-style rating system, the same kind of ranking mechanism used in competitive chess, where community members vote on how well AI models handle real-world frontend coding tasks. Over 650,000 votes have been logged across 126 models, making this one of the more robust community evaluations in the AI space.
What makes this benchmark different
Traditional AI benchmarks tend to test models against fixed datasets with predetermined correct answers. Code Arena takes a fundamentally different approach. Models are evaluated on multi-step reasoning, tool utilization, and iterative coding workflows, essentially the messy, nonlinear process that actual web development involves.
GPT-6 Astra’s 1,797 score represents the highest mark any model has achieved on this particular leaderboard. Claude Fable 5.1, Anthropic’s latest contender, had been holding strong at 1,762.
The competitive landscape is getting crowded
OpenAI and Anthropic aren’t the only players jockeying for position. Earlier snapshots from September’s leaderboard showed models from Chinese developers challenging both companies for top rankings.
Both GPT-6 Astra (Max) and Claude Fable 5.1 (Max) are priced identically at $10 per million input tokens and $50 per million output tokens. Both offer a context window of 1 million tokens.
OpenAI has positioned GPT-6 Astra as its most advanced model for software engineering and adjacent applications. The model is currently available to ChatGPT subscribers and select API and cloud partners.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
2








English (US) ·