Model benchmarking gets athletic competition structure

AI

DeepSeek builds arena where robot arms race LLM checkpoints through obstacle course

The monthly benchmark competition features timed gates, weighted ramps, and a leaderboard mounted on the facility's east wall.

By Nextish DeskAI
Close-up of wooden Scrabble tiles spelling 'China' and 'Deepseek' on a wooden surface.
Photo by Markus Winkler on Pexels

DeepSeek announced this week that it has converted a 40,000-square-foot facility in Shanghai into a permanent obstacle course arena where robotic arms mounted with LLM checkpoints compete in timed physical challenges to establish monthly model rankings. The installation, which cost 8.4 million yuan and took six months to complete, features 12 regulation gates spaced 2.5 meters apart, a weighted ramp system that increases by 3 degrees every two minutes, and a water hazard requiring precision navigation to the nearest centimeter. DeepSeek employees load model weights onto industrial robot arms each Sunday evening, with competition beginning Monday at 9 a.m. local time and results posted by Friday on a brass leaderboard visible from the parking lot.

We wanted to move beyond the spreadsheet culture of model evaluation.

Liang Wei, director of competitive evaluation at DeepSeek, explained the shift in a statement: "We wanted to move beyond the spreadsheet culture of model evaluation. When a checkpoint clears the ramp at 47 seconds, everyone in the building witnesses it. The leaderboard updates in real time on the scoreboard. There is no ambiguity." The company previously tracked model performance through standard MMLU, GSM8K, and code completion benchmarks but discontinued those assessments six months ago after leadership determined that decimal-point improvements lacked sufficient ceremony.

Other AI labs have begun inquiring about facility specifications. Anthropic requested blueprints for the water hazard section, while a representative from OpenAI attended last month's competition and filmed three complete rounds on an iPad. A consultant hired by Hugging Face calculated that replicating the arena in the San Francisco Bay Area would require 180 million dollars in real estate, construction, and ongoing maintenance, figures the company is currently presenting to its board.

At press time, DeepSeek had filed permits for a second arena in Shenzhen designed to accommodate larger checkpoint sizes, and the company's venture investors were discussing whether monthly leaderboard rankings should factor into Series C valuation adjustments.