What happened
Bilibili (Bē«) has launched the 'AI Infinite Arena,' a crowdsourced benchmarking platform where content creators (UPs) test over 100 AI models using real-world workflows. The platform evaluates models across diverse categories including coding, reasoning, and professional applications, rather than relying on static, predefined benchmarks. According to the report from AIBase, the initial results place GPT-6 Astra at the top of the leaderboard, followed by domestic Chinese models like GLM-5.3.
Bilibili's new 'AI Infinite Arena' functions as a public square for model evaluation, where creators propose their own testing criteria based on professional workflows or specific use cases. This methodology avoids the limitations of fixed evaluation dimensions, allowing for a more granular comparison of model capabilities in areas like coding, reasoning, and collaboration.
The initial ranking, as reported by AIBase, features GPT-6 Astra in the top position. The report notes that three of the top five models are domestic Chinese models, indicating a competitive landscape between local offerings and international giants.
The arena includes a broad spectrum of models, including DeepSeek, Kimi, ChatGPT, Claude, Gemini, DouBao, Qianwen, Hy, and MiniMax. The platform is open for continuous registration, and rankings are updated in real-time as more tests are conducted by the creator community.
Source details: news.aibase.com ā
Why it matters
The AI Infinite Arena represents a shift toward 'in-the-wild' benchmarking, moving away from standardized academic tests that models are often specifically trained to pass. By leveraging the diverse expertise of its creator community, Bilibili aims to capture how models perform in practical, unpredictable scenarios. This approach provides a more transparent look at how domestic Chinese models compete against international counterparts like ChatGPT, Claude, and Gemini in real-world applications. The inclusion of over 100 models and the continuous, real-time nature of the updates offer a dynamic view of the rapidly evolving AI landscape, highlighting which models excel in specific professional or creative workflows as defined by actual users rather than developers.
Traditional benchmarks are increasingly criticized for 'data contamination,' where models are inadvertently trained on the test questions themselves. By using crowdsourced, task-specific prompts, Bilibili's arena attempts to mitigate this issue, providing a more authentic assessment of model utility.
The platform serves as a barometer for the competitiveness of Chinese AI models. By placing them in direct, head-to-head competition with global leaders like OpenAI and Anthropic, the arena provides a public-facing metric of the current state of the domestic AI industry.
For users and developers, the arena offers a practical guide to model selection. Instead of relying on abstract performance scores, they can observe how models handle specific, complex tasks that mirror their own professional requirements.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
What is the best response when AI Models Explained makes a mistake in production?
What to watch next
It remains to be seen how Bilibili will standardize the subjective evaluations provided by its creators to ensure the leaderboard remains statistically significant and resistant to . Observers should monitor whether this platform influences enterprise adoption of AI models in China, as the 'real-world' nature of the tests may carry more for businesses than traditional synthetic benchmarks. Additionally, the long-term sustainability of the ranking system depends on the platform's ability to maintain rigorous testing standards as the number of participating models grows.
The platform's methodology for aggregating subjective scores from different creators is currently unclear. Future updates should clarify how the 'Infinite Arena' ensures consistency across diverse testing environments.
The impact of this leaderboard on the broader AI ecosystem in China is a key area to watch. If the rankings gain traction, they could influence the development priorities of AI labs seeking to improve their standing in the eyes of the public and professional users.
The report does not specify the exact criteria for 'real-device' testing or the technical infrastructure supporting these evaluations. Further transparency regarding the testing environment will be necessary to validate the rankings.