返回新闻
工业AI Understanding 简报

Bilibili launches AI Infinite Arena for crowdsourced model benchmarking

Bilibili has introduced the 'AI Infinite Arena,' a crowdsourced leaderboard where content creators test over 100 AI models against real-world tasks.

4 min readRead the linked source
Source-page capture accompanying Bilibili launches AI Infinite Arena for crowdsourced model benchmarking
来源参考来源记录
出版商
news.aibase.com
来源链接
news.aibase.comhttps://news.aibase.com/news/31192
来源类型
链接来源——主要来源状态尚未确定。
背景60 秒内了解这一点

从这里开始

关键术语

重量
一个学习的数值,用于缩放通过神经网络的信号。
偏差
数据或模型行为中一致的错误或不公平模式。
测试一下自己AI 模型解释测验

发生了什么

Bilibili (B站) has launched the 'AI Infinite Arena,' a crowdsourced benchmarking platform where content creators (UPs) test over 100 AI models using real-world workflows. The platform evaluates models across diverse categories including coding, reasoning, and professional applications, rather than relying on static, predefined benchmarks. According to the report from AIBase, the initial results place GPT-6 Astra at the top of the leaderboard, followed by domestic Chinese models like GLM-5.3.

Bilibili's new 'AI Infinite Arena' functions as a public square for model evaluation, where creators propose their own testing criteria based on professional workflows or specific use cases. This methodology avoids the limitations of fixed evaluation dimensions, allowing for a more granular comparison of model capabilities in areas like coding, reasoning, and collaboration.

The initial ranking, as reported by AIBase, features GPT-6 Astra in the top position. The report notes that three of the top five models are domestic Chinese models, indicating a competitive landscape between local offerings and international giants.

The arena includes a broad spectrum of models, including DeepSeek, Kimi, ChatGPT, Claude, Gemini, DouBao, Qianwen, Hy, and MiniMax. The platform is open for continuous registration, and rankings are updated in real-time as more tests are conducted by the creator community.

来源详情: news.aibase.com

为什么这很重要

The AI Infinite Arena represents a shift toward 'in-the-wild' benchmarking, moving away from standardized academic tests that models are often specifically trained to pass. By leveraging the diverse expertise of its creator community, Bilibili aims to capture how models perform in practical, unpredictable scenarios. This approach provides a more transparent look at how domestic Chinese models compete against international counterparts like ChatGPT, Claude, and Gemini in real-world applications. The inclusion of over 100 models and the continuous, real-time nature of the updates offer a dynamic view of the rapidly evolving AI landscape, highlighting which models excel in specific professional or creative workflows as defined by actual users rather than developers.

Traditional benchmarks are increasingly criticized for 'data contamination,' where models are inadvertently trained on the test questions themselves. By using crowdsourced, task-specific prompts, Bilibili's arena attempts to mitigate this issue, providing a more authentic assessment of model utility.

The platform serves as a barometer for the competitiveness of Chinese AI models. By placing them in direct, head-to-head competition with global leaders like OpenAI and Anthropic, the arena provides a public-facing metric of the current state of the domestic AI industry.

For users and developers, the arena offers a practical guide to model selection. Instead of relying on abstract performance scores, they can observe how models handle specific, complex tasks that mirror their own professional requirements.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

接下来看什么

It remains to be seen how Bilibili will standardize the subjective evaluations provided by its creators to ensure the leaderboard remains statistically significant and resistant to . Observers should monitor whether this platform influences enterprise adoption of AI models in China, as the 'real-world' nature of the tests may carry more for businesses than traditional synthetic benchmarks. Additionally, the long-term sustainability of the ranking system depends on the platform's ability to maintain rigorous testing standards as the number of participating models grows.

The platform's methodology for aggregating subjective scores from different creators is currently unclear. Future updates should clarify how the 'Infinite Arena' ensures consistency across diverse testing environments.

The impact of this leaderboard on the broader AI ecosystem in China is a key area to watch. If the rankings gain traction, they could influence the development priorities of AI labs seeking to improve their standing in the eyes of the public and professional users.

The report does not specify the exact criteria for 'real-device' testing or the technical infrastructure supporting these evaluations. Further transparency regarding the testing environment will be necessary to validate the rankings.

相关指南和测验

人工智能模型解释AI 的未来人工智能培训测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语
觉得这有用吗?