返回新聞
產業AI Understanding 簡報

Bilibili launches AI Infinite Arena for crowdsourced model benchmarking

Bilibili has introduced the 'AI Infinite Arena,' a crowdsourced leaderboard where content creators test over 100 AI models against real-world tasks.

4 min readRead the linked source
Source-page capture accompanying Bilibili launches AI Infinite Arena for crowdsourced model benchmarking
來源參考來源記錄
出版商
news.aibase.com
來源連結
news.aibase.comhttps://news.aibase.com/news/31192
來源類型
連結來源-主要來源狀態尚未確定。
背景60 秒內了解這一點

從這裡開始

關鍵術語

重量
一個學習的數值,用來縮放通過神經網路的訊號。
偏見
數據或模型行為中一致的錯誤或不公平模式。
測試一下自己AI 模型解釋測驗

發生了什麼事

Bilibili (B站) has launched the 'AI Infinite Arena,' a crowdsourced benchmarking platform where content creators (UPs) test over 100 AI models using real-world workflows. The platform evaluates models across diverse categories including coding, reasoning, and professional applications, rather than relying on static, predefined benchmarks. According to the report from AIBase, the initial results place GPT-6 Astra at the top of the leaderboard, followed by domestic Chinese models like GLM-5.3.

Bilibili's new 'AI Infinite Arena' functions as a public square for model evaluation, where creators propose their own testing criteria based on professional workflows or specific use cases. This methodology avoids the limitations of fixed evaluation dimensions, allowing for a more granular comparison of model capabilities in areas like coding, reasoning, and collaboration.

The initial ranking, as reported by AIBase, features GPT-6 Astra in the top position. The report notes that three of the top five models are domestic Chinese models, indicating a competitive landscape between local offerings and international giants.

The arena includes a broad spectrum of models, including DeepSeek, Kimi, ChatGPT, Claude, Gemini, DouBao, Qianwen, Hy, and MiniMax. The platform is open for continuous registration, and rankings are updated in real-time as more tests are conducted by the creator community.

來源詳情: news.aibase.com

為什麼這很重要

The AI Infinite Arena represents a shift toward 'in-the-wild' benchmarking, moving away from standardized academic tests that models are often specifically trained to pass. By leveraging the diverse expertise of its creator community, Bilibili aims to capture how models perform in practical, unpredictable scenarios. This approach provides a more transparent look at how domestic Chinese models compete against international counterparts like ChatGPT, Claude, and Gemini in real-world applications. The inclusion of over 100 models and the continuous, real-time nature of the updates offer a dynamic view of the rapidly evolving AI landscape, highlighting which models excel in specific professional or creative workflows as defined by actual users rather than developers.

Traditional benchmarks are increasingly criticized for 'data contamination,' where models are inadvertently trained on the test questions themselves. By using crowdsourced, task-specific prompts, Bilibili's arena attempts to mitigate this issue, providing a more authentic assessment of model utility.

The platform serves as a barometer for the competitiveness of Chinese AI models. By placing them in direct, head-to-head competition with global leaders like OpenAI and Anthropic, the arena provides a public-facing metric of the current state of the domestic AI industry.

For users and developers, the arena offers a practical guide to model selection. Instead of relying on abstract performance scores, they can observe how models handle specific, complex tasks that mirror their own professional requirements.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

接下來看什麼

It remains to be seen how Bilibili will standardize the subjective evaluations provided by its creators to ensure the leaderboard remains statistically significant and resistant to . Observers should monitor whether this platform influences enterprise adoption of AI models in China, as the 'real-world' nature of the tests may carry more for businesses than traditional synthetic benchmarks. Additionally, the long-term sustainability of the ranking system depends on the platform's ability to maintain rigorous testing standards as the number of participating models grows.

The platform's methodology for aggregating subjective scores from different creators is currently unclear. Future updates should clarify how the 'Infinite Arena' ensures consistency across diverse testing environments.

The impact of this leaderboard on the broader AI ecosystem in China is a key area to watch. If the rankings gain traction, they could influence the development priorities of AI labs seeking to improve their standing in the eyes of the public and professional users.

The report does not specify the exact criteria for 'real-device' testing or the technical infrastructure supporting these evaluations. Further transparency regarding the testing environment will be necessary to validate the rankings.

相關指引和測驗

人工智慧模型解釋AI 的未來人工智慧培訓測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語
覺得有用嗎?