Quay lại Tin tức
Công nghiệpAI Understanding tóm tắt

Bilibili launches AI Infinite Arena for crowdsourced model benchmarking

Bilibili has introduced the 'AI Infinite Arena,' a crowdsourced leaderboard where content creators test over 100 AI models against real-world tasks.

4 min readRead the linked source
Source-page capture accompanying Bilibili launches AI Infinite Arena for crowdsourced model benchmarking
Nguồn tham khảoNguồn đã ghi
Nhà xuất bản
news.aibase.com
Liên kết nguồn
news.aibase.comhttps://news.aibase.com/news/31192
Loại nguồn
Nguồn được liên kết - trạng thái nguồn chính chưa được thiết lập.
Bối cảnhHiểu điều này trong 60 giây

Bắt đầu ở đây

Thuật ngữ chính

cân nặng
Một giá trị số đã học để chia tỷ lệ các tín hiệu truyền qua mạng nơ-ron.
thiên vị
Một dạng lỗi hoặc sự không công bằng nhất quán trong dữ liệu hoặc hành vi của mô hình.
Tự kiểm traCâu đố giải thích về mô hình AI

Chuyện gì đã xảy ra

Bilibili (B站) has launched the 'AI Infinite Arena,' a crowdsourced benchmarking platform where content creators (UPs) test over 100 AI models using real-world workflows. The platform evaluates models across diverse categories including coding, reasoning, and professional applications, rather than relying on static, predefined benchmarks. According to the report from AIBase, the initial results place GPT-6 Astra at the top of the leaderboard, followed by domestic Chinese models like GLM-5.3.

Bilibili's new 'AI Infinite Arena' functions as a public square for model evaluation, where creators propose their own testing criteria based on professional workflows or specific use cases. This methodology avoids the limitations of fixed evaluation dimensions, allowing for a more granular comparison of model capabilities in areas like coding, reasoning, and collaboration.

The initial ranking, as reported by AIBase, features GPT-6 Astra in the top position. The report notes that three of the top five models are domestic Chinese models, indicating a competitive landscape between local offerings and international giants.

The arena includes a broad spectrum of models, including DeepSeek, Kimi, ChatGPT, Claude, Gemini, DouBao, Qianwen, Hy, and MiniMax. The platform is open for continuous registration, and rankings are updated in real-time as more tests are conducted by the creator community.

Chi tiết nguồn: news.aibase.com

Tại sao nó quan trọng

The AI Infinite Arena represents a shift toward 'in-the-wild' benchmarking, moving away from standardized academic tests that models are often specifically trained to pass. By leveraging the diverse expertise of its creator community, Bilibili aims to capture how models perform in practical, unpredictable scenarios. This approach provides a more transparent look at how domestic Chinese models compete against international counterparts like ChatGPT, Claude, and Gemini in real-world applications. The inclusion of over 100 models and the continuous, real-time nature of the updates offer a dynamic view of the rapidly evolving AI landscape, highlighting which models excel in specific professional or creative workflows as defined by actual users rather than developers.

Traditional benchmarks are increasingly criticized for 'data contamination,' where models are inadvertently trained on the test questions themselves. By using crowdsourced, task-specific prompts, Bilibili's arena attempts to mitigate this issue, providing a more authentic assessment of model utility.

The platform serves as a barometer for the competitiveness of Chinese AI models. By placing them in direct, head-to-head competition with global leaders like OpenAI and Anthropic, the arena provides a public-facing metric of the current state of the domestic AI industry.

For users and developers, the arena offers a practical guide to model selection. Instead of relying on abstract performance scores, they can observe how models handle specific, complex tasks that mirror their own professional requirements.

Interactive Mechanism

Cơ chế tương tác: Nó thực sự hoạt động như thế nào

Khám phá công nghệ cơ bản đằng sau sự phát triển này một cách tương tác.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Kiểm tra khái niệm tương tác+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

Xem gì tiếp theo

It remains to be seen how Bilibili will standardize the subjective evaluations provided by its creators to ensure the leaderboard remains statistically significant and resistant to . Observers should monitor whether this platform influences enterprise adoption of AI models in China, as the 'real-world' nature of the tests may carry more for businesses than traditional synthetic benchmarks. Additionally, the long-term sustainability of the ranking system depends on the platform's ability to maintain rigorous testing standards as the number of participating models grows.

The platform's methodology for aggregating subjective scores from different creators is currently unclear. Future updates should clarify how the 'Infinite Arena' ensures consistency across diverse testing environments.

The impact of this leaderboard on the broader AI ecosystem in China is a key area to watch. If the rankings gain traction, they could influence the development priorities of AI labs seeking to improve their standing in the eyes of the public and professional users.

The report does not specify the exact criteria for 'real-device' testing or the technical infrastructure supporting these evaluations. Further transparency regarding the testing environment will be necessary to validate the rankings.

Hướng dẫn và câu hỏi liên quan

Giải thích về mô hình AITương lai của AIĐào tạo AIKiểm tra những gì bạn biết — thử một bài kiểm tra AI miễn phíTra cứu một thuật ngữ AI trong bảng thuật ngữ của chúng tôi
Tìm thấy điều này hữu ích?