返回新聞
產品展示AI Understanding 簡報

OpenArt 推出 Arena 進行特定任務的 AI 創意模型排名

OpenArt 推出了一個名為 Arena 的新公共基準,根據電影製作、電子商務和口型同步等特定創意任務,使用創意專業人士的盲目成對評估對 AI 影像和視訊模型進行排名。

5 min readRead the original reporting
Source-provided image accompanying OpenArt launches Arena for task-specific AI creative model rankings
歸因報告來源記錄
出版商
venturebeat.com
來源連結
venturebeat.comhttps://venturebeat.com/orchestration/whats-the-best-ai-model-for-graphic-design-video-ads-lip-sync-and-more-openarts-new-arena-offers-leaderboards-for-different-media-jobs
來源類型
新聞媒體的報道-不是第一方文件。

我們無法獨立確認的內容: 此聲明歸因於指定的商店。我們沒有根據第一方文件對其進行驗證。 (venturebeat.com)

背景60 秒內了解這一點

從這裡開始

關鍵術語

API(應用程式介面)
一種軟體系統向另一個系統發送請求並接收回應的結構化方式。
生成式 AI
產生文字、圖像、音訊、視訊或程式碼等新內容的人工智慧系統。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
測試一下自己AI 模型解釋測驗

發生了什麼事

OpenArt, a San Francisco-based creative AI tooling startup, launched OpenArt Arena, a public benchmarking platform designed to help users select the best AI model for specific multimedia jobs rather than relying on a single overall score. The platform divides model rankings into use-case boards for tasks including filmmaking, e-commerce, graphic design, motion design, video editing, and lip sync. It utilizes blind pairwise evaluations from a pool of creative practitioners and 'tastemakers,' aggregating preferences using the Bradley-Terry statistical model. Initial results show ByteDance’s Seedance 2.5 leading the overall video board, while Alibaba’s Wan 3.0 ranks first in video editing. For image generation, OpenAI’s GPT Image 2 leads in graphic design and editing, while Alibaba’s Seedream 5.0 Pro tops the overall image board. The company plans to publish part of its methodology and prompt sets while keeping active evaluation prompts private to prevent model tuning.

OpenArt, co-founded by former Googlers Coco Mao and John Qiaois, launched OpenArt Arena to answer the question of which AI model is best for specific media jobs. The platform features separate leaderboards for image and video generation, categorized by use cases such as filmmaking, e-commerce, graphic design, motion design, video editing, and lip sync.

The benchmarking process involves generating outputs from curated prompt sets for each board and presenting them to evaluators without model labels. Judges make side-by-side choices, and OpenArt aggregates these preferences using the Bradley-Terry statistical model. The judging pool includes a Creative Expert Council with named practitioners and a larger group of approximately 800 to 1,000 'tastemakers' recruited from users and creative communities.

In the initial video rankings, ByteDance’s proprietary Seedance 2.5 leads the overall board with a score of 1,081, followed by Alibaba’s Wan 3.0 at 1,004 and ByteDance’s Seedance 2.0 at 1,000. Seedance 2.5 also ranks first in film, motion design, and lip sync, while Wan 3.0 narrowly leads in video editing with a score of 1,034 compared to Seedance 2.5’s 1,033.

For image generation, the results are more fragmented. OpenAI’s GPT Image 2 ranks first in graphic design and image editing, while Alibaba’s Seedream 5.0 Pro leads in film-oriented imagery and e-commerce. Seedream 5.0 Pro also tops the overall image board with a score of 1,010, slightly ahead of GPT Image 2. Notably, Alibaba’s Wan 3.0 is the only prominent top-three entrant described as open source, though its current public release does not yet include downloadable model weights for self-hosting.

OpenArt stated it will publish part of its methodology and prompt set at launch but will keep some active evaluation prompts private to reduce the risk of models being tuned to the . The company aims for Arena to become a recognized industry standard for selecting AI models for creative work.

來源詳情: venturebeat.com ↗

為什麼這很重要

This launch addresses a critical pain point for enterprises and creative teams: the difficulty of selecting the optimal AI model among a rapidly expanding and fragmented market. By segmenting performance by specific job function, OpenArt Arena provides more actionable guidance than generic leaderboards, which may obscure a model's strengths in niche areas. This is particularly relevant for businesses adopting , where model selection is often the first major hurdle. The also highlights the trade-offs between proprietary API-based models and open-source options, influencing decisions on data residency, cost, and deployment control. While similar benchmarks exist, OpenArt’s focus on granular, task-specific creative workflows offers a distinct value proposition for professional creative production.

The launch of OpenArt Arena provides a more nuanced approach to evaluating AI creative models by focusing on specific job functions rather than a single overall score. This is significant for enterprises and creative teams that need to route specific production tasks to the most capable model, as a model strong in one area may not be the best in another.

The highlights the practical implications of model selection, including the trade-offs between proprietary API-based models and open-source options. For example, while ByteDance’s Seedance 2.5 leads in video performance, it is a proprietary model, whereas Alibaba’s Wan 3.0 is open source but lacks downloadable weights for self-hosting. This distinction affects deployment control, data residency, and total cost of ownership for businesses.

OpenArt Arena addresses a common challenge for companies adopting : the difficulty of choosing the right model from a rapidly expanding market. By providing task-specific rankings, the platform offers actionable guidance that can help creative teams make informed decisions and potentially improve the efficiency and quality of their AI-assisted production workflows.

The also contributes to the broader conversation about the standardization of AI evaluation. While other platforms like Contra Labs and Artificial Analysis offer similar task-specific rankings, OpenArt’s focus on granular creative workflows and its integration with a commercial creative platform may give it a competitive edge in becoming a widely recognized industry standard.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Model Parameter Size:8B Parameters
VRAM Required5.5 GBGPU memory footprint
Target HardwareMacBook / Single GPUDeployment tier
Privacy100% Air-GappedLocal device capability
Core takeaway: Small, quantized models (3B–8B) now run directly inside smartphones and laptops with complete data privacy, while mammoth 400B+ models remain the domain of datacenter clusters.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

Monitor how OpenArt Arena’s rankings evolve as new models are released and whether the platform becomes a de facto industry standard for creative AI procurement. Watch for the publication of the full methodology and prompt sets, which will allow for independent verification of the results. Additionally, observe how competing benchmarks like Contra Labs’ Human Creativity and Artificial Analysis’ Image Arena respond to OpenArt’s launch, potentially leading to a more standardized approach to evaluating creative AI capabilities.

The publication of OpenArt’s full methodology and prompt sets will be crucial for independent verification of the ’s results. Transparency in the evaluation process will help build trust among users and ensure that the rankings accurately reflect model performance.

The evolution of OpenArt Arena’s rankings as new models are released will be an important indicator of the platform’s relevance and accuracy. Monitoring how the rankings change over time will provide insights into the rapid development of AI creative models and the effectiveness of the in capturing these changes.

The response from competing benchmarks, such as Contra Labs’ Human Creativity and Artificial Analysis’ Image Arena, will be significant. These platforms may adjust their methodologies or launch new features in response to OpenArt’s launch, potentially leading to a more standardized and comprehensive approach to evaluating creative AI capabilities.

The adoption of OpenArt Arena by enterprises and creative professionals will be a key indicator of its success. If the platform becomes a widely used tool for model selection, it could influence the development of AI creative models and the strategies of AI providers seeking to optimize their performance on the .

相關指引和測驗

人工智慧模型解釋AI 的未來什麼是人工智慧?測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?