Назад до новин
ПродуктAI Understanding брифінг

Cognition випускає модель кодування SWE-2 для Devin

MarkTechPost повідомляє, що модель кодування SWE-2 від Cognition доступна в Devin із можливістю вибору рівнів міркувань, але без окремого API чи відкритих ваг.

4 min readRead the linked source
Source-provided image accompanying Cognition releases SWE-2 coding model for Devin
Посилання на джерелоДжерело записано
Видавець
marktechpost.com
Посилання на джерело
marktechpost.comhttps://www.marktechpost.com/2026/09/12/cognition-releases-swe-2-a-kimi-k3-post-trained-coding-model-that-matches-fable-5-1-on-frontiercode-at-64-lower-cost/amp/
Тип джерела
Пов’язане джерело — статус первинного джерела не встановлено.
КонтекстЗрозумійте це за 60 секунд

Почніть тут

Ключові терміни

API (інтерфейс прикладного програмування)
Структурований спосіб для однієї програмної системи надсилати запити до іншої системи та отримувати відповіді від неї.
Навчання з підкріпленням
Навчання за сигналами винагороди, коли агент навчається діям, які максимізують довгострокову віддачу.
Квантування
Перетворення ваг моделі у формати з нижчою точністю, такі як 8- або 4-бітні.
Перевір себеВікторина «Пояснення моделей ШІ».

Що сталося

MarkTechPost reports that Cognition released SWE-2, a coding model post-trained from Moonshot AI’s Kimi K3. The model is available only within Devin, initially through its Desktop and CLI products. Devin Web and Fusion access are reportedly rolling out. MarkTechPost says Cognition reports benchmark and cost results, but those comparisons have not been independently confirmed.

MarkTechPost reports that Cognition, the company behind Devin, released SWE-2 as its most capable coding model to date. According to the outlet, Cognition post-trained Moonshot AI’s Kimi K3, described as a 2.8-trillion-parameter open model, using . MarkTechPost says Cognition reported a 50.0% score on FrontierCode 1.1 Main, within one percentage point of Fable 5.1 at 64% lower cost. These results are company-reported; the source does not provide independent verification.

The model is not offered as an open-weight release or standalone API. MarkTechPost says SWE-2 runs only inside Devin, with Desktop and CLI access available at the time of reporting and Devin Web and Fusion rolling out. The source does not document a price, eligibility requirements, geographic limits, or a general-availability date.

A central change is selectable reasoning effort. MarkTechPost reports that Cognition trained three effort levels in a single reinforcement-learning run and assigned each level a different cost penalty. The outlet also describes claimed improvements over SWE-1.7, including fewer turns before making an initial edit and lower reported cost on FrontierCode. Cognition says the model improved test coverage, tool-use resourcefulness, and verification behavior, but these behavioral claims are not independently established by the source.

Деталі джерела: marktechpost.com ↗

Чому це важливо

SWE-2 represents a meaningful product change because it gives Devin users multiple reasoning-effort choices and reportedly improves coding performance while reducing cost and unnecessary exploration. Its practical reach is limited, however: users cannot deploy the model on their own infrastructure or call it through a standalone API, and pricing is not documented in the source.

The reported release matters because it moves model-level control over reasoning effort into a coding-agent product. If the reported cost and performance differences hold outside Cognition’s evaluations, developers could choose faster or more deliberate behavior according to task difficulty instead of using one fixed setting.

The access model also limits the immediate significance of the release. Organizations that require self-hosting, direct API integration, or control over model weights cannot use SWE-2 on those terms based on the information available. The source does not say whether a future API or broader deployment option is planned.

Контекст оцінки важливий. MarkTechPost каже, що FrontierCode є власним еталонним тестом Cognition і що конкуруючі бали в порівнянні були отримані від оцінки Cognition. Джерело також відзначає значну слабкість у Terminal-Bench 4, де SWE-2, як повідомляється, відстає від інших порівнюваних систем приблизно на 30 балів. Це робить випуск значущим, але не свідченням стабільно кращої продуктивності кодування.

Interactive Mechanism

Інтерактивний механізм: як він насправді працює

Дослідіть технологію, що лежить в основі цієї розробки, в інтерактивному режимі.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Інтерактивна перевірка концепції+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Що дивитися далі

Watch for independent testing of SWE-2 on coding benchmarks, clearer availability and pricing information, and evidence from users about whether its effort settings produce reliable cost-performance tradeoffs in real software projects.

Результати незалежного порівняльного тесту повинні прояснити, чи зберігаються переваги SWE-2 у середовищі кодування, сховищах, мовах і системах оцінювання. Особливу увагу слід приділити Terminal-Bench 4 і тестам, які вимірюють виконані, виправляють зміни, а не лише виконання завдань.

Користувачі повинні стежити за документацією, яка підтверджує, коли Devin Web і Fusion отримують SWE-2, які рівні зусиль доступні в кожному продукті та як виставляється плата за використання. MarkTechPost не надає цін або детальної політики доступу.

Подальше розкриття технічної інформації може допомогти оцінити методи навчання з підкріпленням, якість верифікатора, варіанти квантування та заявлені результати безпеки. MarkTechPost повідомляє, що Cognition повторно провела оцінку надійності, але джерело не встановило, наскільки ці тести репрезентативні для використання робочого кодування, і не відтворило їх висновки.

Пов’язані посібники та вікторини

Пояснення моделей AIАгенти ШІНавчання ШІPrompt EngineeringПеревірте свої знання — пройдіть безкоштовну вікторину зі штучним інтелектомЗнайдіть термін ШІ в нашому глосаріїСлідкуйте за відстеженням випуску моделі AI
Знайшли це корисним?