Kembali ke Berita
ProdukAI Understanding pengarahan

Cognition releases SWE-2 coding model for Devin

MarkTechPost reports that Cognition’s SWE-2 coding model is available inside Devin, with selectable reasoning-effort levels but no standalone API or open weights.

4 min readRead the linked source
Source-provided image accompanying Cognition releases SWE-2 coding model for Devin
Referensi sumberSumber direkam
Penerbit
marktechpost.com
Tautan sumber
marktechpost.comhttps://www.marktechpost.com/2026/09/12/cognition-releases-swe-2-a-kimi-k3-post-trained-coding-model-that-matches-fable-5-1-on-frontiercode-at-64-lower-cost/amp/
Jenis sumber
Sumber tertaut — status sumber utama belum ditetapkan.
KonteksPahami ini dalam 60 detik

Mulai di sini

Istilah-istilah penting

API (Antarmuka Pemrograman Aplikasi)
Cara terstruktur bagi satu sistem perangkat lunak untuk mengirim permintaan dan menerima tanggapan dari sistem lain.
Pembelajaran Penguatan
Pelatihan dengan imbalan memberi sinyal saat agen mempelajari tindakan yang memaksimalkan keuntungan jangka panjang.
Kuantisasi
Mengonversi bobot model ke format presisi lebih rendah seperti 8-bit atau 4-bit.
Test yourselfKuis Penjelasan Model AI

Apa yang terjadi

MarkTechPost reports that Cognition released SWE-2, a coding model post-trained from Moonshot AI’s Kimi K3. The model is available only within Devin, initially through its Desktop and CLI products. Devin Web and Fusion access are reportedly rolling out. MarkTechPost says Cognition reports benchmark and cost results, but those comparisons have not been independently confirmed.

MarkTechPost reports that Cognition, the company behind Devin, released SWE-2 as its most capable coding model to date. According to the outlet, Cognition post-trained Moonshot AI’s Kimi K3, described as a 2.8-trillion-parameter open model, using reinforcement learning. MarkTechPost says Cognition reported a 50.0% score on FrontierCode 1.1 Main, within one percentage point of Fable 5.1 at 64% lower cost. These results are company-reported; the source does not provide independent verification.

The model is not offered as an open-weight release or standalone API. MarkTechPost says SWE-2 runs only inside Devin, with Desktop and CLI access available at the time of reporting and Devin Web and Fusion rolling out. The source does not document a price, eligibility requirements, geographic limits, or a general-availability date.

A central change is selectable reasoning effort. MarkTechPost reports that Cognition trained three effort levels in a single reinforcement-learning run and assigned each level a different cost penalty. The outlet also describes claimed improvements over SWE-1.7, including fewer turns before making an initial edit and lower reported cost on FrontierCode. Cognition says the model improved test coverage, tool-use resourcefulness, and verification behavior, but these behavioral claims are not independently established by the source.

Detail sumber: marktechpost.com

Mengapa itu penting

SWE-2 represents a meaningful product change because it gives Devin users multiple reasoning-effort choices and reportedly improves coding performance while reducing cost and unnecessary exploration. Its practical reach is limited, however: users cannot deploy the model on their own infrastructure or call it through a standalone API, and pricing is not documented in the source.

The reported release matters because it moves model-level control over reasoning effort into a coding-agent product. If the reported cost and performance differences hold outside Cognition’s evaluations, developers could choose faster or more deliberate behavior according to task difficulty instead of using one fixed setting.

The access model also limits the immediate significance of the release. Organizations that require self-hosting, direct API integration, or control over model weights cannot use SWE-2 on those terms based on the information available. The source does not say whether a future API or broader deployment option is planned.

Evaluation context is important. MarkTechPost says FrontierCode is Cognition’s own benchmark and that rival scores in the comparison came from Cognition’s evaluation. The source also notes a substantial weakness on Terminal-Bench 4, where SWE-2 reportedly trails other compared systems by about 30 points. That makes the release consequential but not evidence of uniformly stronger coding performance.

Apa yang harus ditonton selanjutnya

Watch for independent testing of SWE-2 on coding benchmarks, clearer availability and pricing information, and evidence from users about whether its effort settings produce reliable cost-performance tradeoffs in real software projects.

Independent benchmark results should clarify whether SWE-2’s reported advantages persist across coding environments, repositories, languages, and evaluation harnesses. Particular attention should go to Terminal-Bench 4 and to tests that measure completed, correct changes rather than task completion alone.

Users should watch for documentation confirming when Devin Web and Fusion receive SWE-2, which effort levels are available in each product, and how usage is billed. MarkTechPost does not provide pricing or a detailed access policy.

Further technical disclosure could help assess the reinforcement-learning methods, verifier quality, quantization choices, and claimed safety results. MarkTechPost reports that Cognition reran trustworthiness evaluations, but the source does not establish how representative those tests are of production coding use or independently reproduce their findings.

Panduan & kuis terkait

Model AI DijelaskanAgen AIPelatihan AIPrompt EngineeringUji pengetahuan Anda — coba kuis AI gratisCari istilah AI di glosarium kami
Apakah ini berguna?