返回新聞
企業AI Understanding 簡報

福布斯報道自主人工智慧系統正進入企業程式設計和安全領域

《富比士》報導,Blitzy 和 XBOW 正在將日益自主的人工智慧應用於軟體現代化和滲透測試,而 Google 和 OpenAI 正在開發相關的安全系統。兩家公司的性能和部署聲明尚未得到獨立證實。

6 min readRead the original reporting
Source-provided image accompanying Forbes reports autonomous AI systems moving into enterprise coding and security
歸因報告來源記錄
出版商
forbes.com
來源連結
forbes.comhttps://www.forbes.com/sites/sandycarter/2026/08/24/autonomous-ai-outgrows-agents-as-blitzy-xbow-and-google-deliver/
來源類型
新聞媒體的報道-不是第一方文件。

我們無法獨立確認的內容: 此聲明歸因於指定的商店。我們沒有根據第一方文件對其進行驗證。 (forbes.com)

背景60 秒內了解這一點

從這裡開始

關鍵術語

AGI(通用人工智慧)
一個假設的人工智慧系統,可以在許多領域以人類層級執行大多數智力任務。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
測試一下自己AI 代理測驗

發生了什麼事

Forbes reports that autonomous AI systems are being used to pursue extended enterprise tasks with limited or no human intervention. The article highlights Blitzy for legacy-code modernization, XBOW for continuous penetration testing, and related systems from Google and OpenAI.

Forbes reports that autonomous AI is being positioned as a step beyond copilots and short-lived agents. In the article’s framing, copilots assist while a person works, conventional agents handle a task for minutes and return a result for review, while more autonomous systems receive a goal and work for days or weeks before returning completed work. Forbes attributes this distinction partly to Sanjot Malhi, who leads Northzone’s global growth fund and said he had spent nearly two years developing an investment thesis around the technology. The report says Northzone backed Blitzy and XBOW since January, with Blitzy having raised $200 million at a reported $1.4 billion valuation and XBOW having raised $120 million. These financial figures and the characterization of the systems were reported by Forbes and are not independently confirmed here.

Forbes reports that Blitzy applies this approach to enterprise software modernization. The company is described as ingesting hundreds of millions of lines of legacy code and a customer’s compliance policies, then using a stated goal to build new systems. Forbes cites Charles River Development, a State Street company, as a customer using the platform to modernize older code, and says Builders FirstSource tripled its software-development velocity during its first three months on the platform. The article also says OpenAI cited Blitzy in materials about GPT-5.6’s coding capabilities. Forbes presents these customer outcomes as evidence of practical use, but the supplied report does not provide an independent audit, detailed baseline, sample size, error rate, or explanation of how much human review remained in those projects.

Forbes reports that XBOW applies autonomous AI to offensive cybersecurity. Its platform is described as continuously conducting penetration tests against a company’s systems rather than relying only on periodic reviews by scarce human specialists. The article says XBOW reached the top of HackerOne’s U.S. leaderboard in summer 2025, which Forbes describes as the first time the leading hacker was not a human being. Forbes also cites Moderna deputy chief information security officer Farzan Karimi, who reportedly said XBOW found a firewall bypass that he had missed. The report further says XBOW received early access to Anthropic’s Mythos model during Project Glasswing and that Anthropic cited the company’s testing. These claims are attributed to Forbes and the named organizations or executives through Forbes; no independent testing data is supplied.

Forbes places these companies within a broader group of autonomous security and coding systems. The article reports that Horizon3 raised a $250 million Series E at a valuation above $2 billion, and that Google’s Big Sleep, developed by DeepMind and Project Zero, autonomously found 20 vulnerabilities in widely used open-source software. It also says OpenAI incorporated Aardvark, a security-research system, into Codex after testing in which it detected 92% of known flaws. The report describes these developments as signs that AI systems are moving toward sustained work on enterprise goals. However, Forbes does not provide the underlying benchmark protocols, vulnerability disclosures, false-positive rates, remediation results or evidence that the systems operated without meaningful human intervention in every case.

來源詳情: forbes.com ↗

為什麼這很重要

If independently validated, these systems could change how companies handle software development and cybersecurity by shifting AI from short, supervised tasks toward longer-running work. The report also underscores unresolved questions about oversight, auditability, accuracy, cost and accountability.

Forbes’ central claim matters because software development and cybersecurity are both areas where unfinished work can create direct operational consequences. A system that can process a large legacy codebase, interpret compliance constraints and produce a usable modernization project could reduce the time required for work that traditionally depended on systems integrators and long contracts. The reported Builders FirstSource result, if independently verified, would indicate a substantial change in development throughput. The practical question is not whether AI can generate code, but whether it can consistently deliver secure, maintainable and compliant systems over long tasks.

The security examples point to a similarly consequential shift. Continuous automated testing could expand the amount of software examined and reduce dependence on periodic assessments. Finding a vulnerability missed by a human reviewer could be valuable, especially if the finding is reproducible and leads to effective remediation. At the same time, autonomous penetration testing can create risks if its permissions, targets or actions are poorly controlled. Forbes’ account does not establish how these systems are isolated, how customers approve testing boundaries, how findings are validated, or how organizations prevent automated tools from disrupting production systems.

The report also illustrates the difference between investment enthusiasm and demonstrated public performance. Forbes describes Northzone’s Malhi as viewing Blitzy and XBOW as early versions of an “enterprise brain” that retains an organization’s context while models provide capabilities. That is an investment thesis, not an independently established technical category. The article mentions KPMG and Gartner expectations about companies reducing or decommissioning some AI-agent deployments, but those are forecasts rather than evidence about current outcomes. The most meaningful public impact will depend on whether autonomous systems can show reliable results, transparent records and clear accountability across varied customers rather than selected success stories.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
互動式概念檢查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下來看什麼

The key tests are independent evaluations of completed work, security findings, error rates, costs and human oversight. Watch whether companies can document the systems’ actions and reliably assign responsibility when autonomous tools make mistakes.

First, watch for independent evidence behind the performance claims. For Blitzy, that would include details on the projects evaluated, the definition of “tripled” development velocity, defect and security rates, the amount of human review, and whether the resulting systems remained maintainable after deployment. Customer testimony can establish that a system was used, but it does not by itself establish general reliability. Forbes reports the customer claims, while the supplied material does not independently verify them.

For XBOW and related security systems, useful follow-up would include disclosed vulnerabilities, reproducibility, false-positive and false-negative rates, the scope of authorized testing, and the time between discovery and remediation. The reported 92% result for OpenAI’s Aardvark also needs context about the benchmark’s design, the meaning of “known flaws,” and how the system compares with human or automated baselines. Without those details, leaderboard placement and benchmark percentages should not be treated as general measures of real-world security performance.

Finally, organizations adopting these tools will need governance that matches their autonomy. Forbes itself reports concerns about governance gaps and recommends audit trails and clear accountability before companies delegate goals to autonomous systems. Watch whether deployments preserve human approval for high-impact actions, record model and tool activity, limit access to sensitive systems, and provide a responsible operator when work fails. The article’s longer-term predictions about autonomous AI and artificial general intelligence remain uncertain; near-term evidence should come from documented deployments, repeatable evaluations and publicly explainable outcomes.

相關指引和測驗

人工智慧代理人工智慧模型解釋AI 倫理人工智慧培訓測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 資金追蹤器
覺得有用嗎?