Voltar às notícias
EmpresaInstruções AI Understanding

Forbes reports autonomous AI systems moving into enterprise coding and security

Forbes reports that Blitzy and XBOW are applying increasingly autonomous AI to software modernization and penetration testing, while Google and OpenAI are developing related security systems. The companies’ performance and deployment claims have not been independently confirmed.

Por 6 min read
AI-generated editorial illustration accompanying Forbes reports autonomous AI systems moving into enterprise coding and security
A versão curta

Forbes reports that Blitzy and XBOW are applying increasingly autonomous AI to software modernization and penetration testing, while Google and OpenAI are developing related security systems. The companies’ performance and deployment claims have not been independently confirmed.

O que aconteceu

Forbes reports that autonomous AI systems are being used to pursue extended enterprise tasks with limited or no human intervention. The article highlights Blitzy for legacy-code modernization, XBOW for continuous penetration testing, and related systems from Google and OpenAI.

Forbes reports that autonomous AI is being positioned as a step beyond copilots and short-lived agents. In the article’s framing, copilots assist while a person works, conventional agents handle a task for minutes and return a result for review, while more autonomous systems receive a goal and work for days or weeks before returning completed work. Forbes attributes this distinction partly to Sanjot Malhi, who leads Northzone’s global growth fund and said he had spent nearly two years developing an investment thesis around the technology. The report says Northzone backed Blitzy and XBOW since January, with Blitzy having raised $200 million at a reported $1.4 billion valuation and XBOW having raised $120 million. These financial figures and the characterization of the systems were reported by Forbes and are not independently confirmed here.

Forbes reports that Blitzy applies this approach to enterprise software modernization. The company is described as ingesting hundreds of millions of lines of legacy code and a customer’s compliance policies, then using a stated goal to build new systems. Forbes cites Charles River Development, a State Street company, as a customer using the platform to modernize older code, and says Builders FirstSource tripled its software-development velocity during its first three months on the platform. The article also says OpenAI cited Blitzy in materials about GPT-5.6’s coding capabilities. Forbes presents these customer outcomes as evidence of practical use, but the supplied report does not provide an independent audit, detailed baseline, sample size, error rate, or explanation of how much human review remained in those projects.

Forbes reports that XBOW applies autonomous AI to offensive cybersecurity. Its platform is described as continuously conducting penetration tests against a company’s systems rather than relying only on periodic reviews by scarce human specialists. The article says XBOW reached the top of HackerOne’s U.S. leaderboard in summer 2025, which Forbes describes as the first time the leading hacker was not a human being. Forbes also cites Moderna deputy chief information security officer Farzan Karimi, who reportedly said XBOW found a firewall bypass that he had missed. The report further says XBOW received early access to Anthropic’s Mythos model during Project Glasswing and that Anthropic cited the company’s testing. These claims are attributed to Forbes and the named organizations or executives through Forbes; no independent testing data is supplied.

Forbes places these companies within a broader group of autonomous security and coding systems. The article reports that Horizon3 raised a $250 million Series E at a valuation above $2 billion, and that Google’s Big Sleep, developed by DeepMind and Project Zero, autonomously found 20 vulnerabilities in widely used open-source software. It also says OpenAI incorporated Aardvark, a security-research system, into Codex after benchmark testing in which it detected 92% of known flaws. The report describes these developments as signs that AI systems are moving toward sustained work on enterprise goals. However, Forbes does not provide the underlying benchmark protocols, vulnerability disclosures, false-positive rates, remediation results or evidence that the systems operated without meaningful human intervention in every case.

Leia a fonte primária: forbes.com

Por que isso importa

If independently validated, these systems could change how companies handle software development and cybersecurity by shifting AI from short, supervised tasks toward longer-running work. The report also underscores unresolved questions about oversight, auditability, accuracy, cost and accountability.

Forbes’ central claim matters because software development and cybersecurity are both areas where unfinished work can create direct operational consequences. A system that can process a large legacy codebase, interpret compliance constraints and produce a usable modernization project could reduce the time required for work that traditionally depended on systems integrators and long contracts. The reported Builders FirstSource result, if independently verified, would indicate a substantial change in development throughput. The practical question is not whether AI can generate code, but whether it can consistently deliver secure, maintainable and compliant systems over long tasks.

The security examples point to a similarly consequential shift. Continuous automated testing could expand the amount of software examined and reduce dependence on periodic assessments. Finding a vulnerability missed by a human reviewer could be valuable, especially if the finding is reproducible and leads to effective remediation. At the same time, autonomous penetration testing can create risks if its permissions, targets or actions are poorly controlled. Forbes’ account does not establish how these systems are isolated, how customers approve testing boundaries, how findings are validated, or how organizations prevent automated tools from disrupting production systems.

The report also illustrates the difference between investment enthusiasm and demonstrated public performance. Forbes describes Northzone’s Malhi as viewing Blitzy and XBOW as early versions of an “enterprise brain” that retains an organization’s context while models provide capabilities. That is an investment thesis, not an independently established technical category. The article mentions KPMG and Gartner expectations about companies reducing or decommissioning some AI-agent deployments, but those are forecasts rather than evidence about current outcomes. The most meaningful public impact will depend on whether autonomous systems can show reliable results, transparent records and clear accountability across varied customers rather than selected success stories.

O que assistir a seguir

The key tests are independent evaluations of completed work, security findings, error rates, costs and human oversight. Watch whether companies can document the systems’ actions and reliably assign responsibility when autonomous tools make mistakes.

First, watch for independent evidence behind the performance claims. For Blitzy, that would include details on the projects evaluated, the definition of “tripled” development velocity, defect and security rates, the amount of human review, and whether the resulting systems remained maintainable after deployment. Customer testimony can establish that a system was used, but it does not by itself establish general reliability. Forbes reports the customer claims, while the supplied material does not independently verify them.

For XBOW and related security systems, useful follow-up would include disclosed vulnerabilities, reproducibility, false-positive and false-negative rates, the scope of authorized testing, and the time between discovery and remediation. The reported 92% benchmark result for OpenAI’s Aardvark also needs context about the benchmark’s design, the meaning of “known flaws,” and how the system compares with human or automated baselines. Without those details, leaderboard placement and benchmark percentages should not be treated as general measures of real-world security performance.

Finally, organizations adopting these tools will need governance that matches their autonomy. Forbes itself reports concerns about governance gaps and recommends audit trails and clear accountability before companies delegate goals to autonomous systems. Watch whether deployments preserve human approval for high-impact actions, record model and tool activity, limit access to sensitive systems, and provide a responsible operator when work fails. The article’s longer-term predictions about autonomous AI and artificial general intelligence remain uncertain; near-term evidence should come from documented deployments, repeatable evaluations and publicly explainable outcomes.

Guias e questionários relacionados

Agentes de IAModelos de IA explicadosÉtica da IATreinamento de IATeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossário
Achou isso útil?