뉴스로 돌아가기
산업AI Understanding 브리핑

The Register는 AI 공급업체가 추론을 위해 Nvidia GPU를 넘어 다양화하고 있다고 보고합니다.

Register는 Cerebras, Google, Marvell 및 Waymo가 AI 추론 속도, 전력 사용 또는 대기 시간을 개선하기 위해 특수 실리콘을 추구하고 있다고 보고하지만, 기본 공개 및 상업적 배포는 여기에서 독립적으로 확인되지 않습니다.

6 min readRead the original reporting
Source-provided image accompanying The Register reports AI vendors are diversifying beyond Nvidia GPUs for inference
기여 보고녹음된 소스
출판사
theregister.com
소스 링크
theregister.comhttps://www.theregister.com/ai-and-ml/2026/08/24/ai-vendors-are-turning-to-custom-hardware-as-microsoft-winds-back-the-clock-on-windows/5291293
소스 유형
자사 문서가 아닌 뉴스 매체를 통한 보도입니다.

자체적으로는 확인할 수 없었던 내용: 이 소유권 주장은 해당 매장에 귀속됩니다. 당사는 자사 문서와 비교하여 이를 확인하지 않았습니다. (theregister.com)

맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

추론
훈련된 모델이 예측 또는 출력을 생성하는 런타임 단계입니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
정밀도
예측된 긍정 중 실제로 정확한 비율입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

The Register reports a widening move toward custom and alternative AI hardware. Its account covers Cerebras’s modular CS-4 rack system, Marvell’s growing role in custom accelerator design, Google’s continued use of specialized TPUs, and Waymo’s development of a machine-learning accelerator for autonomous vehicles. The report also describes Microsoft restoring several Windows customization and AI-monitoring features, but that is a separate secondary thread.

The Register’s Aug. 24 report centers on what its editors describe as a move beyond Nvidia’s dominant role in AI datacenters. Vendors are increasingly dividing workloads among different kinds of silicon instead of relying on a single accelerator type. It presents , or generating outputs from trained models, as the main pressure point: providers want high token rates for coding assistants and other interactive services while controlling the power and hardware costs of serving requests. The Register’s article is a podcast transcript and analysis, so it should be read as reporting from that outlet rather than independently verified documentation of every technical claim.

The Register reports that Cerebras is moving from large, monolithic systems toward its CS-4 rack-scale design. Earlier Cerebras systems placed a large, power-hungry wafer-scale chip in a self-contained liquid-cooled chassis. The new approach places three chips in a modular rack arrangement, allowing compute components to be replaced without replacing associated power-delivery equipment. The Register says the design doubles clock speed, memory bandwidth and input-output speed, while placing three times as much silicon in a rack. It also says Cerebras has partnerships with AWS and AMD, and that fast token generation has attracted interest for coding assistants and other services.

The report describes a broader custom-silicon ecosystem around hyperscalers: Google has used its own Tensor Processing Units since around 2015 and relied on Broadcom for portions of custom accelerator design, while Marvell is becoming a competitor in custom XPU and connectivity work after building an intellectual-property portfolio through acquisitions. The article frames this as a way for cloud companies to compare suppliers or reduce dependence on a single design partner, but does not document a specific new Google-Marvell contract; assertions about future customer strategies and pricing pressure are commentary, not established developments. Waymo disclosed a custom machine-learning accelerator intended for vehicles because sensor volume and low latency matter when an obstacle appears. The Register says Waymo has relied heavily on field-programmable gate arrays and may seek the greater compute density and easier development of an application-specific integrated circuit, but provides no specifications, testing results, production timetable or safety record. Separately, Microsoft is testing movable Windows 11 taskbars, more customizable context menus and Task Manager visibility into neural-processing-unit workloads.

소스 세부정보: theregister.com ↗

왜 중요한가요?

The shift could make AI infrastructure less dependent on one class of Nvidia GPU and put more emphasis on hardware designed for particular workloads. Faster response times, lower operating costs, better power efficiency and tighter integration with sensors are the practical goals described by The Register. The report does not establish that any alternative has broadly surpassed Nvidia in real-world deployments.

The practical significance is that AI performance is not determined only by the model. The Register’s reporting shows vendors optimizing hardware beneath for particular constraints: token-generation speed for interactive tools, power and cooling for datacenter racks, and latency for vehicle control. If those designs work at scale, companies could choose different accelerators for training, high-volume inference, premium low-latency requests and edge workloads. That would make the AI infrastructure market more specialized and could give customers more alternatives to general-purpose GPU systems.

Cerebras’s reported rack redesign highlights an operational issue that can be missed in headline performance claims. The Register says earlier systems consumed about 15 kilowatts for the chip and bundled compute with power-delivery equipment in a large chassis. A modular rack could make repairs and upgrades less disruptive, even if the silicon remains unusually power-hungry. For datacenter operators, serviceability, cabling, cooling and the ability to replace one failed component may matter as much as peak throughput. The source does not independently verify the reported power or performance figures, so those benefits remain claims requiring technical review.

A wider supplier base could affect bargaining power and investment decisions. The Register suggests that cloud companies often use outside intellectual-property firms for less distinctive parts of custom chips while concentrating internal engineering on features relevant to their services. Competition between Broadcom and Marvell could give hyperscalers another design route, but custom silicon also creates software and support obligations. A chip efficient for one workload may be difficult to program, hard to procure or poorly suited to another. The report offers no cost comparison with Nvidia systems and no evidence that customers are switching at broad scale. Waymo’s example connects custom AI hardware to a safety-sensitive deployment rather than only datacenter economics: low latency matters when a vehicle responds to sensor input, and remote human intervention is not an ideal substitute for immediate onboard processing. That does not demonstrate improved safety; there is no independent evaluation, failure-rate comparison or regulatory assessment. The narrower conclusion is that autonomous vehicles provide a reason to design silicon around sensor processing and response time, while public impact depends on validation in real operating conditions.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The key tests are public technical documentation, independent benchmarks and evidence of scaled customer use. Watch whether Cerebras’s new rack design, Marvell-backed custom accelerators and Waymo’s vehicle silicon reach production, and whether their benefits outweigh software, supply-chain and maintenance costs. Microsoft’s Windows changes should also be judged by public availability and whether they improve user control rather than merely add more system monitoring.

First, look for primary technical material from Cerebras, Waymo, Google, Marvell or their customers. Useful evidence would include architecture documents, power measurements, software support, production status and clearly defined workload tests. The Register’s report supplies a newsworthy map of the companies involved, but does not establish availability, pricing, customer volume or independent performance. Those missing details determine whether the developments are industry moves or mainly engineering plans.

Second, compare like with like. Token-per-second figures can vary with model size, batching, , context length and the number of simultaneous users. A custom accelerator may be excellent for one narrow pattern and less useful for general workloads. Watch for independent tests measuring throughput, response latency, energy per generated token, utilization, reliability and total cost of ownership against contemporary GPU systems. Without that context, claims of faster or cheaper AI remain difficult to evaluate. Third, follow the commercial relationships: The Register reports Cerebras partnerships with AWS and AMD and describes Marvell as entering a field previously dominated by Broadcom and internal hyperscaler teams. Important next evidence will be whether those relationships produce generally accessible services, named deployments or repeat orders. The report’s expectation that AWS or AMD services may use Cerebras hardware is forward-looking commentary, not a confirmed rollout, and should not be treated as one.

Finally, monitor Waymo’s vehicle implementation and Microsoft’s software changes separately. For Waymo, the meaningful questions are whether the accelerator is installed in production vehicles, how it handles sensor workloads, what fallback systems exist and what safety validation has been completed. For Windows, The Register says taskbar movement is in the Release Preview channel and enhanced NPU process visibility is being developed, but public timing and final behavior can change. Neither thread should be presented as complete until the companies publish stable availability and independent users can assess the results.

관련 가이드 및 퀴즈

AI 모델 설명트랜스포머AI의 미래ChatGPT와 LLM알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 자금 추적기를 팔로우하세요
이것이 유용하다고 생각하시나요?