Ascend 960
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- From 8-GPU servers to supernodes: AI data centers for the inference era
Speaking at the 2026 ITValue Summit on September 17, xFusion computing vice president Wang Liangdong said AI has entered an inference-dominated phase in which global token consumption keeps surging and buyers judge infrastructure by cost per token, throughput and cluster stability rather than peak per-card performance. The article cites a range of figures: the 2026 China Development Forum said China's daily token calls topped 140 trillion in March, OpenRouter data showed China accounted for 36% of global token calls in the March 16-22 week, JPMorgan expects Chinese inference token consumption to grow roughly 370x in five years, and IDC projects annual global consumption of 150,000 Peta Tokens by 2030. It argues the 8-GPU server paradigm hits communication bottlenecks under trillion-parameter, high-concurrency inference, pointing to Huawei's Ascend 960 supernode (4,096 cards per node, 8E FP8), Lenovo's Wanquan heterogenous computing platform V5.0 and xFusion's FusionPoD for AI rack-scale designs — while warning that supernodes only suit very large workloads, interconnect standards remain unsettled, and utilization depends on sustained paying demand.
钛媒体 · 🔥 9
Experience and discussion from the community
Share my Ascend 960 experienceAsk about Ascend 960
Nobody has shared their experience with Ascend 960 yet.