
Meta's Custom AMD MI450 Chip Halves Compute Power and Slashes Memory; SemiAnalysis: This Is a Tragedy of Meta's Corporate Culture
Billions of dollars down the drain! Meta has repeatedly made "disastrous" missteps in AI infrastructure: forcing AMD to cripple its most powerful chip and leaving the $2.5 billion acquisition in shambles. Institutions have sharply criticized the short-sighted performance culture and excessive customization, which are stifling hardware-software synergy. This deep-seated organizational crisis is costing Meta dearly
Meta's series of decision-making errors in the field of AI infrastructure are exposing deep-seated problems in its organizational culture at a cost of billions of dollars.
According to the latest disclosure by chip industry research firm SemiAnalysis, Meta is requesting AMD to customize a significantly downgraded MI450X chip—halving the compute units and reducing the HBM memory stack from 12 layers to 8 layers, shrinking memory capacity by nearly two-thirds. SemiAnalysis has directly characterized this decision as "catastrophic" and publicly urged AMD to bypass Meta's infrastructure team and engage directly with Meta's TBD Lab, a superintelligence laboratory, to push for the procurement of the standard MI450X.
This incident is not isolated. SemiAnalysis pointed out that **from the over $2.5 billion Rivos acquisition to the H100 custom server Grand Teton, and then to the GB200 custom solution Ariel, the decision-making pattern of Meta's infrastructure team consistently features excessive engineering, a lack of hardware-software co-design, and short-term political considerations overriding long-term technical rationality." "Meta's infrastructure team needs a cultural reset," SemiAnalysis wrote.

AMD's Most Powerful Chip "Crippled," GenAI Performance Severely Impaired
The AMD MI450X is currently one of the GPUs with the most aggressive silicon engineering specifications on the market: it adopts a 2nm process, hybrid bonding packaging technology, 12-layer HBM4 memory stacking, and the largest CoWoS reticle size on the market, representing the current ceiling for packaging and storage density.

However, according to SemiAnalysis, the custom version ordered by Meta will halve the compute silicon area and reduce the HBM stack from 12-Hi to 8-Hi, effectively compressing two core indicators: compute power and memory bandwidth. The reason given by Meta's infrastructure team is that this configuration is designed for recommendation system (RecSys) workloads, aiming to increase the ratio of CPU to GPU computing.
The problem is that this decision was made before the establishment of TBD Lab, Meta's core large model research team, which has no interest in this chip. SemiAnalysis explicitly stated that compared to NVIDIA's Vera Rubin, the downgraded MI450 holds no appeal for TBD Lab, noting that "TBD will heavily favor Rubin." This means AMD's shipments to Meta will be severely impacted by this decision.
In its report, SemiAnalysis rarely addressed AMD directly: "AMD needs to step up, collaborate directly with the TBD team, and ensure they receive the standard version of the MI450, not this crippled version that is worthless for GenAI."
Rivos Acquisition: Over $2.5 Billion Ends in Shambles
Another typical case of Meta's infrastructure decision-making errors is the completion of the Rivos acquisition in 2024 for over $2.5 billion.
According to SemiAnalysis, few people within Meta's internal chip department truly understood the strategic logic behind this acquisition, and those who initially pushed for the deal have since fallen silent. The prevailing view is that Meta, being cash-rich and seeing the custom chip sector heat up, along with having previously licensed Rivos' intellectual property, led management to believe it was better to acquire the entire company.
However, the deal structure itself planted hidden dangers. The Rivos founding team insisted on selling the entire company as a package, forcing Meta to buy the whole entity and subsequently lay off many departments it did not need. Citing former Meta chip employees, the report stated that the acquisition was led by Yee Jiun Song, Meta's head of silicon, who faced internal opposition during the process, but his own interest waned after the deal was completed. Existing chip team managers viewed Rivos engineers as "free labor," absorbing them into their respective teams, causing the original team structure to rapidly disintegrate.
From a technical perspective, the core value of the acquisition—Rivos' SIMT architecture GPU IP—has vanished with the cancellation of the "Olympus" chip project, which was planned to utilize this technology. The replacement project, "Phoebe," is expected to tape out as early as 2028, but SemiAnalysis noted that internal optimism regarding its feasibility is low.
Personnel losses are also striking. According to SemiAnalysis, about 30% of Rivos employees left in recent layoffs, co-founder Mark Hayter has departed, and several former Rivos personnel joined Nuvacore, a chip startup founded by Gerard Williams, after their first batch of RSUs vested in May this year. SemiAnalysis also revealed that Rivos CEO and co-founder Puneet Kumar may leave within a year or two after all his Meta shares vest.
Grand Teton and Ariel: Replaying the Cost of Design
Meta's obsession with customization is not new. Its H100 custom server, Grand Teton, added an extra switch tray to the standard HGX server, equipped with four Broadcom PCIe switches, 16 SSDs, and eight network cards, aiming to provide more direct-attached storage for each server to save training checkpoints.
However, after actual production deployment, the model teams used this storage far less than expected, and the design was ultimately abandoned. SemiAnalysis pointed out that this is a typical example of the lack of co-design between Meta's hardware and software teams—the infrastructure team paid higher material costs and power consumption for underutilized features, while also increasing dependence on Broadcom, which contradicts Meta's original intention to reduce reliance on NVIDIA's networking equipment.
Entering the Blackwell era, Meta's custom solution "Ariel" continued similar logic. The standard GB200 specification pairs two B200 GPUs with each Grace CPU, whereas Ariel changed this to a one-to-one configuration, halving the number of GPUs. The reason was again to increase the CPU ratio for RecSys workloads.

According to SemiAnalysis calculations, the total cost of ownership (TCO) of the Ariel NVL36x2 solution is 14% higher than the standard GB200 NVL72, and the additional expenditure bought more CPU and DRAM resources, which are exactly what the large model teams do not need. By adopting a cross-rack connection scheme to compensate for the networking gap caused by the reduced number of GPUs, issues with network latency and reliability arose. Meanwhile, the NVL72 backplane, initially considered unstable, has gradually matured, proving Meta's judgment of this technical risk wrong in hindsight.
SemiAnalysis stated that Meta's GB300 servers have returned to standard configurations—this in itself is a negation of the Ariel solution.
Cultural Roots: Short-Term Assessments Override Long-Term Strategy
SemiAnalysis attributes the above problems to the organizational culture of Meta's infrastructure team.
The institution pointed out that Meta's six-month performance review cycle eliminates the bottom 10% to 15% of employees each round, leading teams to pursue short-term visible results rather than long-term technical layouts. Some managers are keen to promote high-visibility but quickly deliverable projects, shifting focus rapidly after delivery, a practice internally known as "window washing." Meanwhile, few people openly challenge superior decisions, making it difficult to correct errors in time.
The supply chain team has limited say in engineering decisions, further exacerbating these issues. SemiAnalysis noted that some suppliers have lowered the priority of Meta's new projects due to frequent changes in direction,转而 prioritizing the needs of Amazon or Google.
SemiAnalysis likened this phenomenon to the expansion history of Meta's Reality Labs—before massive layoffs arrived, billions of dollars were invested in huge engineering teams and R&D projects. Now, as Meta sets out to sell compute power to external customers, the cost of this internal culture will be more directly exposed to the market.
