If xAI defaults on its debt, Apollo Global Management ends up in the GPU rental business. That is in the contract, signed in June 2025, on a five billion dollar debt facility arranged by Morgan Stanley. The lenders have the right to take over Colossus, the company’s 200,000 GPU cluster outside Memphis, and rent it to other AI companies until the loan is repaid.
The interesting question is whether Apollo, or Diameter Capital Partners, or any of the other lenders now financing the AI buildout this way, would want to exercise that right.
The hard question is what they would actually be holding if they did.
A GPU cluster bears little resemblance to a building. Its value at any given moment depends on how it has been provisioned, how it is currently performing, and whether the team that knows its quirks is still there. All of that sits off the lender's balance sheet, beyond the reach of anyone they can call.
This is one of the center problems of the AI infrastructure boom. I cannot determine why no one is talking about it. Tens of billions of dollars in debt is now collateralized by chips whose value depends on operational state, and the operational state is invisible to the people pricing the debt.
This week’s CipherTalk is about what happens to a specific kind of debt when the collateral itself can walk out the door with the operations team.
At the scale these GPU clusters operate, hardware and systems break constantly. Keeping them productive is a craft.
Modern data center GPUs fail at roughly 9% annually. The number traces to Meta’s Llama 3 technical report, which documented 419 unforeseen disruptions across 16,384 H100s over 54 days of training, of which 148 were GPU failures and 72 were HBM3 memory failures. At 200,000 GPUs, that annualized rate works out to approximately 50 GPU failures every day. At xAI’s stated million-GPU target, Epoch AI projects a failure roughly every three minutes. These are not catastrophic events. They are the steady state.
The failure modes that matter for a credit person are the ones that do not look like failures. Silent data corruption (SDC) is the most expensive, where a faulty GPU produces wrong answers without crashing anything, which means a multi-day training run can complete normally and the resulting model weights are quietly poisoned. Cascading failures are the second category, where one bad GPU crashes a training job spread across thousands of others, costing days of compute. Then there are the routine ones: thermal throttle, ECC memory errors, NVLink flap, GPUs falling off the bus.
NVIDIA built NVSentinel because traditional monitoring detects these problems but rarely fixes them. Crusoe built AutoClusters because queue wait time is the largest controllable variable in cluster goodput. Without these tools, remediation timelines run hours to days.
The job of an operations team is to keep all of this in steady state. They know which racks run hot in summer, which cooling loops have been flaky since the last firmware update, which jobs to re-route when a node degrades but has not failed yet. None of that knowledge is written down. It lives in the team.
This is the asset that serves as collateral for tens of billions of dollars in debt and counting.
In the last eighteen months, AI infrastructure went from being financed by corporate debt, to being financed by the chips themselves.
The xAI Colossus 2 SPV is the cleanest example. The structure is roughly $7.5 billion in equity, with up to $2 billion of that contributed by NVIDIA itself, and $12.5 billion in debt. The special purpose vehicle (SPV) purchases NVIDIA GPUs and leases them to xAI on a five-year term. Apollo and Diameter sit on the debt tranche. Valor Equity Partners leads the equity. The debt is collateralized by the chips, not by xAI’s broader balance sheet.
Look at the pricing: xAI’s $5B round was priced at up to 12.5%. CoreWeave's GPU-backed deals priced at roughly 8.5% above the benchmark rate, before terms tightened as lenders got more comfortable with the structure.
If we assume here these are not unsophisticated lenders, then we have to assume they are charging what they think the risk costs. The premium is then, the price of guessing.
The scope is wider than one company. CoreWeave alone holds $18.8 billion in GPU-collateralized debt across multiple SPVs. FluidStack’s $50 billion deal with Anthropic uses a different wrapper, with Google providing a backstop on the lease payments, but the underlying logic is the same.
Every neocloud and most major AI labs are now financed this way.
Every other major asset class that gets used as collateral at this scale has decades of price discovery infrastructure behind it. GPUs have almost none of it.
Aircraft have ISTAT-certified appraisers, a global registry, standardized maintenance logs, ferry pilots, and an active secondary market dating back to the 1970s. Ships have BICA. Cars have NADA. Class A office space has standardized cap rates and vacancy comps. Oil has had a forward curve since the early 1980s.
GPUs have Silicon Data’s H100 Rental Index on Bloomberg terminals, which launched in 2024, and Ornn AI, which raised $5.7 million in October 2025 to build the first regulated exchange for GPU compute derivatives. That is the entire price discovery infrastructure for an asset class now backing tens of billions of dollars in debt.
The price moves underneath all of this are wild. H100 hourly rental rates went from roughly $8 per hour in early 2024 to $1.70 by October 2025, then surged 40% back up to $2.35 by March 2026 on a wave of inference demand nobody had priced in. SemiAnalysis put it bluntly: lenders who used six-year depreciation schedules now look smarter than the analysts who chastised them for being too generous. They were guessing, and they happened to land closer to the right answer than the people calling them reckless. No aircraft lender or shipping lender would underwrite five-year debt against an asset whose price swings like that without a way to hedge it. They would not be allowed to.
CoreWeave’s GPU-backed loans price at roughly 8.5 percentage points above the benchmark rate. For comparison, a typical aircraft loan prices at 1 to 2 points above benchmark, and a commercial mortgage usually sits below that. The extra 6 to 7 points is what lenders charge to bear a risk they cannot measure. There is no GPU futures market, no standardized residual value curve, and no way to lock in a forward rental rate. The premium is is the price of underwriting in the dark.
That spread should compress as the market matures. Hedging instruments will appear. Residual value curves will get more standardized. Secondary markets for used GPUs will deepen. When that happens, the cost of capital for AI infrastructure drops meaningfully, which changes who can build at scale. The companies that benefit are not the ones with the cheapest GPUs today. They are the ones positioned to access cheap debt once the financing infrastructure catches up to the asset class.
The public fight over how fast GPUs depreciate is a tell about how confident the people writing the books actually are.
CoreWeave depreciates GPUs over six years. Nebius, with the same business model and the same hardware, depreciates the same chips over four. AWS, Microsoft, and Google all moved their server useful-life assumptions from three to four years up to six years in 2023, a change that reduced reported depreciation expense by roughly $18 billion annually across $300 billion of combined capex. CoreWeave made the same accounting change in January 2023, before going public, lowering reported expense by hundreds of millions of dollars per year.
NVIDIA announced in 2025 that it is moving from a two-year product cycle to a one-year cycle. The chips backing all of this debt are about to become previous-generation twice as fast.
Michael Burry’s claim is that hyperscalers will cumulatively understate depreciation by approximately $176 billion between 2026 and 2028. He projects Oracle will overstate earnings by roughly 27% and Meta by roughly 21% by 2028. Burry’s motives aside, the math is independently checkable. If the true useful life of frontier-training GPUs is closer to two to four years and the books say six, the gap between paper value and recovery value is real and it is enormous. The recent inference demand surge complicates this. If H100s genuinely have productive life past frontier training, six years may not be wrong. If demand softens again in 2026 or 2027, the writedowns hit at exactly the moment lenders need their collateral to be worth something.
GPU collateral has three different values, and the market is currently pricing only one of them.
Face value is what the SPV says, the purchase price minus straight-line depreciation on whatever schedule the borrower picked. This is the number that determines loan-to-value covenants and the amount of debt the deal can support.
Liquidation value is what a buyer pays in distress. Secondary market data shows moderately-used 2 to 3 year old GPUs trading at 50% to 70% of new pricing under normal conditions. In a default scenario where multiple neoclouds are stressed simultaneously, the buyer pool collapses at the same moment supply spikes, plausibly putting recovery at 30% to 50% of face value in a fire sale.
Going-concern value is what the cluster is worth as a working asset to the next tenant, which depends entirely on whether operational handoff works.
This is where the operational reality from the first section returns. The lender exercising step-in rights inherits a colocation facility owned by someone else, with that facility’s own contracts and constraints. They inherit credentials and topology knowledge that historically lived with the borrower’s operations team, which walked out the door at default. They inherit a market where rental rates already moved 60% in one direction and 40% back the other in eighteen months, with no hedging instrument available. They inherit an asset class where 50 chips a day fail and somebody has to know which racks have been flaky for the last quarter.
The spread between face value and going-concern value is the entire risk that nobody has hedged.
The most telling positions in this market are the ones not being taken.
KKR has been the most aggressive private equity firm in data centers, with the CyrusOne acquisition alongside Global Infrastructure Partners in 2022 for $15 billion, the Global Technical Realty commitment in 2026 for $1.5 billion, and the STT GDC deal in February 2026 for $5.1 billion at a 75% stake. KKR’s digital infrastructure book is a central pillar of $186 billion in real assets. The firm is not in the AIP consortium that bought Aligned Data Centers, not in any xAI SPV, and not in CoreWeave’s debt facilities. KKR owns the buildings, the power, the cooling, and the land, the infrastructure layer that holds value regardless of which AI lab wins or which chip generation dominates.
Peter Thiel sold his entire NVIDIA stake in Q3 2025 and rotated into Apple and Microsoft. The chips are not the durable asset, and the financing structure pricing them as durable will eventually have to reckon with what the chips actually are.
Aircraft became financeable because someone built the registry, the appraisers, and the maintenance logs. Ships became financeable because someone built BICA. The interest premium on these deals exists because no one can answer two basic questions: Is the cluster still working? And: Will it still be working in three years? Answer that, and the cost of building AI infrastructure drops.