The Inspection Illusion, Redux
One sortie test in the Caribbean does not retire a decade of maintenance debt. The Navy is measuring the wrong thing again.
The Inspection Illusion, Redux
One sortie test in the Caribbean does not retire a decade of maintenance debt. The Navy is measuring the wrong thing again.
MD Marine Electric | The Current Conversation
A single February test in the Caribbean Sea is now the headline evidence that the Navy's most expensive warship in history works.¹ Chief of Naval Operations Adm. Daryl Caudle told a roundtable at the Tailhook Symposium in Reno that USS Gerald R. Ford (CVN-78) generated higher sortie rates than a legacy Nimitz-class carrier during a stress test of its Electromagnetic Aircraft Launch System (EMALS) and Advanced Arresting Gear (AAG).¹ The test happened before Ford sailed on to U.S. European Command and the Red Sea.¹
That is the entire public record. One test. One class comparison. One admiral's characterization at an industry symposium, not a certified operational assessment released with methodology and data attached.
This is not a story about whether EMALS works. It may well work, on the day it was tested, under the conditions of that test. This is a story about what the Navy chooses to measure, when it measures it, and what gets left out of the press coverage that follows. Call it what it is: the Inspection Illusion, applied not to a shipyard availability but to a warship's entire developmental history. The habit of substituting a favorable point-in-time snapshot for a defensible trend, then briefing the snapshot as if it settles the argument.
The Test That Got Reported, and the Decade That Didn't
Ford has a documented history that predates this Caribbean test by roughly a decade, and none of that history is erased by one good afternoon of launches and recoveries. The ship's EMALS and AAG systems were the subject of years of developmental struggle, cost growth, and schedule slip that is part of the public record on naval shipbuilding generally. None of that appears in the USNI report on Caudle's remarks, because the remarks themselves didn't address it. The CNO was answering a question about a specific test result. He was not delivering a program history.
The problem is what happens next: a single positive data point, cited by the service's top officer, at a public symposium, becomes the operative narrative. That is the Inspection Illusion in its purest form, and it is a pattern this newsletter has documented before, at the compartment level, on availabilities that closed out clean on paper while systems underneath the paperwork remained degraded.² The mechanism is identical whether the paper is a QAR closeout sheet or a symposium quote: a measurement taken under favorable, bounded conditions gets generalized into a claim about the whole system.
[VERIFY: full transcript or written record of Adm. Caudle's Tailhook remarks, to confirm whether he qualified the comparison by conditions, sortie type, or aircraft mix, source: USNI News or Navy public affairs record of the Tailhook Symposium, Reno, August 2026]
What "Sortie Rate" Actually Measures, and What It Doesn't
Here is the deck-plate detail that the symposium quote skips past. A sortie generation rate test measures the cycle time of launch, recovery, and turnaround under the specific flight deck conditions, aircraft loadout, and crew proficiency present that day. It does not, by itself, measure mean time between failures on the EMALS energy storage subsystem, availability of AAG water twister assemblies under sustained high-tempo cycling, or the maintenance burden generated per hundred launches. Those are the numbers that determine whether a carrier can sustain a sortie rate across a deployment, not whether it can hit one during a discrete test window.
A launch and recovery system that performs well during a single stress test and then requires extensive corrective maintenance afterward has not solved its underlying reliability problem. It has simply deferred the accounting. This is the same structure as a ship that completes an availability on time by pushing discrepancies into a follow-on chit package: the schedule looks clean, the underlying condition does not match the schedule. GAO's own maintenance data supports the general pattern, if not this specific ship: even with dock space available, fewer than 40 percent of ships finish their maintenance availabilities on time.³ A system, or a ship class, that looks good on a single measured event and struggles across a maintenance cycle is not an anomaly in Navy shipbuilding. It is close to the norm.
None of this means Ford's test result was fabricated or misleading on its own terms. It means one test result, reported without the reliability data that would contextualize it, is not evidence of a solved problem. It is evidence of a single successful measurement.
The Inspection Illusion Has a Documented Mechanism
This newsletter has traced the Inspection Illusion through GAO-25-106749 before, and the mechanism there is instructive because it shows how the Navy has, on the record, chosen favorable optics over rigorous inspection when the two came into conflict. In 2020, Navy leadership changed inspection procedures specifically to reduce inspections by almost 50 percent, and did so in order to maintain working relationships with contractors.⁴ That is not an inference. That is GAO's finding, stated plainly: the reduction in scrutiny was a relationship-management decision, not a risk-based one.
Apply that same institutional instinct to a symposium remark about the Navy's most scrutinized and most expensive carrier program. The Ford class has absorbed enormous political and budgetary pressure to demonstrate it was worth building. A CNO citing a favorable sortie comparison at Tailhook, in front of the naval aviation community that has the most riding on Ford's success, is not a neutral act of program reporting. It is a public relations data point, delivered by the officer with the most institutional interest in the program's success, at the venue most receptive to hearing it. That does not make the underlying test result false. It does mean the venue and the messenger both belong in the reader's accounting of how much weight the claim deserves.
The Inspection Illusion works because a single favorable measurement is cheap to produce and expensive to contextualize. Producing the sortie rate test took one stress test, in the Caribbean, on the transit leg of a deployment.¹ Producing the reliability data that would tell you whether that sortie rate is sustainable requires years of maintenance records, failure logs on EMALS energy storage groups and AAG hydraulic and water twister subsystems, and comparison against Nimitz-class baseline data across full deployment cycles, not a single transit leg. The first is a talking point. The second is an audit. The Navy produced the talking point.
Where the Bid-to-Win Trap Meets the Inspection Illusion
There is a structural reason the Navy has an incentive to lead with favorable snapshots rather than reliability trends, and it connects to a pattern this newsletter named earlier in the series: the Bid-to-Win Trap. The Navy's move from cost-plus contracting toward fixed-price, per-ship bidding was intended to increase competition among prime contractors.⁵ But fixed-price structures also change what gets reported publicly and when. A program under fixed-price pressure has a stronger incentive to generate and publicize favorable test results early, because those results become the evidentiary basis for continued funding, follow-on procurement, and political cover, all of which matter more when a contractor's margin depends on the program's survival rather than on cost recovery.
Ford is not built under the fixed-price structure that drives the Bid-to-Win Trap in the destroyer and frigate programs USNI has covered. But the incentive structure surrounding a first-in-class carrier facing years of cost growth is analogous: every favorable data point that can be surfaced publicly reduces political pressure on the program, regardless of whether that data point is representative of sustained performance. $1.84 billion was wasted modernizing four Ticonderoga-class cruisers that were decommissioned before they ever deployed after modernization.⁶ That is what happens when the Navy commits resources on the basis of program-level optimism that does not survive contact with sustained operational reality. Ford's sortie test, taken alone, is not that kind of failure. But the instinct to lead with the good news and let the sustainment data lag behind by years is the same instinct that produced the cruiser write-off.
The Reliability Data the Navy Has Not Released
Here is the trade-level specificity that should anchor any credible assessment of Ford's launch and recovery systems, and it is conspicuously absent from the reporting on Caudle's remarks. EMALS is a linear induction motor system drawing on an energy storage group of rotating machines; AAG uses a water twister braking mechanism with hydraulic damper assemblies to absorb landing energy across a defined deceleration profile. Both systems were designed to reduce manning and increase sortie generation rate relative to the steam catapults and Mark 7 arresting gear on Nimitz-class ships. Both systems also carry documented histories of reliability shortfalls during Ford's post-delivery test and trials period. None of that history is addressed in the USNI report on the Tailhook remarks, because the remarks did not address it.
[VERIFY: current mean time between operational mission failures for EMALS and AAG aboard CVN-78, and how that figure compares to Nimitz-class steam catapult and Mark 7 arresting gear reliability data, source: Navy or DOT&E public reporting on CVN-78 operational test and evaluation]
A single sortie rate test cannot answer the reliability question, and it was never designed to. Treating it as if it does is the Inspection Illusion operating at the fleet level instead of the compartment level: a measurement taken under favorable conditions, generalized into a conclusion the measurement cannot support.
For NAVSEA Program Offices
You have an opportunity here that the Ticonderoga write-off did not have: the chance to publish the reliability trend data alongside the favorable test result, before the narrative hardens into consensus. Releasing MTBF figures for EMALS and AAG, benchmarked honestly against Nimitz-class steam and hydraulic legacy systems across comparable deployment cycles, costs you a news cycle if the numbers are bad. It costs you the program's credibility if you wait until the numbers are forced out by GAO or DOT&E years from now, after billions more have been committed on the strength of a symposium quote.
The alternative to publishing the trend data now is repeating the pattern GAO already documented once: reducing scrutiny to protect a relationship, in this case the relationship between the program and the political coalition that wants Ford to be a success story.⁴ That coalition will still want the story in five years. The fleet will not have the luxury of finding out then that the sortie rate test in the Caribbean was the exception, not the baseline.
Publish the maintenance data. The test result can wait its turn.
References
- USNI News, "USS Gerald R. Ford Sortie Test Exceeded Nimitz-class Performance, Says CNO Caudle," August 24, 2026.
- The Inspection Illusion (MD Marine Electric, The Current Conversation, prior installment).
- GAO finding: even with dock space available, fewer than 40 percent of ships finish maintenance availabilities on time.
- GAO-25-106749.
- USNI News coverage of Navy shift from cost-plus to fixed-price per-ship bidding to increase competition.
- $1.84 billion wasted modernizing four Ticonderoga-class cruisers decommissioned before deploying.

