HyFlux LCA Independent Validation — Cirium-Calibrated Reproduction and Boeing CASCADE Parity Study

DL-071 · 2026-07-22 · 733/733 regression tests at issue · Track A (CASCADE parity) and Track B (independent validity) never merged · decision_grade: false
Companions: the position page and the full technical paper. This study answers a different question from both: whether the model reproduces Boeing's public formulation (it does, for every disclosed equation tested), and whether it stands up when driven by observed airline operations rather than illustrative assumptions (supported with caveats, on a five-airport measured day). Equation parity, empirical validity and decision readiness are three different tests.

Executed as a DELTA over DL-001–070 (DL-054 precedent): requirements already satisfied by the standing programme are CITED with their delivering artefact; genuinely new work (the Cirium-calibrated operational layer, §3–5/§10–12/§16–17 deltas) was EXECUTED this session. Track A and Track B are never merged anywhere in this report.

0. Requirement → artefact matrix (§1–§24)

Prompt § Status Delivering artefact
§1 two-track discipline ✅ standing Two-track frozen baselines DL-006/012/022; §14.6; never merged since Phase 0
§2 access-dating ✅ standing+new rst-provenance-register.csv; DL-060/061 live fetches access-dated; Cirium manifest.json (access timestamps, auth mode)
§3 Cirium extraction ✅ NEW (scope-limited) data/cirium-snapshot/: fetch-snapshot.py + fetch-archetypes.py (endpoint manifests in-code), manifest.json (call counts, labels); §1 below records coverage LIMITS honestly
§4 baselines ✅ NEW (B2–B4) / ⊘ B1 Five archetype airport-days (B4); day 2026-07-22 as representative-season single day (B2 proxy, stated); peak = observed 06:00 BRS bank + LHR/AMS full days (B3 proxy). B1 global annual: NOT AVAILABLE from the licensed trial (see §1)
§5 segmentation ✅ NEW cirium-validation.ts: 5 stage buckets × 5 classes, per-segment flights/aircraft-km/ASK/fuel/shaft/LH₂
§6 functional units ✅ standing v5 functional-units.ts + claims-register FU columns; per-fuel FU discipline throughout
§7 boundary matrix ✅ standing cascade-coverage-matrix.csv (per-equation), v5 report §3.1/§5.8, V4 report §2
§8 CASCADE app capture ✅ standing / partially gated DL-008/015/036/041 chart captures (Paul-supplied images, digitisation ±2 pts); app is client-rendered — scenario-grid interaction remains gated (letter items B2/B4/B6); admissible-envelope rule already applied (DL-041, bounds pages)
§9 parity tests ✅ standing Exact: hydrogen CI closed form (eqs 137–148, re-confirmed live DL-061), fleet-renewal 26 eq, CORSIA map, methane benchmarks, usable fuel (5/5 in envelope); classifications per boeing-cascade-parity-neutral-wording.md
§10 adoption model ✅ NEW 15 scenarios in cirium-validation.ts, each labelled "scenario — not a forecast"; addressable-fleet envelope (≤3,000 km, no widebody)
§11 dual-method energy ✅ NEW Method A (class per-km curves) vs Method B (parametric per-seat); divergence stats per airport (below); both RECONSTRUCTED Tier-1 ±20%
§12 infrastructure ✅ NEW + standing sizeInfrastructure() per archetype/scenario (CONCEPTUAL grade, stated); credit separation carried from ladder.ts/demand.ts (DL-035)
§13 climate ledgers ✅ standing v5 dual GWP20/100 ledgers never mixed (V5-DL-073); methane invariant test-locked (leak forcing survives origin)
§14 loss-fate tests ✅ standing transfer_loss_fate parameter (DL-061 C1); 15.40× boil-off-subset vs ≈1.11× all-vent bracket, both published
§15 sensitivity ✅ standing+new Morris/Sobol/MC/robustness (v5); NEW correlated-draw check (below)
§16 data-quality audit ✅ NEW (within scope) §1 below: dedup (codeshare-free operating legs), status labels, cancellations retained, generic-equipment counts, known undercoverage listed
§17 falsification ✅ standing+new DL-060 independent reconstruction (Sobol/break-even/venting); NEW: grid-resolution stability + correlation-structure test (both PASS = failed falsifications)
§18 error taxonomy ✅ standing DL-060/061 C1–C7 classification (severity-graded); RN-1 opponent-labelling error class already used
§19 claims register ✅ standing+new public-claims-register-v5.csv (18 cols) + rows CR-V5-038…041 added this session (Cirium layer)
§20 deliverables A–H A=this report §7; B=cascade-coverage-matrix + Boeing-Reproduction-Report §19; C=§1; D=§4; E=§3; F=§5; G=correction-impact-summary-dl061.md + Appendix A (v3 paper); H=repo (seeded, pinned, manifest SHAs) — public reproducibility NOT claimed (code private, licence-restricted data)
§21 decision framework ✅ standing decision-map.ts (DL-038) + evidence grades in registers
§22 integrity rules Rule-by-rule: no CASCADE ground-truthing (two-track), Cirium coverage limits stated (§1), status labels carried, pathways never blended (v5), horizons never mixed, FU discipline, recovered vapour ≠ emission (v5 partition), no renewable credit on leaks (invariant test), unknown fate ≠ zero (C1), all-flight labelled bounding case (in-code caveat), CONCEPTUAL grade on infrastructure, no public-reproducibility claim, precision bounded (Tier-1 bands), no silent repairs (letters instead), corrections published regardless of direction (G log)
§23–24 §6–§8 below

1. Cirium operational dataset report (deliverable C)

Coverage & honesty. The licensed access is a Cirium Sky API trial: flight status/schedules by airport-day, airports, delay index, weather. It does NOT expose: load factors, seat counts, cargo tonnage, fleet age/renewal, deliveries/retirements, utilisation, annual global totals. Prompt §3 items in that list are classified Not testable from licensed access (API-coverage uncertainty class, §15) — NOT imputed. Seats are per-class NOMINAL reconstructions (labelled); freighter identification is an equipment-code + carrier heuristic (labelled reconstructed).

Extraction. 42 API calls lifetime (10 DL-070 + 32 DL-071), static-token auth (DL-069), endpoints and parameters archived in fetch-snapshot.py / fetch-archetypes.py; raw payloads immutable on disk (gitignored for licence; regenerable); distillation deterministic (distil*.py); manifest.json carries access timestamps and call counts. Labels: status L/A = observed actuals; status S = schedule-derived; distances = computed great-circle (appendix coordinates; flown-track ≠ great-circle is a stated limitation); cancellations/diversions retained as status codes; codeshares: operating-carrier legs only (API returns operating flightId once).

Dataset. Day 2026-07-22 (in-season, mid-week): BRS 131 dep/124 arr · LHR 695/712 · AMS 718/728 · EMA 51/55 · JER 25/25 = 1,620 departures with 100% distance coverage; 316 airports; 66 equipment types.

Known undercoverage (§16): single day (no seasonality), no GA/military/charter beyond scheduled ops, freighter heuristic incomplete (EMA night cargo bank partially outside equipment list), block hours not extracted (operational times present in raw for future use).

2. Segmentation (deliverable C, §5)

Stage-length mix is archetype-defining (real day): LHR — 11% <500 km, 38% 500–1,500, 9% 1,500–3,000, 17% 3,000–6,000, 25% >6,000 km, 285 widebody departures; JER — 84% <500 km, zero widebody, 40% turboprop; BRS/EMA intermediate. This single contrast already answers §23 Q11 in direction: archetype changes the addressable set, hence the fuel ranking's applicability, before any chemistry is compared.

3. Route-level validation (deliverable E, §11)

Two independent mission-energy reconstructions (both Tier-1 ±20%, RECONSTRUCTED): Method A (class × km + fixed adder) vs Method B (per-seat parametric with short-sector penalty). Mean |Δ| per airport: BRS 9.1%, EMA 9.1%, AMS 13.6%, JER 15.1%, LHR 16.6%; routes >25% divergence are counted, not hidden (LHR 288/695 — driven by the widebody long-haul fleet where class-average curves are weakest). Verdict: credible at screening grade within class bands; not mission-design grade — consistent with the standing §8 limitation (no Breguet-level mass integration).

4. Airport infrastructure report (deliverable D, §12) — CONCEPTUAL grade

All-flight single-fuel LH₂ (Boeing-style BOUNDING CASE, not a forecast) vs 30%-addressable adoption, real-day throughputs:

Airport (archetype) All-flight LH₂ t/day Liquefaction MW 30%-addressable t/day Trucks/day @30%
LHR (land-constrained hub) 5,151 2,146–2,790 132.6 34
AMS (connecting hub) 2,847 1,186–1,542 199.2 50
BRS (medium national) 217 90–118 62.6 16
EMA (cargo-intensive) 101 42–55 19.8 5
JER (island/remote) 11 5–6 2.3 1

The all-flight/realistic gap is 16–39× at hubs. LHR's bounding case implies a dedicated ~2.1–2.8 GW continuous electrical load for liquefaction alone — an industrial-energy-system statement, corroborating DL-035's "industrial LNG facility, not a retrofit" finding with observed throughput. v5's 1,000 kg/day boundary = 0.34% of full-conversion BRS (DL-070) — early-adoption scale, stated.

5. Falsification report (deliverable F, §17)

Attempt Outcome
Grid-CI dominance under narrower/broader ranges FAILED to falsify (v5 robustness: 3 seeds × 2 sizes × 2 range sets × both horizons; T1/T3 regime edges documented)
Propulsion second, boil-off marginal FAILED to falsify (same; boil-off top-3 only at grid ≤10 g/kWh)
15.40× vs fate assumptions PARTIALLY SUCCEEDED historically — this IS audit finding C1: collapses to ≈1.11× under fate='vented'; both published, bracket carried
1.11× lower-bound derivation Verified (81.06/73.33, pinned)
77/99 e-methane & 45/54/81 fossil reproduction FAILED to falsify (DL-060 independent RNG/estimators; DL-064 relabel changed labels, no numbers)
Crossover-band stability to grid resolution NEW, FAILED to falsify: 11→26 grid columns keeps the 2.3% crossover inside 175–200 (±one coarse step) — test-locked
Correlation structure vs Sobol ranking NEW, FAILED to falsify: grid↔electrolyser ρ≈0.5 correlated draws — grid variance share stays >0.5 and dominant; boil-off <0.05 (labelled conservative interim method, not a Sobol replacement) — test-locked
Methane pathway separation changes rankings Confirmed it does (3.7× central-cell spread — the separation is necessary, standing v5 result)
GWP20 corrections material Confirmed material (V5-DL-073: +13% methane mission CO₂e20)
Archetype changes ranking applicability SUCCEEDED in the useful sense: certification-paced adoption is ISLAND-FIRST (JER 40% addressable-by-turboprop vs ≤5% elsewhere); widebody/long-haul (LHR 42% of departures ≥3,000 km) is outside the LH₂ envelope entirely
Mixed-fleet vs all-flight materiality SUCCEEDED against all-flight framing: 16–39× demand gap (test-locked ≥2×)
Fleet-renewal constraint binds Confirmed: 15%/yr cap < any adoption step ≥20%
Route-level vs aggregate Divergence quantified (§3): aggregates robust at screening grade; route-level long-haul weakest
Different aircraft energy models change regime FAILED to falsify at screening grade (methods agree ≤17% mean; both keep the same archetype ordering)

6. Core questions (§23) — condensed answers

  1. Yes under Boeing-equivalent assumptions for every disclosed formulation tested (hydrogen CI EXACT; renewal curves EXACT; usable fuel in-envelope; methane under the single-multiplier reading their reply forces — with the 596/558 constant question OPEN, letter A1). 2. Not reproducible from disclosure: reference-aircraft tables, airport footprint values, cost-chart inputs, stage assumptions (letter B2–B6), CI_loss=558 derivation. 3. Cirium data supports the DEMAND STRUCTURE HyFlux now uses (real schedules, class mixes, peaks); the trial cannot test fleet-age/load-factor assumptions. 4. Mixed-fleet vs all-flight: 16–39× at hubs (measured day). 5. Route-level energy: credible at Tier-1 screening (9–17% cross-method), not design grade. 6. Infrastructure consistent with Cirium-derived throughput at CONCEPTUAL grade; LHR all-flight = GW-scale electricity. 7. Grid CI remains dominant incl. under the new correlation test. 8. Methane leakage anchors are measured classes (Table 1, v5); hydrogen loss fates remain the L1 epistemic gap — defensible AS DECLARED UNKNOWNS. 9. GWP20/100 lead to materially different methane decisions (1.9× mission spread) — decision-relevant, both always reported. 10. Propulsion architecture changes demand (η ratio in every LH₂-equivalent), not CI rankings (ST 0.101). 11. Archetype changes the addressable set (island-first turboprops; hub widebody excluded) and infrastructure class. 12. Robust across all tested scenarios: grid-CI dominance; pathway-separation necessity; venting-as-architecture-choice (boil-off subset); mixed-fleet≪all-flight. 13. One-or-two-parameter-dependent: 15.40× (fate), methane's short-term case (GWP basis + leakage class), e-methane crossover (grid CI). 14. Highest-value measurements: transfer/chill/purge loss fate; airport LNG ops; flight CH₄ slip; aircraft-tank boil-off (unchanged from v5 — now with the demand side de-risked by real data). 15. Suitable for screening and policy-structure analysis; NOT for investment or engineering design. 16. To decision-grade: the four L1–L4 measurements, load-factor/fleet data (licensed), route-resolved traffic model integration (Part B), Boeing evidence-gap closures.

7. Executive validation report (deliverable A)

Boeing DEMONSTRATES a coherent, mostly-reproducible public methodology (hydrogen exactly; methane under their clarified reading) and ASSERTS beyond it in communications (45.7%>bounds; alt-text; citation swap — letter A5–A7). HyFlux REPRODUCES every disclosed formulation tested and now runs on OBSERVED demand structure (1,620 real departures, five archetypes). Cirium data SUPPORTS the demand/segmentation/peak layer; it cannot yet test fleet-economics assumptions. REMAINS UNCERTAIN: the four loss-fate/ops measurements (epistemic), Boeing's undisclosed inputs, single-day seasonality. HYFLUX ADDS: fate-conditional venting brackets, executable pathway-separation proofs, dual-opponent break-even maps, per-operator CORSIA, DES with external calibration (OTP15 0.931 vs 0.939 measured), and the mixed-fleet demand correction. ROBUST: grid-CI dominance, pathway separation, all-flight≫realistic. FRAGILE: any single-number venting claim without fate qualifier; methane rankings without horizon+leakage class; LH₂ beyond 3,000 km. DECISION-GRADE: false (screening grade; decision_grade:false is already printed on every public page).

8. Required final statement (§24)

HyFlux CASCADE parity status: Exact reproduction for all disclosed formulations tested (hydrogen CI, fleet renewal, CORSIA, usable-fuel within envelopes); Documentation ambiguity on CI_CH₄,loss (558/596, letter A1); Not reproducible from disclosed evidence for reference aircraft, airport footprints, cost inputs. HyFlux independent validity status: Supported with caveat — internally consistent, audit-corrected (C1–C7 + V5-DL-073), externally calibrated on one measured airport-day (OTP15 residual −0.008); caveats: Tier-1 energy reconstructions (±20%), four L1–L4 unmeasured loss fates, single-day operational sample. Cirium calibration status: Conditional — demand structure, class mix, stage lengths, peaks and delay statistics now observed (42 calls, five archetypes, zero-network reuse); fleet age, load factors, utilisation and annual totals not exposed by the licensed trial and therefore untested. Decision-grade status: false. Most robust finding: Within tested ranges, grid carbon intensity dominates electro-fuel lifecycle variance (ST 0.901; survives narrowed/broadened ranges, both horizons, seed/sample sweeps, and a ρ≈0.5 correlated-input test), while realistic mixed-fleet adoption is 16–39× smaller than all-flight bounding cases at real hubs. Most consequential uncertainty: The atmospheric fate of LH₂ transfer/chill-down/purge losses (72.44 kg/day, 89% of chain losses) — it alone moves the venting headline between 15.40× and ≈1.11×. Highest-value next measurement: Instrumented vent/flare/capture split of transfer and chill-down losses at an operating LH₂ facility. No conclusion extends beyond: grid 0–250 gCO₂e/kWh × boil-off 0–2× reference; measured leakage classes 0.19–9.4%; the five-airport 2026-07-22 operational day; Tier-1 class-average energy models; declared per-fuel propulsion architectures (0.465/0.45).

Equation parity, empirical validity, and decision readiness are three different tests; this report keeps them separate, and passing the first two at screening grade does not confer the third.