In August 2026, Orbyfy completed a systematic accuracy program across the Earthflow Environmental Intelligence Layer. Every constant in the pipeline was traced to an authoritative source. Every estimate was labeled as an estimate. Every module was re-verified against ground truth, with end-to-end runs across all 48 continental states. Where old values were wrong, they changed — and some changed substantially.
This page reports measured accuracy, not claimed accuracy. A previous edition of this report cited a 100% match rate across 73 ground-truth comparisons. That figure was arithmetically true and structurally misleading: the fields being asserted rarely overlapped the fields most capable of being wrong. We have replaced it with a larger, harder test set — and we publish the bias, the misses, the modules that have no external assertions yet, and the two findings still under investigation. We believe this level of disclosure is what reinsurance-grade transparency requires, and we would rather show you a measured 88% than a hollow 100%.
Of the 7 discrepancies, five trace to a single root cause — a satellite-temperature averaging bug the new accuracy gates caught, root-caused, and fixed the same day (Chapter 2). The remaining two are open findings, published in Chapter 4.
If you received Earthflow output before August 2026, the corrections below are the ones most likely to matter to an underwriting or engineering decision. Each row is a verified before/after from the program record.
| Output Family | What Was Wrong | Correction & Representative Delta |
|---|---|---|
| Hail expected annual loss | The loss formula charged every historical event the break probability of the single largest recorded stone — a systematic overstatement, worst in hail alley. | Each event’s own fragility is now summed over the real event-size distribution. EAL values are 2–7× lower: a corn-belt anchor site corrected from $352,924 to $53,463 /MW/yr; a Chicago-area site from $253,060 to $33,017. |
| Wind design pressures | A return-period conversion factor was applied to map values that are already ultimate-basis design speeds — inflating every design pressure by ≈1.39×. | Double-count removed; design pressures are ≈28% lower (e.g. velocity pressure 41.7 → 29.9 psf at a Texas anchor site). Measured multi-year site wind can now only raise the design value above the code map floor, never lower it. A flat hurricane-zone box was replaced with continuous coastal-distance decay, correcting inland-South sites that read up to 1.9× too high. |
| Design storms | Storm depths came from state-average coefficients — two desert sites 300 km apart reported identical values. | Now a live per-site federal precipitation-frequency point estimate. The two desert sites correctly differ (24.3 vs 22.5 mm); several convective-plains sites rose (68.4 → 87.1 mm 1-hr/100-yr at one Texas site). The two Pacific-Northwest-coast states outside the federal point service’s coverage are honestly labeled degraded, never faked. |
| Flood determinations | An empty map response was silently reported as “minimal zone, high confidence”; total lookup failure guessed a coastal zone; an uncited regional multiplier scaled the score up to 2.09×. | Three-way honesty: mapped-minimal-zone vs not-mapped vs unknown, each with matching confidence. The fabricated fallback zone and the regional multiplier are deleted; the score is now pure evidence. Live verification produced the pipeline’s first true special-flood-hazard-area positive at a Gulf-coast probe point. |
| Lightning strike inputs | Total (cloud + ground) flash density was fed into the ground-strike equation, and a 10-region lookup table shadowed real local gradients. | A gridded satellite optical lightning climatology with bilinear interpolation now provides per-site flash density, converted to ground-flash density via the published continental mean ratio. Risk levels generally shift one tier lower — the old inputs double-counted in-cloud flashes as ground strikes. |
| Carbon offset | A flat, uncited ≈1,500 tons/MW/yr base with guess multipliers. | Modeled site generation × regional grid marginal-emissions rates. Anchor-site values now disperse 824–1,242 tons/MW/yr with the physically correct direction: coal-heavier grids displace more CO₂ per solar MWh. |
| PV bifacial & thermal | Every site silently used fallback inputs: a universal 7.14% bifacial gain, tiered air temperature, 2.0 m/s wind. | Per-site measured surface albedo, temperature, and wind now reach the performance models: bifacial gain disperses 4.8–8.5% across the anchor sites, and each input carries a provenance label (per-site-real vs degraded-fallback). |
| Rainfall erosivity | A synthetic latitude-band estimate polluted the multi-source precipitation mean, and unresolved states silently defaulted to Texas coefficients. | Real sources only; state resolves for every continental coordinate; a published precipitation-only regression covers the no-state case with a degraded label. A Mojave anchor site’s erosivity corrected 21 → 13 as the synthetic overestimate of arid precipitation left the mean. |
Table 1.1 · Principal value corrections from the August 2026 accuracy program. Full enumeration in the Methodology Update Log.
The table below replaces the previous edition’s “100% All-Match (73)” headline. It reports, per module, the signed bias, mean absolute error, and RMSE of Earthflow outputs against the expanded truth set — 158 comparisons at 21 sites in 16 states. Each module also carries a bias gate with a source-uncertainty band; the gates run on every portfolio capture and fail the build when a module drifts.
| Module | n | Bias | MAE | RMSE | Notes |
|---|---|---|---|---|---|
| Design-storm precipitation | 22 | −0.0% | 0.01% | 0.02% | External truths against live federal precipitation-frequency point estimates — the pipeline reproduces the federal point service essentially exactly. |
| Lightning climatology | 24 | 0.0% | 0.0% | 0.0% | Pipeline-consistency truths: verifies the bundled satellite optical lightning climatology is read, interpolated, and converted correctly. Not an independent re-measurement of the climatology itself. |
| Solar resource | 17 | −1.3% | 5.1% | 7.8% | Compared against reference-grade national ground stations; small low bias, within the stations’ own interannual spread. |
| Precipitation / erosivity | 7 | +1.7% | 5.8% | 8.1% | Multi-source precipitation mean after removal of the synthetic latitude-band input (§1.2). |
| Wind design speeds | 19 | +2.2% | 6.4% | 12.5% | Post-correction of the return-period double-count. Small conservative (high-side) bias. |
| Weather-radar station selection | 3 | +3.2% | 3.2% | 5.5% | Station identity and great-circle distance to the national weather-radar network. |
| Tornado | 5 | −5.9% | 11.2% | 16.4% | Event-frequency and design-figure comparisons. |
| Peat / organic-soil fire | 1 | −6.7% | 6.7% | 6.7% | Single external assertion so far — treated as indicative only. |
| Hail | 9 | +6.9% | 7.4% | 11.6% | Event-catalog counts and derived quantities; slight high-side bias, within the ±15% gate reflecting catalog-count uncertainty. |
| Seismic hazard | 10 | +2.6% | 22.6% | 26.2% | Low aggregate bias but wide per-site spread (26% RMSE) — individual hazard values scatter around the reference even though they do not lean one way. Flagged for continued refinement. |
| Satellite land-surface temperature | 7 | +53.5% | 53.5% | 58.5% | Bias gate FAILED (±10% band) on the captured revision — a real bug, found by these gates, fixed the same day. See §2.1. |
Table 2.1 · Measured per-module accuracy, August 2026 capture. 158 comparisons total: 126 MATCH · 13 loose match · 7 discrepancies · 12 missing-field assertions (Chapter 3).
An accuracy gate that has never failed is indistinguishable from an accuracy gate that cannot fail. This one failed, loudly, on a real defect — then confirmed the fix. That is the behavior a reinsurer or engineering reviewer should demand from any vendor’s validation process, and it is why the failure is documented here rather than quietly patched.
The previous edition of this report claimed a 100% match rate across 73 comparisons. The number was real. The problem was coverage: the asserted fields — station identities, catalog counts, resource annuals — almost never overlapped the fields with the most room to be wrong, such as loss formulas, design pressures, and derived scores. A perfect score on an easy test is not evidence of accuracy. We retired the claim.
The replacement test set makes coverage a first-class, published metric: 158 comparisons at 21 sites in 16 states, deliberately extended into loss-bearing and score-bearing field families — and below, the per-module assertion coverage, including the modules that still have zero external assertions. A module listed at 0% has passed internal and cross-module verification but has not yet been independently re-measured; we state that rather than let silence imply otherwise.
| Module | Fields Emitted | Externally Asserted | Coverage | Status |
|---|---|---|---|---|
| Weather-radar station selection | 10 | 2 | 20.0% | Asserted |
| Hail | 27 | 5 | 18.5% | Asserted |
| Lightning | 18 | 2 | 11.1% | Asserted |
| Tornado | 27 | 3 | 11.1% | Asserted |
| Seismic hazard | 47 | 5 | 10.6% | Asserted |
| Wind design speeds | 20 | 2 | 10.0% | Asserted |
| Flood determination | 40 | 2 | 5.0% | Asserted |
| Peat / organic-soil fire | 47 | 2 | 4.3% | Asserted |
| Solar resource | 98 | 3 | 3.1% | Asserted |
| Precipitation | 41 | 1 | 2.4% | Asserted |
| Design-storm precipitation | 156 | 2 | 1.3% | Asserted |
| Satellite land-surface climate | 83 | 1 | 1.2% | Asserted |
| Cross-module & metadata | 302 | 1 | 0.3% | Asserted |
| Frost & winter severity | 15 | 0 | 0.0% | No external assertions yet |
| Financial & performance KPIs | 80 | 0 | 0.0% | No external assertions yet |
| Evapotranspiration | 15 | 0 | 0.0% | No external assertions yet |
| Permitting screens | 18 | 0 | 0.0% | No external assertions yet |
| Road access | 11 | 0 | 0.0% | No external assertions yet |
| Soil erosion (RUSLE) | 68 | 0 | 0.0% | No external assertions yet |
| Soil moisture | 48 | 0 | 0.0% | No external assertions yet |
| Soil characterization | 33 | 0 | 0.0% | No external assertions yet |
Table 3.1 · External-assertion coverage per module, August 2026 truth set. Coverage of a field family is a prerequisite for any accuracy claim about it — a module at 0% carries internal verification only, and we say so.
Ground truth comes from three site classes: federal reference-grade radiometric and surface-measurement ground stations (multi-decade instrument records), the two catastrophic-hail plant sites, and operating utility-scale plants drawn from the national inventory of >10 MW PV facilities in previously unsampled states.
| Site | State | Match | Loose | Discrepancy | Missing Field |
|---|---|---|---|---|---|
| Fighting Jays Solar (operating plant, 2024 hail event) | TX | 15 | 1 | — | — |
| Midway Solar Project (operating plant, 2019 hail event) | TX | 13 | 2 | — | — |
| Federal radiometric reference laboratory · Golden | CO | 13 | 2 | 1 | — |
| Surface-radiation reference station · Table Mountain, Boulder | CO | 3 | 3 | 1 | — |
| Surface-radiation reference station · Desert Rock | NV | 4 | 2 | 1 | — |
| Surface-radiation reference station · Fort Peck | MT | 5 | — | 1 | — |
| Surface-radiation reference station · Goodwin Creek | MS | 6 | 1 | — | — |
| Surface-radiation reference station · Penn State | PA | 3 | 2 | 2 | — |
| Surface-radiation reference station · Sioux Falls | SD | 6 | — | 1 | — |
| Operating utility-scale plant (45.87°N, 120.24°W) | WA | 3 | — | — | 1 |
| Operating utility-scale plant (33.82°N, 115.39°W) | CA | 5 | — | — | 1 |
| Operating utility-scale plant (33.24°N, 112.54°W) | AZ | 5 | — | — | 1 |
| Operating utility-scale plant (38.32°N, 104.40°W) | CO | 5 | — | — | 1 |
| Operating utility-scale plant (34.40°N, 101.60°W) | TX | 5 | — | — | 1 |
| Operating utility-scale plant (29.13°N, 96.26°W) | TX | 5 | — | — | 1 |
| Operating utility-scale plant (41.16°N, 91.17°W) | IA | 5 | — | — | 1 |
| Operating utility-scale plant (39.56°N, 87.98°W) | IL | 5 | — | — | 1 |
| Operating utility-scale plant (42.35°N, 85.06°W) | MI | 5 | — | — | 1 |
| Operating utility-scale plant (30.45°N, 83.19°W) | FL | 5 | — | — | 1 |
| Operating utility-scale plant (38.23°N, 77.78°W) | VA | 5 | — | — | 1 |
| Operating utility-scale plant (39.17°N, 75.95°W) | MD | 5 | — | — | 1 |
Table 3.2 · Per-site comparison outcomes, August 2026 capture. Each of the 12 expansion plants carries one assertion targeting a field name the captured revision did not emit — recorded as a MISSING FIELD failure, not silently dropped.
Two discrepancies from the August 2026 capture remain open and under investigation. We publish them because a validation process that cannot surface an unresolved miss is not a validation process — and because a reader pricing risk from this data deserves to know exactly where the model and its reference expectations currently disagree.
| Finding | Earthflow Output | Reference Expectation | Working Hypothesis |
|---|---|---|---|
| Desert fire-season rating reads low Desert Rock reference station, NV (Mojave) |
Fire risk Low | The federal wildland-fire guide’s regional characterization implies High / Very High for the Mojave fire season | The fire-risk model appears to underrate desert fuel and season dynamics. Under investigation; not quick-fixed, because an untraced correction would repeat the class of error this program exists to remove. |
| Mid-Atlantic hail level reads high Penn State reference station, PA |
Hail level High | Storm-catalog climatology for the area implies Low / Moderate | Hail level thresholds are likely still calibrated to raw event counts rather than severity-weighted climatology. Under investigation. |
Table 4.1 · Open findings, August 2026. Both will be resolved through the same disclosed, gated process as every other change in this program — and this page will be updated when they are.
The truth set itself is also subject to audit. During this program, one golden-record expectation was found to be wrong and corrected: a flood-zone assertion at the Midway site encoded the old “empty map response means minimal zone” convention. A direct map-panel probe proved the point sits outside every digitized flood-map panel — the authority has made no determination there — so the expected value was corrected to not-mapped / unknown. Reference expectations get the same scrutiny as pipeline outputs.
Accuracy at 21 sites answers “is it right where we measured?” Continental verification answers a different question: does the same pipeline behave physically everywhere? Three layers of verification now run at continental scale — one per release program, one per portfolio capture, one on every single deploy.
The August 2026 reference capture ran the full production pipeline end-to-end at 60 sites spanning all 48 continental states: six federal surface-measurement reference stations, 51 operating utility-scale plants drawn from the national >10 MW PV inventory, and three real-land probe points in states with no qualifying plant. All 60 runs succeeded, with the flood, permitting, and carbon modules healthy at 60 of 60 sites. This capture is the reference distribution for the portfolio-percentile KPI and the baseline for future accuracy gates.
Coordinates outside the continental United States are now rejected explicitly with a coverage statement, rather than analyzed on continental fallbacks and reported as a success — a change made after verifying that an Alaska test point had previously run the full pipeline on silent fallback values.
Certain quantities must be monotone along their physical axes at continental scale, no matter what any individual module thinks. The capture is gated on these gradients:
| Gradient Gate | Physical Expectation | Measured on the 60-Site Capture | Verdict |
|---|---|---|---|
| Lightning flash density | Gulf subtropics >> Pacific marine coast | Florida minimum 27.4 > Pacific-coast maximum 1.1 (flashes/km²/yr) | Pass |
| Rainfall erosivity | Humid Southeast >> arid Mountain West | Southeast minimum 305 > Mountain-West maximum 56 | Pass |
| Wind design speed | Hurricane coast >> interior Mountain West | Coastal-Southeast minimum 146 mph > Mountain-West maximum 110 mph | Pass |
Table 5.1 · Continental physical-gradient gates on the August 2026 capture. A pipeline that is merely tuned to its test sites fails these; a physically grounded one cannot.
Between full captures, a fast logic sweep runs on every deploy. In about 15 seconds it verifies, across all 48 states: geographic state resolution (48/48), erosivity-table coverage (48/48), the lightning raster finite and in-range at a point in every state with the correct continental gradient, hail-loss invariants (bounds, monotonicity, empty-catalog behavior), and standing regression guards — the deleted fabricated flood zones stay deleted, the removed regional flood multiplier stays removed, and the wind return-period double-count cannot silently return. The sweep verifies changed logic continent-wide without requiring 48 full pipeline runs per deploy.
Seven fixed anchor coordinates — spanning West Texas, two Mojave/Central-Valley California plants, the Iowa corn belt, the Chicago area, the Nevada desert, and central Florida — are captured live before and after every numeric-change deploy. Each deploy ships with an enumerated accepted-delta list: exactly which fields are expected to move, in which direction, and why. Any field that moves outside the list is investigated before the deploy is accepted.
Earthflow is a Physics AI™ platform: environmental intelligence that starts from established physical and engineering relationships and uses AI where equations alone cannot answer the question. For every continental coordinate, the Environmental Intelligence Layer runs 20 analysis modules and emits roughly 1,190 fields per site. Each module is independent — one upstream outage never blocks the others — and, since this program, every output carries provenance: measured value or labeled estimate, full or degraded status, and the basis it was computed on.
Per standing policy, Earthflow describes each module by the property of its underlying data rather than by vendor or product names; the full provenance chain is available to customers under the platform’s per-field provenance labels.
The previous edition led with this result, and it survives the accuracy program — with one important correction. The classification stands; the loss dollars attached to it in earlier material do not.
The Very High classification at both sites is driven by the long-run climatological event pattern at each coordinate, not by the catastrophic event itself; re-running the pipeline with the event year excluded produces the same classification. It remains, in the truth set, the assertion family with the strongest external outcome evidence.
These seven fixed coordinates anchor the per-deploy harness (§5.4). They were chosen to span the continental extremes the pipeline must get right simultaneously — hail alley, irrigated corn belt, two desert basins, a Great Lakes metro, the Gulf-adjacent Southeast, and the arid Central Valley:
| Anchor | Setting | What It Stresses |
|---|---|---|
| Abilene, TX | Southern-plains hail alley | Hail climatology · convective design storms · high erosivity |
| Solar Star, CA | Antelope Valley desert (operating plant) | Arid precipitation honesty · high solar resource · regional grid rates |
| Topaz, CA | Carrizo Plain (operating plant) | Desert-vs-desert dispersion — must differ from Solar Star, and now does |
| Southeast Iowa | Corn-belt loess | Frost · erosivity state-keying · hail-loss correction magnitude |
| Chicago, IL | Great Lakes metro fringe | Flood-score evidence purity · urban-adjacent flood levels |
| Mesquite, NV | Mojave / Virgin River basin | Arid extremes · real upstream-catchment exposure · measured-wind design floor |
| Central Florida | Gulf-adjacent subtropics | Lightning maximum · hurricane-coast wind · wet-season storms |
Table 6.1 · The golden seven-site deploy harness. Every numeric-change deploy is verified against live captures at these coordinates before it is accepted.
Every constant traced to an authoritative source. Every estimate labeled. Every module measured against independent truth where truth exists — and honestly marked unverified where it does not yet. Every deploy gated, every correction disclosed, every open miss published until it is resolved. That is what this page now reports, and it is the only kind of accuracy claim Earthflow will make going forward.