While excitement is our prevailing sentiment, we suspect puzzlement prevails among subscribers. I thought those guys were dead? long-time subscribers might ask. When and why did I subscribe to this? many of the 700 of you who have signed up since our last post might wonder.
Fair questions, both. To the former: closer than we’d care to admit, but we just kept swimming. As for the latter, we’ve been a bit bemused ourselves — hopefully you’ll be glad you were ensnared in Substack’s recommendation flow.
Today we’re introducing Dory, a deep learning model that predicts per-flight delay distributions and cancellation probabilities with a 24-hour horizon. It’s our first graph-based model conditioned on three-dimensional weather forecasts and reduces error by up to 19% versus a 60-day baseline. Predictions are available at app.aerology.ai as a research preview; to join the waitlist for programmatic access, please email waitlist@aerology.ai. And one practical note: the charts in this post are best viewed on desktop.
Complexity and uncertainty render mental models of air travel disruption inadequate. A traveler can’t foresee the many dependencies that snarl a trip; while an operator can (or should), their span of responsibility creates too many permutations to hold in their mind. Both parties are forced into a reactive posture — travelers can only scavenge for open seats, and operators can only triage.
Dory is the foundation for a more proactive posture, for those who distribute disruption and those who deal with it. Take, for example, September 30th at DFW. That morning’s 06:36 CT Operations Plan from the Command Center noted “no terminal impacts expected for [DFW TRACON],” though carried a reference for possible en-route impacts. As late as the 10:28 CT Operations Plan, no ground delay program (GDP) or ground stop (GS) was referenced for DFW. The 12:31 CT Operations Plan would add a possible DFW GDP or GS for arrivals after 16:00 — owing to an updated TAF, which added a 30% chance for thunderstorms. A ground stop would be published at 14:46 CT and a GDP at 15:57 CT. Arrival delays for 393 flights scheduled to arrive DFW between 15:00–21:59 CT would average 38.3 minutes.
Our 08:36 CT initialization, the last to use Google DeepMind’s 00Z WeatherNext 2 forecast1, called for typical — if slightly elevated — evening arrival delays at DFW. Over the preceding 28 days, actual arrival delays averaged 17.9 minutes between 15:00–21:59 CT; as of 08:36 CT, we were predicting expected arrival delays averaging 23.5 minutes.
Dory, however, would find WeatherNext 2’s fresh 06Z forecast when it was initialized at 08:41 CT: its average expected arrival delay increased to 61.8 minutes for those 393 flights. Nearly four hours before a mention in an ops plan and four hours before any advisory would be issued, Dory was signaling disruption that evening. While its distribution was wider than needed (a delay at or above the 95th percentile occurred for only 1.5% of flights), the distribution was approximately centered. The expected share of flights arriving more than 90 minutes late in the 08:41 CT initialization was 17.7% against 13.5% realized. And the expected share arriving more than 60 minutes late? That matched to within two tenths of a percentage point.
By our 14:01 CT initialization, the first with WeatherNext 2’s 12Z forecast, average expected arrival delay was within 2 minutes of actual — still 45 minutes before the ground stop was issued.
More technically, Dory is an ensemble of graph neural networks (GNNs), whose structure enables each flight’s prediction to draw on the state of the National Airspace System (NAS). Each graph also carries a learned representation of Google DeepMind’s WeatherNext 2 forecasts, themselves an ensemble. Averaging the models improves accuracy, but it isn’t what characterizes uncertainty: the prediction head is natively probabilistic, producing departure and arrival delay distributions that are internally consistent with cancellation probabilities. Each seed totals 8.8 million parameters — though nearly negligible compared to today’s LLMs, that’s not coincidentally on the order of Google DeepMind’s first graph-based, machine-learned weather prediction model2.
Why Dory? It’s a nod to DORI, an artifact from Tim’s previous entrepreneurial effort. We also thought there were some parallels to a certain blue fish: it advances the quest and possesses some surprising capabilities. And, in time, we hope it looks toy-like.
We evaluated Dory on 17 U.S. carriers3 over six held-out weeks, from May 8 to June 18, 2026. The lack of a public benchmark hampered this exercise and, we think, more broadly inhibits progress; we expect to work more on inter-comparison in the future and hope it might foster constructive competition. For now, we compare Dory to three baselines:
The base rate measured over the test window itself. Because it’s computed from the outcomes being forecast, which no forecaster could know in advance, it’s the best possible constant forecast.
A 60-day baseline estimated from outcomes over the trailing window. Rates are computed by a carrier › route › four-hour scheduled-departure block › callsign hierarchy, with sparser levels shrunk closer to their parent. Notably, this baseline does not use lead time.
A learned tabular baseline. We trained CatBoost, a gradient-boosted decision tree framework, on the same per-flight features as Dory, though without weather inputs or network context. Like Dory’s, CatBoost’s forecasts vary by lead time; given the importance of lead time in our evaluation framework, CatBoost is the primary comparator.
We define lead time as the difference between our model’s initialization time4 and a flight’s scheduled departure time. Our 24-hour forecast horizon is also relative to scheduled departure time: our maximum lead time is 24 hours and we make a flight’s first prediction 24 hours before its scheduled departure. In deployment, we re-initialize the model every five minutes with the latest information — so we’ll make hundreds of predictions for a given flight. In evaluation, we’ve downsampled to one randomly chosen initialization per two-hour block.
Scoring composite skill
Because Dory’s forecasts are internally consistent and its arrival-delay distribution is continuous, we can estimate the probability of a composite event: that a flight cancels or arrives more than τ minutes late, for any too-late threshold τ. This convention earns the headline by capturing both forms of flight disruption and bringing the delay distribution into binary scoring. The first such binary scoring tool we reach for is Brier skill score (BSS), which measures the forecast error reduction versus the base rate. A skill score of 0 means no better than the base rate; 1 would be a perfect forecast; 0.5 would represent a 50% reduction in forecast error.
Across six τ values ranging from 15 minutes to 3 hours over our pooled 24-hour horizon, Dory5 reduces forecast error by 18.5–20.9% versus the base rate. Against6 the 60-day baseline, 14.2–19.2%; against CatBoost, 7.1–7.9%. When scored head-to-head on individual days in the held-out window, Dory wins against CatBoost on at least 30 of 42 days (τ = 180 minutes) and up to 42 of 42 days (τ = 15 minutes).
This outperformance holds at individual (τ, lead time) targets: Dory’s skill exceeds CatBoost’s on all 144 targets, and confidently on 141 of them7. In the middle of our horizon (e.g., τ = 120 minutes, lead time = 12 hours) and at lower too-late thresholds (e.g., τ = 30 minutes, lead time = 17 hours), Dory reduces error by more than 8% versus CatBoost.
Unpooling lead time also lets us borrow meteorology’s lead-time gain convention, which fixes a skill level and asks at what lead each model reaches it. At τ = 90 minutes, our first forecasts (made 23 to 24 hours before departure) are as skillful on average as CatBoost’s forecasts made around 9 hours before departure, a lead-time gain of approximately 14 hours (95% interval: 7.5 to 16.3). This gain is what enables a more proactive posture — it provides equivalent information the night before or a shift earlier, when the option set is wider.
Unpooling also reveals an important compositional effect in our evaluations: it requires little skill to forecast disruption after it’s readily apparent. To account for this, we’ve re-scored Brier skill while excluding flights that either Dory or CatBoost assign a disruption probability of 90% or more. Consistent with delays and cancellations typically becoming known close to departure, our measured skill at longer lead times varies little when near-certain disruptions are excluded. For example, at τ = 90 minutes, skill decreases by less than 5% at lead times of 20 hours or more.
At shorter lead times, though, this exclusion can remove up to 11% of rows (τ = 15). These rows represent the base rate’s worst misses — removing them makes the base-rate forecaster’s job easier and mechanically depresses skill for Dory and its comparators. Across the six τ values over our pooled 24-hour horizon, when excluding near-certain disruptions, Dory reduces forecast error by 10.8–16.6% versus the base rate (18.5–20.9% when they’re included). However, Dory’s pooled outperformance against CatBoost is largely unchanged (7.1–7.7% when excluded versus 7.1–7.9% included).
Evaluating composite calibration
We led with Brier skill because it rewards both calibration and conviction. The base rate is instructive — it is perfectly calibrated by construction but never makes a bold call, so it has no skill.
Calibration, however, is what makes a single forecast usable. Skill is a statement about many flights at once; a traveler has one trip and one number. A calibrated forecast lets you take its number at face value: whenever it says 30%, about three in ten of those flights are disrupted.
To measure calibration, we sort predictions into sixteen bins by forecast probability (bins are narrower at low probabilities, where most predictions fall). We then compare each bin’s average forecast with how often its flights were disrupted, represented by the dot in the chart below. A perfectly calibrated forecaster’s dots sit on the dashed diagonal. The observed rate is itself an estimate, so each dot carries a 95% confidence interval8, drawn as a vertical line.
Where that line misses the diagonal, we consider the apparent over- or under-prediction detectable: if the forecast were exactly right, a gap that large would turn up less than 5% of the time. A line entirely above the diagonal means disruption was under-predicted for that bin; below, over-predicted. Where a short line touches the diagonal, we consider the bin calibrated to within the distance from the diagonal to the line’s farther end. A tall line touching the diagonal only supports a more permissive claim: we don’t detect a bias and can only bound its size, if it exists.
Across all 36 combinations of threshold and lead time, 59.3% of Dory's predictions — contained in 26% of bins — are calibrated to within 3.1 percentage points, and half of those to within 0.7 points. CatBoost can make the same claim for 42.3% of its predictions, half of which are within 1 point. In half of bins, holding another 9.2% of Dory’s predictions, the line touches the diagonal but supports only the permissive claim: we can’t detect a bias, but it would be at most 5.0 points for half of those predictions, and at most 6.5 for three-quarters.
We detect a bias in bins containing 31.5% of Dory’s predictions (CatBoost, 39.3%). Dory’s propensity for over-prediction concentrates in low-probability forecast bins, by 0.5 points on average. Conversely, Dory is prone to under-prediction in high-probability bins and, more specifically, at 15- and 30-minute too-late thresholds (by 2.8 points on average).
Scoring Dory’s components
To this point, we’ve evaluated our performance on a composite event — that the flight would arrive too late or cancel. This was intentional, as it highlights Dory’s composability and captures the dual-mode nature of disruption. However, we think scoring components individually is important for understanding composite performance; additionally, we expect there are use cases for its standalone components.
Because our departure and arrival delay predictions are distributions, evaluating them prompts the introduction of one more metric: continuous ranked probability score (CRPS). It rewards a forecast that is both well centered and as narrow as the outcomes allow. Nicely, it’s measured in minutes and lower is better; like Brier skill, it’s most applicable as a comparison between forecasters on the same flights.
Dory’s arrival delay CRPS is 18.4 minutes over our pooled 24-hour horizon, 5.2% lower than CatBoost’s9 (95% CI 4.1 to 6.3). The outperformance holds at every lead hour — where Dory is lower by 3.9–6.0% — and on at least 38 of the 42 held-out days. Scoring delays also offers another comparator: the airlines’ own estimates. Given that airlines don’t post a probabilistic arrival time, we built one from the spread of airlines’ errors10. Against it, Dory’s pooled arrival delay CRPS is 11.1% lower (95% CI 9.5 to 12.7) and lower on all 42 days. The lead-time lens applies here too, where Dory’s lead-time gain at its forecast horizon is 17.5 hours over the airlines (95% CI 13.7 to 19.5) and 12.8 hours over CatBoost (3.8 to 16.9).
And while the composite scoring queried only arrival delays, Dory predicts departure delay distributions as well. Pooled departure delay CRPS is 15.1 minutes, lower than CatBoost’s by 5.0% (3.9 to 6.1); against the airlines’, 9.6% (7.9, 11.2). Our first forecasts’ lead-time gain over CatBoost edges up to 14.1 hours (4.2, 17.2) and is stable versus the airlines’.
Evaluating the cancellation component returns us to a binary outcome, and to the first tool we reached for: Brier skill. To save you a scroll, a skill score of 0 means no more skilled than the base rate and 1, the perfect forecast — though here, only cancellations contribute to the base rate and only cancellation probabilities are scored. Over the pooled horizon, Dory reduces cancellation forecast error by 23.1% versus the base rate (95% CI 18.1% to 27.0%); against CatBoost, 6.9% (0.9% to 11.7%).
Dory scores higher than CatBoost at every lead hour, though the 95% confidence interval on its error reduction includes zero at lead times beyond 15–16 hours. Excluding flights that either Dory or CatBoost assign a cancellation probability of at least 90% disproportionately impacts cancellation skill — it removes 0.3% of forecasts but about one in six cancellations. Dory’s pooled skill score moves from 23.1% to 9.0% when near-certain cancellations are excluded, but its skill relative to CatBoost is again virtually unchanged.
Dory’s performance outside-out-of-time testing
Thus far, our evaluation has focused on the six held-out weeks at the then-end of our corpus. This mirrors the deployed environment, where today and tomorrow are after everything the model saw during training. While epistemically sound, a test window spanning just late spring into early summer does somewhat curb claims about how Dory generalizes across seasons. That said, two other sets of graphs suggest it remains skillful outside May and June.
Dory reduced composite forecast error at every too-late threshold and every lead hour in our validation graphs, where each meteorological season is represented; here, the reference is a still-retrospective base rate, but computed for each validation block of one to three weeks. While Dory did not train on these graphs, they were the basis for architectural decision-making and for when to stop training. This selection effect inflates skill scores by an unmeasured amount, though we think it’s implausible as the source of positive skill in every season, too-late threshold and lead time. We regard Dory’s validation skill as evidence that it doesn’t break outside the test window, not as an indicator of its degree of general skill.
Dory has also been deployed11 for the last month: for the 28 days ending October 4, its composite skill score over the pooled horizon ranges from 14.4% (at τ = 120 minutes, 95% CI 11.3–17.7%) to 17.6% (τ = 15 minutes, 95% CI 15.3–19.8%).
We do observe some degradation in Dory’s deployed skill relative to its held-out skill, most noticeable in cancellation probabilities at longer leads. While some of that gap is within uncertainty, we have identified two material contributors:
An input shift occurring in mid-August, measured at roughly two percentage points of composite skill.
Less disruption, which is where Dory earns its skill. The base rate fell by 6–14% from late spring to fall.
The former we expect to address in Dory’s next version, which we aim to release within a few weeks; that shift, along with Dory’s blindness to runway closures, are the primary reasons we’re calling this a research preview. Skill on calm days plausibly improves as we continue to accumulate more diverse training data.
Even so, we’re finding deployed Dory valuable.
And because Dory’s predictions are internally consistent, they compose not only across a flight’s possible outcomes, but also across multiple flights. This allows us to, for example, predict misconnection probabilities or the chance that either flight in a one-stop trip will cancel. You can try it now at app.aerology.ai.
© 2024-6 Google LLC, whose machine learning models were used to create the experimental data made available under the following license terms https://storage.googleapis.com/weathernext-public/terms-of-use.pdf. This data is intended for experimental modeling only and is not intended, validated, or approved for real world use.
GraphCast, which has 36.7 million parameters. https://arxiv.org/pdf/2212.12794
Operating carriers: Alaska Airlines, American Airlines, Delta Air Lines, Frontier Airlines, JetBlue, Southwest Airlines and United Airlines, plus the regional carriers CommuteAir, Endeavor Air, Envoy Air, GoJet Airlines, Horizon Air, Mesa Airlines, Piedmont Airlines, PSA Airlines, Republic Airways and SkyWest Airlines. Flights in ten non-scheduled flight-number ranges, and one range with a known labeling defect, are excluded.
A forecast’s initialization time is set to the time of the graph’s latest flight data, though there is a few-minute lag from graph snapshot to availability. Notably, this differs from weather prediction, where the lag between initialization and availability is measured in hours.
We apply a post-processing step to the upper tail of Dory’s delay distributions, fitted on data that includes our test window. To keep the held-out figures out-of-sample, we score each half of the window with that step refitted on the other half. The two versions differ by at most about 0.0001 in Brier skill and under 0.01 percentage points in relative CRPS, and cancellation figures, which don’t use this step, are identical. Brier skill intervals are computed on the served version; CRPS and lead-time-gain intervals use a cross-fitted version.
To compare Dory with the 60-day baseline or CatBoost, we compute Dory's Brier skill with that forecast as the reference. Where BSSDory and BSScomp are each forecast’s skill against the base rate:
95% intervals that resample days and exclude zero. The three targets whose intervals include zero are τ = 180 at 23 h, τ = 180 at 22 h and τ = 120 at 23 h, whose CIs are [−0.027, 0.081], [−0.004, 0.095], [−0.003, 0.079], respectively.
A different six weeks, with different weather, would have produced somewhat different rates for the same forecasts. The interval is the wider of one that resamples whole days (flights on the same day share weather and disruption, so they aren't independent) and a standard interval on the bin's predictions alone.
Even a perfectly calibrated forecaster would see about one line in twenty miss the diagonal by chance.
CatBoost predicts each delay distribution as 15 quantiles (from the 1st to the 99th percentile), trained with a pinball loss that, averaged across those levels, approximates CRPS. It is therefore optimized almost directly for the score we report, and we score both models at those same 15 quantiles. Dory, on the other hand, is trained on a different objective and is comparatively constrained in how it shapes its delay distributions. We consider CatBoost advantaged in this setup.
At each lead hour, we add the airline’s posted delay to the distribution of airline estimates’ errors, pooled across carriers, among flights with a similarly sized posted delay, measured on the other half of the test window’s days. We score it at the same 15 quantiles as Dory and CatBoost. This “dressed” forecast improves the airline's arrival CRPS by 22.7% over scoring its posted estimate as a single point, and the result is well calibrated (89.9% of arrivals fell below its 90th percentile).
The model we evaluated was trained without the six test weeks; the deployed model is the same recipe retrained with them included. Live figures score the forecasts as served, including the post-processing step. All of its corrections were fitted before the live window began.

