Virology

Could AI Have Called the Worst Ebola Outbreak in Years?

A new preprint tests whether computational and AI forecasting models could have predicted the 2026 Bundibugyo Ebola epidemic in DR Congo before it became one of the largest on record. Here is what the models got right, what they missed, and why it matters for outbreak science.

For most of the past decade, Ebola forecasting meant one thing: reconstructing what had already happened. Teams counted cases after the fact, drew a curve, and argued about the reproduction number. The 2026 outbreak in the Democratic Republic of the Congo is different. This time, dozens of modelling groups have been publishing live forecasts while the epidemic is still burning, and a new preprint posted on 30 July asks the question everyone in the field has been avoiding: did any of it actually work?

That question sits right at the meeting point of two stories that usually get told separately. One is a filovirus tearing through eastern Congo. The other is the slow arrival of artificial intelligence and computational modelling into frontline outbreak response. The 2026 Bundibugyo epidemic has become the first major Ebola event where those two stories are the same story.

What happened#

The outbreak is caused by Bundibugyo virus, one of the less common members of the Ebola family. According to the WHO Disease Outbreak News, DR Congo confirmed the outbreak in Ituri Province in mid-May 2026, with an estimated index case around 1 April. The European Centre for Disease Prevention and Control reported 4,053 confirmed cases and 1,850 deaths in DR Congo as of 5 August, a case fatality rate near 46 per cent, with the outbreak having spread from a single health zone to dozens across five provinces. Uganda has recorded a small cluster and two deaths, and a single imported case reached France. These figures are moving weekly, so treat any single number as a snapshot rather than a final tally.

That makes this the largest recorded Bundibugyo outbreak by a wide margin. Previous Bundibugyo events in Uganda in 2007 and in the DR Congo in 2012 stayed in the low hundreds. The current one has already blown past them and, on present trends, is closing on the scale of the region's worst Ebola epidemics.

Against that backdrop, the 30 July preprint (which has not been peer-reviewed) did something unusual. Rather than publish yet another forecast, its authors ran a "rolling-origin" evaluation: they replayed the outbreak week by week using only the surveillance data available at each point, generated forecasts as if in real time, and then scored them against what actually happened. In plain terms, they graded the crystal balls after the fact, using only what each ball could have seen.

How outbreak forecasting actually works#

The forecasting looks at recent case counts and extends the curve, assuming the near future resembles the near past. It knows nothing about biology. It is often surprisingly hard to beat over short horizons.

In the middle sit mechanistic models. A branching-process or compartmental model encodes how the disease spreads: how many people each case infects, how long between infections, how isolation cuts transmission. The US Centers for Disease Control and Prevention used a branching-process model for its published projections. Under poor isolation of 20 per cent of cases, it put the chance of exceeding 20,000 cases within three months at 65 per cent. Push isolation to 70 per cent and only about one simulation in twenty crossed 10,000. The model does not predict a single future. It maps how the future bends depending on what responders do.

At the far end sit Bayesian and machine-learning approaches that fit these mechanisms to messy real-world data while carrying explicit uncertainty. The recalibrated stochastic model published in The Lancet Infectious Diseases estimated a central reproduction number of about 1.71 and put the probability of spillover into Uganda at 94 per cent and into South Sudan at 69 per cent. A network model using human displacement-tracking data forecast where importations into Uganda were most likely, and a separate preprint mapped importation risk into Europe under different expansion scenarios. None of these is a chatbot. They are statistical engines, increasingly wrapped in machine-learning methods, feeding on real surveillance streams.

What the July preprint found#

The headline result is modest and, for that reason, believable. Routine national situation reports, the ordinary bureaucratic paperwork of an outbreak, were good enough to support genuinely useful short-term forecasts. You did not need exotic data to see a week or two ahead.

The nuances matter more than the headline. A saturating growth model, which assumes the epidemic is about to level off, performed badly and the authors advise against it. A plain recent-trend baseline and a recalibrated Bayesian model each captured something the other missed. Crucially, a transparent combination of the two beat either one alone. That is a familiar lesson from weather forecasting: ensembles of complementary models tend to outperform any single clever model, because their errors partly cancel.

But there is a limit the authors are honest about. Short-term forecasts, out to a couple of weeks, were useful for planning bed and staffing levels. Longer projections carried so much uncertainty that presenting them as point predictions would mislead. The value was in the range, not the single line down the middle.

Why this matters#

For outbreak science, this is a quiet turning point. Ebola has historically been forecast in the past tense. Seeing an ecosystem of live models, then a rigorous audit of whether they worked, is closer to how meteorology matured from folklore into a real forecasting discipline. If the finding holds up through peer review, it means low-income health systems, sitting on nothing more than daily situation reports, can still generate forecasts good enough to guide where isolation beds go next week. That is a far more useful claim than any promise of a magic predictive AI.

For public health, the CDC projections carry the sharper message. The gap between 20 per cent and 70 per cent case isolation is the difference between a contained outbreak and a catastrophe of tens of thousands. The models do not save anyone. They quantify how much the response choices matter, which is exactly the kind of argument that moves budgets and staff.

For the broader AI-in-biology field, the outbreak is a reality check on hype. The genuinely novel AI tools are arriving elsewhere in the response. Genomic surveillance now leans on machine-assisted phylogenetics through platforms like Nextstrain, and sequencing showed the 2026 virus is a distinct lineage from a fresh zoonotic spillover rather than a resurgence of an old one. On the countermeasures side, generative protein models are being tested against filovirus and influenza targets, though not yet at outbreak speed.

Critical analysis#

The strengths of this work are its honesty and its method. Grading forecasts using only the data available at the time is the correct way to test them, and it is depressingly rare. The conclusion that a simple ensemble beats a single clever model is robust and reproducible across fields.

The limitations are real. The evaluation is a preprint, not yet peer-reviewed, and its findings should not be read as settled science. It rests on surveillance data that, in an active conflict zone, is incomplete and delayed. Case counts lag, deaths are missed, and the denominator is always uncertain. A model fed bad numbers forecasts confidently wrong. The related work on outbreak size exists precisely because nobody trusts the raw counts: an independent Bayesian re-analysis has tried to estimate the true epidemic size behind the reported figures.

The unresolved questions are harder. Nobody has shown that a live forecast changed an operational decision in this outbreak and improved the outcome. Producing an accurate number and having a strained ministry of health act on it are different things. There is also a competing view worth taking seriously: some experienced epidemiologists argue that in fast-moving filovirus outbreaks, the binding constraint is never the forecast but the response capacity, so effort spent refining models is effort not spent on isolation beds and contact tracing. That critique has force.

Expert perspective#

Set this against the 2014 to 2016 West African Ebola epidemic and the shift is clear. A decade ago, high-profile forecasts badly overshot the outbreak, projecting more than a million cases in one widely reported scenario, and the miss damaged the credibility of modelling for years. What is different now is not that the models are magically better. It is that the community is measuring its own accuracy in the open, publishing misses alongside hits, and combining models rather than betting on a single model.

It also separates two things that get lumped together as AI. Forecasting the epidemic curve is mostly statistics and mechanistic modelling with machine learning bolted on. Designing an antibody or reading a genome at scale is where deep learning does something people could not do by hand. Both are useful. Only the second is new, and neither is a substitute for a functioning health system.

Key takeaways#

  1. The 2026 Bundibugyo outbreak is the largest on record for this virus and, per ECDC, had caused roughly 1,850 deaths in DR Congo by early August, with a case fatality rate near 46 per cent.

  2. A new preprint shows that routine surveillance data supported useful short-term forecasts, and that a simple ensemble of two models outperformed either model alone.

  3. CDC projections show that the response, not the virus, drives the outcome: a case isolation rate of 20 versus 70 per cent is the difference between containment and disaster.

  4. Most outbreak "AI" here is statistical and mechanistic modelling. The genuinely novel AI in genomics and protein design is evident elsewhere in the response.

  5. Forecasts are only as good as the surveillance feeding them, and no one has yet shown a live forecast changed the course of this outbreak.

Frequently asked questions#

Is Bundibugyo virus the same as the Ebola everyone knows? It is in the same family. The virus most people mean by "Ebola" is Zaire ebolavirus. Bundibugyo is a separate species, historically less transmissible and less deadly, though this outbreak is testing that assumption.

Did AI predict this outbreak? Not in the sense of raising an alarm before it started. Once it was under way, computational and machine-learning models produced short-term forecasts that a recent evaluation found genuinely useful, with important caveats.

How reliable are the death and case figures? Treat them as evolving estimates. Counts come from an active conflict zone, lag reality, and are revised often. That is why separate studies try to estimate the true epidemic size behind the reported numbers.

Is there a vaccine or treatment? There is no licensed vaccine or therapy specific to Bundibugyo virus. Existing Ebola vaccines target the Zaire species and cross-protect weakly. Trials of Bundibugyo-specific candidates, including an mRNA vaccine and a viral-vector vaccine, began in 2026, with results still months away.

What is a rolling-origin evaluation? It is a way of testing forecasts fairly. You replay history step by step, let the model see only the data available at each point, make a forecast, and score it against what actually happened. It stops models from cheating with hindsight.

Should other countries be worried? Regional spillover risk is real, and models put the probability of spread to neighbouring countries at a moderate to high level. The risk to distant countries with strong health systems remains low, and the single case reported in France was contained.

References#

World Health Organization. Ebola disease caused by Bundibugyo virus, Democratic Republic of the Congo (Disease Outbreak News).

European Centre for Disease Prevention and Control. Ebola disease outbreak in the Democratic Republic of the Congo and Uganda.

US Centers for Disease Control and Prevention. Modeled Scenario Projections for the Ebola Disease Outbreak Caused by Bundibugyo Virus, 2026. MMWR.

How Predictable Was the 2026 Bundibugyo Virus Disease Outbreak? A Rolling-Origin Evaluation. medRxiv, 28 July 2026. Preprint (not peer reviewed).

Size of the 2026 Ebola outbreak and risk of cross-border spillover from Bundibugyo virus in Ituri Province: a recalibrated stochastic modelling study. The Lancet Infectious Diseases

Network-based modelling of Bundibugyo Ebola virus disease importation and spread in Uganda. medRxiv. Preprint (not peer reviewed).

Shifting patterns of importation risk of Bundibugyo Ebola virus disease to Europe. medRxiv. Preprint (not peer reviewed).

Epiforecasts. Estimating the current size of the 2026 DRC Bundibugyo virus outbreak. Preprint / ongoing analysis (not peer reviewed).

Emergence of a Bundibugyo virus variant in the 2026 outbreak in the Democratic Republic of the Congo and Uganda. Nature Medicine.

Nextstrain. Real-time tracking of pathogen evolution.

Gavi. The world's first Bundibugyo Ebola vaccine has entered human trials.

Related observations

Adjacent work from the same lines of enquiry.

Ebola's 100-Day Test: The Bundibugyo Vaccine Sprint

The Bundibugyo Ebola outbreak is now the largest in DRC history. With no licensed vaccine, three candidates have reached human trials in weeks — a real-world stress test of the 100 Days Mission, AI-assisted drug discovery and outbreak modelling.