Joshua C. Macdonald
Computational Scientist · Johns Hopkins Bloomberg School of Public Health
International Vaccine Access Center
Department of International Health
Johns Hopkins Bloomberg School of Public Health
Baltimore, MD
My work determines what actions to take, what experiments to run, and what measurements are worth collecting in systems where interventions are costly and uncertainty is unavoidable. I build the computational machinery — generative models, scientific machine learning methods, inference algorithms, forecasting pipelines, and design and evaluation tools — that makes this possible for partially observed systems across health, environmental, and earth sciences; concretely, the infectious disease tools contribute multi-model evidence used for public-health planning, while other tools support wildlife disease surveillance design in sub-Saharan Africa and public health resource allocation. Prediction alone is not enough; the structural assumptions buried in every model must be exposed, tested, and defended before anyone acts on the output.
Currently I am an Assistant Scientist at the International Vaccine Access Center at Johns Hopkins (postdoctoral scholar there 2024–2026), where I develop operational infectious disease forecasting and decision support systems for ACCIDDA, a CDC Center for Forecasting and Outbreak Analytics collaboration. This work contributes scenario projections to the CDC Flu Scenario Modeling Hub, which supports public health resource allocation and intervention timing. I also collaborate with the Jolles Lab at Oregon State University on cross-scale modeling of Crimean–Congo hemorrhagic fever (CCHF) in wildlife and livestock systems, designing surveillance strategies that balance information gain against field cost. Previously I was a Zuckerman STEM Leadership Fellow at Tel Aviv University, working with Yoav Ram on Bayesian dimensionality reduction for incomplete cultural and genetic datasets.
I hold a PhD in Mathematics from the University of Louisiana at Lafayette and a BA in Archaeology from UNC Greensboro — a combination that continues to shape how I think about inference from incomplete records.
The common structure
The systems I work on share a recurring architecture: a latent process we care about, an observation process that maps and distorts it, and the data we actually get. Whether the latent state is disease prevalence in a wildlife herd, nutrient cycling in a marine ecosystem, or cultural trait frequencies in a human population, the decision-relevant state is rarely fully observed. The observation assumptions and validation requirements remain domain-specific. The figure below maps this architecture across four domains.
This is not just a conceptual analogy. The mathematical structure is shared: each domain requires a generative model that encodes how hidden states produce observables, an inference engine that inverts the observation process under uncertainty, and a decision layer that translates posterior beliefs into actionable recommendations. Building this infrastructure so that it transfers across domains — rather than rebuilding it from scratch for each application — is the central goal of my research program.
The operational tools I build reflect this shared structure and compose into a stack — from model specification to numerical integration to orchestration to evaluation:
- Model specification. OP System parses restricted model expressions, lowers them through a typed intermediate representation, and compiles vectorized right-hand sides that run with NumPy, JAX, or PyTorch arrays.
- Numerical integration. OP Engine combines portable Array-API numerical kernels with backend-native execution providers for explicit, IMEX, fully implicit, and stochastic integration.
- Orchestration. FlepiMoP2 drives configuration-defined campaigns over locations and scenarios, connecting systems, engines, parameters, processes, jobs, and persistence providers to operational modeling workflows.
- Design and evaluation. trade-study uses simulators with known ground truth to score competing configurations — model formulations, solver choices, measurement strategies, or any design decision — against ground truth using proper scoring rules and multi-objective Pareto optimization, so that decisions validated on benchmarks transfer to real data.
- Dimensionality reduction. vbpca-py and pp-eigentest recover latent population structure from incomplete datasets with full posterior uncertainty and principled rank determination.
The stack currently runs forward — specify a model, integrate it, orchestrate campaigns over it, evaluate the results — but the loop back to real, observed data is not yet closed by shared infrastructure: fitting a specified model’s parameters to observed surveillance data today happens as one-off, project-specific code rather than a reusable layer of the stack. Closing that loop is the near-term priority.
Where this is going
The research program moves along three axes. First, operationalize decisions: determine what intervention to deploy, what experiment to run, and what to measure next — then package these recommendations into tested, documented, open-source software with CI pipelines so that collaborators and decision-makers can act on model output they have reason to trust. Second, harden the inference: build reusable tools for Bayesian rank selection (pp-eigentest), multi-objective design and evaluation (trade-study), missing-data-native dimensionality reduction (vbpca-py), and — the next piece — parameter calibration against real observational data, so the same operator-partitioned models used for forward simulation can be fit to it directly. Third, generalize: extend the partially observed decision framework to new domains and new classes of systems — multi-host zoonoses, marine ecosystems, cultural evolution, spatial processes.
Current analyses include dengue antibody-dependent enhancement, outcome-independent structural scores for scientific predictions, COVID-19 observation-noise sensitivity, diphtheria outbreak vaccination, and CCHF surveillance and control design. Their working repositories remain private while the analyses and manuscripts mature.
The question that ties it all together: what should we do, given what we can’t observe?
Software
| Tool | What it does | Status |
|---|---|---|
| op_system | Restricted model parser, typed IR, and Array-API-polymorphic compiler | Active development · docs |
| op_engine | Array-API numerical kernels with backend-native explicit, IMEX, implicit, and stochastic execution | Developed for CDC-funded scenario modeling · docs |
| flepimop2 | Configuration-driven orchestration for forecasting & scenario analysis | Active development · CDC scenario modeling |
| trade-study | Design & evaluation: score competing configurations against known ground truth via simulators, scoring rules, Pareto optimization, stacking | Released on PyPI · docs |
| vbpca-py | Variational Bayesian PCA for incomplete data with full posterior uncertainty | Released on PyPI · JOSS submission in preparation |
| pp-eigentest | Posterior predictive eigenvalue testing for signal rank determination | Private pre-release · methods paper in preparation, extending arXiv:2409.12129 |
Methods & Expertise
These capabilities are deployed to design interventions, optimize surveillance strategies, and determine what experiments and measurements are worth their cost.
Modeling. Generative models (ODE/PDE/stochastic/hybrid); compartmental and agent-based models; scientific machine learning (physics-embedded surrogates, lawful learning); multivariate data analysis and dimensionality reduction; stability, bifurcation, and sensitivity analysis; parameter identifiability and inverse problems.
Inference. Bayesian hierarchical models; variational and simulation-based inference; data assimilation; uncertainty quantification; model calibration and validation; proper scoring rules and forecast evaluation.
Scientific computing. Python, Julia, C++, Stan, R, MATLAB. High-performance numerical solvers; operator splitting; reproducible workflows (Git, CI/CD, containerization); data pipelines and exploratory analysis.
news
| Jul 01, 2026 | Mini-symposium talk at SMB 2026 in Graz, Austria — “Decision-Support Modeling for One Health Pathogens: Using Mechanistic Models for Surveillance and Forecast Design.” |
|---|---|
| Apr 09, 2026 | Presented FlepiMoP2 and the Operator-Partitioned Simulation Stack at the Insight Net Third Annual Meeting tools workshop, Friday Center, Chapel Hill, NC. |
| Mar 30, 2026 | Invited seminar at Woods Hole Oceanographic Institution — “Cross-Scale Feedback Motifs: Structure-Preserving Models and Computational Tools for Complex Systems.” |
| Mar 23, 2026 | Invited seminar at UMBC — “Decision-Support Modeling for One Health Pathogens: Using Mechanistic Models for Surveillance and Forecast Design.” |
| Jul 13, 2025 | Presented at SMB 2025 in Edmonton — “Recovering Ecological Geometry: A Trait- and Depth-Structured IPDE Model of Plankton Dynamics.” |