Joshua C. Macdonald

Computational Scientist · Johns Hopkins Bloomberg School of Public Health

prof_pic.jpg

International Vaccine Access Center

Department of International Health

Johns Hopkins Bloomberg School of Public Health

Baltimore, MD

My work determines what actions to take, what experiments to run, and what measurements are worth collecting in systems where interventions are costly and uncertainty is unavoidable. I build the computational machinery — generative models, scientific machine learning methods, inference algorithms, forecasting pipelines, and design and evaluation tools — that makes this possible for partially observed systems across health, environmental, and earth sciences; concretely, the infectious disease tools contribute multi-model evidence used for public-health planning, while other tools support wildlife disease surveillance design in sub-Saharan Africa and public health resource allocation. Prediction alone is not enough; the structural assumptions buried in every model must be exposed, tested, and defended before anyone acts on the output.

Currently I am an Assistant Scientist at the International Vaccine Access Center at Johns Hopkins (postdoctoral scholar there 2024–2026), where I develop operational infectious disease forecasting and decision support systems for ACCIDDA, a CDC Center for Forecasting and Outbreak Analytics collaboration. This work contributes scenario projections to the CDC Flu Scenario Modeling Hub, which supports public health resource allocation and intervention timing. I also collaborate with the Jolles Lab at Oregon State University on cross-scale modeling of Crimean–Congo hemorrhagic fever (CCHF) in wildlife and livestock systems, designing surveillance strategies that balance information gain against field cost. Previously I was a Zuckerman STEM Leadership Fellow at Tel Aviv University, working with Yoav Ram on Bayesian dimensionality reduction for incomplete cultural and genetic datasets.

I hold a PhD in Mathematics from the University of Louisiana at Lafayette and a BA in Archaeology from UNC Greensboro — a combination that continues to shape how I think about inference from incomplete records.

Shared research stack connecting model specification, numerical execution, orchestration, design, evaluation, and inference to operational and emerging scientific domains.
Reusable computational layers support domain-specific models and decisions; labels distinguish operational, active-research, and planned work.

The common structure

The systems I work on share a recurring architecture: a latent process we care about, an observation process that maps and distorts it, and the data we actually get. Whether the latent state is disease prevalence in a wildlife herd, nutrient cycling in a marine ecosystem, or cultural trait frequencies in a human population, the decision-relevant state is rarely fully observed. The observation assumptions and validation requirements remain domain-specific. The figure below maps this architecture across four domains.

Four examples of latent scientific processes passing through domain-specific observation models to produce epidemiological, environmental, earth-system, and human data.
Across domains, a partially observed latent process is connected to data through an explicit observation model.

This is not just a conceptual analogy. The mathematical structure is shared: each domain requires a generative model that encodes how hidden states produce observables, an inference engine that inverts the observation process under uncertainty, and a decision layer that translates posterior beliefs into actionable recommendations. Building this infrastructure so that it transfers across domains — rather than rebuilding it from scratch for each application — is the central goal of my research program.

The operational tools I build reflect this shared structure and compose into a stack — from model specification to numerical integration to orchestration to evaluation:

  • Model specification. OP System parses restricted model expressions, lowers them through a typed intermediate representation, and compiles vectorized right-hand sides that run with NumPy, JAX, or PyTorch arrays.
  • Numerical integration. OP Engine combines portable Array-API numerical kernels with backend-native execution providers for explicit, IMEX, fully implicit, and stochastic integration.
  • Orchestration. FlepiMoP2 drives configuration-defined campaigns over locations and scenarios, connecting systems, engines, parameters, processes, jobs, and persistence providers to operational modeling workflows.
  • Design and evaluation. trade-study uses simulators with known ground truth to score competing configurations — model formulations, solver choices, measurement strategies, or any design decision — against ground truth using proper scoring rules and multi-objective Pareto optimization, so that decisions validated on benchmarks transfer to real data.
  • Dimensionality reduction. vbpca-py and pp-eigentest recover latent population structure from incomplete datasets with full posterior uncertainty and principled rank determination.

The stack currently runs forward — specify a model, integrate it, orchestrate campaigns over it, evaluate the results — but the loop back to real, observed data is not yet closed by shared infrastructure: fitting a specified model’s parameters to observed surveillance data today happens as one-off, project-specific code rather than a reusable layer of the stack. Closing that loop is the near-term priority.

Where this is going

The research program moves along three axes. First, operationalize decisions: determine what intervention to deploy, what experiment to run, and what to measure next — then package these recommendations into tested, documented, open-source software with CI pipelines so that collaborators and decision-makers can act on model output they have reason to trust. Second, harden the inference: build reusable tools for Bayesian rank selection (pp-eigentest), multi-objective design and evaluation (trade-study), missing-data-native dimensionality reduction (vbpca-py), and — the next piece — parameter calibration against real observational data, so the same operator-partitioned models used for forward simulation can be fit to it directly. Third, generalize: extend the partially observed decision framework to new domains and new classes of systems — multi-host zoonoses, marine ecosystems, cultural evolution, spatial processes.

Current analyses include dengue antibody-dependent enhancement, outcome-independent structural scores for scientific predictions, COVID-19 observation-noise sensitivity, diphtheria outbreak vaccination, and CCHF surveillance and control design. Their working repositories remain private while the analyses and manuscripts mature.

Research program organized around generalizing to new scientific systems, hardening inference and evaluation, and operationalizing policy and surveillance decisions.
The program links generalization, methodological hardening, and operational decision support through observations and policy feedback.

The question that ties it all together: what should we do, given what we can’t observe?


Software

Tool What it does Status
op_system Restricted model parser, typed IR, and Array-API-polymorphic compiler Active development · docs
op_engine Array-API numerical kernels with backend-native explicit, IMEX, implicit, and stochastic execution Developed for CDC-funded scenario modeling · docs
flepimop2 Configuration-driven orchestration for forecasting & scenario analysis Active development · CDC scenario modeling
trade-study Design & evaluation: score competing configurations against known ground truth via simulators, scoring rules, Pareto optimization, stacking Released on PyPI · docs
vbpca-py Variational Bayesian PCA for incomplete data with full posterior uncertainty Released on PyPI · JOSS submission in preparation
pp-eigentest Posterior predictive eigenvalue testing for signal rank determination Private pre-release · methods paper in preparation, extending arXiv:2409.12129

Methods & Expertise

These capabilities are deployed to design interventions, optimize surveillance strategies, and determine what experiments and measurements are worth their cost.

Modeling. Generative models (ODE/PDE/stochastic/hybrid); compartmental and agent-based models; scientific machine learning (physics-embedded surrogates, lawful learning); multivariate data analysis and dimensionality reduction; stability, bifurcation, and sensitivity analysis; parameter identifiability and inverse problems.

Inference. Bayesian hierarchical models; variational and simulation-based inference; data assimilation; uncertainty quantification; model calibration and validation; proper scoring rules and forecast evaluation.

Scientific computing. Python, Julia, C++, Stan, R, MATLAB. High-performance numerical solvers; operator splitting; reproducible workflows (Git, CI/CD, containerization); data pipelines and exploratory analysis.

news

Jul 01, 2026 Mini-symposium talk at SMB 2026 in Graz, Austria — “Decision-Support Modeling for One Health Pathogens: Using Mechanistic Models for Surveillance and Forecast Design.”
Apr 09, 2026 Presented FlepiMoP2 and the Operator-Partitioned Simulation Stack at the Insight Net Third Annual Meeting tools workshop, Friday Center, Chapel Hill, NC.
Mar 30, 2026 Invited seminar at Woods Hole Oceanographic Institution — “Cross-Scale Feedback Motifs: Structure-Preserving Models and Computational Tools for Complex Systems.”
Mar 23, 2026 Invited seminar at UMBC — “Decision-Support Modeling for One Health Pathogens: Using Mechanistic Models for Surveillance and Forecast Design.”
Jul 13, 2025 Presented at SMB 2025 in Edmonton — “Recovering Ecological Geometry: A Trait- and Depth-Structured IPDE Model of Plankton Dynamics.”