Research spanning machine learning, statistical genetics, and biostatistics. A complete publication record is available through Google Scholar and the CV.
Selected Research
Machine Learning
Machine learning methods that create, evaluate, or leverage learned phenotypes for valid scientific discovery.
Synthetic surrogates improve power for genome-wide association studies of partially missing phenotypes in population biobanks
Joint modeling of observed phenotypes and learned surrogate outcomes can improve association power while preserving valid inference.
DeepNull: Modeling non-linear covariate effects improves phenotype prediction and association power
Nonlinear covariate adjustment with machine learning improves phenotype prediction and can increase power in genetic association studies.
EmbedGEM: A framework to evaluate the utility of embeddings for genetic discovery
A framework for evaluating learned representations through both their genetic association strength and their clinical relevance.
Large-scale machine learning-based phenotyping significantly improves genomic discovery for optic nerve head morphology
Deep-learning-derived optic nerve phenotypes substantially expand genetic discovery for glaucoma-related morphology.
Inference of chronic obstructive pulmonary disease with deep learning on raw spirograms identifies new genetic loci and improves risk models
Deep learning on raw lung-function traces improves COPD phenotyping, genetic discovery, and risk prediction.
Unsupervised representation learning improves genomic discovery for lung function and respiratory disease prediction
Self-supervised representations of spirograms capture lung-function variation that supports discovery and respiratory-disease prediction.
Selected Research
Statistical Genetics
Statistical methods for robust, scalable discovery across common and rare genetic variation.
An allelic-series rare-variant association test for candidate-gene discovery
A gene-level rare-variant test designed to identify allelic series with increasingly large effects for increasingly deleterious variants.
A Scalable Framework for Identifying Allelic Series from Summary Statistics
A summary-statistics framework that extends allelic-series testing to large cohorts and meta-analytic settings.
Operating Characteristics of the Rank-Based Inverse Normal Transformation for Quantitative Trait Analysis in Genome-Wide Association Studies
A robust omnibus testing strategy for quantitative-trait association when regression residuals are non-normal.
Pitfalls in performing genome-wide association studies on ratio traits
An analysis of how ratio phenotypes can induce misleading associations and complicate interpretation in genome-wide studies.
Selected Research
Biostatistics
Statistical methods for interpretable inference in clinical studies, event-time analysis, and large-scale testing.
Nonparametric estimation of the total treatment effect with multiple outcomes in the presence of terminal events
A nonparametric estimand that integrates treatment effects across multiple outcomes while accounting for terminal events.
Practical Recommendations on Quantifying and Interpreting Treatment Effects in the Presence of Terminal Competing Risks: A Review
Practical guidance for choosing interpretable treatment-effect summaries when terminal competing risks are present.
Testing a Large Number of Composite Null Hypotheses Using Conditionally Symmetric Multidimensional Gaussian Mixtures in Genome-Wide Studies
A multidimensional Gaussian-mixture approach for testing many composite null hypotheses while controlling false discoveries.