Computational biologist | Statistics, meta-omics, biomarker discovery, and reproducible workflows
Hi, I'm Yixuan 👋 — a computational biologist and Ph.D. candidate in Bioinformatics at North Carolina State University, advised by Benjamin Callahan, with M.S. and B.S. training in Statistics.
My work combines statistical modeling, reproducible workflows, and scientific software to study complex biological data. My research spans metaproteomics, microbiome science, genome assembly, and molecular evolution, with particular interest in how measurement and study design shape biological inference.
- Quantitative methods: Statistical inference, study design, causal inference, compositional data analysis, differential expression and abundance, machine learning
- Biological data: Metagenomics, metaproteomics, PacBio HiFi sequencing, amplicon sequencing, RNA-seq
- Scientific computing: Python, R/Bioconductor, Nextflow, containers (Docker, Singularity), scalable computing (HPC/Slurm, AWS), SQL, Parquet, MuData
An ongoing project to harmonize public human gut metaproteomics data and support reproducible research on proteins of unknown function.
With Benjamin Callahan and Karen R. Muñana
Reproducible analyses for a household-matched case-control study of gut microbiome alterations in 98 dogs. The study evaluated community structure and six differential-abundance methods, identifying household as the dominant source of microbiome variation.
With Lina Quesada and Benjamin Callahan · Manuscript in preparation
A modular Nextflow DSL2 workflow for recovering target eukaryotic genomes from highly contaminated PacBio HiFi reads. It integrates assembly, multi-stage decontamination, quality control, reproducible HPC execution, and RNA-seq-supported protein validation.
An R/Bioconductor package for protein-language-model-informed representation distances, structure-aware sample comparison, and biomarker discovery. Bioconductor submission in progress.
With Jeff Thorne and Xiang Ji
Statistical modeling of interlocus gene conversion, natural selection, and paralog homogenization. This work extended the MG94 codon framework with an IGC component and evaluated competing evolutionary hypotheses using maximum-likelihood estimation and likelihood-ratio tests.
Implemented a DIRECT-based optimizer with an 18.5× median speedup (up to 82.7×) while maintaining interval Jaccard ≥ 0.8 across all nine benchmark cases. TrIdent is an R/Bioconductor package for detecting, classifying, and characterizing active transduction events from sequencing-coverage patterns.
- Yang, Y., Nettifee, J., Azcarate-Peril, M. A., Muñana, K. R., & Callahan, B. (2026). Gut microbiome alterations in canine idiopathic epilepsy: a pairwise case-control study. Animal Microbiome. doi:10.1186/s42523-026-00594-1
- Yang, Y., Xu, T., Conant, G. C., Kishino, H., Thorne, J. L., & Ji, X. (2023). Interlocus gene conversion, natural selection, and paralog homogenization. Molecular Biology and Evolution, 40, msad198. doi:10.1093/molbev/msad198
Google Scholar · ORCID · LinkedIn · Email
Always happy to chat about computational biology, statistics, multi-omics, and reproducible research.


