Jun Ho Yoon
Jun Ho Yoon

Jun Ho Yoon

Machine Learning Engineer, Search Ranking at Uber · San Francisco

I’m a machine learning engineer at Uber, working on search ranking. I received my Ph.D. in Computational Biology from Carnegie Mellon University’s School of Computer Science, where I developed scalable methods for Gaussian processes and sparse graphical models.

Interests Probabilistic & statistical modeling Graphs Applications

Research

Graphical models of a gene network perturbed by cis-acting and trans-acting eQTLs, and the sum-difference models derived from them
Figure 1. Overview of CiTruss: a gene network (blue) perturbed by cis-acting (red) and trans-acting (orange) eQTLs (a–c), and the sum-difference model derived from each (d–f).
CiTrussbioRxiv 2023

Learning Gene Networks Under SNP Perturbation Using SNP and Allele-Specific Expression Data

A statistical framework for simultaneously learning a gene network and the cis-acting and trans-acting expression quantitative trait loci (eQTLs) that perturb this network, given population allele-specific expression and SNP data.

Many existing methods identified eQTLs with single-gene, single-SNP association analysis and placed them against known gene networks for functional interpretation. CiTruss instead reconstructs a gene network perturbed by eQTLs, using a multi-level conditional Gaussian graphical model: trans-acting eQTLs perturb the expression of both alleles in the gene network at the top level, and cis-acting eQTLs perturb the expression of each allele at the bottom level. A transformation of this model allows efficient learning for large-scale human data.

  • New insights into the genetics of gene regulation from GTEx and LG×SM advanced intercross line mouse data for multiple tissue types
  • Gene networks made of local subnetworks over proximally located genes and global subnetworks over genes scattered across the genome
Abstract

Allele-specific expression quantification from RNA-seq reads provides opportunities to study the control of gene regulatory networks by cis-acting and trans-acting genetic variants. Many existing methods performed a single-gene and single-SNP association analysis to identify expression quantitative trait loci (eQTLs), and placed the eQTLs against known gene networks for functional interpretation. Instead, we view eQTL data as a capture of the effects of perturbation of gene regulatory system by a large number of genetic variants and reconstruct a gene network perturbed by eQTLs. We introduce a statistical framework called CiTruss for simultaneously learning a gene network and cis-acting and trans-acting eQTLs that perturb this network, given population allele-specific expression and SNP data. CiTruss uses a multi-level conditional Gaussian graphical model to model trans-acting eQTLs perturbing the expression of both alleles in gene network at the top level and cis-acting eQTLs perturbing the expression of each allele at the bottom level. We derive a transformation of this model that allows efficient learning for large-scale human data. Our analysis of the GTEx and LG×SM advanced intercross line mouse data for multiple tissue types with CiTruss provides new insights into genetics of gene regulation. CiTruss revealed that gene networks consist of local subnetworks over proximally located genes and global subnetworks over genes scattered across genome, and that several aspects of gene regulation by eQTLs such as the impact of genetic diversity, pleiotropy, tissue-specific gene regulation, and local and long-range linkage disequilibrium among eQTLs can be explained through these local and global subnetworks.

Surface plots of a multi-task function decomposed into shared and specific components
Figure 2. Doubly mixed-effects GP (left) and translated mixed-effects GP (right).
DMGPAISTATS 2022Oral · 2.6% of submissions

Doubly Mixed-Effects Gaussian Process Regression

Multi-task Gaussian process regression that separates how inputs affect outputs into parts shared across tasks and samples and parts specific to each.

Instead of the tensor product used in most multi-task GPs, the kernel combines task and sample covariance functions with the direct sum and the Kronecker sum. The resulting family of mixed-effects GPs, including doubly and translated mixed-effects GPs, models complex task relationships while decomposing input effects into four components. A stochastic variational inference method makes these models efficient to fit, and also significantly reduces the cost of inference for existing mixed-effects GPs.

  • Fixed effects shared across tasks and across samples, plus random effects specific to each task and each sample
  • Higher test accuracy and interpretable decompositions on simulated and real-world data
Abstract

We address the problem of multi-task regression with Gaussian processes (GPs) with the goal of obtaining a decomposition of input effects on outputs into components shared across or specific to tasks and samples. We propose a family of mixed-effects GPs, including doubly and translated mixed-effects GPs, that performs such a decomposition, while also modeling the complex task relationships. Instead of the tensor product widely used in multi-task GPs, we use the direct sum and Kronecker sum for Cartesian product to combine task and sample covariance functions. With this kernel, the overall input effects on outputs decompose into four components: fixed effects shared across tasks and across samples and random effects specific to each task and to each sample. We describe an efficient stochastic variational inference method for our proposed models that also significantly reduces the cost of inference for the existing mixed-effects GPs. On simulated and real-world data, we demonstrate higher test accuracy and interpretable decomposition from our approach.

Two small graphs combined by the Kronecker sum into a grid-structured graph, compared with the denser Kronecker product
Figure 3. Kronecker sum (left) vs. Kronecker product (right).
Block matrix diagram showing how the sparse Kronecker-sum structure collapses into two small matrices for gradient computation
Figure 4. Collapsing the sparse Kronecker-sum structure to compute gradients.
EiGLassoUAI 2020JMLR 2022100–1000× faster

Scalable Sparse Kronecker-Sum Inverse Covariance Estimation

A highly scalable estimator for sparse Kronecker-sum inverse covariance matrices, the Cartesian product of a sample graph and a feature graph, for data with dependencies among both samples and features.

Existing methods for this problem did not scale beyond a few hundred features and samples, and unidentifiable parameters made estimation difficult. EiGLasso combines Newton’s method with an eigendecomposition of the sample and feature graphs to exploit the Kronecker-sum structure, and approximates the Hessian from the same eigendecomposition to cut computation further. It also introduces a simple approach to estimating the unidentifiable parameters that generalizes existing methods.

  • Two to three orders of magnitude faster than existing methods on simulated and real-world data
  • Quadratic convergence with the exact Hessian, linear convergence with the approximate Hessian
Abstract

In many real-world data, complex dependencies are present both among samples and among features. The Kronecker sum or the Cartesian product of two graphs, each modeling dependencies across features and across samples, has been used as an inverse covariance matrix for a matrix-variate Gaussian distribution, as an alternative to Kronecker-product inverse covariance matrix due to its more intuitive sparse structure. However, the existing methods for sparse Kronecker-sum inverse covariance estimation are limited in that they do not scale to more than a few hundred features and samples and that the unidentifiable parameters pose challenges in estimation. In this paper, we introduce EiGLasso, a highly scalable method for sparse Kronecker-sum inverse covariance estimation, based on Newton’s method combined with eigendecomposition of the sample and feature graphs to exploit the Kronecker-sum structure. EiGLasso further reduces computation time by approximating the Hessian matrix based on the eigendecomposition of the two graphs. EiGLasso achieves quadratic convergence with the exact Hessian and linear convergence with the approximate Hessian. We describe a simple new approach to estimating the unidentifiable parameters that generalizes the existing methods. On simulated and real-world data, we demonstrate that EiGLasso achieves two to three orders-of-magnitude speed-up, compared to the existing methods.

Publications

  1. 2023

    Learning Gene Networks Under SNP Perturbation Using SNP and Allele-Specific Expression Data

    Jun Ho Yoon, Seyoung Kim

    bioRxiv preprint, 2023. doi:10.1101/2023.10.23.563661

  2. 2022

    Doubly Mixed-Effects Gaussian Process Regression

    Jun Ho Yoon, Daniel P. Jeong, Seyoung Kim

    Proceedings of the 25th International Conference on Artificial Intelligence and Statistics (AISTATS), PMLR 151:6893–6908, 2022.

    Oral · 2.6% of submissions
  3. 2022

    EiGLasso for Scalable Sparse Kronecker-Sum Inverse Covariance Estimation

    Jun Ho Yoon, Seyoung Kim

    Journal of Machine Learning Research, 23(110):1–39, 2022.

  4. 2020

    EiGLasso: Scalable Estimation of Cartesian Product of Sparse Inverse Covariance Matrices

    Jun Ho Yoon, Seyoung Kim

    Proceedings of the 36th Conference on Uncertainty in Artificial Intelligence (UAI), PMLR 124:1248–1257, 2020.

Software

Education

  1. Ph.D. in Computational Biology

    School of Computer Science, Carnegie Mellon University

    2023
  2. B.S. in Computer Science

    Columbia University

    2017