Applications of synthetic null controls to enhance omics data analysis, from gene expression to sequences
Skip to main content
eScholarship
Open Access Publications from the University of California

UCLA

UCLA Electronic Theses and Dissertations bannerUCLA

Applications of synthetic null controls to enhance omics data analysis, from gene expression to sequences

Abstract

A pervading question in statistics is whether observed differences are meaningful or due to noise. P values estimate the chance of seeing such an extreme difference if the null hypothesis is true. In simple settings, such as testing whether or not a coin is fair, simulating the null hypothesis for estimation of the P value is straightforward, but more complex applications may be harder to describe. We propose that using synthetic null data as a concrete realization of the null hypothesis can be broadly used to enhance analysis. As proof of concept, I describe three applications explored during my Ph.D.

My first two projects focus on applications in high throughput genomics data, such as single-cell RNA-sequencing (scRNA-seq). A common step in analysis is data visualization, which requires moving from a high-dimensional space, such as the PCA space, to a 2-dimensional (2D) plane. Embedding methods may impart some data distortion, yet there are few ways to assess the fidelity of the visualization. In my first project, we introduce scDEED (single-cell dubious embedding detector). The key idea is that each cell’s neighbors, defined pre- and post-embedding, should appear similar in the 2D-visualization. For contrast, scDEED uses permutation to remove cell-cell relationships, creating a synthetic null dataset. By using the null dataset as a comparison, scDEED can better assess the target (original) dataset, identifying which cells are reliably embedded near their high-dimensional neighbors and which are not. In this first application, we show how scDEED can be used to better understand visualizations output by the embedding method.

Visualization is commonly followed by clustering and post-clustering differential expression (DE) analysis to identify subpopulations, such as patient cohorts or cell types. However, this method requires double-dipping: the same data is used for both cluster separation and DE analysis. This can result in spurious discoveries and an inflated false discovery rate. Moreover, this double-dipping issue is present in many genomic pipelines, not only scRNA-seq. To resolve this problem, our lab developed ClusterDE. ClusterDE uses a synthetic null dataset to replicate the null hypothesis: only one true cluster exists. Using this dataset as a contrast, ClusterDE separates meaningful and spurious discoveries in the target dataset. In my second project, I show that ClusterDE can identify better cell type markers in scRNA-seq data compared to standard double-dipping techniques and extend the application to microbiome data.

In the third project, I examine aspects of translation regulation, particularly the presence of upstream open reading frames (uORFs). Translation is the process of making proteins from mRNA and is highly regulated by sequence elements such as uORFs, categorized here as a canonical start codon (AUG) followed by a canonical termination codon (UGA/UAA/UAG) prior to the main protein coding sequence. We questioned whether evolutionary pressure could account for aspects of uORF repressiveness, such as their length. For comparison with uORFs, we used several methods of null control to create random sequences that replicate structural requirements and known nucleotide biases of mRNA. Applied to a wide range of phylogenies, we present evidence that uORFs are shorter than expected and that evolutionary pressure may have contributed to reduce the repressive effect of uORFs.

In summary, we propose that the principle of a synthetic null control has broad applications and present these applications as proof of concept.