Skip to main content
eScholarship
Open Access Publications from the University of California

Department of Statistics, UCLA

Open Access Policy Deposits bannerUCLA

This series is automatically populated with publications deposited by UCLA Department of Statistics researchers in accordance with the University of California’s open access policies. For more information see Open Access Policy Deposits and the UC Publication Management System.

Cover page of Introduction to Special Edition: The Future of the Textbook

Introduction to Special Edition: The Future of the Textbook

(2013)

A brief overview of the papers and commentaries in this special edition.

Cover page of A Statistical Analysis of Santa Barbara Ambulance Response in 2006: Performance Under Load

A Statistical Analysis of Santa Barbara Ambulance Response in 2006: Performance Under Load

(2009)

Ambulance response times in Santa Barbara County for 2006 are analyzed using point process techniques, including kernel intensity estimates and K-functions. Clusters of calls result in significantly higher response times, and this effect is quantified. In particular, calls preceded by other calls within 20 km and within the previous hour are significantly more likely to result in violations. This effect appears to be especially pronounced within semi-rural neighborhoods.

[WestJEM. 2009;10:42-47.]

Cover page of Bayesian Inference for Spatially‐Temporally Misaligned Data Using Predictive Stacking

Bayesian Inference for Spatially‐Temporally Misaligned Data Using Predictive Stacking

(2026)

ABSTRACT Air pollution remains a major environmental risk factor that is often associated with adverse health outcomes. However, quantifying and evaluating its effects on human health is challenging due to the complex nature of exposure data. Recent technological advances have led to the collection of various indicators of air pollution at increasingly high spatial‐temporal resolutions (e.g., daily averages of pollutant levels at spatial locations referenced by latitude‐longitude). However, health outcomes are typically aggregated over several spatial‐temporal coordinates (e.g., annual prevalence for a county) to comply with survey regulations. This article develops a Bayesian hierarchical model to analyze such spatially‐temporally misaligned exposure and health outcome data. We develop Bayesian predictive stacking for spatially and temporally misaligned data to optimally combine inference from multiple predictive spatial‐temporal models. Stacking allows us to avoid iterative estimation algorithms such as Markov chain Monte Carlo that struggle due to convergence issues inflicted by the presence of weakly identified parameters. We apply our proposed method to study the effects of ozone on asthma in the state of California.

Cover page of Nonstationary Spatial Process Models with Spatially Varying Covariance Kernels

Nonstationary Spatial Process Models with Spatially Varying Covariance Kernels

(2026)

Building spatial process models that capture nonstationary behavior while delivering computationally efficient inference is challenging. Nonstationary spatially varying kernels (see, e.g., Paciorek, 2003) offer flexibility and richness, but computation is impeded by high-dimensional parameter spaces resulting from spatially varying process parameters. Matters are exacerbated if the number of locations recording measurements is massive. With limited theoretical tractability, obviating computational bottlenecks requires synergy between model construction and algorithm development. We build a class of scalable nonstationary spatial process models using spatially varying covariance kernels. We implement a Bayesian modeling framework using Hybrid Monte Carlo with nested interweaving. We conduct experiments on synthetic data sets to explore model selection and parameter identifiability, and assess inferential improvements accrued from nonstationary modeling. We illustrate strengths and pitfalls with a data set on remote sensed normalized difference vegetation index.

Cover page of Fairness-Aware Kidney Exchange and Kidney Paired Donation

Fairness-Aware Kidney Exchange and Kidney Paired Donation

(2026)

The kidney paired donation (KPD) program provides an innovative solution to overcome incompatibility challenges in kidney transplants by matching incompatible donor-patient pairs and facilitating kidney exchanges. To address unequal access to transplant opportunities, there are two widely used fairness criteria: group fairness and individual fairness. However, these criteria do not consider protected patient features, which refer to characteristics legally or ethically recognized as needing protection from discrimination, such as race and gender. Motivated by the calibration principle in machine learning, we introduce a new fairness criterion: the matching outcome should be conditionally independent of the protected feature, given the sensitization level. We integrate this fairness criterion as a constraint within the KPD optimization framework and propose a computationally efficient solution using linearization strategies and column-generation methods. Theoretically, we analyze the associated price of fairness using random graph models. Empirically, we compare our fairness criterion with group fairness and individual fairness through both simulations and a real-data example.

Cover page of Causal Machine Learning: A Deductive–Inductive Framework for Sociological Research

Causal Machine Learning: A Deductive–Inductive Framework for Sociological Research

(2026)

Causal explanation is central to sociological research, shaping both theoretical development and empirical inquiry. This paper argues that causal machine learning—which integrates deductive identification strategies with inductive estimation techniques—offers an analytical approach for modeling complex, nonlinear social processes within the potential outcomes framework. We argue that causal machine learning operates through an iterative feedback loop: Theoretical assumptions guide flexible estimation, which inductively uncovers complex heterogeneities and nonlinearities, and these discoveries subsequently refine and expand sociological knowledge. Drawing on a systematic review of recent sociological research (2014–2024), we highlight how causal machine learning is advancing work in three key areas: causal effect heterogeneity, causal mediation analysis, and time-varying causal inference. These developments expand the methodological tool kit available to sociologists and strengthen the discipline’s ability to test, refine, and extend theories of social explanation. We conclude by outlining emerging directions, including high-dimensional causal inference and generative artificial intelligence, that are opening new methodological frontiers in causal machine learning for sociology.

Cover page of SnakeAltPromoter Facilitates Differential Alternative Promoter Analysis.

SnakeAltPromoter Facilitates Differential Alternative Promoter Analysis.

(2026)

Background: Alternative promoter usage contributes to isoform diversity and gene regulation in mammals but remains difficult to study at scale. Cap Analysis of Gene Expression precisely maps transcription start sites, but its cost limits large-scale application. Alternatively, ProActiv, Salmon, and DEXSeq can be utilized with widely available RNA sequencing (RNA-seq) data to infer promoter activity. However, there is currently no framework available to automate the generation of reproducible results for these methods. Results: SnakeAltPromoter, a scalable end-to-end Snakemake workflow, has been developed to automate alternative promoter analysis from raw RNA-seq data. The workflow performs quality control, alignment, and promoter quantification using 3 complementary RNA-seq analysis methods (junction-based, transcript-based, and first-exon-based), followed by promoter classification and differential activity or usage analysis. SnakeAltPromoter supports both command-line and graphical user interface usage and utilizes standardized modules to enhance reproducibility. A built-in benchmarking module compares promoter activities inferred from RNA-seq data to matched Cap Analysis of Gene Expression data to evaluate quantification performance. Our analyses revealed robust and complementary performance profiles among the evaluated methods across tissues and cell types, highlighting the value of a unified framework for promoter analysis. Conclusions: To our knowledge, SnakeAltPromoter is the first unified and reproducible framework that combines scalable execution and guided method selection for RNA-seq-based promoter analysis. By standardizing and integrating existing RNA-seq-based promoter analysis tools, it provides researchers with a robust and accessible platform to investigate promoter-level regulation and alternative promoter activity and will enhance the research value of large, public RNA-seq repositories. Code is freely available at https://github.com/YidanSunResearchLab/SnakeAltPromoter.git.

Cover page of CellScope: high-performance cell atlas workflow with tree-structured representation

CellScope: high-performance cell atlas workflow with tree-structured representation

(2025)

Single-cell sequencing enables comprehensive profiling of individual cells, revealing cellular heterogeneity and function with unprecedented resolution. However, current analysis frameworks lack the ability to simultaneously explore and visualize cellular hierarchies at multiple biological levels. To address these limitations, we present CellScope, a promising framework for constructing high-resolution cell atlases at multiple clustering levels. CellScope employs a two-stage manifold fitting process for gene selection and noise reduction, followed by agglomerative clustering, and integrates UMAP visualization with hierarchical clustering to intuitively represent cellular relationships simultaneously at multiple levels—such as cell lineage, cell type, and cell subtype levels. Compared to established pipelines such as Seurat and Scanpy, CellScope comprehensively improves clustering performance, visualization clarity, computational efficiency, and algorithm interpretability, while reducing dependence on hyperparameters across a multitude of single-cell datasets. Most importantly, it can reveal biological insights that other contemporary methods are unable to detect, thereby deepening our understanding of cellular heterogeneity and function, and potentially informing disease research.

Spatial transcriptomics iterative hierarchical clustering (stIHC): A novel method for identifying spatial gene co‐expression modules

(2025)

Abstract Recent advancements in spatial transcriptomics (ST) technologies allow researchers to simultaneously measure RNA expression levels for hundreds to thousands of genes while preserving spatial information within tissues, providing critical insights into spatial gene expression patterns, tissue organization, and gene functionality. However, existing methods for clustering spatially variable genes (SVGs) into co‐expression modules often fail to detect rare or unique spatial expression patterns. To address this, we present spatial transcriptomics iterative hierarchical clustering (stIHC), a novel method for clustering SVGs into co‐expression modules, representing groups of genes with shared spatial expression patterns. Through three simulations and applications to ST datasets from technologies such as 10x Visium, 10x Xenium, and Spatial Transcriptomics, stIHC outperforms clustering approaches used by popular SVG detection methods, including SPARK, SPARK‐X, MERINGUE, and SpatialDE. Gene ontology enrichment analysis confirms that genes within each module share consistent biological functions, supporting the functional relevance of spatial co‐expression. Robust across technologies with varying gene numbers and spatial resolution, stIHC provides a powerful tool for decoding the spatial organization of gene expression and the functional structure of complex tissues.