Skip to main content
eScholarship
Open Access Publications from the University of California

This series is automatically populated with publications deposited by UC Riverside Bourns College of Engineering Computer Science and Engineering Department researchers in accordance with the University of California’s open access policies. For more information see Open Access Policy Deposits and the UC Publication Management System.

Cover page of <i>Babesia hegotelforum</i> sp. nov., a zoonotic <i>Babesia</i> species previously referred to as <i>Babesia sp</i>. <i>MO1</i>.

Babesia hegotelforum sp. nov., a zoonotic Babesia species previously referred to as Babesia sp. MO1.

(2026)

A zoonotic Babesia species previously referred to as Babesia sp. MO1 is formally described and named here as Babesia hegotelforum sp. nov. This taxon is distinct from Babesia divergens based on genome-wide sequence divergence, phylogenetic placement, host associations, and clinical presentation. The parasite infects erythrocytes of humans, and eastern cottontail rabbits (Sylvilagus floridanus), and is transmitted by Ixodes dentatus. The holotype consists of a Giemsa-stained thin blood smear and cryopreserved infected erythrocytes from the cloned isolate BML-Bh-B12 at ≤10 passages in continuous in vitro culture. Paratype material includes five additional clones (BML-Bh-H1, BML-Bh-F12, BML-Bh-H6, BML-Bh-A3, and BML-Bh-F1) derived from BEI Resources strain NR-50441, along with the original mixed isolate NR-50441. This species description meets the requirements of the International Code of Zoological Nomenclature and establishes Babesia hegotelforum sp. nov. as a distinct species of clinical and epidemiological significance in North America.

Cover page of ENVnet provides a global molecular resource of dissolved organic matter

ENVnet provides a global molecular resource of dissolved organic matter

(2026)

Dissolved organic matter (DOM) is a central component of Earth’s carbon cycle and one of the planet’s most chemically diverse pools, yet the molecular structures of its constituents remain largely unresolved. This limitation has hindered our ability to link DOM composition to microbial processes and ecosystem function. Here we present ENVnet, a global molecular repository built from tandem mass spectrometry data collected across 13 terrestrial and aquatic environment types, including 419 newly generated samples that expand publicly available DOM metabolomics data and cover previously underrepresented environments. By computationally deconvolving chimeric mass spectra, a longstanding challenge in environmental metabolomics, we recover high-quality fragmentation data for >22,000 distinct molecular features (defined by a specific precursor mass and fragmentation pattern). Using ENVnet, we uncover conserved and environment-specific molecular patterns in DOM composition and underlying biogeochemical processes. We also use molecular features encoded in ENVnet to train predictive models of DOM persistence, allowing molecular-level assessment of microbial turnover in independent systems.

Cover page of Comparative structural analysis of protein complexes with SPICE.

Comparative structural analysis of protein complexes with SPICE.

(2026)

Computational tools for studying the structure of protein complexes are essential for providing mechanistic insights into protein-protein interactions and therapeutic drug design. Here, we present SPICE (Structural Protein Interaction Complex Evaluator), a web-based platform that allows structural biologists to perform rapid, modular analyses of protein complexes directly from Protein Data Bank (PDB) structures. SPICE allows users to define and execute analysis workflows via an intuitive web interface, reducing analysis times from minutes to seconds. The platform offers a broad range of analytical capabilities, including (i) detection of hydrogen bonds, salt bridges, and disulfide bonds; (ii) protein-protein interface mapping; and (iii) computation of solvent accessibility, van der Waals energetics, and other key geometric descriptors. SPICE further provides interactive 3D visualization and supports comparative analyses across multiple complexes, enabling the study of mutational effects and binding variants. The tool is freely available at https://spice.cs.ucr.edu (no registration required).

Cover page of SHICEDO: single-cell Hi-C data enhancement with reduced over-smoothing

SHICEDO: single-cell Hi-C data enhancement with reduced over-smoothing

(2025)

MOTIVATION: Single-cell Hi-C (scHi-C) technologies have significantly advanced our understanding of the 3D genome organization. However, scHi-C data are often sparse and noisy, leading to substantial computational challenges in downstream analyses. RESULTS: In this study, we introduce SHICEDO, a novel deep-learning model specifically designed to enhance scHi-C contact matrices by imputing missing or sparsely captured chromatin contacts through a generative adversarial framework. SHICEDO leverages the unique structural characteristics of scHi-C matrices to derive customized features that enable effective data enhancement. Additionally, the model incorporates a channel-wise attention mechanism to mitigate the over-smoothing issue commonly associated with scHi-C enhancement methods. Through simulations and real-data applications, we demonstrate that SHICEDO outperforms the state-of-the-art methods, achieving superior quantitative and qualitative results. Moreover, SHICEDO enhances key structural features in scHi-C data, thus enabling more precise delineation of chromatin structures such as A/B compartments, TAD-like domains, and chromatin loops. AVAILABILITY AND IMPLEMENTATION: SHICEDO is publicly available at https://github.com/wmalab/SHICEDO.

Cover page of Prediction of DNA Methylation With Long-Range State-Space Models

Prediction of DNA Methylation With Long-Range State-Space Models

(2025)

The prediction of DNA methylation from the primary DNA sequence allows one to impute the methylation status of cytosines with insufficient sequencing coverage. Various deep learning models have been proposed in the literature, including transformer-based models and convolutional neural networks. In this study, we investigate the performance of long-range state-space models based on the Hyena architecture on the task of DNA methylation prediction on six plant species. First, we train the HyenaDNA framework to obtain a genome-wide foundation model for each species. Then, we fine-tune these foundation models using the sequence data surrounding the methylated or unmethylated cytosines. Extensive experimental results show that our model predicts DNA methylation with higher accuracy than state-of-the-art methods in the literature.

Cover page of Fine-tuned protein language model identifies antigen-specific B cell receptors from immune repertoires

Fine-tuned protein language model identifies antigen-specific B cell receptors from immune repertoires

(2025)

Abstract Scalable identification of antigen-specific antibodies from whole immune repertoire V(D)J sequences is a central challenge in biomedical engineering. We show that protein language models (PLMs) fine-tuned on antibody heavy-chain sequences can directly predict antigen specificity from unselected immune repertoires. We assessed our model, Antigen Specificity Predictor (ASPred), against SARS-CoV-2, influenza, and HIV-AIDS antigens, observing comparable predictive performance. In the whole immune repertoire V(D)J sequences of mice immunized with the SARS-CoV-2 spike protein’s receptor-binding domain (RBD), ASPred identified antibody sequences specific to RBD. Several candidate sequences were validated, including one as a heavy chain-only nanobody with 20.7 nM dissociation constant. Molecular dynamics simulations supported the predicted interactions at coarse-grained and atomic levels. Benchmarking against Barcode-Enabled Antigen Mapping (BEAM) of B cell receptor sequence data had highly significant overlaps with ASPred predictions, suggesting scalability. The predicted SARS-CoV-2 binders differed substantially from training sequences, demonstrating generalization beyond sequence memorization. Together, we establish that heavy chain antibody sequences encode sufficient information for PLMs to infer specificity, offering a scalable framework for antibody discovery with broad applications.

Cover page of Towards optimal selection of ultra-deep sequencing reads for de novo genome assembly

Towards optimal selection of ultra-deep sequencing reads for de novo genome assembly

(2025)

When sequencing a new genome, it is common practice to expect that 30-50× sequencing depth will be sufficient for a complete and highly contiguous assembly. With the rapid decrease in the cost of sequencing DNA, on small genomes it is not uncommon to have excessive sequencing data, sometimes exceeding 1000× sequencing depth (which we call ultra-deep). Because ultra-deep sequencing data significantly degrades the quality of the final assembly (for reasons not entirely clear to us), one faces the problem of how to select a subsample of the data for optimal assembly. The optimal read selection problem for genome assembly is largely unexplored. Here we first show that this problem is related to the minimum tiling path (MTP) problem which is known to be NP-hard. Then, we propose a heuristic (called AWinK) based on single-copy k-mer to select a subset of ultra-deep sequencing reads that maximizes the genomic coverage. Our experiments on both synthetic and real ultra-deep sequencing data demonstrate that AWinK can approximate the minimum tiling path in obtaining highly contiguous, accurate, and complete genome assembly. Compared to other six read selection strategies, subsets of reads chosen with AWinK produced assemblies that had the highest genome fraction and sequence identity.

Cover page of Author Correction: A universal language for finding mass spectrometry data patterns

Author Correction: A universal language for finding mass spectrometry data patterns

(2025)

Correction to: Nature Methodshttps://doi.org/10.1038/s41592-025-02660-z, published online 12 May 2025. This article was originally published under standard Springer Nature license (© The Author(s), under exclusive licence to Springer Nature America, Inc.). It is now available as an open-access paper under a Creative Commons Attribution 4.0 International license, © The Author(s). The error has been corrected in the HTML and PDF versions of the article.