Skip to main content
eScholarship
Open Access Publications from the University of California

LBL Publications

Lawrence Berkeley National Laboratory (Berkeley Lab) has been a leader in science and engineering research for more than 70 years. Located on a 200 acre site in the hills above the Berkeley campus of the University of California, overlooking the San Francisco Bay, Berkeley Lab is a U.S. Department of Energy (DOE) National Laboratory managed by the University of California. It has an annual budget of nearly $480 million (FY2002) and employs a staff of about 4,300, including more than a thousand students.

Berkeley Lab conducts unclassified research across a wide range of scientific disciplines with key efforts in fundamental studies of the universe; quantitative biology; nanoscience; new energy systems and environmental solutions; and the use of integrated computing as a tool for discovery. It is organized into 17 scientific divisions and hosts four DOE national user facilities. Details on Berkeley Lab's divisions and user facilities can be viewed here.

Cover page of Characterizing Lossless GPU Data Compression Across AMD CDNA and RDNA Architectures

Characterizing Lossless GPU Data Compression Across AMD CDNA and RDNA Architectures

(2027)

Data movement remains a major bottleneck in high-performance computing workflows, making GPU-accelerated compression an increasingly important optimization. While NVIDIA platforms benefit from mature compression libraries such as nvCOMP, the performance of practical lossless compression on AMD GPUs remains undercharacterized. This paper addresses that gap by porting LZ4, Snappy, and Cascaded from CUDA to HIP and evaluating them across four AMD GPUs spanning three CDNA generations (MI50, MI210, MI300X) and RDNA 3 (RX 7900 XT), using both synthetic datasets and seismic simulation data. Our study makes three contributions. First, we provide a cross-generational characterization of lossless GPU compression on recent AMD accelerators. Second, we present a functional CUDA-to-HIP port of three widely used algorithms, establishing an open baseline for future AMD-specific optimization. Third, we analyze how architectural differences shape compression and decompression behavior across algorithms and datasets. The results show that the MI300X achieves up to 11×$$11\times $$ higher decompression throughput than the MI50, while RDNA 3 is competitive for compression workloads characterized by irregular memory accesses. We also show that transfers dominate end-to-end execution time on all evaluated platforms, indicating that the main benefits of GPU compression are most likely to emerge in GPU-resident workflows. For realistic seismic floating-point data, lossless compression ratios remain modest (1.03-1.04×$$1.03-1.04\times $$), suggesting that error-bounded lossy compression is a promising direction for future work.

Cover page of Thermo-hydro-mechanical analysis of subsurface ice-based thermal energy storage

Thermo-hydro-mechanical analysis of subsurface ice-based thermal energy storage

(2026)

Ice-based thermal energy storage systems are widely utilized for cooling and managing peak electrical demand globally, offering daily or weekly storage capabilities for both individual homes and larger office buildings. However, scaling these systems for district-level cooling or integrating them with renewable energy sources presents challenges, especially in accommodating larger volumes and addressing seasonal storage requirements in densely populated urban areas. This paper proposes a novel solution by evaluating subsurface ice-based thermal energy storage, in which the underground is subjected to seasonal freeze/thaw cycles. However, these cycles may influence ground behavior, affecting pore pressure and inducing ground movement. To systematically investigate these challenges, we enhance the TOUGH-FLAC simulator by integrating water/ice phase change capabilities and updating the effective stress–strain constitutive relation. Both modifications are validated against analytical solutions or experimental data. Through numerical simulations spanning a decade with ten seasonal freeze/thaw cycles, we evaluate the performance and long-term stability of a generic subsurface ice-based thermal energy storage system, considering factors such as ground permeability, freezing pipe spacing, freeze/thaw damage, and glycol solution temperature. The simulations indicate that ice formation induces pore pressure variations that drive seasonal surface heave and settlement, controlled by ground permeability, pipe spacing, and glycol solution temperature, along with tensile and localized shear deformation around freeze pipes. This highlights the need for accurate ground property characterization and geomechanical analysis for subsurface ice-based thermal energy storage.

Robust electron counting for direct electron detectors with the Back-propagation counting method

(2026)

Electron microscopy (EM) is a foundational tool for directly assessing the structure of materials. Recent advances in direct electron detectors have improved signal-to noise ratios via single-electron counting. However, accurately counting electrons at high flux remains challenging. We developed a new method of electron counting for direct electron detectors, Back-Propagation Counting (BPC). BPC uses machine learning techniques designed for mathematical operations on large tensors but does not require large training datasets. In synthetic data, we show BPC is able to count multiple electron strikes per pixel and is robust to increasing occupancy. In experimental data, frames counted with BPC are shown to reconstruct diffraction peaks corresponding to individual nanoparticles with relatively higher intensity and produce images with improved contrast when compared to a standard counting method. Together, these results show that BPC excels in experiments where pixels see a high flux of electron irradiation such as in situ TEM movies and diffraction.

Cover page of Estimating the impact of tariff-driven behind-the-meter storage operation on distribution grid investments

Estimating the impact of tariff-driven behind-the-meter storage operation on distribution grid investments

(2026)

Increasing growth of distributed solar photovoltaics (PV) and electric vehicles (EV) can strain local distribution networks and require costly upgrades. Distributed battery storage, often deployed alongside PV, can be used to mitigate those costs, depending on how batteries are operated. This study evaluates the potential deferral value of distributed battery storage across a range of tariff structures, focusing on the rate structures most commonly available to residential customers today and related variants. Deferrals are evaluated with a least-cost distribution grid expansion optimization model to identify requirements on line reconductoring, transformer upgrades, and voltage regulator installations under each tariff. Results show that TOU rates and net billing tariffs can yield meaningful deferral value, depending on specific tariff structure features. Under the best performing tariff structure tested, storage produced a median annualized deferral value of $7.18 per kW of storage capacity ( kW S ) across all feeders in the sample, though deferral values were considerably larger for feeders with peak loads that coincide with utility system peak, i.e., timing of TOU peak period. In contrast, under an unrestricted TOU design with no restrictions on grid charging or discharging, the median deferral value was $0/ kW S illustrating the critical importance of tariff structure details.

Cover page of A sharp interface method for two-phase incompressible flow with surface tension

A sharp interface method for two-phase incompressible flow with surface tension

(2026)

We present a sharp-interface method for resolved two-phase incompressible viscous flow based on embedded boundary finite volume discretizations of the individual phase domains. The numerical algorithm is a fractional step method, in which incompressibility is enforced at the half-step and the full-step by solving coupled Hodge projections that respect the jump boundary conditions in pressure and pressure gradient at the interface. Surface tension enters through these boundary conditions. The viscous term in the momentum equations is solved implicitly using a Crank-Nicholson time discretization, respecting the jump conditions on velocity and velocity gradient. The method is implemented with block-structured adaptive mesh refinement. We demonstrate stable behavior in 2D and 3D, with the velocity converging in all norms as measured using Richardson extrapolation. The method successfully models spherical cap bubble shapes and velocities, showing good agreement with a range of experiments.

Cover page of Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

(2026)

Device-to-device variability in experimental noise critically impacts reproducibility, especially in automated, high-throughput systems like additive manufacturing farms. While manageable in small labs, such variability can escalate into serious risks at larger scales, such as architectural 3D printing, where noise may cause structural or economic failures. This contribution presents a noise-aware decision-making algorithm that quantifies and models device-specific noise profiles to manage variability adaptively. It uses distributional analysis and pairwise divergence metrics with clustering to choose between single-device and robust multi-device Bayesian optimization strategies. Unlike conventional methods that assume homogeneous devices or enforce generic robustness, the proposed framework explicitly determines whether shared optimization across devices is appropriate based on the degree of inter-device noise heterogeneity. This enables improved performance, reproducibility, and efficiency. An experimental case study involving three nominally identical 3D printers (same brand, model, and close serial numbers) demonstrates reduced redundancy, lower resource usage, and improved reliability, along with improved convergence stability and solution quality through the selection of the appropriate optimization strategy based on the degree of inter-device noise heterogeneity. Overall, this framework establishes a general approach for precision- and resource-aware optimization in scalable, automated experimental platforms, demonstrated here on a representative multi-device 3D printing case study.

Demonstration and analysis of volumetric additive manufacturing via sub-orbital spaceflight testing

(2026)

Computed Axial Lithography (CAL) represents a significant advancement in the emerging field of Volumetric Additive Manufacturing (VAM). CAL addresses key limitations of traditional photopolymer additive manufacturing technologies, by eliminating the need for layering and support structures. Unlike conventional methods, CAL prints components by illuminating all points within a desired geometry simultaneously, using tomographic reconstruction to form the object in a single step. This unique approach eliminates the relative motion between the object and the precursor material, enabling faster printing speeds and reducing the waste associated with support structures. However, CAL parts require post-processing steps before they can be utilized. CAL's core attributes make it particularly suited for In-Space Manufacturing (ISM), due to its fast fabrication times, wide breadth of materials it can use, and minimized footprint. CAL has been successfully demonstrated in microgravity during parabolic flight experiments. However to fully validate and understand CAL's behaviour in microgravity, all manufacturing and post-processing steps must be integrated. In June 2024, we conducted SpaceCAL Mission 3, testing this entire workflow on a suborbital flight aboard Virgin Galactic's SpaceShipTwo. During ∼140 s of microgravity, the system autonomously manufactured and post-processed four parts using PEGDA700 resin. Post-flight analysis showed that 2/4 parts were recognisable, while others were distorted due to bubble formation from residual water droplets, off-axis optical aberrations, and non-uniform solvent rinsing. Despite these limitations, this study represents the first integrated CAL workflow in space, providing an initial experimental demonstration and analysis for closed-loop in-space manufacturing.

Cover page of Insights from a coupled thermo-hydro-mechanical analysis of a layered high-temperature thermal energy storage reservoir

Insights from a coupled thermo-hydro-mechanical analysis of a layered high-temperature thermal energy storage reservoir

(2026)

Coupled thermal-hydraulic-mechanical (THM) modeling is applied to investigate the performance of a seasonal high-temperature aquifer thermal energy storage operation based on data and conditions from current site investigations at the Geostorage Forsthaus pilot project in Bern (Switzerland). The model includes subhorizontal sand lenses of various lengths and dips that are embedded in a low permeability clay matrix. Thermal energy storage is simulated by seasonal injection and withdrawal of hot (up to 90 °C) water from a main well, with reservoir pressure regulated by two auxiliary wells at a distance of about 70 m from the main well. The results show how targeted injection into deeper permeable storage formations, along with active deep well pressure control, can effectively minimize geomechanical impact and the potential risk of damaging subsurface storage and sealing formations, or even surface facilities. With such pressure control, the subsurface mechanical responses are dominated by thermal strain and stress, which can be monitored with subsurface fiber optics. The study demonstrates how coupled THM modeling can be applied for the design of a safe and efficient thermal energy storage operation, and how subsurface fiber optic monitoring can be applied for performance confirmation, allowing for more confident operational forecasting.

Cover page of Fast solvers for tokamak fluid models with PETSc

Fast solvers for tokamak fluid models with PETSc

(2026)

Multigrid (MG) is widely recognized as a highly effective solver for the model problem, the Laplacian, but textbook MG fails on most problems of interest. MG methods have been applied to complex, real-world applications with careful consideration of the physical model and discretization. This work develops the first step in applying MG methods to science and engineering relevant magnetohydrodynamics (MHD) tokamak models in the M3D-C1 (https://m3dc1.pppl.gov) fusion energy science code. The semi-implicit time integrator in M3D-C1 is composed of many linear solves. The implicit advance of the momentum equation is the most challenging and is the focus of this work. The current production solver in M3D-C1 is a block Jacobi (BJ) preconditioner within a Krylov solver, where blocks group degrees of freedom on planes of constant toroidal coordinate. BJ convergence degrades as the number of planes increases due to the spectral properties of the matrix preconditioned with BJ. The partially magnetic field-aligned, regular toroidal grid structure in M3D-C1 is amenable to semi-coarsening geometric MG in the toroidal direction. This paper develops such a solver and demonstrates competitive performance on a runaway electron model of a SPARC (https://cfs.energy/technology/sparc) disruption, and superior robustness on a stellarator model on which the BJ solver fails to converge.