Skip to main content
eScholarship
Open Access Publications from the University of California

The Department of Statistics at UCLA coordinates undergraduate and graduate statistics teaching and research within the College of Letters and Sciences. We teach a large number of undergraduates and we have a substantial graduate program. Our research and teaching have a strong emphasis on computational and applied statistics.

Cover page of Stratified Multiple Imputation Estimation for Complex Surveys

Stratified Multiple Imputation Estimation for Complex Surveys

(2026)

This research presents an approach to conducting multiple imputation of the R&D expenditures of businesses. It could be titled: don’t forget the past when imputing the present.

The approach uses all past years of survey data to learn about the attrition patterns of businesses. This gives a criterion for imputing the data. Based on that criterion, all years of data are imputed at once, each year’s imputation gaining strength from the other year’s information.

Cover page of Converting Statistical Literacy Resources to Data Science Resources

Converting Statistical Literacy Resources to Data Science Resources

(2023)

Data Science is considered a pseudonym for handling big data, machine learning, statistics, computing and mathematics. It is no uncommon for learners to think that all that requires a radical change in their education and even a change in the name of their Statistics major. However, it is not too difficult for a program promoting statistical literacy to at the same time use its resources as data science resources. It takes translation, some acquaintance with what all practitioners of data science usually do, and a willingness to edit the resources to make learners feel that they are immersed in the data science world.

Cover page of CGM and insulin pump data to introduce classical and machine learning time series analysis concepts to students

CGM and insulin pump data to introduce classical and machine learning time series analysis concepts to students

(2021)

The case study engages students and makes them use the tools they know to investigate a complex process. At the same time, they learn basic time-series concepts using only their intro stats tools.

Cover page of Women mathematicians in data-centric occupations (with a context)

Women mathematicians in data-centric occupations (with a context)

(2020)

Texas State University, Women Doing Math and Talk Math 2 Me Joint Statistics Seminar

  • 1 supplemental PDF
Cover page of Sensitivity of Econometric Estimates to Item Non-response Adjustment

Sensitivity of Econometric Estimates to Item Non-response Adjustment

(2016)

Non-response in establishment surveys is a very important problem that can bias results of statistical analysis. The bias can be considerable when the survey data is used to do multivariate analysis that involve several variables with different response rates, which can reduce the effective sample size considerably. Fixing the non-response, however, could potentially cause other econometric problems. This paper uses an operational approach to analyze the sensitivity of results of multivariate analysis to multiple imputation procedures applied to the U.S. Census Bureau/NSF‘s Business Research and Development and Innovation Survey (BRDIS) to address item non-response. Multiple imputation is first applied using data from all survey units and periods for which there is data, presenting scenario 1. A scenario 2 involves separate imputation for units that have participated in the survey only once and those that repeat. Scenario 3 involves no imputation. Sensitivity analysis is done by comparing the model estimates and their standard errors, and measures of the additional uncertainty created by the imputation procedure. In all cases, unit non-response is addressed by using the adjusted weights that accompany BRDIS micro data. The results suggest that substantial benefit may be derived from multiple imputation, not only because it helps provide more accurate measures of the uncertainty due to item non-response but also because it provides alternative estimates of effect sizes and population totals.