Skip to main content
eScholarship
Open Access Publications from the University of California

School of Information

Open Access Policy Deposits bannerUC Berkeley

Open Access Policy Deposits

This series is automatically populated with publications deposited by UC Berkeley School of Information researchers in accordance with the University of California’s open access policies. For more information see Open Access Policy Deposits and the UC Publication Management System.

Cover page of The Silicon Valley-Hsinchu Connection: Technical Communities and Industrial Upgrading

The Silicon Valley-Hsinchu Connection: Technical Communities and Industrial Upgrading

(2001)

Silicon Valley in California and the Hsinchu-Taipei region of Taiwan are among the most frequently cited ‘miracles’of the information technology era. The dominant accounts of these successes treat them in isolation, focusing either on free markets, multinationals or the state. This paper argues that the dynamism of these regional economies is attributable to their increasing interdependencies. A community of US-educated Taiwanese engineers has coordinated a decentralized process of reciprocal industrial upgrading by transferring capital, skill, and know-how and by facilitating collaboration between specialist producers in the two regions. This case underscores the significance of technical communities and their institutions in diffusing ideas and organizing production at the global as well as the local level.

Cover page of Michael K. Buckland -- Bibliography

Michael K. Buckland -- Bibliography

(2026)

Bibliography of publications.

Cover page of Clifford Lynch at Berkeley

Clifford Lynch at Berkeley

(2025)

abstract: Clifford Lynch is known for his long tenure as executive director of the Coalition for Networked Information (CNI). Here, two other achievements are summarized. From 1979 until he moved to CNI in 1997, Clifford was responsible for developing and implementing library infrastructure for the multicampus University of California system, including MELVYL, a highly innovative, user-oriented online replacement for card catalogs and its extension to provide access to medical and other bibliographical resources. To support it and other applications, he and others built an intercampus network that evolved into the university's Internet node. In addition, for more than three decades he also team-taught the Friday Afternoon Seminar, a weekly Berkeley campus colloquium series featuring a wide range of research reports.

Cover page of Categorical misalignment: Making autism(s) in big data biobanking

Categorical misalignment: Making autism(s) in big data biobanking

(2025)

The opaque relationship between biology and behavior is an intractable problem for psychiatry, and it increasingly challenges longstanding diagnostic categorizations. While various big data sciences have been repeatedly deployed as potential solutions, they have so far complicated more than they have managed to disentangle. Attending to categorical misalignment, this article proposes one reason why this is the case: Datasets have to instantiate clinical categories in order to make biological sense of them, and they do so in different ways. Here, I use mixed methods to examine the role of the reuse of big data in recent genomic research on autism spectrum disorder (ASD). I show how divergent regimes of psychiatric categorization are innately encoded within commonly used datasets from MSSNG and 23andMe, contributing to a rippling disjuncture in the accounts of autism that this body of research has produced. Beyond the specific complications this dynamic introduces for the category of autism, this paper argues for the necessity of critical attention to the role of dataset reuse and recombination across human genomics and beyond.

Cover page of Philip Olin Keeney: Checklist of Writings

Philip Olin Keeney: Checklist of Writings

(2025)

Checklist of writings by Philip Olin Keeney (1891-1962).

A Postgenomic Quilt: The Turn to Endophenotypes and Deep Phenotyping in Human Genetics

(2025)

How do experts stitch together seemingly irreconcilable scientific objects? This paper examines the way “endophenotypes” and “deep phenotyping” are deployed to bridge the infamously recalcitrant genotype–phenotype divide. The endophenotype concept emerged in mid-20th-century evolutionary biology/entomology and was briefly adopted in schizophrenia genetics in 1972 before virtually disappearing for a generation. Decades later, in the “postgenomic” 2000s, the concept exploded in medical—and especially psychiatric—genetics in the United States and United Kingdom. Endophenotypes now refer to phenotypic observations like subclinical biomarkers or traits that are both more fine-grained than disease categories and yield stronger associations with genetics. We argue that the endophenotype concept functions as an “epistemic quilting device” that allows experts to forge research programs across disjunct scientific ontologies and fields, creating new wholes that respond to historically specific needs. But they are also destabilizing existing medical categories and foundational concepts in genetics like penetrance and recessivity, with radical implications for “precision medicine.” Today, endophenotypes and the more loosely defined “deep phenotyping” have been integrated into the infrastructures of postgenomic research in both rare disease and big-data genomics, quilting together genetics and other fields interested in human illness and difference even as it disrupts them.

The cost of (data) community: error and repair in data processing pipelines

(2025)

This essay closely examines the ongoing development and use of an open-source software tool commonly used in microbiome research in order to make three interlocking contributions. First, I identify data cleaning as a set of richly epistemic practices which are functionally inextricable from data analysis. This means that ‘good’ data cleaning decisions are not necessarily universal, but must be made suitably for the specific analytic purposes intended by later users of shared data. Second, I examine how the repair and modularity of data processing software can offer data users a variety of cleaning choices in order to enable divergent analytic goals. In doing so, this software facilitates the development of epistemically diverse data communities. Finally, turning to a high-profile paper retraction which hinged on a data processing error, I explore how repair at different points in the research production process can serve to enable or constrain the growth of data communities. Through this analysis, I argue that error serves as a site of productive negotiation over the interpretive flexibility of shared data, and that repair plays a critical role in stabilizing data community membership.

X under Musk’s leadership: Substantial hate and no reduction in inauthentic activity

(2025)

Numerous studies have reported an increase in hate speech on X (formerly Twitter) in the months immediately following Elon Musk's acquisition of the platform on October 27th, 2022; relatedly, despite Musk's pledge to "defeat the spam bots," a recent study reported no substantial change in the concentration of inauthentic accounts. However, it is not known whether any of these trends endured. We address this by examining material posted on X from the beginning of 2022 through June 2023, the period that includes Musk's full tenure as CEO. We find that the increase in hate speech just before Musk bought X persisted until at least May of 2023, with the weekly rate of hate speech being approximately 50% higher than the months preceding his purchase, although this increase cannot be directly attributed to any policy at X. The increase is seen across multiple dimensions of hate, including racism, homophobia, and transphobia. Moreover, there is a doubling of hate post "likes," indicating increased engagement with hate posts. In addition to measuring hate speech, we also measure the presence of inauthentic accounts on the platform; these accounts are often used in spam and malicious information campaigns. We find no reduction (and a possible increase) in activity by these users after Musk purchased X, which could point to further negative outcomes, such as the potential for scams, interference in elections, or harm to public health campaigns. Overall, the long-term increase in hate speech, and the prevalence of potentially inauthentic accounts, are concerning, as these factors can undermine safe and democratic online environments, and increase the risk of offline harms.

Cover page of People are poorly equipped to detect AI-powered voice clones

People are poorly equipped to detect AI-powered voice clones

(2025)

As generative artificial intelligence (AI) continues its ballistic trajectory, everything from text to audio, image, and video generation continues to improve at mimicking human-generated content. Through a series of perceptual studies, we report on the realism of AI-generated voices in terms of identity matching and naturalness. We find human participants cannot consistently identify recordings of AI-generated voices. Specifically, participants perceived the identity of an AI-generated voice to be the same as its real counterpart approximately $$80\%$$ of the time, and correctly identified a voice as AI generated only about $$60\%$$ of the time.

Cover page of Abortion access barriers shared in “r/abortion” after Roe: a qualitative analysis of a Reddit community post-Dobbs decision leak in 2022

Abortion access barriers shared in “r/abortion” after Roe: a qualitative analysis of a Reddit community post-Dobbs decision leak in 2022

(2024)

With drastic changes to abortion policy, the months following the Dobbs leak and subsequent decision in 2022 were a uniquely uncertain and difficult time for abortion access in the United States. To understand experiences of challenges to abortion access during that time, we used a hybrid inductive and deductive thematic coding approach to analyse descriptions of barriers and their impacts shared in an abortion subreddit (r/abortion). A simple random sample of 10% of posts was obtained from those shared from 02 May 2022 through 23 December 2022; comments were purposively sampled during the coding process. In this sample of submissions (n = 523 posts, 88 comments), people described structural barriers identified in past research, including state abortion bans and gestational limits, high costs, limited appointment availability, and long travel required. Posters also commonly described known social barriers, including limited social support and abortion stigma. Several impactful barriers not well-described in past research emerged inductively, including wait time for receiving mail-ordered abortion medication, low credibility of online ordering platforms, and concerns about legal risks of accessing abortion or related medical care. The most common consequences of experiencing barriers were adverse mental health outcomes, delayed access to care, and being compelled to self-manage their abortion because of access barriers. This analysis provides timely insights into the experiences and impacts of abortion access barriers in a group of people with a range of engagement with clinical abortion care, lived experiences, and points in their abortion processes, with public health implications for mental health and abortion access.