Skip to main content
Download PDF
- Main
DS SERVE: A Framework for Efficient and Scalable Neural Retrieval
- Liu, Jinjian;
- Wang, Yichuan;
- Lyu, Xinxi;
- Shao, Rulin;
- Gonzalez, Joseph E;
- Zaharia, Matei;
- Min, Sewon
Published Web Location
https://doi.org/10.1609/aaai.v40i48.42363Abstract
We present DS SERVE, a framework that transforms large-scale text datasets—comprising half a trillion tokens—into a high-performance neural retrieval system. DS SERVE offers both a web interface and API endpoints, achieving low latency with modest memory overhead on a single node. The framework also supports inference-time tradeoffs between latency, accuracy, and result diversity. We anticipate that DS SERVE will be broadly useful for a range of applications such as large-scale retrieval-augmented generation (RAG), training data attribution, training a search agent, and beyond.
Many UC-authored scholarly publications are freely available on this site because of the UC's open access policies. Let us know how this access is important for you.