- Main
Titian: Data Provenance Support in Spark.
- Interlandi, Matteo;
- Shah, Kshitij;
- Tetali, Sai Deep;
- Gulzar, Muhammad Ali;
- Yoo, Seunghyun;
- Kim, Miryung;
- Millstein, Todd;
- Condie, Tyson
Published Web Location
https://doi.org/10.14778/2850583.2850595Abstract
Debugging data processing logic in Data-Intensive Scalable Computing (DISC) systems is a difficult and time consuming effort. Today's DISC systems offer very little tooling for debugging programs, and as a result programmers spend countless hours collecting evidence (e.g., from log files) and performing trial and error debugging. To aid this effort, we built Titian, a library that enables data provenance-tracking data through transformations-in Apache Spark. Data scientists using the Titian Spark extension will be able to quickly identify the input data at the root cause of a potential bug or outlier result. Titian is built directly into the Spark platform and offers data provenance support at interactive speeds-orders-of-magnitude faster than alternative solutions-while minimally impacting Spark job performance; observed overheads for capturing data lineage rarely exceed 30% above the baseline job execution time.
Many UC-authored scholarly publications are freely available on this site because of the UC's open access policies. Let us know how this access is important for you.