Spotify's RAP: Data Lakes Now Serve Real-Time Needs

Alps Wang

Alps Wang

Aug 13, 2026 · 1 views

Unlocking Data Lake Potential

Spotify's introduction of Random Access Parquet (RAP) is a compelling solution to a long-standing challenge: enabling low-latency, key-based point queries directly on data lakes, which have historically been optimized for batch analytical scans. By introducing an external indexing layer over Parquet files, RAP effectively bridges the gap between cost-effective, massive data lake storage and the immediate access requirements of online services and AI applications. This approach cleverly avoids the prohibitive cost and complexity of replicating petabytes of data into operational databases, a common workaround that introduces data staleness and maintenance overhead. The architecture's compatibility with existing Apache Iceberg tables and Parquet files is a significant advantage, facilitating adoption within existing data ecosystems.

The innovation lies in its pragmatic approach to optimizing for point lookups without sacrificing the analytical capabilities of the data lake. Techniques like data sorting, record grouping, interleaved column layouts, and covering indexes are not entirely novel individually, but their integration into an indexing framework specifically for data lakes is noteworthy. The ability to support secondary indexes and various indexing strategies (hash, sorted) further enhances its flexibility. This is particularly impactful for AI agents that require rapid, granular data retrieval for tasks like personalization or real-time decision-making. The architecture's ability to serve both analytical workloads and latency-sensitive applications from a single source of truth is a major step towards unifying data platforms and reducing operational complexity.

However, potential limitations could include the overhead associated with maintaining the external index, especially for rapidly changing datasets. The complexity of managing this index layer, while designed to be append-only, might still introduce operational considerations. Furthermore, the effectiveness of the storage layout optimizations is highly dependent on the query patterns. While RAP aims to minimize file access, the initial cost of organizing data for optimal locality could be a factor. The success of this approach will also hinge on the performance and scalability of the index builder and query resolution mechanisms under extreme load. Nonetheless, for organizations grappling with the cost and latency of serving data from their lakes, RAP presents a highly attractive and technically sound solution.

Key Points

  • Spotify developed Random Access Parquet (RAP), an external indexing layer for data lakes.
  • RAP enables low-latency point queries directly on data lake data (Parquet files).
  • This avoids replicating data into operational databases, saving costs and complexity.
  • It supports online services and AI applications needing immediate, granular data access.
  • RAP is compatible with Apache Iceberg tables and existing Parquet files.
  • Techniques like data sorting, record grouping, and interleaved columns optimize query performance.
  • Secondary indexes and hash/sorted indexes enhance query flexibility.
  • The solution reduces query planning and metadata traversal overhead for point lookups.

Article Image


📖 Source: Spotify Builds External Index to Enable Low Latency Point Queries on Its Data Lake

Related Articles

Comments (0)

No comments yet. Be the first to comment!