presto Archives | Page 2 of 11

Speed Up Uber’s Presto with Alluxio

March 4, 2022

This talk covers how Uber’s Presto team implements the cache invalidation and dashboard for Alluxio’s Local Cache. Liang Chen will also share his experience using a customized cache filter to resolve the performance degradation due to a large working set.

Tags: alluxio day, local cache, performance, presto, uber

Using Consistent Hashing in Presto to Improve Caching Data Locality in Dynamic Clusters

February 2, 2022 By Rongrong Zhong

Running Presto with Alluxio is gaining popularity in the community. It avoids long latency reading data from remote storage by utilizing SSD or memory to cache hot dataset close to Presto workers. Presto supports hash-based soft affinity scheduling to enforce that only one or two copies of the same data are cached in the entire cluster, which improves cache efficiency by allowing more hot data cached locally. The current hashing algorithm used, however, does not work well when cluster size changes. This article introduces a new hashing algorithm for soft affinity scheduling, consistent hashing, to address this problem.

Improve Presto Architectural Decisions with Shadow Cache

October 12, 2021

This talk describes the design of shadow cache, a lightweight component to track the working set size of Alluxio cache. Shadow cache can keep track of the working set size over the past window dynamically, and is implemented by a series of bloom filters. We’ve deployed the shadow cache in Facebook Presto and leverage the result to understand the system bottleneck and help with routing design decisions.

Tags: alluxio day, architecture, cache, facebook, presto, shadow cache

Improving Presto performance with Alluxio at TikTok

June 24, 2021

Nowadays it is not straightforward to integrate Alluxio with popular query engines like Presto on existing Hive data. Solutions proposed by the community like Alluxio Catalog Service or Transparent URI brings unnecessary pressure on Alluxio masters when querying files should not be cached.

Tags: alluxio day, cache layer, hive, presto, tiktok

RaptorX: Building a 10X Faster Presto with hierarchical cache

June 24, 2021

RaptorX is an internal project name aiming to boost query latency significantly beyond what vanilla Presto is capable of. For this session, we introduce the hierarchical cache work including Alluxio data cache, fragment result cache, etc. Cache is the key building block for RaptorX.

Tags: alluxio day, disaggregated storage, facebook, presto, raptorx

Building a high-performance data lake analytics engine at Alibaba Cloud with Presto+Alluxio

April 27, 2021

Data Lake Analytics(DLA) is a large scale serverless data federation service on Alibaba Cloud. One of its serverless analytics engine is based on Presto. The DLA Presto engine supports a variety of data sources and is widely used in different application scenarios in the cloud. In this session, we will talk about the system architecture of DLA Presto engine, as well as the challenges and solutions. In particular, we will introduce the use of alluxio local cache to solve performance issues on OSS data sources caused by access delay and OSS bandwidth limitation. We will discuss the principle of alluxio local cache and some improvements we have made.

Tags: alibaba, alluxio day, data lake analytics, local cache, performance, presto

Tag: presto