“Zero-Copy” Hybrid Cloud for Data Analytics – Strategy, Architecture and Benchmark Report
This whitepaper details how to leverage a public cloud, such as Amazon AWS, Google GCP, or Microsoft Azure to scale analytic workloads directly on data on-premises without copying and synchronizing the data into the cloud. We will show an example of what it might look like to run on-demand Presto and Hive with Alluxio in the public cloud using on-prem HDFS. We will also show how to set up and execute performance benchmarks in two geographically dispersed Amazon EMR clusters along with a summary of our findings.
Tags: aws, azure, data analytics, emr, gcp, hdfs, hive, hybrid cloud, presto, public cloud, zero copy