On Demand Video

Community Office Hour: Accelerating Hive with Alluxio on S3


Many organizations are leveraging Hive to run big data analytics on public cloud. However, reading and writing data to S3 directly can result in slow and inconsistent performance. Alluxio is a data orchestration layer for the cloud, and in this use case it caches data for S3, ensuring high and predictable performance as well as reduced network traffic.

In this Office Hour we’ll go over:

  • Bazaarvoice’s use case leveraging Apache Spark, Hive, and Alluxio on S3
  • How to set up Hive with Alluxio such that Hive jobs can seamlessly read from and write to S3
  • Open Session for discussion on any topics such as solving the separation of compute and storage problem, and more

Bin Fan is the founding engineer of Alluxio, Inc. and the PMC member of Alluxio open source project. Prior to Alluxio, he worked for Google where he won the Technical Infrastructure Award. Bin received his Ph.D. in Computer Science from Carnegie Mellon University working on distributed systems

Nakkul Sreenivas is a software engineer at Alluxio. Prior to Alluxio, Nakkul worked as a consultant where he built and supported an entirely open source Hadoop platform for financial services clients.

Questions? Slack with the speakers, users, and many other community members!
Welcome to join Alluxio Global Online Meetup Group to attend online meetups like this!


Presentation slides: