Resource Hub

Presentation

Presentation

Bay Area Meetup: Interactive Analytics in the Cloud with Presto and Alluxio

ALLUXIO BAY AREA MEETUP

This talk describes a stack to combine Presto, Alluxio, and Cloud object storage systems (e.g.,AWS S3) for high-concurrent and low-latency SQL queries on big data on the cloud. Presto, an open-source distributed SQL engine, is widely recognized for its low-latency queries, high concurrency, and native ability to query multiple data sources. Alluxio is an open-source data orchestration that brings data closer to compute and provides a unified data access layer at in-memory speeds. Presto can use Alluxio as a distributed caching tier on top of S3 for the hot data to query, avoiding reading data repeatedly from the cloud.

This talk covers:

The architecture of Presto, its separation of compute and storage, cloud-readiness, recent advancements in the project such as Cost-Based Optimizer and Kubernetes Support.
An overview of Alluxio’s key concepts, architecture and data flow,
Presto and Alluxio production use cases running hundreds of nodes, including ING Bank, JD.com, and NetEase Games.

Presentation

Presentation

Austin Meetup: Efficient Data Engineering with Apache Spark, Hive, and Alluxio on S3

Cloud, Data, & Orchestration – Austin Meetup

At Bazaarvoice, a software-as-a-service digital marketing company, the data engineering team is tasked to handle data at massive Internet-scale to serve over 1,900 of the biggest internet retailers and brands.

We built our data pipelines all in the cloud using Apache Spark and Hive on AWS EC2 accessing data in S3. AWS enables us to scale “out” the infrastructure capacity effortlessly to keep up with the Internet-scale data and web traffic, but scaling out also exposes certain limitations like the ability to further scale “up”. While this cloud native stack is scalable and elastic we experience performance limitations, because data access is limited by the network bandwidth, and this is exacerbated for workloads that involve repeated queries.

To address the data access challenges, we leverage Alluxio, an open source data orchestration system for analytics in the cloud. Alluxio serves as a transparent caching layer for hot and warm data, such that Hive and Spark jobs are able to access all data transparently in S3. We have seen 10x performance acceleration of Spark and Hive jobs on S3 with Alluxio.

Blog

Blog

Four Different Ways to Write to Alluxio

Alluxio is a new layer on top of under storage systems that can not only improve raw I/O performance but also enables applications flexible options to read, write and manage files. This article focuses on describing different ways to write files to Alluxio, realizing the tradeoffs in performance, consistency, and also the level of fault tolerance compared to HDFS.

On Demand Videos

On Demand Videos

Tech Talk: Accelerating Spark with Kubernetes

Blog

Blog

Creating Grafana Dashboards to Visualize Alluxio Metrics

Monitoring metrics is highly important to operate distributed systems in production. Alluxio collects metrics using the Codahale Metrics Library on I/O throughput, RPC throughput, and resource usage. Alluxio metrics are shown in its webUI, but are also available through a REST endpoint or exportable to several third-party sinks in a time-series manner (see docs).

On Demand Videos

On Demand Videos

Tech Talk: Alluxio 2.0 Deep Dive – Simplifying data access for cloud workloads

Blog

Blog

Accelerating Writeintensive Data Workloads on AWS S3

Alluxio is an open-source data orchestration system widely used to speed up data-intensive workloads in the cloud. Alluxio v2.0 introduced Replicated Async Write to allow users to complete writes to Alluxio file system and return quickly with high application performance, while still providing users with peace of mind that data will be persisted to the chosen under storage like S3 in the background.

On Demand Videos

On Demand Videos

Bay Area Meetup: Alluxio 2.0 Deep Dive and Near Real-time Analytics with Spark

ALLUXIO BAY AREA MEETUP

‍

Blog

Blog

Recap AWS Summit New York

Alluxio is a proud sponsor and exhibitor at the AWS Summit in New York. If you weren't able to attend, here are the highlights

Presentation

Presentation

Scalable Filesystem Metadata Services with RocksDB

Alluxio maintainer and founding engineer Calvin Jia presents on Scalable Filesystem Metadata Services with RocksDB at the RocksDB meetup at Twitter.

Alluxio provides a unified namespace where you can mount multiple different storage systems and access them through the same API. To serve the file system requests to operate on all the files and directories in this namespace, Alluxio masters must handle the file system metadata at a scale of all mounted systems combined. We are writing several engineering blogs describing the design and implementation of Alluxio master to address this scalability challenge. This is the first article focusing on metadata storage and service, particularly how to use RocksDB as an embedded persistent key-value store to encode and store the file system inode tree with high performance.

Presentation

Presentation

Alluxio New York Meetup: Accelerating Analytical Workloads for Public & Hybrid Clouds

ALLUXIO NEW YORK MEETUP

The most innovative organizations like Uber, Twitter, and others have moved to disaggregated stacks – a separate tier for computational frameworks like Spark and Presto and a separate tier for Storage. And the need for more compute flexibility is making users move towards hybrid clouds.

In this meetup, Dipti and HY presented a new approach to hybrid analytical workloads using Alluxio, an open source data orchestration layer, which sits between compute and storage layer. Applications like Apache Spark or TensorFlow can then seamlessly access multiple disparate data sources with consistent performance using data locality and abstraction that the data orchestration tier brings.

Haoyuan Li (H.Y.), Alluxio
Haoyuan is the Founder and CTO of Alluxio. He graduated with a Computer Science Ph.D. from the AMPLab at UC Berkeley. At the AMPLab, he co-created and led Alluxio (formerly Tachyon), an open source virtual distributed file system. Before UC Berkeley, he got a M.S. from Cornell University and a B.S. from Peking University, all in Computer Science.

Dipti Borkar, Alluxio
Dipti Borkar is the VP of Product & Marketing at Alluxio with over 15 years experience in data and database technology across relational and non-relational. Prior to Alluxio, Dipti was VP of Product Marketing at Kinetica and Couchbase. Dipti holds a M.S. in Computer Science from the UC San Diego, and an MBA from the Haas School of Business at UC Berkeley.

Blog

Blog

The Practice of Alluxio in Ctrip RealTime Computing Platform

Today, real-time computation platform is becoming increasingly important in many organizations. In this article, we will describe how ctrip.com applies Alluxio to accelerate the Spark SQL real-time jobs and maintain the jobs’ consistency during the downtime of our internal data lake (HDFS). In addition, we leverage Alluxio as a caching layer to dramatically reduce the workload pressure on our HDFS NameNode.

On Demand Videos

On Demand Videos

Democratizing Data Orchestration | Haoyuan (H.Y.) Li – Alluxio

TFiR – Open Source & Emerging Technologies

On Demand Videos

On Demand Videos

Tech Talk: Accelerate and Scale Big Data Analytics with Disaggregated Compute and Storage

Blog

Blog

Orchestrating Data for the Cloud World with Alluxio 2.0

Today, I’m thrilled to announce the GA of Alluxio 2.0.0, Alluxio’s biggest release to date (see our Release Notes & Release Blog) with over 900 commits.

Blog

Blog

Getting Started with the AlluxioPresto Sandbox

The Alluxio-Presto sandbox is a docker application featuring installations of MySQL, Hadoop, Hive, Presto, and Alluxio. The sandbox lets you easily dive into an interactive environment where you can explore Alluxio, run queries with Presto, and see the performance benefits of using Alluxio in a big data software stack.

Blog

Blog

2.0 is here! Embrace silos orchestrate data accelerate innovation

Here in New York, at the AWS Summit, we are super excited to announce that Alluxio 2.0 is here, our most major release since the Alluxio launch. A couple months ago, we released 2.0 Preview - which included some of the capabilities, but 2.0 now includes even more, to continue building on to our data orchestration approach for the cloud.

Blog

Blog

Turn cloud storage or HDFS into your local file system for faster AI model training with TensorFlow

This article aims to provide a different approach to help connect and make distributed files systems like HDFS or cloud storage systems look like a local file system to data processing frameworks: the Alluxio POSIX API. To explain the approach better, we used the TensorFlow + Alluxio + AWS S3 stack as an example.