alluxio engineering Archives | Page 5 of 9

Scalable Filesystem Metadata Services with RocksDB

July 22, 2019

Alluxio maintainer and founding engineer Calvin Jia presents on Scalable Filesystem Metadata Services with RocksDB at the RocksDB meetup at Twitter.

Tags: alluxio engineering, meetup, metadata management, performance, scale, storage, unified namespace

Democratizing Data Orchestration | Haoyuan (H.Y.) Li – Alluxio

July 18, 2019

TFiR – Open Source & Emerging Technologies In this interview we spoke to Haoyuan (H.Y.) Li, Founder, Chairman and CTO of Open Source Alluxio, a company that is democratizing data in the cloud.

Tags: alluxio engineering, analytics, compute, compute storage separation, data, data orchestration

Cybersecurity and fraud detection at ING Bank using Presto & Alluxio on S3

Alluxio Global Online Meetup * August 1, 2019

This event features leading financial services company ING Bank’s user story on how they leverage open source technologies like Presto and Alluxio with S3.

Getting Started with the Alluxio-Presto Sandbox

July 11, 2019 By Zac Blanco

The Alluxio-Presto sandbox is a docker application featuring installations of MySQL, Hadoop, Hive, Presto, and Alluxio. The sandbox lets you easily dive into an interactive environment where you can explore Alluxio, run queries with Presto, and see the performance benefits of using Alluxio in a big data software stack.

Accelerating Analytical Workloads for Public & Hybrid Clouds

New York Meetup * July 10, 2019

In this meetup, Dipti and HY will present a new approach to hybrid analytical workloads using Alluxio, an open source data orchestration layer, which sits between compute and storage layer. Applications like Apache Spark or TensorFlow can then seamlessly access multiple disparate data sources with consistent performance using data locality and abstraction that the data orchestration tier brings.

Turn cloud storage or HDFS into your local file system for faster AI model training with TensorFlow

July 3, 2019 By Lu Qiu and Bin Fan

This article aims to provide a different approach to help connect and make distributed files systems like HDFS or cloud storage systems look like a local file system to data processing frameworks: the Alluxio POSIX API. To explain the approach better, we used the TensorFlow + Alluxio + AWS S3 stack as an example.

How do you partition Hive Table across storage systems using Alluxio?

Today when we create a Hive table, it is a common technique to partition the table across different values and ranges to improve query performance and reduce maintenance cost. However, Hive can not access a single table directly using a single query with the data of this Hive table across different mediums of storage and … Continued

Scalable Metadata Service in Alluxio: Storing Billions of Files

May 10, 2019 By Andrew Audibert

We are writing several engineering blogs describing the design and implementation of Alluxio master to address this scalability challenge. This is the first article focusing on metadata storage and service, particularly how to use RocksDB as an embedded persistent key-value store to encode and store the file system inode tree with high performance.
Alluxio serves its metadata from a single active master as the primary and potentially multiple standby master for high availability. The master handles all metadata requests and uses a write-ahead log to journal all changes so that we can recover from crashes. The log is typically written to shared storage like HDFS for persistence and availability. Standby masters read the write-ahead log to keep their own state up-to-date. If the primary master dies, one of the standbys can quickly take over for it.

Moving From Apache Thrift to gRPC: A Perspective From Alluxio

April 13, 2019 By Gokturk Gezer and Bin Feng

As part of the Alluxio 2.0 release, we have moved our RPC framework from Apache Thrift to gRPC. In this article, we will talk about the reasons behind this change as well as some lessons we learned along the way.
In Alluxio 1.x, the RPC communication between clients and servers is built mostly on top of Apache Thrift. Thrift enabled us to define Alluxio service interface in simple IDL files and implement client binding using native Java interfaces generated by Thrift compiler. However, we faced several challenges as we continued developing new features and improvements for Alluxio.

Tag: alluxio engineering