performance Archives | Page 2 of 13

Speed up Large-scale ML/DL Offline Inference Jobs with Alluxio at Microsoft Bing

January 6, 2022 By Binyang Li and Qianxi Zhang

Running inference at scale is challenging. In this blog, we will share our observations and the practice to use Alluxio to speed up the I/O performance for large-scale ML/DL offline inference at Microsoft Bing.

Metadata Synchronization in Alluxio: Design, Implementation and Optimization

December 14, 2021 By David Zhu

Metadata synchronization (sync) is a core feature in Alluxio that keeps files and directories consistent with their source of truth in under storage systems, thus making it simple for users to reason the data retrieved from Alluxio. Meanwhile, understanding the internal process is important in order to tune the performance. This article describes the design and the implementation in Alluxio to keep metadata synchronized.

Accelerating Machine Learning / Deep Learning in the Cloud: Architecture and Benchmark

December 7, 2021

This whitepaper introduces how to speed up end-to-end distributed training in the cloud using Alluxio to accelerate data access. With the help of Alluxio, loading data from cloud storage, training and caching data can be done in a transparent and distributed way as a part of the training process. This whitepaper also demonstrates how to set up and benchmark the end-to-end performance of the training process, along with a comparison of other popular approaches.

Tags: benchmark, cache, cloud, data orchestration, deep learning, distributed training, machine learning, performance, storage

Speeding up TensorFlow and PyTorch with Alluxio

September 9, 2021

The Alluxio core engineering team re-designed things to come up with a more efficient and transparent way for users to leverage data orchestration through the POSIX interface. This enables much better performance for ML workloads where data is accessed via the POSIX interface.

Tags: data orchestration, fuse, ml, performance, POSIX, pytorch, tensorflow

Design and Implementation of Alluxio POSIX Support

August 31, 2021

Applications like Tensorflow, PyTorch can access data through Alluxio FUSE service without modifying any code just like accessing their local file systems by Unix/Linux POSIX API. This article describes the design and implementation of Alluxio FUSE service, its current status and future plans.

Tags: alluxio engineering, fuse, performance, POSIX

Building a high-performance data lake analytics engine at Alibaba Cloud with Presto+Alluxio

April 27, 2021

Data Lake Analytics(DLA) is a large scale serverless data federation service on Alibaba Cloud. One of its serverless analytics engine is based on Presto. The DLA Presto engine supports a variety of data sources and is widely used in different application scenarios in the cloud. In this session, we will talk about the system architecture of DLA Presto engine, as well as the challenges and solutions. In particular, we will introduce the use of alluxio local cache to solve performance issues on OSS data sources caused by access delay and OSS bandwidth limitation. We will discuss the principle of alluxio local cache and some improvements we have made.

Tags: alibaba, alluxio day, data lake analytics, local cache, performance, presto

Accelerate Analytics and ML in the Hybrid Cloud Era

September 23, 2020

Many companies we talk to have on premises data lakes and use the cloud(s) to burst compute. Many are now establishing new object data lakes as well. As a result, running analytics such as Hive, Spark, Presto and machine learning are experiencing sluggish response times with data and compute in multiple locations. We also know there is an immense and growing data management burden to support these workflows.

Tags: analytics, data lake, data management, data orchestration, hybrid cloud, machine learning, performance, webinar

Optimizing Latency-Sensitive Queries for Presto at Facebook: A Collaboration Between Presto & Alluxio

IDEAS online webinar * September 13, 2020

For many latency-sensitive SQL workloads, Presto is often bound by retrieving distant data. In this talk, Rohit Jain, James Sun from Facebook and Bin Fan from Alluxio will introduce their teams’ collaboration on adding a local on-SSD Alluxio cache inside Presto workers to improve unsatisfied Presto latency.

Tag: performance