Community Office Hour: Hands-on with Alluxio Structured Data Management

January 14, 2020

Bin Fan

VP of Technology

Alluxio

Gene Pang

PMC Maintainer & founding member

Alluxio

ALLUXIO COMMUNITY OFFICE HOUR

Users deploy Alluxio in a wide range of use cases from analytics to AI platforms, for Alluxio’s unified access to data and transparent caching for acceleration. However, many frameworks are SQL engines, like Presto, Apache Spark SQL, or Apache Hive, and consume data structured as tables of rows and columns. Since Alluxio is commonly used as a filesystem of files and directories, there is a mismatch between how Alluxio exposes data (files, directories), and how SQL engines deal with data (tables, rows, columns). This gap creates various challenges and inefficiencies.

Therefore, in the Alluxio 2.1 release, we introduce Alluxio Structured Data Management, which is a new set of services that enables structured data applications to interact with data more efficiently. The new services include the catalog service and a transformation service, which all work together to bridge the gap between storage and SQL engines and enable physical data independence.

In this office hour, we introduce the concepts and components of Alluxio Structured Data Management, and go through a demo with Presto.

In this Office Hour we’ll go over:

Introduction and motivation of Alluxio Structured Data Management
Overview of the different services of Alluxio Structured Data Management in Alluxio 2.1
A demo of using Alluxio Structured Data Management with Presto

ALLUXIO COMMUNITY OFFICE HOUR

In this office hour, we introduce the concepts and components of Alluxio Structured Data Management, and go through a demo with Presto.

In this Office Hour we’ll go over:

Introduction and motivation of Alluxio Structured Data Management
Overview of the different services of Alluxio Structured Data Management in Alluxio 2.1
A demo of using Alluxio Structured Data Management with Presto

Video:

Slides:

Hands-on with Alluxio Structured Data Management from Alluxio, Inc.

‍

Videos:

Presentation Slides:

Community Office Hour: Hands-on with Alluxio Structured Data Management from Alluxio, Inc.

Video:

Slides:

Hands-on with Alluxio Structured Data Management from Alluxio, Inc.

‍

Videos:

Presentation Slides:

Community Office Hour: Hands-on with Alluxio Structured Data Management from Alluxio, Inc.

Complete the form below to access the full overview:

Videos

AI/ML Infra Meetup | LLM Agents and Implementation Challenges

In this talk, Pritish Udgata from Adobe provides a comprehensive overview of implementation challenges and solutions for LLM agents.

Topic include:

CoT vs RAG vs Agentic AI
Anatomy of an agent
Single Agent with MCP
Multi Agents with A2A
Implementation Challenges and Solutions

August 14, 2025

Product Update: Alluxio AI 3.7 Now with Sub-Millisecond Latency

Watch this on-demand video to learn about the latest release of Alluxio Enterprise AI. In this webinar, discover how Alluxio AI 3.7 eliminates cloud storage latency bottlenecks with breakthrough sub-millisecond performance, delivering up to 45× faster data access than S3 Standard without changing your code. Alluxio AI 3.7 is also packed with new features designed to supercharge your AI infrastructure while keeping your data secure.Key highlights include:

Alluxio Ultra Low Latency Caching for Cloud Storage
Role-Based Access Control (RBAC) for S3 Access
5X Faster Cache Preloading with Alluxio Distributed Cache Preloader
FUSE Non-Disruptive Upgrade
Other New Features for Alluxio Admins

August 13, 2025

Optimizing Tiered Storage for Low-Latency Real-Time Analytics at AI Scale

Real-time OLAP databases are optimized for speed and often rely on tightly coupled storage-compute architectures using disks or SSDs. Decoupled architectures, which use cloud object storage, introduce an unavoidable tradeoff: cost efficiency at the expense of performance. This makes them unsuitable for databases that need to provide low-latency, real-time analytics, especially the new wave of LLM-powered dashboards, retrieval-augmented generation (RAG), and vector-embedding searches that thrive only when fresh data is milliseconds away. Can we achieve both cost efficiency and performance?

In this talk, we’ll explore the engineering challenges of extending Apache Pinot—a real-time OLAP system—onto cloud object storage while still maintaining sub-second P99 latencies.

We’ll dive into how we built an abstraction in Apache Pinot to make it agnostic to the location of data. We’ll explain how we can query data directly from the cloud (without needing to download the entire dataset, as with lazy-loading) while achieving sub-second latencies. We’ll cover the data fetch and optimization strategies we implemented, such as pipelining fetch and compute, prefetching, selective block fetches, index pinning, and more. We'll also share our latest work about integration with open table formats like iceberg, and how we will continue to achieve fast analytics directly on parquet files by implementing all the same techniques that apply to tiered storage.

‍

July 15, 2025

Sign-up for a Live Demo or Book a Meeting with a Solutions Engineer

Request a demo