1 of 5

File Systems

This section contains a collection of short guides to show you how to import from a Pinot supported file system.

FileSystem is an abstraction provided by Pinot to access data in distributed file systems (DFS).

Pinot uses distributed file systems for the following purposes:

Batch Ingestion Job - To read the input data (CSV, Avro, Thrift, etc.) and to write generated segments to DFS

Amazon S3

You can enable Amazon S3 Filesystem backend by including the plugin pinot-s3 .

By default Pinot loads all the plugins, so you can just drop this plugin there. Also, if you specify -Dplugins.include, you need to put all the plugins you want to use, e.g. pinot-json, pinot-avro , pinot-kafka-2.0...

You can also configure the S3 filesystem using the following options:

Each of these properties should be prefixed by pinot.[node].storage.factory.s3. where node is either controller or server depending on the config

e.g.

S3 Filesystem supports authentication using the . The credential provider looks for the credentials in the following order -

Environment Variables - AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY (RECOMMENDED since they are recognized by all the AWS SDKs and CLI except for .NET), or AWS_ACCESS_KEY and AWS_SECRET_KEY (only recognized by Java SDK)
Java System Properties - aws.accessKeyId and aws.secretKey

You can also specify the accessKey and secretKey using the properties. However, this method is not secure and should be used only for POC setups.

Examples

Job spec

Controller config

Server config

Minion config

Azure Data Lake Storage

This guide shows you how to import data from files stored in Azure Data Lake Storage Gen2 (ADLS Gen2)

You can enable the Azure Data Lake Storage using the plugin pinot-adls. In the controller or server, add the config -

Azure Blob Storage provides the following options -

accountName : Name of the azure account under which the storage is created
accessKey : access key required for the authentication
fileSystemName

Each of these properties should be prefixed by pinot.[node].storage.factory.class.adl2. where node is either controller or server depending on the config

e.g.

Examples

Job spec

Controller config

Server config

Minion config

HDFS

This guide shows you how to import data from HDFS.

You can enable the Hadoop DFS using the plugin pinot-hdfs. In the controller or server, add the config:

HDFS implementation provides the following options -

hadoop.conf.path : Absolute path of the directory containing hadoop XML configuration files such as hdfs-site.xml, core-site.xml .
hadoop.write.checksum : create checksum while pushing an object. Default is false

Each of these properties should be prefixed by pinot.[node].storage.factory.class.hdfs. where node is either controller or server depending on the config

The kerberos configs should be used only if your Hadoop installation is secured with Kerberos. Please check on how to generate Kerberos security identification.

You will also need to provide proper Hadoop dependencies jars from your Hadoop installation to your Pinot startup scripts.

Push HDFS segment to Pinot Controller

To push HDFS segment files to Pinot controller, you just need to ensure you have proper Hadoop configuration as we mentioned in the previous part. Then your remote segment creation/push job can send the HDFS path of your newly created segment files to the Pinot Controller and let it download the files.

For example, the following curl requests to Controller will notify it to download segment files to the proper table:

Examples

Job spec

Standalone Job:

Hadoop Job:

Controller config

Server config

Minion config

Google Cloud Storage

This guide shows you how to import data from GCP (Google Cloud Platform).

You can enable the using the plugin pinot-gcs. In the controller or server, add the config -

Azure Data Lake Storage

This guide shows you how to import data from files stored in Azure Data Lake Storage Gen2 (ADLS Gen2)

You can enable the Azure Data Lake Storage using the plugin pinot-adls. In the controller or server, add the config -

-Dplugins.dir=/opt/pinot/plugins -Dplugins.include=pinot-adls

Azure Blob Storage provides the following options -

accountName : Name of the azure account under which the storage is created
accessKey : access key required for the authentication
fileSystemName

Each of these properties should be prefixed by pinot.[node].storage.factory.class.adl2. where node is either controller or server depending on the config

e.g.

Examples

Job spec

Controller config

Server config

Minion config

Amazon S3

You can enable Amazon S3 Filesystem backend by including the plugin pinot-s3 .

-Dplugins.dir=/opt/pinot/plugins -Dplugins.include=pinot-s3

You can also configure the S3 filesystem using the following options:

Each of these properties should be prefixed by pinot.[node].storage.factory.s3. where node is either controller or server depending on the config

e.g.

S3 Filesystem supports authentication using the . The credential provider looks for the credentials in the following order -

Environment Variables - AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY (RECOMMENDED since they are recognized by all the AWS SDKs and CLI except for .NET), or AWS_ACCESS_KEY and AWS_SECRET_KEY (only recognized by Java SDK)
Java System Properties - aws.accessKeyId and aws.secretKey

You can also specify the accessKey and secretKey using the properties. However, this method is not secure and should be used only for POC setups.

Examples

Job spec

Controller config

Server config

Minion config

HDFS

This guide shows you how to import data from HDFS.

You can enable the Hadoop DFS using the plugin pinot-hdfs. In the controller or server, add the config:

-Dplugins.dir=/opt/pinot/plugins -Dplugins.include=pinot-hdfs

HDFS implementation provides the following options -

hadoop.conf.path : Absolute path of the directory containing hadoop XML configuration files such as hdfs-site.xml, core-site.xml .
hadoop.write.checksum : create checksum while pushing an object. Default is false

Each of these properties should be prefixed by pinot.[node].storage.factory.class.hdfs. where node is either controller or server depending on the config

The kerberos configs should be used only if your Hadoop installation is secured with Kerberos. Please check on how to generate Kerberos security identification.

You will also need to provide proper Hadoop dependencies jars from your Hadoop installation to your Pinot startup scripts.

Push HDFS segment to Pinot Controller

For example, the following curl requests to Controller will notify it to download segment files to the proper table:

File Systems

Amazon S3

hashtagExamples

hashtagJob spec

hashtagController config

hashtagServer config

hashtagMinion config

Azure Data Lake Storage

hashtagExamples

hashtagJob spec

hashtagController config

hashtagServer config

hashtagMinion config

HDFS

hashtagPush HDFS segment to Pinot Controller

hashtagExamples

hashtagJob spec

hashtagController config

hashtagServer config

hashtagMinion config

Google Cloud Storage

Azure Data Lake Storage

hashtagExamples

hashtagJob spec

hashtagController config

hashtagServer config

hashtagMinion config

Amazon S3

hashtagExamples

hashtagJob spec

hashtagController config

hashtagServer config

hashtagMinion config

HDFS

hashtagPush HDFS segment to Pinot Controller

hashtagExamples

hashtagJob spec

hashtagController config

hashtagServer config

hashtagMinion config

Google Cloud Storage

File Systems

hashtagExamples

hashtagJob spec

hashtagController config

hashtagServer config

hashtagMinion config

hashtagSupported File Systems

hashtagEnabling a File System

Examples

Job spec

Controller config

Server config

Minion config

Examples

Job spec

Controller config

Server config

Minion config

Push HDFS segment to Pinot Controller

Examples

Job spec

Controller config

Server config

Minion config

Examples

Job spec

Controller config

Server config

Minion config

Examples

Job spec

Controller config

Server config

Minion config

Push HDFS segment to Pinot Controller

Examples

Job spec

Controller config

Server config

Minion config

Examples

Job spec

Controller config

Server config

Minion config

Supported File Systems

Enabling a File System