LogoLogo
release-1.2.0
release-1.2.0
  • Introduction
  • Basics
    • Concepts
      • Pinot storage model
      • Architecture
      • Components
        • Cluster
          • Tenant
          • Server
          • Controller
          • Broker
          • Minion
        • Table
          • Segment
            • Deep Store
            • Segment threshold
            • Segment retention
          • Schema
          • Time boundary
        • Pinot Data Explorer
    • Getting Started
      • Running Pinot locally
      • Running Pinot in Docker
      • Quick Start Examples
      • Running in Kubernetes
      • Running on public clouds
        • Running on Azure
        • Running on GCP
        • Running on AWS
      • Create and update a table configuration
      • Batch import example
      • Stream ingestion example
      • HDFS as Deep Storage
      • Troubleshooting Pinot
      • Frequently Asked Questions (FAQs)
        • General
        • Pinot On Kubernetes FAQ
        • Ingestion FAQ
        • Query FAQ
        • Operations FAQ
    • Import Data
      • From Query Console
      • Batch Ingestion
        • Spark
        • Flink
        • Hadoop
        • Backfill Data
        • Dimension table
      • Stream ingestion
        • Ingest streaming data from Apache Kafka
        • Ingest streaming data from Amazon Kinesis
        • Ingest streaming data from Apache Pulsar
        • Configure indexes
      • Stream ingestion with Upsert
      • Segment compaction on upserts
      • Stream ingestion with Dedup
      • Stream ingestion with CLP
      • File Systems
        • Amazon S3
        • Azure Data Lake Storage
        • HDFS
        • Google Cloud Storage
      • Input formats
        • Complex Type (Array, Map) Handling
        • Ingest records with dynamic schemas
      • Reload a table segment
      • Upload a table segment
    • Indexing
      • Bloom filter
      • Dictionary index
      • Forward index
      • FST index
      • Geospatial
      • Inverted index
      • JSON index
      • Native text index
      • Range index
      • Star-tree index
      • Text search support
      • Timestamp index
    • Release notes
      • 1.1.0
      • 1.0.0
      • 0.12.1
      • 0.12.0
      • 0.11.0
      • 0.10.0
      • 0.9.3
      • 0.9.2
      • 0.9.1
      • 0.9.0
      • 0.8.0
      • 0.7.1
      • 0.6.0
      • 0.5.0
      • 0.4.0
      • 0.3.0
      • 0.2.0
      • 0.1.0
    • Recipes
      • Connect to Streamlit
      • Connect to Dash
      • Visualize data with Redash
      • GitHub Events Stream
  • For Users
    • Query
      • Querying Pinot
      • Query Syntax
        • Aggregation Functions
        • Cardinality Estimation
        • Explain Plan (Single-Stage)
        • Explain Plan (Multi-Stage)
        • Filtering with IdSet
        • GapFill Function For Time-Series Dataset
        • Grouping Algorithm
        • JOINs
        • Lookup UDF Join
        • Querying JSON data
        • Transformation Functions
        • Window aggregate
        • Funnel Analysis
      • Query Options
      • Multi stage query
        • Operator Types
          • Aggregate
          • Filter
          • Join
          • Intersect
          • Leaf
          • Literal
          • Mailbox receive
          • Mailbox send
          • Minus
          • Sort or limit
          • Transform
          • Union
          • Window
        • Understanding Stages
        • Explain
        • Stats
      • User-Defined Functions (UDFs)
    • APIs
      • Broker Query API
        • Query Response Format
      • Controller Admin API
      • Controller API Reference
    • External Clients
      • JDBC
      • Java
      • Python
      • Golang
    • Tutorials
      • Use OSS as Deep Storage for Pinot
      • Ingest Parquet Files from S3 Using Spark
      • Creating Pinot Segments
      • Use S3 as Deep Storage for Pinot
      • Use S3 and Pinot in Docker
      • Batch Data Ingestion In Practice
      • Schema Evolution
  • For Developers
    • Basics
      • Extending Pinot
        • Writing Custom Aggregation Function
        • Segment Fetchers
      • Contribution Guidelines
      • Code Setup
      • Code Modules and Organization
      • Dependency Management
      • Update documentation
    • Advanced
      • Data Ingestion Overview
      • Ingestion Aggregations
      • Ingestion Transformations
      • Null value support
      • Use the multi-stage query engine (v2)
      • Troubleshoot issues with the multi-stage query engine (v2)
      • Advanced Pinot Setup
    • Plugins
      • Write Custom Plugins
        • Input Format Plugin
        • Filesystem Plugin
        • Batch Segment Fetcher Plugin
        • Stream Ingestion Plugin
    • Design Documents
      • Segment Writer API
  • For Operators
    • Deployment and Monitoring
      • Set up cluster
      • Server Startup Status Checkers
      • Set up table
      • Set up ingestion
      • Decoupling Controller from the Data Path
      • Segment Assignment
      • Instance Assignment
      • Rebalance
        • Rebalance Servers
        • Rebalance Brokers
        • Rebalance Tenant
      • Separating data storage by age
        • Using multiple tenants
        • Using multiple directories
      • Pinot managed Offline flows
      • Minion merge rollup task
      • Consistent Push and Rollback
      • Access Control
      • Monitoring
      • Tuning
        • Real-time
        • Routing
        • Query Routing using Adaptive Server Selection
        • Query Scheduling
      • Upgrading Pinot with confidence
      • Managing Logs
      • OOM Protection Using Automatic Query Killing
    • Command-Line Interface (CLI)
    • Configuration Recommendation Engine
    • Tutorials
      • Authentication
        • Basic auth access control
        • ZkBasicAuthAccessControl
      • Configuring TLS/SSL
      • Build Docker Images
      • Running Pinot in Production
      • Kubernetes Deployment
      • Amazon EKS (Kafka)
      • Amazon MSK (Kafka)
      • Monitor Pinot using Prometheus and Grafana
      • Performance Optimization Configurations
  • Configuration Reference
    • Cluster
    • Controller
    • Broker
    • Server
    • Table
    • Ingestion
    • Schema
    • Ingestion Job Spec
    • Monitoring Metrics
    • Functions
      • ABS
      • ADD
      • ago
      • EXPR_MIN / EXPR_MAX
      • arrayConcatDouble
      • arrayConcatFloat
      • arrayConcatInt
      • arrayConcatLong
      • arrayConcatString
      • arrayContainsInt
      • arrayContainsString
      • arrayDistinctInt
      • arrayDistinctString
      • arrayIndexOfInt
      • arrayIndexOfString
      • ARRAYLENGTH
      • arrayRemoveInt
      • arrayRemoveString
      • arrayReverseInt
      • arrayReverseString
      • arraySliceInt
      • arraySliceString
      • arraySortInt
      • arraySortString
      • arrayUnionInt
      • arrayUnionString
      • AVGMV
      • Base64
      • caseWhen
      • ceil
      • CHR
      • codepoint
      • concat
      • count
      • COUNTMV
      • COVAR_POP
      • COVAR_SAMP
      • day
      • dayOfWeek
      • dayOfYear
      • DISTINCT
      • DISTINCTAVG
      • DISTINCTAVGMV
      • DISTINCTCOUNT
      • DISTINCTCOUNTBITMAP
      • DISTINCTCOUNTHLLMV
      • DISTINCTCOUNTHLL
      • DISTINCTCOUNTBITMAPMV
      • DISTINCTCOUNTMV
      • DISTINCTCOUNTRAWHLL
      • DISTINCTCOUNTRAWHLLMV
      • DISTINCTCOUNTRAWTHETASKETCH
      • DISTINCTCOUNTTHETASKETCH
      • DISTINCTSUM
      • DISTINCTSUMMV
      • DIV
      • DATETIMECONVERT
      • DATETRUNC
      • exp
      • FIRSTWITHTIME
      • FLOOR
      • FrequentLongsSketch
      • FrequentStringsSketch
      • FromDateTime
      • FromEpoch
      • FromEpochBucket
      • FUNNELCOUNT
      • FunnelCompleteCount
      • FunnelMaxStep
      • FunnelMatchStep
      • Histogram
      • hour
      • isSubnetOf
      • JSONFORMAT
      • JSONPATH
      • JSONPATHARRAY
      • JSONPATHARRAYDEFAULTEMPTY
      • JSONPATHDOUBLE
      • JSONPATHLONG
      • JSONPATHSTRING
      • jsonextractkey
      • jsonextractscalar
      • LAG
      • LASTWITHTIME
      • LEAD
      • length
      • ln
      • lower
      • lpad
      • ltrim
      • max
      • MAXMV
      • MD5
      • millisecond
      • min
      • minmaxrange
      • MINMAXRANGEMV
      • MINMV
      • minute
      • MOD
      • mode
      • month
      • mult
      • now
      • percentile
      • percentileest
      • percentileestmv
      • percentilemv
      • percentiletdigest
      • percentiletdigestmv
      • percentilekll
      • percentilerawkll
      • percentilekllmv
      • percentilerawkllmv
      • quarter
      • regexpExtract
      • regexpReplace
      • remove
      • replace
      • reverse
      • round
      • ROW_NUMBER
      • rpad
      • rtrim
      • second
      • SEGMENTPARTITIONEDDISTINCTCOUNT
      • sha
      • sha256
      • sha512
      • sqrt
      • startswith
      • ST_AsBinary
      • ST_AsText
      • ST_Contains
      • ST_Distance
      • ST_GeogFromText
      • ST_GeogFromWKB
      • ST_GeometryType
      • ST_GeomFromText
      • ST_GeomFromWKB
      • STPOINT
      • ST_Polygon
      • strpos
      • ST_Union
      • SUB
      • substr
      • sum
      • summv
      • TIMECONVERT
      • timezoneHour
      • timezoneMinute
      • ToDateTime
      • ToEpoch
      • ToEpochBucket
      • ToEpochRounded
      • TOJSONMAPSTR
      • toGeometry
      • toSphericalGeography
      • trim
      • upper
      • Url
      • UTF8
      • VALUEIN
      • week
      • year
      • yearOfWeek
      • Extract
    • Plugin Reference
      • Stream Ingestion Connectors
      • VAR_POP
      • VAR_SAMP
      • STDDEV_POP
      • STDDEV_SAMP
    • Dynamic Environment
  • Reference
    • Single-stage query engine (v1)
    • Multi-stage query engine (v2)
    • Troubleshooting
      • Troubleshoot issues with the multi-stage query engine (v2)
      • Troubleshoot issues with ZooKeeper znodes
  • RESOURCES
    • Community
    • Team
    • Blogs
    • Presentations
    • Videos
  • Integrations
    • Tableau
    • Trino
    • ThirdEye
    • Superset
    • Presto
    • Spark-Pinot Connector
  • Contributing
    • Contribute Pinot documentation
    • Style guide
Powered by GitBook
On this page

Was this helpful?

Edit on GitHub
Export as PDF
  1. For Users
  2. Query
  3. Multi stage query

Stats

Learn more about multi-stage stats and how to use them to improve your queries.

Multi-stage stats are more complex but also more expressive than single-stage stats. While in single-stage stats Apache Pinot returns a single set of statistics for the query, in multi-stage stats Apache Pinot returns a set of statistics for each operator of the query execution.

These stats can be seen when using Pinot controller UI by running the query and clicking on the Show JSON format button. Then the whole JSON response will be shown and the multi-stage stats will be in a field called stageStats. Different drivers may provide different ways to see the stats.

For example the following query:

SELECT playerName, teamName
FROM baseballStats_OFFLINE as playerStats
JOIN dimBaseballTeams_OFFLINE AS teams
    ON playerStats.teamID = teams.teamID
LIMIT 10

Returns the following stageStats:

{
    "type": "MAILBOX_RECEIVE",
    "executionTimeMs": 222,
    "emittedRows": 10,
    "fanIn": 3,
    "rawMessages": 4,
    "deserializedBytes": 1688,
    "upstreamWaitMs": 651,
    "children": [
      {
        "type": "MAILBOX_SEND",
        "executionTimeMs": 210,
        "emittedRows": 10,
        "stage": 1,
        "parallelism": 3,
        "fanOut": 1,
        "rawMessages": 4,
        "serializedBytes": 338,
        "children": [
          {
            "type": "SORT_OR_LIMIT",
            "executionTimeMs": 585,
            "emittedRows": 10,
            "children": [
              {
                "type": "MAILBOX_RECEIVE",
                "executionTimeMs": 585,
                "emittedRows": 10,
                "fanIn": 3,
                "inMemoryMessages": 4,
                "rawMessages": 8,
                "deserializedBytes": 1775,
                "deserializationTimeMs": 1,
                "upstreamWaitMs": 1480,
                "children": [
                  {
                    "type": "MAILBOX_SEND",
                    "executionTimeMs": 397,
                    "emittedRows": 30,
                    "stage": 2,
                    "parallelism": 3,
                    "fanOut": 3,
                    "inMemoryMessages": 4,
                    "rawMessages": 8,
                    "serializedBytes": 1108,
                    "serializationTimeMs": 2,
                    "children": [
                      {
                        "type": "SORT_OR_LIMIT",
                        "executionTimeMs": 379,
                        "emittedRows": 30,
                        "children": [
                          {
                            "type": "TRANSFORM",
                            "executionTimeMs": 377,
                            "emittedRows": 5092,
                            "children": [
                              {
                                "type": "HASH_JOIN",
                                "executionTimeMs": 376,
                                "emittedRows": 5092,
                                "timeBuildingHashTableMs": 167,
                                "children": [
                                  {
                                    "type": "MAILBOX_RECEIVE",
                                    "executionTimeMs": 206,
                                    "emittedRows": 10000,
                                    "fanIn": 1,
                                    "inMemoryMessages": 4,
                                    "rawMessages": 21,
                                    "deserializedBytes": 649374,
                                    "deserializationTimeMs": 3,
                                    "downstreamWaitMs": 5,
                                    "upstreamWaitMs": 390,
                                    "children": [
                                      {
                                        "type": "MAILBOX_SEND",
                                        "executionTimeMs": 94,
                                        "emittedRows": 97889,
                                        "stage": 3,
                                        "parallelism": 1,
                                        "fanOut": 3,
                                        "inMemoryMessages": 4,
                                        "rawMessages": 20,
                                        "serializedBytes": 649076,
                                        "serializationTimeMs": 17,
                                        "children": [
                                          {
                                            "type": "LEAF",
                                            "table": "baseballStats_OFFLINE",
                                            "executionTimeMs": 75,
                                            "emittedRows": 97889,
                                            "numDocsScanned": 97889,
                                            "numEntriesScannedPostFilter": 195778,
                                            "numSegmentsQueried": 1,
                                            "numSegmentsProcessed": 1,
                                            "numSegmentsMatched": 1,
                                            "totalDocs": 97889,
                                            "threadCpuTimeNs": 19888000
                                          }
                                        ]
                                      }
                                    ]
                                  },
                                  {
                                    "type": "MAILBOX_RECEIVE",
                                    "executionTimeMs": 163,
                                    "emittedRows": 51,
                                    "fanIn": 1,
                                    "inMemoryMessages": 2,
                                    "rawMessages": 4,
                                    "deserializedBytes": 2330,
                                    "downstreamWaitMs": 14,
                                    "upstreamWaitMs": 162,
                                    "children": [
                                      {
                                        "type": "MAILBOX_SEND",
                                        "executionTimeMs": 17,
                                        "emittedRows": 51,
                                        "stage": 4,
                                        "parallelism": 1,
                                        "fanOut": 3,
                                        "inMemoryMessages": 1,
                                        "rawMessages": 4,
                                        "serializedBytes": 2092,
                                        "children": [
                                          {
                                            "type": "LEAF",
                                            "table": "dimBaseballTeams_OFFLINE",
                                            "executionTimeMs": 62,
                                            "emittedRows": 51,
                                            "numDocsScanned": 51,
                                            "numEntriesScannedPostFilter": 102,
                                            "numSegmentsQueried": 1,
                                            "numSegmentsProcessed": 1,
                                            "numSegmentsMatched": 1,
                                            "totalDocs": 51,
                                            "threadCpuTimeNs": 1919000,
                                            "systemActivitiesCpuTimeNs": 4677167
                                          }
                                        ]
                                      }
                                    ]
                                  }
                                ]
                              }
                            ]
                          }
                        ]
                      }
                    ]
                  }
                ]
              }
            ]
          }
        ]
      }
    ]
  }

Each node in the tree represents an operation that is executed and the tree structure form is similar (but not equal) to the logical plan of the query that can be obtained with the EXPLAIN PLAN command.

PreviousExplainNextUser-Defined Functions (UDFs)

Was this helpful?

As you can see, each operator has a type and the stats carried on the node depend on that type. You can learn more about each operator types and their stats in the section.

Operator Types