Meaning
Statistical summary structure used to estimate the value of specific percentiles in high volume data streams. The t-digest quantile calculation uses a series of centroids to summarize large amounts of data while maintaining high accuracy at the edges of the distribution. It is particularly effective for calculating p99 latencies in high volume systems.
This structure is more memory efficient than storing every individual data point.
Sketch Precision
The precision of the estimate depends on the number of clusters used to represent the data. More clusters provide better accuracy but require more memory. Most implementations allow the user to tune this parameter based on their specific needs for detail and resource usage.
Resource Constraint
Because it uses a fixed amount of memory, the algorithm is suitable for inclusion in monitoring agents and edge devices. It can process millions of values per second without consuming much CPU. This makes it a standard tool for real time performance auditing.
Metric Aggregation
Multiple summaries can be merged together to provide a global view of performance across many servers. This merging process is mathematically sound and does not lose the accuracy of the underlying quantiles. It enables a unified view of a distributed system from individual local measurements.