Podfit | vertical pod right-sizing — Self-hosted 2.0
PerfectScale Podfit provides comprehensive insights into the health and costs of your cluster and its components, helping you quickly pinpoint areas requiring attention along with data-driven, actionable recommendations to streamline and enhance your optimization process.

Podfit screen
Cluster overview and telemetry
The overview and telemetry section delivers a comprehensive summary of performance risks, costs, and waste insights for the selected cluster, along with identified optimization opportunities you can quickly achieve with PerfectScale. This view enables a quick evaluation of your cluster's overall health and efficiency, pinpointing configuration issues and empowering you to streamline and enhance your optimization process effectively.
Overview section

Podfit overview
Cluster selector allows for dynamic switching between clusters, enabling seamless management and monitoring of multi-cluster environment.
Tenant - the account name (PerfectScale in the example above).
Optimization Policy - displays the optimization policy of the selected cluster. Optimization Policy allows you to specify how your resources should be allocated in order to support the individual needs of your workloads. Define the policies that best suit your environment and business goals, depending on whether you want to maximize cost savings or provide extra headroom to maintain the resilience of mission-critical services.
- MaxSavings - maximum cost savings, the best for non-production environments
- Balanced (default) - optimally balances cost and resiliency
- ExtraHeadroom - the best fit for latency-sensitive environments
- MaxHeadroom - keeps the environment above the highest spikes
If a policy is set through the exporter when installing the PerfectScale Agent, it cannot be modified in the UI afterward. You can still change the policy by upgrading the exporter with the new value, or you can return it to the default by upgrading the exporter without specifying any value (this will also enable the option to change the custom time window through the UI).
Timeframe allows you to adjust the period for reviewing metrics, enabling a focused analysis for a specific time range.
Export allows you to easily download your data as a .csv file, enabling smooth analysis and effortless sharing.
Telemetry section
The telemetry section provides a comprehensive overview of aggregated data for the selected cluster, offering key insights into the cluster's health and efficiency. This helps you evaluate the cluster's performance easily and identifies opportunities for optimization, giving you a clear view of its overall status.

Podfit telemetry section
Current Risks shows the total risks identified within the cluster for the selected period. This value is dynamic and updates based on the filters applied in the workload table.
Unused Resources provides insights into the resources within the cluster that are not being effectively utilized:
- Pod Waste displays the total cost of wasted resources within the cluster. Clicking on this metric will direct you to the workload waste report, offering a detailed visual breakdown of the workloads contributing to the waste. This allows you to quickly identify the most impactful areas requiring attention, enabling more efficient optimization.
- Node Idle indicates the total cost of unutilized node space. By clicking on this metric, you'll be navigated to a comprehensive view of the cluster at the infrastructure level. This view provides valuable insights into the behavior of different node groups and types, enabling you to optimize the underlying infrastructure for your workloads effectively.
Cost & Expected Optimized Cost is a powerful widget that offers insights into the total costs incurred compared to the actual resource utilization. This information helps you evaluate whether the cluster is well-balanced, over-provisioned, or under-provisioned. Additionally, the widget provides a Recommended Cost, reflecting the potential savings achievable through PerfectScale's recommendations, ensuring your cluster operates efficiently and cost-effectively. Clicking on this metric will direct you to the cluster cost report for further investigation.
Negative savings indicate an under-provisioned environment.
CPU/Memory Utilization Over Time provides a comprehensive visual representation of resource allocation, requests, and usage trends within your cluster. Tracking these metrics over a specified timeframe allows you to analyze historical data to understand how resource dynamics have changed, compare actual usage with allocated and requested resources, and identify utilization patterns.
- Used - p99 of utilization
- Requested - p99 of combined requests of all the workloads
- Allocated - p99 of available cluster compute
Workloads table
The Workload table provides a detailed overview of all the workloads running in your cluster. Each row represents a specific workload and its containers, including critical metrics like cost, waste, and potential cost increase due to under-provisioned resources. This view will help you quickly identify workloads that are misaligned with resource demands, highlighting optimization opportunities and areas at risk that require attention. With dynamic filtering and sorting options, you can easily focus on specific namespaces, labels, or workloads, making it easier to prioritize optimization tasks and run clusters efficiently.
Workloads are sets of pods of a
Deployment,StatefulSet,DaemonSet,Job,or custom resource CRD (for example -Runner,SparkJob,etc)
Hover over the column name to view hints.

Workloads table
Filtering resiliency issues
Status indicates workloads at risk. Workloads could be easily filtered by the resiliency risk level or particular indicator.
Risk indicators are dynamic, i.e., the presence of OOM indicator in the list means that at least one workload experienced an out-of-memory event in a given timeframe.
The dot count is a visual indicator of risk levels, with three levels: Low, Medium, and High (three dots represent the High-risk level).
Hollow dots indicate a muted workload, while shaded dots indicate the presence of a workflow ticket in progress.

Status
Automation status
This column shows the current automation status of each workload. You can quickly filter the data by automation status, prioritizing and focusing on the most relevant workloads for further investigation.
Multiselect is available.

Automation status
| Status | I | Description |
|---|---|---|
| Active | ![]() | Once the configuration is completed, automation will be indicated as successfully enabled. |
| Limited by Rule | ![]() | Indicates that automation is intentionally restricted by a specific rule. Learn more about resource allocation constraints here. |
| Delayed | ![]() | If the defined CRD maintenance window causes time constraints, the execution of recommendations will be postponed. |
| Disabled | ![]() | The merged CRD will disable automation for the workload. For example, if the cluster-level configuration enables automation while the namespace-level configuration disables it, the namespace-level configuration takes precedence, resulting in disabled automation for the particular workloads within the cluster. |
| Stopped | ![]() | PerfectScale will forcibly stop the automation. For example, to prevent your environment from recursive resource increases, such as those resulting from memory leaks. |
Type
This column identifies the workload type (e.g., Deployment, StatefulSet). You can use filtering, sorting, and multi-select options to tailor the data display, making focusing on specific workload types easier.
Namespace
The namespace column shows the namespace of each workload. You can apply filtering, sorting, and multi-select options to customize the data display, allowing you to focus on specific namespaces.
If PerfectScale does not detect any workload in the Namespaces for 7 consecutive days, those Namespaces will be consolidated into a separate Namespace __deleted-namespaces__.
Running Hours
The workload running hours column indicates the total duration each workload, including its replicas, has been actively running in the cluster during the selected period. You can use the sorting option to arrange the data in your preferred order.
Cost/h
The workload cost per hour column indicates the total hourly expense of the workload. You can use the sorting option to arrange the data in your preferred order.
Cost
The workload total cost column shows the total expense of the workload for the selected period, considering both its hourly cost and the duration it has been actively running. You can easily identify the most costly workloads in the cluster with a single click using the sorting option.
Waste
The workload waste column indicates the historical waste caused by over-provisioned resources allocated to a workload. You can easily identify the most wasteful workloads in the cluster with a single click using the sorting option.
Savings Opportunity
The savings opportunity displays the expected total savings achievable by applying data-driven workload right-sizing recommendations.
Risk Mitigation
The risk mitigation column shows the projected rise in workload cost based on PerfectScale’s recommendations, indicating that the workload is under-provisioned. This helps you predict the cost adjustments required to maintain cluster stability.
Container
The container column lists the containers associated with each workload. You can use filtering options to display the data for a specific container(s). Multi-select is available.
View Customization
Easily jump between Recommendations, Labels and Policies, and HPA views using the switcher above the table.

Podfit view customization
Recommendations view
The Recommendations Table offers clear insights into necessary workload resource adjustments to maintain the cluster's stability and cost efficiency. To access more information, click on the workload. This will open up a Zoom-in window that provides a comprehensive breakdown.
| Name | Description |
|---|---|
CPU Request | PerfectScale guidelines for CPU Request. |
CPU Limit | PerfectScale guidelines for CPU Limit. |
Memory Request | PerfectScale guidelines for Memory Request. |
Memory Limit | PerfectScale guidelines for Memory Limit. |
If one or more resources have reached their CRD-defined size constraints, the recommendations will not be executed. In this case, the Limited by Rule indicator, along with an explanatory tooltip, will be displayed near the recommendations.

Learn more about resource allocation constraints here.
To customize your recommendations view, use the Resource Change View drop-down menu.

Recommendations formats
- Detailed - to display the changes made to resources (shows both the previous and new values).
- Total Impact in Units - to display changes made to resources as an absolute number, factoring in replica count.
- Single Instance Impact in Units - to display changes made to resources as an absolute number.
- Single Instance Impact in % - to display changes made to resources in a percentage format.
When the recommendation view is set to Total Impact in Units, the resource change impact summary is available. This view provides a clear understanding of the effect of total resource adjustments, enabling seamless evaluation of the optimization process.

Resource change summary




