> For the complete documentation index, see [llms.txt](https://docs.perfectscale.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.perfectscale.io/visibility-and-optimization/gpu-optimization.md).

# GPU optimization

Reduce cloud GPU costs with real-time utilization visibility and insights

{% hint style="info" %}
PerfectScale now only supports NVIDIA Data Center GPU Manager (DCGM).
{% endhint %}

{% hint style="info" %}
**GPU** support is available starting with the **exporter version 1.0.55**.\
**GPU memory** support is available starting with the **exporter version 1.1.11**.
{% endhint %}

PerfectScale delivers exceptional **GPU** and **GPU memory** utilization visibility to monitor and optimize GPU resources within your Kubernetes clusters. This feature helps teams identify optimization opportunities, reduce resource waste, and improve overall K8s efficiency.

{% hint style="warning" %}
To enable GPU visibility support in PerfectScale, the **NVIDIA DCGM exporter** should be installed. Additionally, specific configuration parameters should be set when deploying or upgrading the PerfectScale agent. Learn more [here](/getting-started/how-to-onboard-a-cluster.md#gpu-support).
{% endhint %}

When PerfectScale detects active GPU resources within a cluster, it automatically enables GPU and GPU memory widgets and utilization insights in the UI.&#x20;

<figure><img src="https://1573387604-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FABMqnYtsOO44JmQTVSnn%2Fuploads%2FLVBFalZhcQo2t7p5WIq8%2FGPU%20widgets.png?alt=media&amp;token=96a86104-88f0-4969-b3ec-ad1c64ebc981" alt=""><figcaption><p>GPU utilization widgets</p></figcaption></figure>

This view provides detailed GPU usage and allocation efficiency metrics, helping you identify savings opportunities and drive further data-driven Kubernetes optimization.

## PodFit GPU visibility

<figure><img src="https://1573387604-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FABMqnYtsOO44JmQTVSnn%2Fuploads%2FKczemX17MUjUZFtxN3dX%2FGPU%20switch%20(1).png?alt=media&amp;token=80f4b245-1820-47ed-a58e-b79bd1e7d7b4" alt=""><figcaption><p>GPU view</p></figcaption></figure>

To quickly identify GPU-allocated workloads in **PodFit**, switch to the **GPU view** by clicking the GPU tab from the view selector, as shown below. Once there, you will see the GPU and GPU memory utilization metrics per container, as well as the GPU scheduler type, for example, KAI, and the sharing type, for example, TimeSlicing. Sort the table by GPU usage to bring all GPU-consuming workloads to the top.

Clicking on a specific workload opens the detailed workload view - Zoom-in window. This panel provides in-depth information about the workload’s current state and behavior, along with historical data on resource allocation and utilization over time. It includes GPU and GPU memory utilization metrics as well as other key workload metrics. Learn more about zoom-in capabilities [here](/visibility-and-optimization/podfit-vertical-pod-right-sizing.md#detailed-workload-analysis).

<figure><img src="https://1573387604-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FABMqnYtsOO44JmQTVSnn%2Fuploads%2FaRMF1o6LcCNE828w7wI3%2FGroup%201000002511.png?alt=media&amp;token=92490cd9-3176-4441-876a-32ac154cb6b0" alt=""><figcaption><p>Workload GPU utilization widgets</p></figcaption></figure>

## NodeFit GPU visibility

To see detailed GPU usage across your infrastructure, go to **NodeFit**. The GPU chart shows how much of your GPUs are being used versus how much was requested, making it easy to spot inefficiencies and find ways to optimize.

<figure><img src="https://1573387604-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FABMqnYtsOO44JmQTVSnn%2Fuploads%2FO92jtG2OCdt2Dg250Iud%2FInfrafit%20GPU.png?alt=media&amp;token=64c8852c-4b9d-4fc3-9ab2-1c6eac849fc9" alt=""><figcaption><p>Node group GPU utilization</p></figcaption></figure>

This view helps you quickly evaluate the difference between requested GPU and GPU memory resources and actual usage, making it easy to pinpoint underutilized or idle GPU capacity across your clusters.&#x20;

By clicking on the specific node group, you will get a granular breakdown of individual instances within that group, along with key metrics for each one.

<figure><img src="https://1573387604-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FABMqnYtsOO44JmQTVSnn%2Fuploads%2F6s9xXQcxCcQKqs8OPTIb%2FNode%20type%20GPU.png?alt=media&amp;token=dbd8d615-7663-4efc-a7b9-4e655519ac3c" alt=""><figcaption><p>Node type GPU utilization</p></figcaption></figure>

Clicking on a specific instance will display a list of workloads running on that machine, allowing for deeper investigation and analysis.
