gputop

Derived metrics

Derived metrics are computed by gputop from vendor readings. They are tagged with source derived, and each is documented with its inputs, formula, assumptions and limitations. When an input is unavailable, the derived value is unavailable; gputop never fills gaps with guesses.

All windows use the fast-tier samples kept by the derive tracker (up to 1024 per GPU).

VRAM utilization and headroom

Power fraction

State

Allocation and idle-but-allocated GPUs

Unused allocated capacity (GPU-equivalents)

Utilization imbalance and stragglers

Designed for distributed workloads, where one slow GPU slows every rank.

Efficiency score

A utilization-based efficiency indicator, deliberately not “GPU utilization renamed”.

Fleet summary

Counts (busy, active, idle, unavailable, allocated, idle-allocated, throttled), sums (power, power limit, VRAM), averages (utilization, temperature, health) and maxima (temperature) over available GPUs. The health average counts unavailable GPUs as 0.