How discovery works
The Collector discovers entities, relationships, and connectivity between components on each monitored host, and exposes the results on a Prometheus-compatible endpoint for collection.
Discovery cycle
Discovery of entities, relationships, and connectivity between components runs when the Collector starts, and then every 30 minutes afterward. The Collector caches the discovered data and returns it with every poll, so each polling cycle models the environment accurately. Additional discoveries run in the background and don't block subsequent calls to /metrics.
Discovery requirements
To perform discovery, the nv-hostengine process must be running on the host and reachable from the running container. The Collector uses the NVIDIA Data Center GPU Manager (DCGM) API to connect to nv-hostengine and scrape metrics and topology data. The Collector also uses the nvidia-smi command to scrape performance and configuration data.
Program discovery
Program discovery is the exception to the discovery cycle described above: it runs continuously in the background rather than on a 30-minute cycle.
Every 10 seconds, the Collector scrapes GPUs to determine which programs are running. When it detects a program running on a GPU, it caches the process ID and checks it every 10 seconds afterward to capture CPU, memory, GPU, and energy utilization. The Collector aggregates this data and reports it to Infrastructure Observability (IO) at 1-minute intervals.
To map a program back to a specific pod, container name, and program name, the Collector requires the host's /proc directory to be mounted to the container.