Installation prerequisites
The NVIDIA AI Factory Observability (AIFO) integration uses two components: the Collector and the Translator. The Collector runs as a container on each GPU host that you want to monitor and collects metrics from that host. The Translator runs on Infrastructure Observability (IO). It either passively receives metrics from the Collector (push model) or actively polls the Collector for metrics (pull model), and then sends the data to AIFO.
AIFO supports multiple deployment options, each with its own prerequisites. For Kubernetes environments, deploy the Linux Collector with OpenTelemetry (OTel). In this push model configuration, OTel scrapes the metrics and pushes them to AIFO, so GPU hosts are discovered automatically when they start.
You can also deploy the Linux Collector with AIFO only. In this pull model, you must either manually provide the list of GPU hosts that AIFO can access or use your own automation to discover hosts.
For Windows environments, deploy the Windows Collector with AIFO. In this pull model setup, AIFO remotely scrapes data from Windows GPU hosts, similar to the Linux AIFO-only configuration.
Translator requirements
Infrastructure Observability 2026.4.1 or later is installed and configured.
You have purchased an IO Wisdom Pack license that supports the integration, and it's available for upload.
Tip
For complete Linux OS monitoring, install and configure the Linux OS integration on each GPU host.
Collector requirements
The Collector is installed as a container on the target GPU host. When it runs, it must run in privileged mode, with access to all GPUs, and with the /proc directory mounted. The required configuration is shipped in one of several forms for installation in your Kubernetes environment, for example as a Helm chart. For installation details, see the Helm chart prerequisites.
The nv-hostengine process is running and reachable from the container.
The Collector is reachable from the IO instance that collects the data.
When the Collector runs, you must provide the following information to the Collector's container using environment variables:
AIDC_HOSTNAMESpecifies the hostname of the host that the Collector runs on. Ensures consistent reporting of host information in the Collector payloads.
NVHENGINE_ADDRSpecifies the IP address or hostname of the host where nv-hostengine runs, so the Collector can access the nv-hostengine service.
NVHENGINE_PORTSpecifies the port that nv-hostengine runs on (default
5555), so the Collector can access the nv-hostengine service.
Set the following additional container flags when you run the container:
--gpus allProvides access to all GPUs so they can be queried.
--privilegedGrants permission to access process and GPU information owned by other users.
-v /proc:/proc2Mounts the host's
/procdirectory so process information can be resolved from the PID in the GPU.
Verify system requirements
Before you install the integration, verify that your environment meets the following requirements:
Infrastructure Observability (IO) 2026.4.1 or later is installed and configured.
A valid AIFO Wisdom Pack license that supports GPU monitoring is available.
A supported version of the NVIDIA Gateway (AIFO NVIDIA Gateway) is installed and configured.
SSH access is available to the NVIDIA Gateway.
The SSH user has at least read-only privileges.
TCP port 22, or your configured SSH port, is accessible.