Skip to main content

Installation prerequisites

The NVIDIA AI Factory Observability (AIFO) integration uses two components: the Collector and the Translator. The Collector runs as a container on each GPU host that you want to monitor and collects metrics from that host. The Translator runs on Infrastructure Observability (IO). It either passively receives metrics from the Collector (push model) or actively polls the Collector for metrics (pull model), and then sends the data to AIFO.

AIFO supports multiple deployment options, each with its own prerequisites. For Kubernetes environments, deploy the Linux Collector with OpenTelemetry (OTel). In this push model configuration, OTel scrapes the metrics and pushes them to AIFO, so GPU hosts are discovered automatically when they start.

You can also deploy the Linux Collector with AIFO only. In this pull model, you must either manually provide the list of GPU hosts that AIFO can access or use your own automation to discover hosts.

For Windows environments, deploy the Windows Collector with AIFO. In this pull model setup, AIFO remotely scrapes data from Windows GPU hosts, similar to the Linux AIFO-only configuration.

Translator requirements

  • Infrastructure Observability 2026.4.1 or later is installed and configured.

  • You have purchased an IO Wisdom Pack license that supports the integration, and it's available for upload.

Tip

For complete Linux OS monitoring, install and configure the Linux OS integration on each GPU host.

Collector requirements

The Collector is installed as a container on the target GPU host. When it runs, it must run in privileged mode, with access to all GPUs, and with the /proc directory mounted. The required configuration is shipped in one of several forms for installation in your Kubernetes environment, for example as a Helm chart. For installation details, see the Helm chart prerequisites.

  • The nv-hostengine process is running and reachable from the container.

  • The Collector is reachable from the IO instance that collects the data.

When the Collector runs, you must provide the following information to the Collector's container using environment variables:

  • AIDC_HOSTNAME

    Specifies the hostname of the host that the Collector runs on. Ensures consistent reporting of host information in the Collector payloads.

  • NVHENGINE_ADDR

    Specifies the IP address or hostname of the host where nv-hostengine runs, so the Collector can access the nv-hostengine service.

  • NVHENGINE_PORT

    Specifies the port that nv-hostengine runs on (default 5555), so the Collector can access the nv-hostengine service.

Set the following additional container flags when you run the container:

  • --gpus all

    Provides access to all GPUs so they can be queried.

  • --privileged

    Grants permission to access process and GPU information owned by other users.

  • -v /proc:/proc2

    Mounts the host's /proc directory so process information can be resolved from the PID in the GPU.

Verify system requirements

Before you install the integration, verify that your environment meets the following requirements:

  • Infrastructure Observability (IO) 2026.4.1 or later is installed and configured.

  • A valid AIFO Wisdom Pack license that supports GPU monitoring is available.

  • A supported version of the NVIDIA Gateway (AIFO NVIDIA Gateway) is installed and configured.

  • SSH access is available to the NVIDIA Gateway.

  • The SSH user has at least read-only privileges.

  • TCP port 22, or your configured SSH port, is accessible.