GPU missing from nvidia-smi, GPU fallen off bus, driver error, Xid/SXid/NVRM, PCIe or enumeration issue | Standard manual collection. Include affected GPU index/UUID/serial, Xid/SXid if known, and the failure timestamp. |
| GPU ECC, SRAM/DRAM, row remapper, page retirement, memory error, or suspected GPU memory fault | Standard manual collection. Include affected GPU index/UUID/serial, observed ECC or row-remapper alert, and workload impact. |
| NVLink, NVSwitch, Fabric Manager, IMEX, NVLSM, OpenSM, Fabric state, training, Xid 145, or Xid 149 issue | Standard manual collection from each affected compute host. For NVL systems, include switch-side NVOS tech-support or NVSwitch logs when available or requested. |
| Thermal, fan, power, throttling, PSU, or voltage issue | Standard manual collection. Include the NVDebug Redfish/BMC bundle when BMC access is available, plus facility alarms, sensor exports, or photos/screenshots when useful. |
| Unexpected reboot, node power loss, power cycle, or suspected BMC event | Standard manual collection after the host is back online. Include the NVDebug Redfish/BMC bundle when BMC access is available. Include exact reboot or outage time if known. |
| NVMe, disk I/O, filesystem, Linux MD RAID, or hardware RAID issue | Standard manual collection. Include the affected device path, mount point, or volume if known. Nscale Support may request storage-specific command output after initial triage. |
| System memory, CPU, MCE, EDAC, RAS, OOM, or host-level thermal issue | Standard manual collection. Include any visible machine-check, EDAC, OOM, DIMM slot, CPU socket, or sensor alert details. |
| Ethernet switch device issue | Device logs, CLI command output, interface state and error counters, environmental or hardware alerts, and screenshots when available. Include switch asset, site, component, current state, and impact. |
| Ethernet link issue | Endpoint A and Port A, Endpoint B and Port B if known, plus command output for both sides where available: net show interface <swp interface number> detail, net show interface pluggables, and sudo l1-show <swp interface number>. Include link state, counters, cleaning or optics replacement attempts, and impact. |
| InfiniBand switch device issue | Device logs, UFM or fabric health evidence, port state and error counters, environmental or hardware alerts, and screenshots when available. Include switch asset, site, component, current state, and impact. |
| InfiniBand link issue | Endpoint A and Port A, Endpoint B and Port B if known, UFM evidence, and mlxlink -d lid <lid> -p <physical_port_no> -c output. Note whether the port was isolated in UFM and whether perfquery -x <lid> <logical_port> was run as part of your normal process. |
| Rack issue | Rack ID, affected rack unit/side/bay/PDU, photos or screenshots, related device logs, power or environmental evidence, and any impacted host logs. |
| Facility, cooling, rPDU, or widespread power issue | Site, affected racks/rows/clusters/devices, start time, current state, alarm screenshots, rPDU or facility exports, photos when useful, and sample host logs from impacted nodes. |