NVIDIA-Certified Professional: AI Operations Practice Test
El examen NVIDIA-Certified Professional: AI Operations valida las habilidades necesarias para implementar, gestionar y optimizar cargas de trabajo de IA en infraestructuras NVIDIA en entornos de producción. Este examen evalúa la capacidad del candidato para configurar y mantener pilas de software NVIDIA AI Enterprise, incluyendo el uso de NVIDIA GPU Operator, particionamiento MIG (Multi-Instance GPU) y Triton Inference Server para la entrega escalable de modelos. Cubre la monitorización y observabilidad con herramientas como DCGM (Data Center GPU Manager) y Prometheus, así como la orquestación de clústeres utilizando Kubernetes para pipelines de IA.
Preguntas de Muestra
Prueba algunas preguntas para ver cómo es el examen completo.
After installing Run:ai via BCM on a new Kubernetes cluster, the `runai list nodes` command shows all GPUs as "unallocated" even though `kubectl describe node` shows the nvidia.com/gpu resources and the nodes are Ready. Training workloads submitted with fractional GPU requests (0.5) stay pending. What must be done to make the GPUs visible and allocatable to Run:ai?
A team wants to run both a latency-sensitive real-time inference service (SLO 40 ms P99) and a high-throughput batch embedding job on the same H100 without the batch job starving the real-time service. The platform uses Kubernetes and the GPU Operator. Which resource configuration satisfies the isolation requirement with minimal waste?
An organization wants to charge back AI project teams for GPU-hours. They are running both Slurm and Run:ai on the same BCM-managed cluster. Which combination of tools provides accurate, auditable per-project GPU utilization data that can be fed into an external billing system?
A Slurm cluster managed by BCM has two partitions: "ai-training" (8 H100 nodes, QOS "high") and "ai-inference" (4 H100 nodes, QOS "normal"). A user submits a job with `#SBATCH --partition=ai-training --qos=normal --gres=gpu:4`. The job remains pending with reason "QOSGrpCpuLimit". What is the most probable cause?
After a BCM-orchestrated update of the software image on the "h100" category, several nodes report that the nvidia-fabricmanager service is in a failed state with "Failed to initialize NVSwitch". Base View marks the nodes degraded. The update included a new version of the fabric manager package. What is the correct next action?
Plan de Estudio
Cada dominio está ponderado para coincidir con el examen de certificación real, por lo que una simulación de práctica completa predice tu resultado.