NVIDIA-Certified Professional: AI Operations Practice Test
NVIDIA-Certified Professional: AI Operations परीक्षा उन कौशलों को मान्य करती है जो NVIDIA अवसंरचना पर उत्पादन वातावरण में AI कार्यभार को तैनात, प्रबंधित और अनुकूलित करने के लिए आवश्यक हैं। यह परीक्षा उम्मीदवार की NVIDIA AI Enterprise सॉफ़्टवेयर स्टैक्स को कॉन्फ़िगर और बनाए रखने की क्षमता का परीक्षण करती है, जिसमें NVIDIA GPU Operator, MIG (Multi-Instance GPU) विभाजन, और Triton Inference Server का उपयोग शामिल है। यह DCGM (Data Center GPU Manager) और Prometheus जैसे उपकरणों के साथ निगरानी और अवलोकन, और AI पाइपलाइनों के लिए Kubernetes का उपयोग करके क्लस्टर ऑर्केस्ट्रेशन को कवर करती है।
नमूना प्रश्न
पूरी परीक्षा कैसी है देखने के लिए कुछ प्रश्न आज़माएं।
After installing Run:ai via BCM on a new Kubernetes cluster, the `runai list nodes` command shows all GPUs as "unallocated" even though `kubectl describe node` shows the nvidia.com/gpu resources and the nodes are Ready. Training workloads submitted with fractional GPU requests (0.5) stay pending. What must be done to make the GPUs visible and allocatable to Run:ai?
A team wants to run both a latency-sensitive real-time inference service (SLO 40 ms P99) and a high-throughput batch embedding job on the same H100 without the batch job starving the real-time service. The platform uses Kubernetes and the GPU Operator. Which resource configuration satisfies the isolation requirement with minimal waste?
An organization wants to charge back AI project teams for GPU-hours. They are running both Slurm and Run:ai on the same BCM-managed cluster. Which combination of tools provides accurate, auditable per-project GPU utilization data that can be fed into an external billing system?
A Slurm cluster managed by BCM has two partitions: "ai-training" (8 H100 nodes, QOS "high") and "ai-inference" (4 H100 nodes, QOS "normal"). A user submits a job with `#SBATCH --partition=ai-training --qos=normal --gres=gpu:4`. The job remains pending with reason "QOSGrpCpuLimit". What is the most probable cause?
After a BCM-orchestrated update of the software image on the "h100" category, several nodes report that the nvidia-fabricmanager service is in a failed state with "Failed to initialize NVSwitch". Base View marks the nodes degraded. The update included a new version of the fabric manager package. What is the correct next action?
परीक्षा ब्लूप्रिंट
प्रत्येक डोमेन वास्तविक सर्टिफिकेशन परीक्षा के अनुसार भारित है, इसलिए एक पूर्ण अभ्यास सिमुलेशन आपके परिणाम की भविष्यवाणी करता है।