Databricks Certified Machine Learning Associate Practice Test
Databricks Certified Machine Learning Associate के लिए अपना आत्मविश्वास बढ़ाएँ। अवधारणाओं का अभ्यास करें, उत्तरों को समझें और हर सवाल के साथ अपना ज्ञान मज़बूत करें।
एक नमूना सवाल आज़माएँपरीक्षा का परिचय और विवरण
The Databricks Certified Machine Learning Associate certification validates foundational expertise in implementing machine learning workflows on the Databricks Lakehouse Platform. This professional credential demonstrates a candidate's ability to use Databricks' integrated tools-including MLflow, AutoML, and Spark MLlib-to build, track, register, and deploy models at scale. It covers the complete ML lifecycle from feature engineering and model training to evaluation, responsible AI practices, and production deployment. Earning this certification signals to employers a verified, practical skill set in leveraging a unified platform for data and AI, which is critical for developing efficient, reproducible, and collaborative machine learning solutions in enterprise environments. It is designed for data scientists and ML engineers beginning their journey with Databricks, establishing a solid foundation for advanced specialization.
नमूना प्रश्न
अभ्यास कैसे काम करता है, यह जानने के लिए एक उत्तर चुनें और व्याख्या देखें।
A team runs Databricks AutoML for a loan-default model using bureau attributes, application channel, income bands, prior delinquency counts, and a late-arriving label. The review requires editable preprocessing code, trial metrics, and a defensible reason for the selected estimator before any registry promotion. Which artifact should they inspect first? In this scenario, the team is comparing the result against a simple baseline from last quarter. Select the best answer.
A reviewer compares 60 MLflow runs for a claims severity model. The required champion is the run with the lowest validation log loss among runs tagged candidate=true. Which approach is most direct? In this scenario, the source Delta table is shared by two downstream model teams. Select the best answer.
A team uses Databricks ML Runtime for a demand forecast instead of a standard runtime. The administrator asks what practical advantage justifies that choice for basic ML work. What is the best answer? In this scenario, the data includes one late-arriving column that is excluded from training. Select the best answer.
After validating a challenger for a card-fraud classifier, the MLOps owner wants production code to keep loading the logical champion without changing the model URI. What registry action is best? In this scenario, the pipeline must be rerunnable after a feature backfill. Select the best answer.
A team runs Databricks AutoML for a patient no-show model using appointment history, clinic, reminder channel, travel distance, and several missing demographics. The review requires editable preprocessing code, trial metrics, and a defensible reason for the selected estimator before any registry promotion. Which artifact should they inspect first? In this scenario, the workspace uses Unity Catalog grants rather than notebook-only access controls. Select the best answer.
करियर के अवसर और वेतन
जब तक लोकल रेंज न दिखे, ये US मार्केट के आंकड़े हैं।
इस परीक्षा में क्या शामिल है
पढ़ाई की योजना बनाने के लिए प्रकाशित डोमेन भार का उपयोग करें। अभ्यास के परिणाम आपके सर्टिफिकेशन परीक्षा के स्कोर का अनुमान नहीं हैं।
01Databricks Machine Learning
This domain covers the operational aspects of machine learning on Databricks, focusing on MLOps strategies, the advantages of using ML runtimes, and how AutoML facilitates model and feature selection processes for efficient production workflows.
विषय
- Identify the best practices of an MLOps strategy
- Identify the advantages of using ML runtimes
- Identify how AutoML facilitates model/feature selection.
- Identify the advantages AutoML brings to the model development process
- Identify the benefits of creating feature store tables at the account level in Unity Catalog in Databricks vs at the workspace level
- Create a feature store table in Unity Catalog
- Write data to a feature store table
- Train a model with features from a feature store table.
- Score a model using features from a feature store table.
- Describe the differences between online and offline feature tables
- Identify the best run using the MLflow Client API.
- Manually log metrics, artifacts, and models in an MLflow Run.
- Identify information available in the MLFlow UI
- Register a model using the MLflow Client API in the Unity Catalog registry
- Identify benefits of registering models in the Unity Catalog registry over the workspace registry
- Identify scenarios where promoting code is preferred over promoting models and vice versa
- Set or remove a tag for a model
- Promote a challenger model to a champion model using aliases
सीखने के उद्देश्य
- Identify the best practices of an MLOps strategy
- Identify the advantages of using ML runtimes
- Identify how AutoML facilitates model/feature selection.
- Identify the advantages AutoML brings to the model development process
- Identify the benefits of creating feature store tables at the account level in Unity Catalog in Databricks vs at the workspace level
- Create a feature store table in Unity Catalog
- Write data to a feature store table
- Train a model with features from a feature store table.
- Score a model using features from a feature store table.
- Describe the differences between online and offline feature tables
- Identify the best run using the MLflow Client API.
- Manually log metrics, artifacts, and models in an MLflow Run.
- Identify information available in the MLFlow UI
- Register a model using the MLflow Client API in the Unity Catalog registry
- Identify benefits of registering models in the Unity Catalog registry over the workspace registry
- Identify scenarios where promoting code is preferred over promoting models and vice versa
- Set or remove a tag for a model
- Promote a challenger model to a champion model using aliases
02Model Development
This domain addresses the technical process of developing machine learning models, including selecting the appropriate algorithm for specific scenarios, identifying methods to mitigate data imbalance in training sets, and comparing estimators and transformers.
विषय
- Use ML foundations to select the appropriate algorithm for a given model scenario
- Identify methods to mitigate data imbalance in training data
- Compare estimators and transformers
- Develop a training pipeline
- Use Hyperopt's fmin operation to tune a model's hyperparameters
- Perform random or grid search or Bayesian search as a method for tuning hyperparameters.
- Parallelize single node models for hyperparameter tuning
- Describe the benefits and downsides of using cross-validation over a train-validation split.
- Perform cross-validation as a part of model fitting.
- Identify the number of models being trained in conjunction with a grid-search and cross-validation process.
- Use common classification metrics: F1, Log Loss, ROC/AUC, etc
- Use common regression metrics: RMSE, MAE, R-squared, etc.
- Choose the most appropriate metric for a given scenario objective
- Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions
- Assess the impact of model complexity and the bias variance tradeoff on model performance
सीखने के उद्देश्य
- Use ML foundations to select the appropriate algorithm for a given model scenario
- Identify methods to mitigate data imbalance in training data
- Compare estimators and transformers
- Develop a training pipeline
- Use Hyperopt's fmin operation to tune a model's hyperparameters
- Perform random or grid search or Bayesian search as a method for tuning hyperparameters.
- Parallelize single node models for hyperparameter tuning
- Describe the benefits and downsides of using cross-validation over a train-validation split.
- Perform cross-validation as a part of model fitting.
- Identify the number of models being trained in conjunction with a grid-search and cross-validation process.
- Use common classification metrics: F1, Log Loss, ROC/AUC, etc
- Use common regression metrics: RMSE, MAE, R-squared, etc.
- Choose the most appropriate metric for a given scenario objective
- Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions
- Assess the impact of model complexity and the bias variance tradeoff on model performance
03Data Processing
This domain focuses on data preparation tasks using Spark DataFrames, including computing summary statistics with dbutils, removing outliers based on standard deviation or IQR, and creating effective visualizations for both categorical and continuous features.
विषय
- Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries
- Remove outliers from a Spark DataFrame based on standard deviation or IQR
- Create visualizations for categorical or continuous features
- Compare two categorical or two continuous features using the appropriate method
- Compare and contrast imputing missing values with the mean or median or mode value
- Impute missing values with the mode, mean, or median value
- Use one-hot encoding for categorical features
- Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate.
- Identify scenarios where log scale transformation is appropriate
सीखने के उद्देश्य
- Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries
- Remove outliers from a Spark DataFrame based on standard deviation or IQR
- Create visualizations for categorical or continuous features
- Compare two categorical or two continuous features using the appropriate method
- Compare and contrast imputing missing values with the mean or median or mode value
- Impute missing values with the mode, mean, or median value
- Use one-hot encoding for categorical features
- Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate.
- Identify scenarios where log scale transformation is appropriate
04Model Deployment
This domain covers the deployment of models into production environments using various serving patterns, including batch, realtime, and streaming methods, while also focusing on deploying custom models to endpoints and performing batch inference using pandas.
विषय
- Identify the differences and advantages of model serving approaches: batch, realtime, and streaming
- Deploy a custom model to a model endpoint
- Use pandas to perform batch inference
- Identify how streaming inference is performed with Delta Live Tables
- Deploy and query a model for realtime inference
- Split data between endpoints for realtime interference
सीखने के उद्देश्य
- Identify the differences and advantages of model serving approaches: batch, realtime, and streaming
- Deploy a custom model to a model endpoint
- Use pandas to perform batch inference
- Identify how streaming inference is performed with Delta Live Tables
- Deploy and query a model for realtime inference
- Split data between endpoints for realtime interference