Databricks Certified Machine Learning Associate Practice Test

100 preguntas disponibles

Gana confianza para Databricks Certified Machine Learning Associate. Practica los conceptos, comprende las respuestas y refuerza tus conocimientos pregunta a pregunta.

Probar una pregunta
Prueba 5 preguntas gratis
No necesitas cuenta. Una cuenta gratuita incluye 20 preguntas de este examen.
Examen de certificación
1 hora 30 minutos Límite de Tiempo
Associate Nivel
Tu práctica
100 Preguntas de práctica
1 hora 40 minutos Tiempo de Práctica
Prueba 5 preguntas gratis
No necesitas cuenta. Una cuenta gratuita incluye 20 preguntas de este examen.
El banco 100 Preguntas de práctica verificadas con los objetivos oficiales.
Explora los temas del examen
Databricks100 preguntas de práctica
Temario verificadoVerificado con Databricks official objectivesMetadatos verificados 2026-09-17Cómo verificamos

Descripción y detalles del examen

The Databricks Certified Machine Learning Associate certification validates foundational expertise in implementing machine learning workflows on the Databricks Lakehouse Platform. This professional credential demonstrates a candidate's ability to use Databricks' integrated tools-including MLflow, AutoML, and Spark MLlib-to build, track, register, and deploy models at scale. It covers the complete ML lifecycle from feature engineering and model training to evaluation, responsible AI practices, and production deployment. Earning this certification signals to employers a verified, practical skill set in leveraging a unified platform for data and AI, which is critical for developing efficient, reproducible, and collaborative machine learning solutions in enterprise environments. It is designed for data scientists and ML engineers beginning their journey with Databricks, establishing a solid foundation for advanced specialization.

Preguntas de Muestra

Elige una respuesta y consulta la explicación para ver cómo funciona la práctica.

Databricks Machine Learning

A team runs Databricks AutoML for a loan-default model using bureau attributes, application channel, income bands, prior delinquency counts, and a late-arriving label. The review requires editable preprocessing code, trial metrics, and a defensible reason for the selected estimator before any registry promotion. Which artifact should they inspect first? In this scenario, the team is comparing the result against a simple baseline from last quarter. Select the best answer.

Databricks Machine Learning

A reviewer compares 60 MLflow runs for a claims severity model. The required champion is the run with the lowest validation log loss among runs tagged candidate=true. Which approach is most direct? In this scenario, the source Delta table is shared by two downstream model teams. Select the best answer.

Databricks Machine Learning

A team uses Databricks ML Runtime for a demand forecast instead of a standard runtime. The administrator asks what practical advantage justifies that choice for basic ML work. What is the best answer? In this scenario, the data includes one late-arriving column that is excluded from training. Select the best answer.

Databricks Machine Learning

After validating a challenger for a card-fraud classifier, the MLOps owner wants production code to keep loading the logical champion without changing the model URI. What registry action is best? In this scenario, the pipeline must be rerunnable after a feature backfill. Select the best answer.

Databricks Machine Learning

A team runs Databricks AutoML for a patient no-show model using appointment history, clinic, reminder channel, travel distance, and several missing demographics. The review requires editable preprocessing code, trial metrics, and a defensible reason for the selected estimator before any registry promotion. Which artifact should they inspect first? In this scenario, the workspace uses Unity Catalog grants rather than notebook-only access controls. Select the best answer.

Oportunidades profesionales y salario

Salario medio: $120,230mercado de EE. UU.– Data Scientists

Fuente: BLS Occupational Employment and Wage Statistics, May 2025 -- Data Scientists (SOC 15-2051), US national. Occupation median, not a certification salary. (2025)

Data Scientists

Los rangos son cifras del mercado de EE. UU. salvo que se muestre un rango local.

Qué temas cubre este examen

Usa las ponderaciones publicadas de los dominios para planificar tu estudio. Los resultados de práctica no predicen tu puntuación en el examen de certificación.

01Databricks Machine Learning

38%

This domain covers the operational aspects of machine learning on Databricks, focusing on MLOps strategies, the advantages of using ML runtimes, and how AutoML facilitates model and feature selection processes for efficient production workflows.

Temas

  • Identify the best practices of an MLOps strategy
  • Identify the advantages of using ML runtimes
  • Identify how AutoML facilitates model/feature selection.
  • Identify the advantages AutoML brings to the model development process
  • Identify the benefits of creating feature store tables at the account level in Unity Catalog in Databricks vs at the workspace level
  • Create a feature store table in Unity Catalog
  • Write data to a feature store table
  • Train a model with features from a feature store table.
  • Score a model using features from a feature store table.
  • Describe the differences between online and offline feature tables
  • Identify the best run using the MLflow Client API.
  • Manually log metrics, artifacts, and models in an MLflow Run.
  • Identify information available in the MLFlow UI
  • Register a model using the MLflow Client API in the Unity Catalog registry
  • Identify benefits of registering models in the Unity Catalog registry over the workspace registry
  • Identify scenarios where promoting code is preferred over promoting models and vice versa
  • Set or remove a tag for a model
  • Promote a challenger model to a champion model using aliases

Objetivos de aprendizaje

  • Identify the best practices of an MLOps strategy
  • Identify the advantages of using ML runtimes
  • Identify how AutoML facilitates model/feature selection.
  • Identify the advantages AutoML brings to the model development process
  • Identify the benefits of creating feature store tables at the account level in Unity Catalog in Databricks vs at the workspace level
  • Create a feature store table in Unity Catalog
  • Write data to a feature store table
  • Train a model with features from a feature store table.
  • Score a model using features from a feature store table.
  • Describe the differences between online and offline feature tables
  • Identify the best run using the MLflow Client API.
  • Manually log metrics, artifacts, and models in an MLflow Run.
  • Identify information available in the MLFlow UI
  • Register a model using the MLflow Client API in the Unity Catalog registry
  • Identify benefits of registering models in the Unity Catalog registry over the workspace registry
  • Identify scenarios where promoting code is preferred over promoting models and vice versa
  • Set or remove a tag for a model
  • Promote a challenger model to a champion model using aliases

02Model Development

31%

This domain addresses the technical process of developing machine learning models, including selecting the appropriate algorithm for specific scenarios, identifying methods to mitigate data imbalance in training sets, and comparing estimators and transformers.

Temas

  • Use ML foundations to select the appropriate algorithm for a given model scenario
  • Identify methods to mitigate data imbalance in training data
  • Compare estimators and transformers
  • Develop a training pipeline
  • Use Hyperopt's fmin operation to tune a model's hyperparameters
  • Perform random or grid search or Bayesian search as a method for tuning hyperparameters.
  • Parallelize single node models for hyperparameter tuning
  • Describe the benefits and downsides of using cross-validation over a train-validation split.
  • Perform cross-validation as a part of model fitting.
  • Identify the number of models being trained in conjunction with a grid-search and cross-validation process.
  • Use common classification metrics: F1, Log Loss, ROC/AUC, etc
  • Use common regression metrics: RMSE, MAE, R-squared, etc.
  • Choose the most appropriate metric for a given scenario objective
  • Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions
  • Assess the impact of model complexity and the bias variance tradeoff on model performance

Objetivos de aprendizaje

  • Use ML foundations to select the appropriate algorithm for a given model scenario
  • Identify methods to mitigate data imbalance in training data
  • Compare estimators and transformers
  • Develop a training pipeline
  • Use Hyperopt's fmin operation to tune a model's hyperparameters
  • Perform random or grid search or Bayesian search as a method for tuning hyperparameters.
  • Parallelize single node models for hyperparameter tuning
  • Describe the benefits and downsides of using cross-validation over a train-validation split.
  • Perform cross-validation as a part of model fitting.
  • Identify the number of models being trained in conjunction with a grid-search and cross-validation process.
  • Use common classification metrics: F1, Log Loss, ROC/AUC, etc
  • Use common regression metrics: RMSE, MAE, R-squared, etc.
  • Choose the most appropriate metric for a given scenario objective
  • Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions
  • Assess the impact of model complexity and the bias variance tradeoff on model performance

03Data Processing

19%

This domain focuses on data preparation tasks using Spark DataFrames, including computing summary statistics with dbutils, removing outliers based on standard deviation or IQR, and creating effective visualizations for both categorical and continuous features.

Temas

  • Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries
  • Remove outliers from a Spark DataFrame based on standard deviation or IQR
  • Create visualizations for categorical or continuous features
  • Compare two categorical or two continuous features using the appropriate method
  • Compare and contrast imputing missing values with the mean or median or mode value
  • Impute missing values with the mode, mean, or median value
  • Use one-hot encoding for categorical features
  • Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate.
  • Identify scenarios where log scale transformation is appropriate

Objetivos de aprendizaje

  • Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries
  • Remove outliers from a Spark DataFrame based on standard deviation or IQR
  • Create visualizations for categorical or continuous features
  • Compare two categorical or two continuous features using the appropriate method
  • Compare and contrast imputing missing values with the mean or median or mode value
  • Impute missing values with the mode, mean, or median value
  • Use one-hot encoding for categorical features
  • Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate.
  • Identify scenarios where log scale transformation is appropriate

04Model Deployment

12%

This domain covers the deployment of models into production environments using various serving patterns, including batch, realtime, and streaming methods, while also focusing on deploying custom models to endpoints and performing batch inference using pandas.

Temas

  • Identify the differences and advantages of model serving approaches: batch, realtime, and streaming
  • Deploy a custom model to a model endpoint
  • Use pandas to perform batch inference
  • Identify how streaming inference is performed with Delta Live Tables
  • Deploy and query a model for realtime inference
  • Split data between endpoints for realtime interference

Objetivos de aprendizaje

  • Identify the differences and advantages of model serving approaches: batch, realtime, and streaming
  • Deploy a custom model to a model endpoint
  • Use pandas to perform batch inference
  • Identify how streaming inference is performed with Delta Live Tables
  • Deploy and query a model for realtime inference
  • Split data between endpoints for realtime interference

Detalles del Examen Machine Learning Associate | $200 USD | 1 hora 30 minutos

Código del Examen Machine Learning Associate
Proveedor Databricks
Costo del Examen $200 USD
Límite de Tiempo 1 hora 30 minutos
Disponible En
EnglishJapanesePortuguese (Brazil)Korean

Preguntas Frecuentes

¿Quién es el candidato ideal para el examen Databricks Certified Machine Learning Associate?

¿Cuál es el formato del examen y cómo se entrega?

¿En qué se diferencia esta certificación de la Databricks Certified Data Engineer Associate?

¿Cuáles son las áreas clave del esquema del examen en las que debo enfocarme?

¿Cuál es la mejor manera de prepararse para este examen basado en el rendimiento?