DASCA Certified Data Scientist Exam Practice Test
Preparation for DASCA CDS covering data science methodology, machine learning algorithms, statistical modeling, Python/R programming, big data tools, and data visualization. Administered by DASCA. Key domains include Data Science Foundation, Deep Learning Foundation, Machine Learning Associate and Machine Learning Expert.
Preguntas de Muestra
Prueba algunas preguntas para ver cómo es el examen completo.
A logistics carrier is preparing the winter 2025 dispatch redesign for 1.4 billion route events and weather joins from three providers. The case notes that weather joins are missing for rural lanes and drivers need recommendations within two minutes, and dispatch wants better ETA accuracy without destabilizing daily route planning. fleet branding is changing on the mobile app. The validation note says the holdout was collected after a policy change and before a marketing campaign; the test plan must preserve temporal order and site boundaries. the evidence sample contains 1267 records from 6 operating units. The workstream is SHAP explanations. Which explanation approach best fits a tabular gradient-boosting credit model?
A steering committee at A regional hospital network is comparing options for 37 facilities and two imaging vendors. The case file notes that FHIR feeds arrive 45 minutes late twice a week and a rural site changed triage workflow, and the clinical safety board needs an auditable go/no-go decision. Ignore the side issue that an executive also asks whether the dashboard can use the corporate color palette. The latest slice shows the model card must document assumptions before approval; the service must support both batch scoring and a real-time API; the evidence sample contains 1441 records from 3 operating units. Regarding transformer sequence modeling, Why can a transformer handle long text better than a simple RNN baseline?
A B2B software company has a production analytics issue involving product telemetry, support cases, and renewal records. In the FY2025 renewal forecast, enterprise customers use private deployments and success teams update records late, and finance wants a forecast that sales leaders will trust for pipeline planning, but a new logo package is being released in the same quarter. Current measurements show the champion has 0.82 ROC AUC but weak calibration in a high-risk segment; privacy counsel will not approve raw-record sharing across business units; the evidence sample contains 2195 records from 11 operating units. For master data management, the committee asks: Which action best reduces customer-identity errors across systems?
After a post-release review, A B2B software company documents a concern in product telemetry, support cases, and renewal records. The review finds that enterprise customers use private deployments and success teams update records late, and finance wants a forecast that sales leaders will trust for pipeline planning. The team also hears that a new logo package is being released in the same quarter. The model evidence shows a shadow deployment shows stable throughput but a shifted score distribution; privacy counsel will not approve raw-record sharing across business units; the evidence sample contains 1499 records from 5 operating units. In this analytics adoption decision, Which leadership action most reduces the risk that a technically sound model will not be used?
A pharmaceutical research group is preparing the 2025 interim analysis for multi-site trial data with lab values, notes, and imaging features. The case notes that one site changed lab equipment mid-study and the endpoint committee receives delayed labels, and the sponsor asks for explainable evidence before changing enrollment strategy. one investigator also requested a different chart style. The validation note says the validation set contains delayed labels from two independent review teams; privacy counsel will not approve raw-record sharing across business units. the evidence sample contains 3355 records from 6 operating units. The workstream is vector database retrieval. Which retrieval design best reduces unsupported answers in a RAG system?