Back to Blog

Data Scientist Interview Questions: Stats, ML, SQL, Cases

Published December 14, 2025
Updated October 4, 2026Technical Tips4 min read

By

Last updated October 4, 2026

Data Scientist Interview Questions: Stats, ML, SQL, Cases

Data scientist interviews usually combine statistics and machine learning fundamentals, SQL and coding, a case or product question, and behavioral questions. The mix depends on the role: product-analytics roles weight experimentation and SQL; machine learning roles weight modeling and system design. This page lists the common questions, what they test, and what a strong answer includes.

How the process usually runs

  1. Recruiter screen: background and role fit.
  2. Technical screen: SQL and/or Python, sometimes a statistics question.
  3. Onsite loop: statistics and ML, a case study, a project deep dive, and behavioral questions.

Statistics and experimentation

  • "Explain a p-value to a non-technical person." Strong answer: the probability of seeing a result at least this extreme if there were no real effect. It is not the probability that the hypothesis is true.
  • "How would you design an A/B test for a new checkout button?" Strong answer: define the metric and the unit of randomization, estimate the sample size from the minimum effect worth detecting, fix the duration before starting (avoid peeking), and check that the groups are balanced.
  • "The test shows a significant lift but you are not sure to trust it. Why?" Strong answer: novelty effects, multiple comparisons, sample ratio mismatch, a non-representative period, or a metric that does not capture the harm.
  • "What is the difference between correlation and causation?" Strong answer: give an example and say what would be needed to support a causal claim (an experiment or a well-justified quasi-experimental design).

Machine learning fundamentals

  • "Explain the bias-variance trade-off." High bias means a model too simple to capture the pattern (underfits); high variance means it fits noise (overfits). Regularization, more data, and simpler models reduce variance; more features or complexity reduces bias.
  • "How do you handle class imbalance?" Choose metrics that fit (precision, recall, PR-AUC rather than accuracy), then consider resampling, class weights, or threshold tuning, and say what the business cost of each error is.
  • "Precision versus recall: when do you prefer each?" Prefer recall when missing a positive is costly (disease screening) and precision when false alarms are costly (spam filtering of important mail).
  • "How do you prevent overfitting?" Hold-out or cross-validation, regularization, early stopping, simpler models, more data, and checking for leakage.
  • "Why might a model perform well offline and poorly in production?" Data leakage, distribution shift, training-serving skew, or a metric that does not match the product outcome.

SQL and coding

  • Window functions: running totals, ranking within groups, the previous row's value.
  • Joins and their effect on row counts (many-to-many joins that multiply rows).
  • Aggregations with filters, and finding the top N per group.
  • Basic Python or pandas: grouping, merging, handling missing values.

Example question: find the second-highest salary per department. A strong answer ranks within each department with DENSE_RANK() over a partition, filters to rank two, and discusses ties. See also SQL interview questions.

Case and product questions

  • "How would you detect fraudulent transactions?" Clarify the cost of errors, define labels, choose features, pick metrics for imbalance, discuss a threshold, monitoring and feedback loops.
  • "Subscriptions dropped last month. How do you investigate?" Check data first, segment, look at changes and seasonality, form hypotheses, and say which analyses would confirm them.

Worked answer: "Walk me through a project"

Outline of a strong answer: Start with the business question and who would use the answer. Describe the data (source, size, quality problems you found). Explain the approach and why you chose it over one alternative. Say how you evaluated it and what the baseline was. State the result in terms the stakeholder cares about, what limited it, and what you would do next. Be ready to defend each choice.

Behavioral questions

  • Tell me about a time your analysis was wrong or incomplete.
  • Describe explaining a technical result to a non-technical stakeholder.
  • Tell me about a project that did not lead to a decision, and what you did.

Common mistakes

  • Reciting definitions without examples or consequences.
  • Reporting accuracy on an imbalanced problem.
  • Starting to model before defining the business question.
  • Never naming a limitation of your result.

Practice plan

Spend a week each on statistics and experimentation, ML fundamentals, and SQL; then run mock interviews where you explain a past project aloud. The data analyst practice page has related questions, and data engineer questions cover the pipeline side.

Frequently asked questions

How much math is required?

Enough to explain statistical concepts clearly and apply them: distributions, hypothesis testing, regression, and probability. Be able to explain, not only compute.

Do I need to know deep learning?

For machine learning roles, yes at a conceptual level. For analytics-focused roles, experimentation and SQL matter more.

How do I prepare for the case study?

Practice structuring an ambiguous problem: define the goal, the data you would need, the method, and how you would evaluate and communicate the result.

Share:
#TechnicalTips#InterviewPrep#CareerGrowth