Practice Interview AI Start practicing

Data Analyst Interview Questions

Prepare for SQL, statistics, metrics, investigation cases, and stakeholder questions by practicing clear explanations, not only final answers.

Updated October 2026 ยท 13 min read

Data analyst interviews test more than whether you can write a query. A hiring team needs to know how you define a business question, work safely with imperfect data, explain uncertainty, and help someone make a decision. The best preparation is to practice your reasoning out loud.

Use this list to identify the type of question you need to rehearse. For technical prompts, narrate the table grain, assumptions, and validation checks before you talk through syntax. For case prompts, state a structured investigation plan before you start naming causes.

SQL interview questions

SQL interviews may be live exercises, written prompts, or verbal walkthroughs. In every format, explain the unit of analysis first. Are you working with one row per order, event, account, or customer? That choice controls whether a join or aggregation creates a trustworthy answer.

Common prompts cover joins, deduplication, window functions, funnels, cohorts, dates, and NULL handling. Interviewers often care as much about how you find a double count as they do about the completed query.

Example: most recent row per customer

SELECT customer_id, status, updated_at
FROM (
  SELECT
    customer_id,
    status,
    updated_at,
    ROW_NUMBER() OVER (
      PARTITION BY customer_id
      ORDER BY updated_at DESC
    ) AS row_num
  FROM customer_status_history
) AS ranked
WHERE row_num = 1;

This query ranks each customer's history from newest to oldest, then keeps the first row. In an interview, explain what happens if two rows have the same timestamp and whether a secondary sort field is needed. Also say whether customer_status_history truly contains one record per status change, since the answer depends on its grain.

Example: conversion through a funnel

SELECT
  COUNT(DISTINCT CASE WHEN event_name = 'viewed_signup' THEN user_id END) AS viewed,
  COUNT(DISTINCT CASE WHEN event_name = 'completed_signup' THEN user_id END) AS completed
FROM product_events
WHERE event_time >= DATE '2026-09-01'
  AND event_time < DATE '2026-10-01';

This counts unique users who reached each event in a period. It is a starting point, not a complete funnel. A stronger explanation asks whether completion occurred after the view, whether events have reliable user identifiers, and whether the denominator should be restricted to users eligible to sign up.

Statistics and experimentation questions

Statistics prompts test whether you can match evidence to the confidence of a decision. Be prepared to explain a p-value in plain language, distinguish statistical significance from practical impact, and describe what a good experiment needs before results arrive.

For an A/B test, start with a specific hypothesis and a primary metric. Define guardrails before launch, such as errors, cancellations, or support contacts. Discuss the population, random assignment, expected test duration, and the minimum change worth detecting. If the sample is biased or a result is noisy, say what conclusion is unsafe and what you would do next.

A useful answer does not overclaim. A p-value does not tell you whether an effect is important to users or the business. Correlation does not establish that one metric caused another. Explain the limitation, then describe the additional evidence you would seek.

Metrics and dashboards questions

Metric questions ask whether you can make measurement useful for a real audience. A weekly executive dashboard should not be a catalog of everything the company tracks. It should show the few measures that describe progress, their trend, the reason a movement matters, and the area that needs attention.

When defining a metric, write the numerator, denominator, population, time window, and exclusions. This prevents two teams from reporting different versions of the same concept. For a product metric, connect the measure to a user action that indicates value, then add guardrails for quality and unintended behavior.

Expect questions about north star metrics, active user definitions, retention, conversion, dashboard design, anomalies, and reconciling conflicting reports. A strong answer considers data lineage and tracking changes before treating a surprising number as a real shift in behavior.

How to validate an analysis before sharing it

Interviewers may ask how you know an analysis is correct because data work is only useful when people can trust it. Start with the source: confirm the tables are current, the filters match the stated population, and the time zone or date boundary is appropriate. Compare key totals with an existing trusted report when one exists, but do not assume that report is correct without understanding its definition.

Then test the logic at a smaller scale. Inspect a handful of records, check for duplicate keys after joins, and make sure the numerator and denominator use the same eligible population. If a result is surprising, try to reproduce it with a simpler query or a different aggregation. Explain these checks in an interview, even if the prompt only asks for a final number. They show that you can distinguish a plausible result from a reliable one.

Finally, communicate uncertainty directly. State which assumptions are confirmed, which are provisional, and what additional data would reduce the risk of a wrong decision. This is not hesitation. It gives a stakeholder a clear view of what the analysis can support now.

Case and investigation questions

Case-style data analyst questions usually begin with a changed number: signups fell, churn rose, revenue missed forecast, or a launch affected a funnel. The goal is to show an investigation that is fast enough to be useful and careful enough not to send a team after the wrong cause.

Start by restating the metric. Confirm its definition, reporting freshness, and whether the apparent change is outside a normal range. Then segment the change in a way that could distinguish causes. Useful cuts include platform, release version, new versus returning users, channel, geography, plan, or a specific funnel step.

Turn the most informative cuts into hypotheses. A 20% signup drop overnight might come from broken instrumentation, an outage, a release, a change in paid traffic, or a browser-specific problem. Prioritize by likely impact and speed to validate. Finish with what you would communicate now, what you would test next, and what decision follows from each possible result.

Behavioral and stakeholder questions

Analysts work through other people. Behavioral questions examine whether you can clarify an ambiguous request, push back on a misleading metric, correct an error openly, and translate an analysis for an audience that does not use your tools.

Use the STAR method to organize your stories. Set the context briefly, make your responsibility clear, describe what you personally did, and give a specific result. If the result was not positive, explain the correction and what you changed in your process. The behavioral interview question bank offers more prompts for building a story library.

Full list of data analyst interview questions

Answer technical questions as if you are teaching a careful teammate. For case questions, give the interviewer your investigation plan before diving into the data. For stories, show the decision your work enabled.

  1. Explain the difference between an INNER JOIN and a LEFT JOIN, using a business example.

    Tests join semantics, awareness of missing records, and ability to explain a query without jargon.

    Practice
  2. How would you prevent double counting when joining orders to a table with multiple order tags?

    Tests grain awareness, join diagnosis, and a safe approach to aggregating after a many-to-many join.

    Practice
  3. When would you use a window function instead of GROUP BY?

    Tests whether you can preserve row detail while calculating rankings, running totals, or comparisons.

    Practice
  4. How would you deduplicate records while keeping the most recent row for each customer?

    Tests partitioning logic, ordering, and a clear explanation of the chosen deduplication rule.

    Practice
  5. How do NULL values affect comparisons, counts, and joins in SQL?

    Tests accurate NULL reasoning and practical safeguards against quietly excluding or miscounting rows.

    Practice
  6. How would you write a query to calculate monthly retention by signup cohort?

    Tests table grain, date logic, cohort definitions, and explanation of the resulting metric.

    Practice
  7. How would you calculate conversion through a three-step signup funnel?

    Tests event ordering, user-level deduplication, denominator choice, and treatment of incomplete journeys.

    Practice
  8. A dashboard query became slow after a new table was added. How would you investigate?

    Tests systematic debugging, understanding of joins and filters, and communication with data engineering partners.

    Practice
  9. What does a p-value tell you, and what does it not tell you?

    Tests precise statistical reasoning and avoidance of treating a p-value as proof of practical importance.

    Practice
  10. How would you design an A/B test for a new onboarding flow?

    Tests hypothesis definition, assignment, primary metrics, guardrails, and decisions before reviewing results.

    Practice
  11. What factors affect the sample size needed for an experiment?

    Tests understanding of baseline rate, minimum detectable effect, variance, power, and test duration.

    Practice
  12. Give an example of sampling bias in product data and how you would address it.

    Tests recognition of unrepresentative data, impact on conclusions, and a practical mitigation approach.

    Practice
  13. A metric is correlated with retention. How would you avoid claiming it causes retention?

    Tests causal caution, confounders, experiment thinking, and appropriate language for observational results.

    Practice
  14. How would you choose a north star metric for a marketplace?

    Tests connection between a metric and user value, balancing both sides, and avoiding vanity metrics.

    Practice
  15. What belongs on a weekly executive dashboard for a subscription product?

    Tests audience awareness, metric hierarchy, trends, context, and restraint in dashboard design.

    Practice
  16. Two teams report different active-user numbers. How would you reconcile them?

    Tests metric governance, definition comparison, data lineage, and clear communication of a final standard.

    Practice
  17. What guardrail metrics would you use when optimizing checkout conversion?

    Tests consideration of user harm, refunds, support load, and longer-term outcomes beyond conversion.

    Practice
  18. How would you decide whether a sudden dashboard spike is real or a data issue?

    Tests validation steps, source checks, segmentation, and measured communication while facts are uncertain.

    Practice
  19. Signups dropped 20% overnight. How do you investigate?

    Tests a structured root-cause process, data validation, segmentation, hypotheses, and a prioritized next step.

    Practice
  20. A new-user retention metric fell after a release. What would you do?

    Tests release awareness, cohort analysis, alternative explanations, and a recommendation based on evidence.

    Practice
  21. Revenue is below forecast, but traffic is steady. How would you break down the problem?

    Tests decomposition across conversion, pricing, mix, and data quality before selecting an action.

    Practice
  22. A leader asks for a dashboard by tomorrow with no defined question. How do you respond?

    Tests requirement gathering, scope management, useful first delivery, and expectation setting.

    Practice
  23. How would you identify which customers are most at risk of churn?

    Tests outcome definition, feature selection, segmentation, validation, and an actionable handoff to a partner.

    Practice
  24. An experiment improves conversion but increases customer complaints. What would you recommend?

    Tests tradeoff analysis, guardrail interpretation, further investigation, and a clear decision framework.

    Practice
  25. A stakeholder asks for a metric that cannot be measured from current data. What do you do?

    Tests honest communication, alternative proxies, instrumentation planning, and a useful next decision.

    Practice
  26. Tell me about a time you changed a stakeholder's mind with data.

    Tests influence, context, understandable analysis, ownership, and a concrete business or product result.

    Practice
  27. Tell me about a time you solved an ambiguous analysis problem.

    Tests problem framing, proactive communication, assumptions, and progress despite incomplete information.

    Practice
  28. Tell me about a mistake in your analysis and what you did next.

    Tests accountability, data quality habits, transparent correction, and learning that improved later work.

    Practice
  29. Tell me about a cross-functional project where priorities conflicted.

    Tests collaboration, tradeoff communication, personal contribution, and an outcome for the team.

    Practice
  30. Tell me about a time you explained a complex analysis to a nontechnical audience.

    Tests audience adaptation, clarity, decision focus, and evidence that the explanation was understood.

    Practice

Sample data analyst answers

Sample case answer

Prompt: Signups dropped 20% overnight. How do you investigate?

First, I would verify the metric definition, data refresh, and event tracking so we do not mistake a reporting issue for a product problem. I would compare the affected day with recent days and the same weekday, then split the change by platform, country, acquisition channel, app version, and signup funnel step. If the drop is isolated to one browser or release version, I would review errors and recent changes with engineering. If it is channel-specific, I would check campaign traffic and landing-page changes. I would share an early status noting what is confirmed and what is still a hypothesis. My next action would be the fastest check that separates a tracking failure from a real conversion failure, then I would recommend a rollback, fix, or deeper analysis based on that result.

Sample stakeholder answer

Prompt: Tell me about a time you changed a stakeholder's mind with data.

In a prior role, a marketing partner wanted to move budget toward a channel because its reported conversion rate looked highest. I checked the metric definition and found the channel was using a shorter attribution window than the others. I rebuilt the comparison with the same window, added customer quality and refund rates, and summarized the result in a one-page view rather than leading with query details. The channel still performed well for initial conversion, but it had a lower rate of completed activation. I recommended a smaller test allocation with an activation guardrail. The partner agreed, and the team used the same definition for later channel reviews. The key was showing the decision impact, not only pointing out an inconsistency.

How to practice explaining your work out loud

Silent problem solving can hide gaps in your reasoning. Speak through each step as you would in an interview. Start with the question the analysis must answer, name the data you expect to need, identify the grain and possible data-quality risks, then explain the calculation and decision. This makes it easier to catch a missing denominator or an unsupported conclusion.

  1. Choose a question type

    Use SQL for query reasoning, statistics for experimental judgment, a case for investigation structure, or a behavioral question for stakeholder communication.

  2. Practice under time pressure

    Open the data analyst practice track and answer by voice or text while the timer runs.

  3. Check the feedback

    Look for missing structure, evidence, ownership, or an unclear recommendation. Rewrite one sentence that would help a nontechnical listener follow you.

  4. Repeat with a stricter explanation

    Answer again, this time naming assumptions, validation steps, and the evidence that would change your conclusion.

For broader technical preparation, review the technical interview questions. Pair that with the behavioral interview questions and the STAR method guide so both your analysis and your working style are ready to discuss.

Frequently asked questions

What questions are asked in a data analyst interview?

Data analyst interviews commonly include SQL reasoning, statistics and experimentation, metric definitions, dashboard judgment, case-style investigation, and behavioral questions about stakeholders and communication.

How do I prepare for a data analyst SQL interview?

Practice explaining your query before writing it. State the table grain, the join keys, the denominator, how you handle duplicates and NULLs, and how you would validate the result. Review technical interview questions for more practice.

Do data analyst interviews require coding?

Many roles test SQL and may ask you to explain data logic. The exact tools vary. Focus on the reasoning behind joins, aggregations, window functions, data quality, and the business question being answered.

How should I answer a data analyst case study?

Confirm the metric and time period, check whether the change is real, segment the data, form hypotheses, validate the highest-value explanations, and recommend a next step. Explain what evidence would change your conclusion.

What statistics should a data analyst know for interviews?

Be ready to discuss experiment design, p-values, sample size, sampling bias, correlation versus causation, confidence intervals, and practical significance. Explain concepts plainly and connect them to a decision.