Stewardship • Catalog & Lineage • Data Quality • MDM • Privacy • Access • 2026

Data Governance Interview Questions

Most data governance interviews start with the basics, like owners versus stewards and what a catalog is for, then test data quality rules, lineage, master data, privacy classification and access control. Many also ask how tools such as Collibra, Microsoft Purview or Unity Catalog fit, and nearly all include a messy scenario, like two teams reporting different revenue. This page is written for anyone facing a data governance round, whether you're moving in from analytics or engineering or already run a governance programme. Each question shows what the interviewer is checking, the shape of a strong answer and a short answer to say out loud. Swap in your own stories.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Questions for freshers 8 questions

Foundations

Easy Technical round Fresher, Mid-level Practice question

1. In your own words, what is data governance, and how is it different from data management?

What the interviewer is really testing:
Whether you see governance as decision rights and accountability, not as a tool or a team that owns all the data.
Answer frame:
Governance vs Management
Governancewho decides, who is accountable, and which rules apply to data.
Managementthe day-to-day work of storing, moving, cleaning and securing data under those rules.

Why it matters: people trust and reuse data because definitions, owners and controls are clear.

Sample spoken answer:

“I think of data governance as the rules of the road for data. It decides who owns each data set, who can make decisions about it, what the standard definitions are, and which policies apply, like quality thresholds, privacy handling and access. Data management is the actual driving: engineers building pipelines, admins running databases, analysts cleaning data. Governance doesn't do that work itself; it sets the decisions and checks they're followed. A simple way I put it is that governance answers who decides and how we'll know it's right, and management does the work. When governance is working, people spend less time arguing about which number is correct and more time using the data, and the company can show a regulator or auditor that sensitive data is handled properly.”

Red flag to avoid:

Describing governance as buying a catalog tool, or as a central team that controls every data change.

They may ask next:
  • Who should a governance programme report to, and why does that choice matter?
  • How would you explain the value of governance to a sales leader in one sentence?
Say it in 60 seconds

Ownership & Stewardship

Easy Technical round Fresher, Mid-level Practice question

2. What's the difference between a data owner, a data steward and a data custodian?

What the interviewer is really testing:
Whether you know the core roles and that accountability sits with the business, not with IT.
Answer frame:
Owner vs Steward vs Custodian
Ownera senior business person accountable for the data: approves access, sets quality expectations, signs off definitions.
Stewardthe hands-on expert who maintains definitions, triages quality issues and answers questions.
Custodianusually IT or engineering; runs the systems, backups and technical controls.
Sample spoken answer:

“The data owner is accountable. It's usually a business leader, say the head of sales for customer account data, who decides who gets access, what good quality means and which definition is official. They don't do the daily work. The data steward does. That's someone who knows the data deeply, keeps the glossary terms and quality rules up to date, investigates issues and is the first person people ask. The custodian looks after the technical side: the database team or platform engineers who store the data, run backups and apply the access controls the owner approved. The key point is that ownership sits with the business, because they understand what the data means and carry the risk if it's wrong. IT keeps it safe and running, but shouldn't be deciding what 'active customer' means.”

Red flag to avoid:

Saying IT owns the data because it sits in their systems.

They may ask next:
  • What would you do if a data owner never has time to approve anything?
  • Can one person be both owner and steward for a small data set?
Say it in 60 seconds

Metadata & Glossary

Easy Technical round Fresher, Mid-level Practice question

3. How do a business glossary, a data dictionary and a data catalog differ, and how do they connect?

What the interviewer is really testing:
Whether you can separate business meaning from technical structure and see how a catalog links the two.
Answer frame:
Glossary vs Dictionary vs Catalog
Glossarybusiness terms and their agreed meaning, like 'active customer', with an owner.
Dictionarytechnical detail of tables and columns: names, types, allowed values.
Cataloga searchable inventory of data assets that links glossary terms to the physical columns, plus owners, lineage and usage.
Sample spoken answer:

“A business glossary is about meaning. It says what 'active customer' or 'net revenue' means in plain language, who owns that definition, and any rules around it. A data dictionary is technical. It describes a table's columns, their data types, formats and allowed values. A data catalog is the searchable inventory that pulls it together: it lists the data assets across our systems, links each column to the glossary term it represents, and shows the owner, lineage, quality score and who uses it. So if an analyst searches for 'active customer', a good catalog shows the approved definition and then points to the exact tables and columns that hold it. Without that link, you get a nice glossary that nobody can connect to real data, or a list of tables nobody understands.”

Red flag to avoid:

Treating the three as the same thing, or saying a catalog is just a list of tables.

They may ask next:
  • Who should write glossary definitions, the governance team or the business?
  • How would you handle two glossary terms that mean nearly the same thing?
Say it in 60 seconds
Easy Technical round Fresher, Mid-level Practice question

4. What are the main types of metadata, and can you give a real example of each?

What the interviewer is really testing:
Whether you know metadata goes beyond column names and why each type helps a governance programme.
Answer frame:

Technical: schemas, data types, table names, file locations.

Business: definitions, owners, sensitivity class, approved uses.

Operational: load times, row counts, job runs, freshness, access and usage logs.

Sample spoken answer:

“I usually group metadata into three types. Technical metadata describes structure: that orders has a column order_date of type date, where the table lives and how it's partitioned. Business metadata describes meaning and accountability: that order_date means the date the customer placed the order, not when it shipped, that the owner is the head of e-commerce, and that the table is classed as internal. Operational metadata describes what happens to the data: when it last loaded, how many rows arrived, whether the job failed, and who queried it last week. Governance needs all three. Technical metadata tells you what exists, business metadata tells you what it means and who to ask, and operational metadata tells you whether you can trust it today and who might be affected if something breaks.”

Red flag to avoid:

Saying metadata is only column names and types.

They may ask next:
  • Which of these can be collected automatically, and which need people?
  • How would you use usage metadata to decide which data sets to govern first?
Say it in 60 seconds

Catalog & Lineage

Easy Technical round Fresher, Mid-level Practice question

5. What is data lineage, and what's the difference between table-level and column-level lineage?

What the interviewer is really testing:
Whether you understand lineage and its uses, like impact analysis, root cause and audit.
Answer frame:

Definition: where data came from, what transformed it and where it goes next.

Table vs column: table-level shows which tables feed which; column-level shows exactly which fields build a given field.

Uses: impact analysis before a change, root cause when a number is wrong, proof for auditors.

Sample spoken answer:

“Data lineage is the map of how data flows: from the source system, through each transformation, into the tables and reports people use. Table-level lineage tells me that the revenue dashboard reads from a sales summary table, which is built from orders and refunds. Column-level lineage goes further and shows that the net_revenue field is calculated from order_amount minus refund_amount, and which source columns those came from. I use lineage three ways. Before a change, for impact analysis: if a source field is renamed, what breaks downstream? When a number looks wrong, to trace it back to the step that caused it. And for audit, to prove where a reported figure or a piece of personal data came from. Column-level is much more useful, but it's also harder to capture completely.”

Red flag to avoid:

Describing lineage as a diagram someone draws once in a slide deck.

They may ask next:
  • A source system is renaming a field next month. How would you use lineage to judge the impact?
  • Who uses lineage most in your experience, engineers, analysts or auditors, and does that change what you capture?
Say it in 60 seconds

Data Quality

Easy Technical round Fresher, Mid-level Practice question

6. What are the common data quality dimensions? Give me an example of a failure for each.

What the interviewer is really testing:
Whether you can name the dimensions and turn each into something you could actually test.
Answer frame:

Content: accuracy, completeness, validity.

Consistency and uniqueness: same value across systems, no duplicate records.

Timeliness: data is fresh enough for its use.

Sample spoken answer:

“The six I use most are accuracy, completeness, validity, consistency, uniqueness and timeliness. Accuracy is whether the value matches reality, like a customer's phone number actually reaching them. Completeness is whether required values are there, say ten thousand orders with no delivery postcode. Validity is whether values follow the rules: a date of birth in the future or a country code that isn't on the approved list. Consistency means the same fact agrees across systems, so a customer isn't 'closed' in billing and 'active' in the CRM. Uniqueness means no duplicates, like the same supplier set up three times. Timeliness is whether data arrives in time for its use: yesterday's sales that only land at noon are no good for the 9 a.m. meeting. Naming them helps turn a vague 'the data is bad' into a rule you can measure.”

Red flag to avoid:

Listing the names with no example of how you'd test any of them.

They may ask next:
  • Which dimension is hardest to measure automatically, and why?
  • How would you decide which dimensions matter most for a given data set?
Say it in 60 seconds

Master Data

Easy Technical round Fresher, Mid-level Practice question

7. What's the difference between master data, reference data and transactional data?

What the interviewer is really testing:
Whether you can classify data correctly, since each type is governed differently.
Answer frame:
Master data vs Reference data vs Transactional data
Master datacore business entities shared across systems: customers, products, suppliers, employees.
Reference datasmall, controlled code lists: country codes, currencies, units, status values.
Transactional dataevents that reference them: orders, payments, shipments.
Sample spoken answer:

“Master data describes the core things the business deals with, like customers, products, suppliers and locations. It changes slowly and is used by many systems, so getting one agreed version matters a lot. Reference data is the set of allowed values used to categorise other data: country codes, currency codes, order status values, units of measure. It's small and often comes from outside standards, like ISO country codes, but it has to be the same everywhere or joins and reports break. Transactional data records events: an order, a payment, a delivery. It's high volume and each record points to master and reference data, so an order points to a customer, a product and a currency. Governance-wise, master data needs owners and match rules, reference data needs a single controlled list with change approval, and transactional data needs quality checks at volume.”

Red flag to avoid:

Calling everything in the database master data.

They may ask next:
  • Where does something like a price list fit?
  • Who should approve adding a new value to a shared status code list?
Say it in 60 seconds

Privacy & Classification

Easy Technical round Fresher, Mid-level Practice question

8. How would you set up a data classification scheme for a company, and what levels would you use?

What the interviewer is really testing:
Whether you know a simple, usable scheme and that each level must drive concrete handling rules.
Answer frame:

Levels: a few clear ones, often public, internal, confidential, restricted.

Handling rules: each level maps to controls: who can access, encryption, sharing, retention.

Apply it: tag at column or data set level, owners confirm, highest level wins when mixed.

Sample spoken answer:

“I'd keep it to four levels, because people can't apply ten. Public is anything already published, like the website. Internal is the default for business data that isn't meant outside but wouldn't do much harm if leaked, like an org chart. Confidential is data that could hurt the company or people if exposed, like customer contact details or contracts. Restricted is the most sensitive: things like health data, government ID numbers, card numbers or passwords. A scheme only matters if each level changes what happens, so I'd attach handling rules: restricted data is encrypted, masked by default, access is approved by the owner and reviewed regularly, and it can't be copied to personal devices. I'd tag at column level where possible, have owners confirm the tags, and when a table mixes levels, the table takes the highest one.”

Red flag to avoid:

Listing levels with no handling rules attached to them.

They may ask next:
  • Who should decide the classification of a new data set?
  • How would you handle derived data, like an aggregate built from restricted records?
Say it in 60 seconds

Questions for every level 4 questions

Data Quality

Medium Technical round Fresher, Mid-level, Senior Practice question

9. How would you turn a business expectation like 'every customer must have a valid country' into a data quality rule you can measure?

What the interviewer is really testing:
Whether you can go from a vague expectation to a precise, testable rule with a threshold and an owner.
Answer frame:

Define precisely: which table, which column, what counts as valid, which records are in scope.

Measure: a query that counts passing and failing rows, run on a schedule.

Act: a threshold agreed with the owner, and a clear route for failures.

Sample spoken answer:

“I'd first pin down the words with the data owner. Which customers count: all, or only active ones? What is 'valid': not empty and on the approved country code reference list? Then the rule becomes measurable: for active customers, country_code is not null and exists in the reference table. I'd write it as a query that counts total rows and failing rows, so we get a pass rate, and run it on every load. Next, a threshold agreed with the owner, say above ninety-nine in a hundred is green, and a response for red: an alert to the steward, with the failing records listed so they can be fixed at the source. Finally the rule goes into the catalog next to the column, with its owner and threshold, so anyone using the data can see how it's doing.”

Code:
-- Completeness and validity of country for active customers
SELECT
  COUNT(*) AS total_rows,
  SUM(CASE WHEN c.country_code IS NULL THEN 1 ELSE 0 END) AS missing_country,
  SUM(CASE WHEN c.country_code IS NOT NULL AND r.code IS NULL
           THEN 1 ELSE 0 END) AS invalid_country
FROM customers c
LEFT JOIN ref_country r ON r.code = c.country_code
WHERE c.status = 'ACTIVE';
Red flag to avoid:

Writing the rule without agreeing scope and thresholds with the business, or having no plan for what happens when it fails.

They may ask next:
  • Should a failing rule stop the data from loading, or just raise an alert?
  • How do you set a threshold when the owner says it must be perfect?
Say it in 60 seconds
Medium Behavioral round Fresher, Mid-level, Senior Practice question

10. Tell me about a data quality issue you traced to its root cause. What was really going on, and what did you change?

What the interviewer is really testing:
Whether you dig past the symptom to process causes, and fix things so they stay fixed.
Answer frame:

Symptom: what people noticed and why it mattered.

Trace: how you followed the data back, using lineage, samples and people.

Fix: the change at the source, plus a rule so you'd see it again early.

Sample spoken answer:

“In a previous role, our customer churn report showed a sudden jump in closed accounts one month. People were worried. I pulled a sample of the closed accounts and noticed most were closed and reopened on the same day. Using the lineage in our catalog, I followed the status field back to the CRM, then talked to the service team. It turned out a new process for changing a customer's billing address closed the old account and opened a new one, because that was the only way the screen allowed it. So they weren't churned customers at all. The fix had two parts. The service team and the CRM owner changed the process so addresses could be updated in place. And I added a quality rule that flagged accounts closed and reopened within a week, reported to the steward. Churn went back to normal and leadership trusted the report again.”

Red flag to avoid:

A story that ends with patching the report and never touching the source process.

They may ask next:
  • How did you correct the months that had already been reported?
  • What would have happened if you'd only fixed the report logic?
Say it in 60 seconds

Access Control

Medium Technical round Fresher, Mid-level, Senior Practice question

11. What's the difference between role-based and attribute-based access control, and why do governance teams often push for attributes?

What the interviewer is really testing:
Whether you understand how access scales and how classification tags can drive policy.
Answer frame:

RBAC: access granted to roles or groups, like 'finance analysts can read the finance schema'.

ABAC: a policy checks attributes of the user, the data and the context, like region or sensitivity tag.

Why ABAC scales: one policy on a tag covers every new table carrying it, instead of hundreds of role grants.

Sample spoken answer:

“Role-based access control grants permissions to a role or group, and people get access by joining it. It's simple to understand and audit, but as the data estate grows you end up with hundreds of roles, one per team per region per data set, and nobody can say what anyone can see. Attribute-based access control writes policies on attributes instead: the user's department and region, the data's classification tag, maybe the purpose. So one policy can say restricted columns are masked unless the user is in HR, and it applies automatically to every new table tagged restricted. That's why governance teams like it: classification becomes the thing that drives access. In practice most companies use both. Roles for broad access, attributes for sensitive data and row-level rules like 'managers only see their own region'.”

Red flag to avoid:

Saying one model is always better, or not connecting attributes to data classification.

They may ask next:
  • What goes wrong with ABAC if the tags themselves are wrong?
  • How would you audit who can actually see a given column?
Say it in 60 seconds
Medium Technical round Fresher, Mid-level, Senior Practice question

12. How should access to sensitive data be requested, approved and reviewed over time?

What the interviewer is really testing:
Whether you know least privilege is a lifecycle: grant, justify, expire and recertify.
Answer frame:

Request: through one process, stating the data set and the purpose.

Approve: the data owner approves, access goes to a group, with an end date where possible.

Review: regular recertification by owners, and automatic removal when people move or leave.

Sample spoken answer:

“I'd have one front door for access requests, ideally from the catalog page of the data set, where the person says what they need and why. The data owner, or a steward they've delegated to, approves it, and the grant goes to a group rather than to the individual, so it's easy to see and remove. For restricted data I'd add an end date, so access expires unless it's renewed. Over time access creeps: people change teams and keep old permissions. So I'd run a recertification every quarter or half-year where owners see who has access to their data and confirm or remove each person. Joiner, mover and leaver changes from HR should remove access automatically. And everything is logged, so we can answer an auditor's question about who could see this data on a given date.”

Red flag to avoid:

Granting access directly to individuals with no expiry and no review.

They may ask next:
  • How do you stop recertification turning into owners clicking approve on everything?
  • What would you do about service accounts that have broad access?
Say it in 60 seconds

Questions for experienced candidates 18 questions

Foundations

Medium Technical round Mid-level, Senior Practice question

13. How would you measure whether a data governance programme is actually working?

What the interviewer is really testing:
Whether you measure outcomes the business cares about, not just activity like the number of glossary terms.
Answer frame:

Outcomes: fewer incidents, less time spent reconciling, faster audits, faster access to data.

Health: quality scores on critical data, share of critical assets with an owner and certified description.

Adoption: catalog use, repeat users, requests handled on time.

Sample spoken answer:

“I'd avoid counting activity, like how many terms we've added, because that can go up while nothing improves. I'd track three layers. First, business outcomes: the number of data incidents that reached a report, how long month-end reconciliation takes, how quickly someone gets approved access to a data set, and how smoothly the last audit went. Second, health of the critical data: quality scores on the key elements over time, and how many of them have a named owner, a steward and a certified definition. Third, adoption: how many people use the catalog each month and come back, and whether access and issue requests are handled within the agreed time. I'd set a baseline before we start, report a short scorecard to the governance council each quarter, and tie at least one number to something a leader already worries about.”

Red flag to avoid:

Only naming activity metrics like the count of policies written or tables scanned.

They may ask next:
  • Which single metric would you show the executive sponsor, and why?
  • How do you get a baseline for time spent reconciling numbers?
Say it in 60 seconds
Hard Technical round Senior Practice question

14. Centralised, federated or hybrid: how would you choose a governance operating model for a large company?

What the interviewer is really testing:
Whether you can design who does what across a large organization and match it to culture and maturity.
Answer frame:

Centralised: one team sets and enforces everything; consistent but slow and a bottleneck.

Federated: domains own their data and rules within a few shared standards; scales but needs mature domains.

Hybrid: a small central team owns standards, tooling and shared definitions; domains own their data and stewards.

Sample spoken answer:

“A centralised model puts one team in charge of standards, definitions and approvals. It's consistent and good for a small company or a heavily regulated start, but it becomes a bottleneck and business teams stop feeling responsible. A federated model pushes ownership to domains like sales, finance and supply chain, with only a thin set of shared rules. That scales and fits ideas like data mesh, where standards are built into the platform, but it needs domains that are mature enough to own their data. Most large companies I'd steer toward a hybrid. A small central office owns the policies, the tooling, the shared terms like customer and revenue, and reporting to a governance council. Each domain has an owner and stewards who do the actual work. I'd start more central while maturity is low and hand more to domains as they prove they can run it.”

Red flag to avoid:

Choosing one model as always right without looking at company size, regulation and maturity.

They may ask next:
  • What decisions must stay central even in a federated model?
  • How would you set up a governance council so it doesn't become a talking shop?
Say it in 60 seconds
Hard Situational round Senior Practice question

15. You're the first data governance hire at a mid-sized company with no programme at all. What would you do in your first 90 days?

What the interviewer is really testing:
Whether you start small, tie governance to a real business pain and build momentum instead of writing a big framework nobody uses.
Answer frame:

Listen: meet leaders and data users; find the pains, risks and regulatory pressure.

Pick a pilot: one domain or one critical report with a willing sponsor and owner.

Show value: a few definitions, quality rules and access fixes, measured and shared.

Sample spoken answer:

“In the first month I'd mostly listen. I'd meet the executives, the data team and the heaviest data users, and ask where data costs them time, money or sleep: numbers that don't match, audits, customer complaints, slow access. I'd also check what regulation we're under. In month two I'd pick one pilot with a sponsor who cares, maybe the monthly revenue report or customer data. For that domain I'd name an owner and a steward, agree the key definitions, set a handful of quality rules on the critical fields and tidy up who has access. I'd keep policy light: a short charter, the roles, and how decisions get made. In month three I'd show results against the baseline, set up a small council, and plan the next two domains. The goal is that people ask to be next, not that I've written a thick framework.”

Red flag to avoid:

Starting with a long policy document or a tool purchase before understanding the business problems.

They may ask next:
  • What would you do if leadership wants a company-wide catalog rollout in the first month?
  • How would you pick the pilot if three leaders all want to go first?
Say it in 60 seconds

Ownership & Stewardship

Medium Situational round Mid-level, Senior Practice question

16. Nobody wants to own the supplier data, and quality keeps getting worse. How would you get a real owner and steward in place?

What the interviewer is really testing:
Whether you can assign accountability through influence and a clear, small ask rather than a policy nobody reads.
Answer frame:

Show the pain: tie the bad data to a cost the business already feels, like late payments or duplicate suppliers.

Pick by process: the leader whose process creates or depends most on the data is the natural owner.

Make it small: a clear, time-boxed role description and a sponsor to back it.

Sample spoken answer:

“First I'd find where bad supplier data is already hurting someone. Maybe invoices bounce because bank details are wrong, or we have the same supplier three times and lose volume discounts. I'd put a few real examples in front of the procurement and finance leads. Then I'd propose ownership based on who creates the data, which is usually procurement, since they onboard suppliers. The ask would be small and specific: the owner approves the definition of an active supplier, agrees two or three quality rules and reviews a monthly score. The steward would be someone on their team who already fixes these records informally, and we'd make that official with a bit of their time set aside. I'd also get our executive sponsor to confirm it, so it isn't just me asking. Then I'd show early progress so the role feels worth it.”

Red flag to avoid:

Assigning an owner by email with no business case, or making the governance team the owner by default.

They may ask next:
  • What if two departments both claim ownership of the same data?
  • How much of a steward's time would you ask for, and how would you justify it?
Say it in 60 seconds
Medium Behavioral round Mid-level, Senior Practice question

17. Tell me about a time you got a business team to take real ownership of their data when they saw it as IT's job.

What the interviewer is really testing:
Whether you can influence without authority and make ownership feel useful rather than extra work.
Answer frame:

Situation: the data, the team and why they pushed back.

Action: how you showed the cost, shrank the ask and gave them control.

Result: what changed, with a before and after.

Sample spoken answer:

“At my last company, the marketing team saw campaign and lead data as IT's problem, even though they set up every campaign. Reports kept splitting one campaign into several because of inconsistent naming, so nobody trusted the lead numbers. The marketing head's view was, fix it in the database. I sat with one of her managers and showed how much of their weekly reporting time went on cleaning names by hand, and that the leads from their biggest campaign were being undercounted. That got attention. Then I made the ask small: they'd own a naming convention for campaigns, and one person would be steward for an hour a week. I built a simple check that flagged badly named campaigns the day they were created. Within two months the manual cleaning stopped, and the marketing head started presenting the lead numbers herself. That's when I knew they owned it.”

Red flag to avoid:

A story where you force ownership through an escalation, or quietly do all the work yourself.

They may ask next:
  • What would you have done if the marketing head had still said no?
  • How did you keep the steward role going once the first push was over?
Say it in 60 seconds

Metadata & Glossary

Hard Situational round Mid-level, Senior Practice question

18. Sales and finance both show 'revenue' to the leadership team, and the numbers never match. How would you sort this out?

What the interviewer is really testing:
Whether you can resolve a definition conflict through facts and decision rights, keeping both valid measures without picking a winner by politics.
Answer frame:

Find the real differences: timing, refunds, discounts, tax, currency, which deals count.

Decide who decides: the owner of the term, often finance for reported revenue, with the council as tie-breaker.

Name and publish: one official term, other measures renamed, definitions in the glossary, certified sources.

Sample spoken answer:

“I'd first reconcile the two numbers line by line for one month. Usually both are right for their purpose. Sales might count bookings when a contract is signed, at list price, while finance counts recognised revenue after discounts, refunds and tax, following accounting rules. Once the differences are clear, it stops being an argument about who's wrong. Then decision rights: the owner of the official 'revenue' term is usually finance, because it ties to the accounts, and our governance council breaks ties if needed. Sales doesn't lose its measure. It gets renamed to something honest, like 'bookings', with its own definition. Both go in the glossary with owners and calculation rules, each points to a certified data set, and dashboards get relabelled. I'd also add a small reconciliation view that shows how bookings turn into revenue, so the leadership team sees why the two differ.”

Red flag to avoid:

Picking one number as correct straight away, or forcing one definition on everyone without a decision owner.

They may ask next:
  • What if the head of sales refuses to rename their number?
  • How would you stop a third definition from appearing next quarter?
Say it in 60 seconds

Catalog & Lineage

Medium Technical round Mid-level, Senior Practice question

19. The company bought a data catalog a year ago and hardly anyone uses it. How would you turn that around?

What the interviewer is really testing:
Whether you know catalogs fail on content and habit, not software, and can drive adoption with a focus.
Answer frame:

Diagnose: look at search logs and ask analysts why they skip it.

Curate the top assets: describe, certify and assign owners for the most-used data sets first.

Put it in the flow: link it from dashboards and tools people already use, and make it the answer to common questions.

Sample spoken answer:

“I'd start by finding out why. Usually it's because the catalog is full of scanned tables with no descriptions, so searching it doesn't answer anything. I'd pull usage logs to find the twenty or so data sets people actually query most, and get their owners and stewards to write clear descriptions, link the glossary terms and mark the trusted ones as certified. That small, high-value slice is worth more than ten thousand empty entries. Then I'd put the catalog where people already are: links from dashboards back to the certified source, and a habit in the analyst channel of answering 'where do I find X' with a catalog link. I'd track searches, certified asset views and repeat users each month. And I'd retire or hide junk entries, because noise is what teaches people the catalog is useless.”

Red flag to avoid:

Proposing a company-wide training session and a mandate, with no curation of content.

They may ask next:
  • What does 'certified' mean to you, and who is allowed to certify a data set?
  • How would you keep descriptions from going stale once they're written?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

20. Your automated lineage looks great in the demo but has holes in production. Where do the gaps usually come from, and how would you close them?

What the interviewer is really testing:
Whether you have real experience with lineage tools and know their blind spots.
Answer frame:

Common gaps: hand-written scripts, spreadsheets, dynamic SQL, file drops and tools the scanner can't read.

Prioritise: close gaps on critical data elements and regulated reports first.

Fill in: connectors or parsers where possible, manual lineage with an owner where not.

Sample spoken answer:

“Automated lineage is only as good as what the tool can read. The holes usually come from custom Python or shell scripts, stored procedures with dynamic SQL, spreadsheets someone emails around, files dropped on a shared folder, and third-party tools with no connector. The result is a lineage graph that just stops, or a report that seems to appear from nowhere. I wouldn't try to fix everything. I'd pick the critical data elements, the numbers in board reports and regulatory returns, and trace them end to end. Where there's a connector or parser we're not using, turn it on. Where there isn't, document the step manually in the catalog with an owner and a review date. And where the gap is a spreadsheet in the middle of a critical flow, that's often a sign the process itself should change, not just the lineage.”

Red flag to avoid:

Assuming the tool captures everything, or trying to map every data set before any critical one is done.

They may ask next:
  • How would you keep manually added lineage from going out of date?
  • Which data would you call critical data elements, and who decides?
Say it in 60 seconds

Master Data

Hard Technical round Mid-level, Senior Practice question

21. Walk me through the main MDM implementation styles. When would you choose a registry approach over a centralised hub?

What the interviewer is really testing:
Whether you know the styles and the trade-off between low disruption and strong control.
Answer frame:

Registry: sources keep their records; the hub holds matching keys and builds a view when asked.

Consolidation: the hub builds a golden record for reporting, but doesn't write back.

Coexistence and centralised: the golden record is synced back to sources, or the hub becomes where records are created.

Sample spoken answer:

“There are four common styles. Registry leaves the data in the source systems and only keeps a cross-reference in the hub: these three records are the same customer. It's the least disruptive and quick to start, but you don't really fix the sources. Consolidation copies records into the hub and builds a golden record, mainly for analytics, without sending it back. Coexistence also builds the golden record but syncs it back to the sources, so operational systems get cleaner too. Centralised, sometimes called transactional, makes the hub the only place records are created and changed, with approval workflows, which gives the most control but changes how people work. I'd pick registry when there are many systems we can't change easily and the first goal is a single view. I'd move toward centralised when bad master data is causing real operational pain, like failed shipments or payments.”

Red flag to avoid:

Knowing only one style, or suggesting a centralised hub for every case without mentioning how disruptive it is.

They may ask next:
  • Which style would you pick for supplier data in a company with three different ERP systems?
  • What organizational changes does a centralised hub need to succeed?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

22. How does match and merge work in MDM, and how do survivorship rules decide what goes into the golden record?

What the interviewer is really testing:
Whether you understand matching risk and can define survivorship per attribute, not per record.
Answer frame:

Match: standardise first, then compare with exact and fuzzy rules to score likely duplicates.

Review: auto-merge above a high score, send the grey zone to stewards, keep apart below.

Survivorship: per attribute: most trusted source, most recent, most complete.

Sample spoken answer:

“Matching starts with standardising: same case, trimmed spaces, standard address formats, phone numbers in one format. Then match rules compare records, some exact like a tax ID, some fuzzy like a similar name at the same postcode, and give each pair a score. I'd set two thresholds: above the high one we merge automatically, in the middle a steward reviews, and below we keep them apart. False merges are worse than missed ones, because you might combine two real customers and show one their neighbour's orders. Once records are grouped, survivorship picks the best value for each attribute, not the whole record. So the legal name might come from the finance system as the most trusted source, the email from the most recently updated record, and the phone from whichever is complete and valid. We keep the cross-reference to every source record so merges can be undone.”

Red flag to avoid:

Picking one whole source record as the winner, or never mentioning the risk of false merges.

They may ask next:
  • How would you tune the thresholds, and who signs off the trade-off?
  • What happens when a merge turns out to be wrong six months later?
Say it in 60 seconds

Privacy & Classification

Medium Technical round Mid-level, Senior Practice question

23. You've been asked to find all the personal data spread across hundreds of databases and file shares. How would you go about it?

What the interviewer is really testing:
Whether you can combine automated discovery with human review and focus on risk.
Answer frame:

Scan: automated classifiers look at column names and sample values for patterns like emails, IDs, card numbers.

Confirm: stewards review matches, since scanners produce false positives and miss free text.

Record and act: tag it in the catalog, then apply masking, access and retention rules.

Sample spoken answer:

“I'd start with automated discovery, because hundreds of sources can't be checked by hand. Most catalog and governance tools can scan sources and flag likely personal data from column names and from patterns in sample values, like email formats, phone numbers, card numbers that pass a checksum, or national ID formats. I'd prioritise sources by risk: customer-facing systems, HR and anything shared with third parties first. Scanners get it wrong both ways. A column called notes can hold phone numbers and names, and a column called id might look like a card number but isn't. So stewards review the results and confirm the tags. Every confirmed field goes into the catalog with its classification, and that drives the next steps: masking, tighter access, retention rules and an entry in the record of processing if the law where we operate needs one.”

Red flag to avoid:

Relying only on column names, or only on a one-off manual survey.

They may ask next:
  • How would you deal with personal data hidden in free-text fields or PDFs?
  • How often would you re-scan, and why?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

24. Which privacy principles shape how you govern personal data, and how do you turn them into controls?

What the interviewer is really testing:
Whether you know the common principles that most privacy laws share and can make them concrete, while staying neutral about jurisdiction.
Answer frame:

Principles: lawful basis, purpose limitation, data minimisation, accuracy, storage limitation, security.

Controls: purpose tags on data sets, collect only needed fields, retention schedules, access and masking.

Rights: processes to find, export, correct and delete a person's data on request.

Sample spoken answer:

“Most privacy laws share a core set of principles, whether it's the GDPR in Europe or similar laws elsewhere, though the details differ, so I always check the specifics with legal. The ones that shape my work are: have a valid reason for using the data, use it only for the purpose it was collected for, collect no more than needed, keep it accurate, keep it no longer than needed, and protect it. Governance turns those into controls. Data sets carry a purpose tag, and a new use needs a check before access is granted. New collections go through a review asking whether each field is needed. Retention schedules are set per data type and enforced by deletion jobs. And we need a reliable way to find everything about one person across systems, which is where the catalog and lineage pay off, so we can answer access and deletion requests on time.”

Red flag to avoid:

Treating privacy as only a security problem, or naming one country's law as if it applied everywhere.

They may ask next:
  • Is pseudonymised data still personal data?
  • How do you handle a deletion request when the person's data is in last month's backups?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

25. How would you design a data retention policy and make sure data is really deleted when its time is up?

What the interviewer is really testing:
Whether you can balance legal minimums and maximums, and whether you know deletion must be proven, not just written down.
Answer frame:

Schedule: per data type, the legal minimum, business need and a maximum, agreed with legal and owners.

Enforce: automated deletion or anonymisation jobs, including copies and extracts.

Exceptions and proof: legal holds pause deletion; logs show what was deleted and when.

Sample spoken answer:

“I'd build a retention schedule by data type with legal and the data owners. Some records have a minimum period, like financial records that tax rules say must be kept for a number of years, and personal data has a maximum: no longer than we need it. Each entry says the period, what starts the clock, like account closure, and what happens at the end, delete or anonymise. Writing it down is the easy part. The hard part is copies: extracts in analysts' folders, test environments, old backups. So I'd use lineage to find where each data type flows, automate deletion jobs on the main stores, and set backups to expire on a known cycle. Legal holds must be able to pause deletion for specific records. And I'd keep deletion logs, because an auditor will ask for proof, not just the policy.”

Red flag to avoid:

Saying 'keep everything forever, storage is cheap' or ignoring copies and backups.

They may ask next:
  • What would you do about personal data in a test environment copied from production?
  • When would you anonymise instead of delete?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

26. Teams want to feed company data into AI models and assistants. What changes for data governance?

What the interviewer is really testing:
Whether you can extend existing controls to AI uses without hype, focusing on purpose, sensitivity, access and provenance.
Answer frame:

Purpose and consent: is this data allowed to be used for this new purpose?

Sensitivity and access: classified data needs the same controls when it flows into prompts, indexes or training sets.

Provenance: record which data built which model or index, and keep it current.

Sample spoken answer:

“The basics don't change, but the risks move. First, purpose: data collected for billing may not be allowed for training a model, so a new AI use needs the same purpose check as any new use, and contracts with customers might restrict it. Second, sensitivity and access. If a document search assistant indexes everything on the file shares, it can show people content they could never open directly, so it must respect the same permissions, and restricted data should be kept out or masked. Third, provenance: we need lineage from source data to each training set, model or search index, so we can answer what it learned from, and remove data when a deletion request or retention rule applies. Quality matters more too, because a model repeats bad data confidently. I'd add AI uses to the existing review and catalog, not build a separate process from scratch.”

Red flag to avoid:

Treating AI as out of scope for governance, or proposing a blanket ban with no path to approved use.

They may ask next:
  • How would you handle a deletion request for data that's already in a trained model?
  • Who should approve a new AI use of customer data?
Say it in 60 seconds

Access Control

Medium Technical round Mid-level, Senior Practice question

27. How do column masking and row-level security help you share a table safely? Give an example of each.

What the interviewer is really testing:
Whether you can share data with more people by controlling what each sees, instead of making separate copies.
Answer frame:

Column masking: sensitive columns show a masked or partial value unless the user is allowed.

Row filters: users only see rows that match their attributes, like their own region.

Why: one governed table instead of many copies with different content.

Sample spoken answer:

“Both let many people use the same table while each sees only what they should. Column masking changes what a sensitive column returns based on who's asking. An HR group sees full ID numbers, while everyone else sees a masked value, or maybe just the last four digits. Row-level security filters rows: a regional sales manager querying the orders table only gets rows for their region, without needing a separate table. The big win is fewer copies. Without these, teams make a trimmed-down extract for each audience, and every extract is another place personal data can leak and another thing to delete later. Many platforms support both natively. In Databricks Unity Catalog, for example, you attach a masking function to a column and it applies to every query, whatever tool the user connects with.”

Code:
-- Unity Catalog style: mask a column for everyone outside HR
CREATE FUNCTION hr.mask_national_id(id STRING)
  RETURN CASE WHEN is_account_group_member('hr_team') THEN id
              ELSE '*****' END;

ALTER TABLE hr.employees
  ALTER COLUMN national_id SET MASK hr.mask_national_id;
Red flag to avoid:

Solving it by creating a separate copy of the table for every audience.

They may ask next:
  • Is masking enough for data you want to share with an outside partner?
  • How would you test that a row filter really hides the rows it should?
Say it in 60 seconds
Medium Behavioral round Mid-level, Senior Practice question

28. Tell me about a time a governance control you put in place slowed people down too much. How did you find out, and what did you change?

What the interviewer is really testing:
Whether you treat governance as a service that must be easy to follow, and can admit and fix your own mistakes.
Answer frame:

The control: what it was and why it made sense at the time.

The signal: how you learned it was hurting, with evidence.

The change: how you kept the protection but removed the friction.

Sample spoken answer:

“After a small leak of customer data at my last company, I set up a rule that every access request to customer tables needed the data owner's personal approval. It felt safe. A few months later I noticed analysts were exporting old extracts instead of asking for access, and requests were taking over a week because the owner was travelling. So my control had made things less safe, not more. I looked at the requests and saw most were for the same few data sets for standard reporting. We changed it: a pre-approved group for the common reporting data, with personal columns masked by default, and owner approval only for unmasked access. The owner delegated routine approvals to a steward. Requests dropped to a day, the stray extracts stopped, and the sensitive columns were better protected than before. I learned to check how a control actually changes behaviour, not just what it says.”

Red flag to avoid:

Never having changed a control, or telling a story where people are blamed for working around it.

They may ask next:
  • How do you now check whether a new control is working before rolling it out widely?
  • Who did you involve before changing the approval rule?
Say it in 60 seconds

Tools

Medium Technical round Mid-level, Senior Practice question

29. How would you compare tools like Collibra, Microsoft Purview and Databricks Unity Catalog when choosing one for a company?

What the interviewer is really testing:
Whether you choose a tool from the problem and the data estate, and understand what each type of tool is strongest at.
Answer frame:

Start from needs: what problems, which platforms, who the users are, business or technical.

Know their angle: business-led catalog and workflows, a governance suite tied to one cloud and office ecosystem, or governance built into one data platform.

Test: a short proof of concept on real, messy sources and real users.

Sample spoken answer:

“I'd start with the problem and the estate, not the tool. Are we mainly trying to get business definitions and stewardship working, or to control access and lineage inside our data platform? And where does the data live? Broadly, Collibra is strong as a business-led catalog with a glossary, stewardship roles and approval workflows across many sources. Microsoft Purview fits well when a company is mostly on Azure and Microsoft 365, since it scans sources into a data map, classifies sensitive data and ties in with sensitivity labels. Unity Catalog is built into Databricks, so it governs access, lineage and auditing for data and AI assets on that platform very tightly, but it isn't meant to catalog everything outside it. Many companies combine a platform-level tool with an enterprise catalog. I'd shortlist two, run a proof of concept on real sources with real stewards, and check connectors, cost and effort to run.”

Red flag to avoid:

Picking a tool by brand before defining the problem, or claiming one tool does everything equally well.

They may ask next:
  • What would make you use two governance tools together, and how would you avoid duplicated work?
  • What would you test in the proof of concept that a vendor demo won't show you?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

30. How does Databricks Unity Catalog organise and secure data? Walk me through the object model and how permissions work.

What the interviewer is really testing:
Whether you've worked with governance built into a data platform and understand hierarchy, inheritance and auditing.
Answer frame:

Hierarchy: a metastore holds catalogs; catalogs hold schemas; schemas hold tables, views, volumes, functions and models.

Permissions: granted in SQL to groups; privileges on a parent are inherited by its children.

Beyond grants: lineage and audit logs captured by the platform, tags, row filters and column masks.

Sample spoken answer:

“Unity Catalog gives you a hierarchy with a three-level name for data: catalog, schema, then table, so something like sales.orders.daily_totals. Above catalogs sits the metastore, which is the top container for a region. Schemas can hold tables and views, and also volumes for files, functions and machine learning models, so data and AI assets are governed in one place. Permissions are granted with SQL, like GRANT SELECT on a schema to a group, and they're inherited downwards, so a grant on a catalog applies to everything in it. That makes it sensible to organise catalogs by domain or environment. It also captures lineage for queries and jobs that run on the platform, keeps audit logs of who accessed what, and supports tags, row filters and column masks for finer control. I'd grant to groups, never individuals, and keep production catalogs locked down.”

Code:
GRANT USE CATALOG ON CATALOG sales TO `sales_analysts`;
GRANT USE SCHEMA ON SCHEMA sales.orders TO `sales_analysts`;
GRANT SELECT ON SCHEMA sales.orders TO `sales_analysts`;
Red flag to avoid:

Granting privileges to individual users, or not knowing that grants on a parent flow down to its children.

They may ask next:
  • How would you lay out catalogs for development, test and production?
  • What can't Unity Catalog lineage see?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card