Most data governance interviews start with the basics, like owners versus stewards and what a catalog is for, then test data quality rules, lineage, master data, privacy classification and access control. Many also ask how tools such as Collibra, Microsoft Purview or Unity Catalog fit, and nearly all include a messy scenario, like two teams reporting different revenue. This page is written for anyone facing a data governance round, whether you're moving in from analytics or engineering or already run a governance programme. Each question shows what the interviewer is checking, the shape of a strong answer and a short answer to say out loud. Swap in your own stories.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Foundations
| Governance | who decides, who is accountable, and which rules apply to data. |
|---|---|
| Management | the day-to-day work of storing, moving, cleaning and securing data under those rules. |
Why it matters: people trust and reuse data because definitions, owners and controls are clear.
“I think of data governance as the rules of the road for data. It decides who owns each data set, who can make decisions about it, what the standard definitions are, and which policies apply, like quality thresholds, privacy handling and access. Data management is the actual driving: engineers building pipelines, admins running databases, analysts cleaning data. Governance doesn't do that work itself; it sets the decisions and checks they're followed. A simple way I put it is that governance answers who decides and how we'll know it's right, and management does the work. When governance is working, people spend less time arguing about which number is correct and more time using the data, and the company can show a regulator or auditor that sensitive data is handled properly.”
Describing governance as buying a catalog tool, or as a central team that controls every data change.
Ownership & Stewardship
| Owner | a senior business person accountable for the data: approves access, sets quality expectations, signs off definitions. |
|---|---|
| Steward | the hands-on expert who maintains definitions, triages quality issues and answers questions. |
| Custodian | usually IT or engineering; runs the systems, backups and technical controls. |
“The data owner is accountable. It's usually a business leader, say the head of sales for customer account data, who decides who gets access, what good quality means and which definition is official. They don't do the daily work. The data steward does. That's someone who knows the data deeply, keeps the glossary terms and quality rules up to date, investigates issues and is the first person people ask. The custodian looks after the technical side: the database team or platform engineers who store the data, run backups and apply the access controls the owner approved. The key point is that ownership sits with the business, because they understand what the data means and carry the risk if it's wrong. IT keeps it safe and running, but shouldn't be deciding what 'active customer' means.”
Saying IT owns the data because it sits in their systems.
Metadata & Glossary
| Glossary | business terms and their agreed meaning, like 'active customer', with an owner. |
|---|---|
| Dictionary | technical detail of tables and columns: names, types, allowed values. |
| Catalog | a searchable inventory of data assets that links glossary terms to the physical columns, plus owners, lineage and usage. |
“A business glossary is about meaning. It says what 'active customer' or 'net revenue' means in plain language, who owns that definition, and any rules around it. A data dictionary is technical. It describes a table's columns, their data types, formats and allowed values. A data catalog is the searchable inventory that pulls it together: it lists the data assets across our systems, links each column to the glossary term it represents, and shows the owner, lineage, quality score and who uses it. So if an analyst searches for 'active customer', a good catalog shows the approved definition and then points to the exact tables and columns that hold it. Without that link, you get a nice glossary that nobody can connect to real data, or a list of tables nobody understands.”
Treating the three as the same thing, or saying a catalog is just a list of tables.
Technical: schemas, data types, table names, file locations.
Business: definitions, owners, sensitivity class, approved uses.
Operational: load times, row counts, job runs, freshness, access and usage logs.
“I usually group metadata into three types. Technical metadata describes structure: that orders has a column order_date of type date, where the table lives and how it's partitioned. Business metadata describes meaning and accountability: that order_date means the date the customer placed the order, not when it shipped, that the owner is the head of e-commerce, and that the table is classed as internal. Operational metadata describes what happens to the data: when it last loaded, how many rows arrived, whether the job failed, and who queried it last week. Governance needs all three. Technical metadata tells you what exists, business metadata tells you what it means and who to ask, and operational metadata tells you whether you can trust it today and who might be affected if something breaks.”
Saying metadata is only column names and types.
Catalog & Lineage
Definition: where data came from, what transformed it and where it goes next.
Table vs column: table-level shows which tables feed which; column-level shows exactly which fields build a given field.
Uses: impact analysis before a change, root cause when a number is wrong, proof for auditors.
“Data lineage is the map of how data flows: from the source system, through each transformation, into the tables and reports people use. Table-level lineage tells me that the revenue dashboard reads from a sales summary table, which is built from orders and refunds. Column-level lineage goes further and shows that the net_revenue field is calculated from order_amount minus refund_amount, and which source columns those came from. I use lineage three ways. Before a change, for impact analysis: if a source field is renamed, what breaks downstream? When a number looks wrong, to trace it back to the step that caused it. And for audit, to prove where a reported figure or a piece of personal data came from. Column-level is much more useful, but it's also harder to capture completely.”
Describing lineage as a diagram someone draws once in a slide deck.
Data Quality
Content: accuracy, completeness, validity.
Consistency and uniqueness: same value across systems, no duplicate records.
Timeliness: data is fresh enough for its use.
“The six I use most are accuracy, completeness, validity, consistency, uniqueness and timeliness. Accuracy is whether the value matches reality, like a customer's phone number actually reaching them. Completeness is whether required values are there, say ten thousand orders with no delivery postcode. Validity is whether values follow the rules: a date of birth in the future or a country code that isn't on the approved list. Consistency means the same fact agrees across systems, so a customer isn't 'closed' in billing and 'active' in the CRM. Uniqueness means no duplicates, like the same supplier set up three times. Timeliness is whether data arrives in time for its use: yesterday's sales that only land at noon are no good for the 9 a.m. meeting. Naming them helps turn a vague 'the data is bad' into a rule you can measure.”
Listing the names with no example of how you'd test any of them.
Master Data
| Master data | core business entities shared across systems: customers, products, suppliers, employees. |
|---|---|
| Reference data | small, controlled code lists: country codes, currencies, units, status values. |
| Transactional data | events that reference them: orders, payments, shipments. |
“Master data describes the core things the business deals with, like customers, products, suppliers and locations. It changes slowly and is used by many systems, so getting one agreed version matters a lot. Reference data is the set of allowed values used to categorise other data: country codes, currency codes, order status values, units of measure. It's small and often comes from outside standards, like ISO country codes, but it has to be the same everywhere or joins and reports break. Transactional data records events: an order, a payment, a delivery. It's high volume and each record points to master and reference data, so an order points to a customer, a product and a currency. Governance-wise, master data needs owners and match rules, reference data needs a single controlled list with change approval, and transactional data needs quality checks at volume.”
Calling everything in the database master data.
Privacy & Classification
Levels: a few clear ones, often public, internal, confidential, restricted.
Handling rules: each level maps to controls: who can access, encryption, sharing, retention.
Apply it: tag at column or data set level, owners confirm, highest level wins when mixed.
“I'd keep it to four levels, because people can't apply ten. Public is anything already published, like the website. Internal is the default for business data that isn't meant outside but wouldn't do much harm if leaked, like an org chart. Confidential is data that could hurt the company or people if exposed, like customer contact details or contracts. Restricted is the most sensitive: things like health data, government ID numbers, card numbers or passwords. A scheme only matters if each level changes what happens, so I'd attach handling rules: restricted data is encrypted, masked by default, access is approved by the owner and reviewed regularly, and it can't be copied to personal devices. I'd tag at column level where possible, have owners confirm the tags, and when a table mixes levels, the table takes the highest one.”
Listing levels with no handling rules attached to them.
Data Quality
Define precisely: which table, which column, what counts as valid, which records are in scope.
Measure: a query that counts passing and failing rows, run on a schedule.
Act: a threshold agreed with the owner, and a clear route for failures.
“I'd first pin down the words with the data owner. Which customers count: all, or only active ones? What is 'valid': not empty and on the approved country code reference list? Then the rule becomes measurable: for active customers, country_code is not null and exists in the reference table. I'd write it as a query that counts total rows and failing rows, so we get a pass rate, and run it on every load. Next, a threshold agreed with the owner, say above ninety-nine in a hundred is green, and a response for red: an alert to the steward, with the failing records listed so they can be fixed at the source. Finally the rule goes into the catalog next to the column, with its owner and threshold, so anyone using the data can see how it's doing.”
-- Completeness and validity of country for active customers
SELECT
COUNT(*) AS total_rows,
SUM(CASE WHEN c.country_code IS NULL THEN 1 ELSE 0 END) AS missing_country,
SUM(CASE WHEN c.country_code IS NOT NULL AND r.code IS NULL
THEN 1 ELSE 0 END) AS invalid_country
FROM customers c
LEFT JOIN ref_country r ON r.code = c.country_code
WHERE c.status = 'ACTIVE';
Writing the rule without agreeing scope and thresholds with the business, or having no plan for what happens when it fails.
Symptom: what people noticed and why it mattered.
Trace: how you followed the data back, using lineage, samples and people.
Fix: the change at the source, plus a rule so you'd see it again early.
“In a previous role, our customer churn report showed a sudden jump in closed accounts one month. People were worried. I pulled a sample of the closed accounts and noticed most were closed and reopened on the same day. Using the lineage in our catalog, I followed the status field back to the CRM, then talked to the service team. It turned out a new process for changing a customer's billing address closed the old account and opened a new one, because that was the only way the screen allowed it. So they weren't churned customers at all. The fix had two parts. The service team and the CRM owner changed the process so addresses could be updated in place. And I added a quality rule that flagged accounts closed and reopened within a week, reported to the steward. Churn went back to normal and leadership trusted the report again.”
A story that ends with patching the report and never touching the source process.
Access Control
RBAC: access granted to roles or groups, like 'finance analysts can read the finance schema'.
ABAC: a policy checks attributes of the user, the data and the context, like region or sensitivity tag.
Why ABAC scales: one policy on a tag covers every new table carrying it, instead of hundreds of role grants.
“Role-based access control grants permissions to a role or group, and people get access by joining it. It's simple to understand and audit, but as the data estate grows you end up with hundreds of roles, one per team per region per data set, and nobody can say what anyone can see. Attribute-based access control writes policies on attributes instead: the user's department and region, the data's classification tag, maybe the purpose. So one policy can say restricted columns are masked unless the user is in HR, and it applies automatically to every new table tagged restricted. That's why governance teams like it: classification becomes the thing that drives access. In practice most companies use both. Roles for broad access, attributes for sensitive data and row-level rules like 'managers only see their own region'.”
Saying one model is always better, or not connecting attributes to data classification.
Request: through one process, stating the data set and the purpose.
Approve: the data owner approves, access goes to a group, with an end date where possible.
Review: regular recertification by owners, and automatic removal when people move or leave.
“I'd have one front door for access requests, ideally from the catalog page of the data set, where the person says what they need and why. The data owner, or a steward they've delegated to, approves it, and the grant goes to a group rather than to the individual, so it's easy to see and remove. For restricted data I'd add an end date, so access expires unless it's renewed. Over time access creeps: people change teams and keep old permissions. So I'd run a recertification every quarter or half-year where owners see who has access to their data and confirm or remove each person. Joiner, mover and leaver changes from HR should remove access automatically. And everything is logged, so we can answer an auditor's question about who could see this data on a given date.”
Granting access directly to individuals with no expiry and no review.
Foundations
Outcomes: fewer incidents, less time spent reconciling, faster audits, faster access to data.
Health: quality scores on critical data, share of critical assets with an owner and certified description.
Adoption: catalog use, repeat users, requests handled on time.
“I'd avoid counting activity, like how many terms we've added, because that can go up while nothing improves. I'd track three layers. First, business outcomes: the number of data incidents that reached a report, how long month-end reconciliation takes, how quickly someone gets approved access to a data set, and how smoothly the last audit went. Second, health of the critical data: quality scores on the key elements over time, and how many of them have a named owner, a steward and a certified definition. Third, adoption: how many people use the catalog each month and come back, and whether access and issue requests are handled within the agreed time. I'd set a baseline before we start, report a short scorecard to the governance council each quarter, and tie at least one number to something a leader already worries about.”
Only naming activity metrics like the count of policies written or tables scanned.
Centralised: one team sets and enforces everything; consistent but slow and a bottleneck.
Federated: domains own their data and rules within a few shared standards; scales but needs mature domains.
Hybrid: a small central team owns standards, tooling and shared definitions; domains own their data and stewards.
“A centralised model puts one team in charge of standards, definitions and approvals. It's consistent and good for a small company or a heavily regulated start, but it becomes a bottleneck and business teams stop feeling responsible. A federated model pushes ownership to domains like sales, finance and supply chain, with only a thin set of shared rules. That scales and fits ideas like data mesh, where standards are built into the platform, but it needs domains that are mature enough to own their data. Most large companies I'd steer toward a hybrid. A small central office owns the policies, the tooling, the shared terms like customer and revenue, and reporting to a governance council. Each domain has an owner and stewards who do the actual work. I'd start more central while maturity is low and hand more to domains as they prove they can run it.”
Choosing one model as always right without looking at company size, regulation and maturity.
Listen: meet leaders and data users; find the pains, risks and regulatory pressure.
Pick a pilot: one domain or one critical report with a willing sponsor and owner.
Show value: a few definitions, quality rules and access fixes, measured and shared.
“In the first month I'd mostly listen. I'd meet the executives, the data team and the heaviest data users, and ask where data costs them time, money or sleep: numbers that don't match, audits, customer complaints, slow access. I'd also check what regulation we're under. In month two I'd pick one pilot with a sponsor who cares, maybe the monthly revenue report or customer data. For that domain I'd name an owner and a steward, agree the key definitions, set a handful of quality rules on the critical fields and tidy up who has access. I'd keep policy light: a short charter, the roles, and how decisions get made. In month three I'd show results against the baseline, set up a small council, and plan the next two domains. The goal is that people ask to be next, not that I've written a thick framework.”
Starting with a long policy document or a tool purchase before understanding the business problems.
Ownership & Stewardship
Show the pain: tie the bad data to a cost the business already feels, like late payments or duplicate suppliers.
Pick by process: the leader whose process creates or depends most on the data is the natural owner.
Make it small: a clear, time-boxed role description and a sponsor to back it.
“First I'd find where bad supplier data is already hurting someone. Maybe invoices bounce because bank details are wrong, or we have the same supplier three times and lose volume discounts. I'd put a few real examples in front of the procurement and finance leads. Then I'd propose ownership based on who creates the data, which is usually procurement, since they onboard suppliers. The ask would be small and specific: the owner approves the definition of an active supplier, agrees two or three quality rules and reviews a monthly score. The steward would be someone on their team who already fixes these records informally, and we'd make that official with a bit of their time set aside. I'd also get our executive sponsor to confirm it, so it isn't just me asking. Then I'd show early progress so the role feels worth it.”
Assigning an owner by email with no business case, or making the governance team the owner by default.
Situation: the data, the team and why they pushed back.
Action: how you showed the cost, shrank the ask and gave them control.
Result: what changed, with a before and after.
“At my last company, the marketing team saw campaign and lead data as IT's problem, even though they set up every campaign. Reports kept splitting one campaign into several because of inconsistent naming, so nobody trusted the lead numbers. The marketing head's view was, fix it in the database. I sat with one of her managers and showed how much of their weekly reporting time went on cleaning names by hand, and that the leads from their biggest campaign were being undercounted. That got attention. Then I made the ask small: they'd own a naming convention for campaigns, and one person would be steward for an hour a week. I built a simple check that flagged badly named campaigns the day they were created. Within two months the manual cleaning stopped, and the marketing head started presenting the lead numbers herself. That's when I knew they owned it.”
A story where you force ownership through an escalation, or quietly do all the work yourself.
Metadata & Glossary
Find the real differences: timing, refunds, discounts, tax, currency, which deals count.
Decide who decides: the owner of the term, often finance for reported revenue, with the council as tie-breaker.
Name and publish: one official term, other measures renamed, definitions in the glossary, certified sources.
“I'd first reconcile the two numbers line by line for one month. Usually both are right for their purpose. Sales might count bookings when a contract is signed, at list price, while finance counts recognised revenue after discounts, refunds and tax, following accounting rules. Once the differences are clear, it stops being an argument about who's wrong. Then decision rights: the owner of the official 'revenue' term is usually finance, because it ties to the accounts, and our governance council breaks ties if needed. Sales doesn't lose its measure. It gets renamed to something honest, like 'bookings', with its own definition. Both go in the glossary with owners and calculation rules, each points to a certified data set, and dashboards get relabelled. I'd also add a small reconciliation view that shows how bookings turn into revenue, so the leadership team sees why the two differ.”
Picking one number as correct straight away, or forcing one definition on everyone without a decision owner.
Catalog & Lineage
Diagnose: look at search logs and ask analysts why they skip it.
Curate the top assets: describe, certify and assign owners for the most-used data sets first.
Put it in the flow: link it from dashboards and tools people already use, and make it the answer to common questions.
“I'd start by finding out why. Usually it's because the catalog is full of scanned tables with no descriptions, so searching it doesn't answer anything. I'd pull usage logs to find the twenty or so data sets people actually query most, and get their owners and stewards to write clear descriptions, link the glossary terms and mark the trusted ones as certified. That small, high-value slice is worth more than ten thousand empty entries. Then I'd put the catalog where people already are: links from dashboards back to the certified source, and a habit in the analyst channel of answering 'where do I find X' with a catalog link. I'd track searches, certified asset views and repeat users each month. And I'd retire or hide junk entries, because noise is what teaches people the catalog is useless.”
Proposing a company-wide training session and a mandate, with no curation of content.
Common gaps: hand-written scripts, spreadsheets, dynamic SQL, file drops and tools the scanner can't read.
Prioritise: close gaps on critical data elements and regulated reports first.
Fill in: connectors or parsers where possible, manual lineage with an owner where not.
“Automated lineage is only as good as what the tool can read. The holes usually come from custom Python or shell scripts, stored procedures with dynamic SQL, spreadsheets someone emails around, files dropped on a shared folder, and third-party tools with no connector. The result is a lineage graph that just stops, or a report that seems to appear from nowhere. I wouldn't try to fix everything. I'd pick the critical data elements, the numbers in board reports and regulatory returns, and trace them end to end. Where there's a connector or parser we're not using, turn it on. Where there isn't, document the step manually in the catalog with an owner and a review date. And where the gap is a spreadsheet in the middle of a critical flow, that's often a sign the process itself should change, not just the lineage.”
Assuming the tool captures everything, or trying to map every data set before any critical one is done.
Master Data
Registry: sources keep their records; the hub holds matching keys and builds a view when asked.
Consolidation: the hub builds a golden record for reporting, but doesn't write back.
Coexistence and centralised: the golden record is synced back to sources, or the hub becomes where records are created.
“There are four common styles. Registry leaves the data in the source systems and only keeps a cross-reference in the hub: these three records are the same customer. It's the least disruptive and quick to start, but you don't really fix the sources. Consolidation copies records into the hub and builds a golden record, mainly for analytics, without sending it back. Coexistence also builds the golden record but syncs it back to the sources, so operational systems get cleaner too. Centralised, sometimes called transactional, makes the hub the only place records are created and changed, with approval workflows, which gives the most control but changes how people work. I'd pick registry when there are many systems we can't change easily and the first goal is a single view. I'd move toward centralised when bad master data is causing real operational pain, like failed shipments or payments.”
Knowing only one style, or suggesting a centralised hub for every case without mentioning how disruptive it is.
Match: standardise first, then compare with exact and fuzzy rules to score likely duplicates.
Review: auto-merge above a high score, send the grey zone to stewards, keep apart below.
Survivorship: per attribute: most trusted source, most recent, most complete.
“Matching starts with standardising: same case, trimmed spaces, standard address formats, phone numbers in one format. Then match rules compare records, some exact like a tax ID, some fuzzy like a similar name at the same postcode, and give each pair a score. I'd set two thresholds: above the high one we merge automatically, in the middle a steward reviews, and below we keep them apart. False merges are worse than missed ones, because you might combine two real customers and show one their neighbour's orders. Once records are grouped, survivorship picks the best value for each attribute, not the whole record. So the legal name might come from the finance system as the most trusted source, the email from the most recently updated record, and the phone from whichever is complete and valid. We keep the cross-reference to every source record so merges can be undone.”
Picking one whole source record as the winner, or never mentioning the risk of false merges.
Privacy & Classification
Scan: automated classifiers look at column names and sample values for patterns like emails, IDs, card numbers.
Confirm: stewards review matches, since scanners produce false positives and miss free text.
Record and act: tag it in the catalog, then apply masking, access and retention rules.
“I'd start with automated discovery, because hundreds of sources can't be checked by hand. Most catalog and governance tools can scan sources and flag likely personal data from column names and from patterns in sample values, like email formats, phone numbers, card numbers that pass a checksum, or national ID formats. I'd prioritise sources by risk: customer-facing systems, HR and anything shared with third parties first. Scanners get it wrong both ways. A column called notes can hold phone numbers and names, and a column called id might look like a card number but isn't. So stewards review the results and confirm the tags. Every confirmed field goes into the catalog with its classification, and that drives the next steps: masking, tighter access, retention rules and an entry in the record of processing if the law where we operate needs one.”
Relying only on column names, or only on a one-off manual survey.
Principles: lawful basis, purpose limitation, data minimisation, accuracy, storage limitation, security.
Controls: purpose tags on data sets, collect only needed fields, retention schedules, access and masking.
Rights: processes to find, export, correct and delete a person's data on request.
“Most privacy laws share a core set of principles, whether it's the GDPR in Europe or similar laws elsewhere, though the details differ, so I always check the specifics with legal. The ones that shape my work are: have a valid reason for using the data, use it only for the purpose it was collected for, collect no more than needed, keep it accurate, keep it no longer than needed, and protect it. Governance turns those into controls. Data sets carry a purpose tag, and a new use needs a check before access is granted. New collections go through a review asking whether each field is needed. Retention schedules are set per data type and enforced by deletion jobs. And we need a reliable way to find everything about one person across systems, which is where the catalog and lineage pay off, so we can answer access and deletion requests on time.”
Treating privacy as only a security problem, or naming one country's law as if it applied everywhere.
Schedule: per data type, the legal minimum, business need and a maximum, agreed with legal and owners.
Enforce: automated deletion or anonymisation jobs, including copies and extracts.
Exceptions and proof: legal holds pause deletion; logs show what was deleted and when.
“I'd build a retention schedule by data type with legal and the data owners. Some records have a minimum period, like financial records that tax rules say must be kept for a number of years, and personal data has a maximum: no longer than we need it. Each entry says the period, what starts the clock, like account closure, and what happens at the end, delete or anonymise. Writing it down is the easy part. The hard part is copies: extracts in analysts' folders, test environments, old backups. So I'd use lineage to find where each data type flows, automate deletion jobs on the main stores, and set backups to expire on a known cycle. Legal holds must be able to pause deletion for specific records. And I'd keep deletion logs, because an auditor will ask for proof, not just the policy.”
Saying 'keep everything forever, storage is cheap' or ignoring copies and backups.
Purpose and consent: is this data allowed to be used for this new purpose?
Sensitivity and access: classified data needs the same controls when it flows into prompts, indexes or training sets.
Provenance: record which data built which model or index, and keep it current.
“The basics don't change, but the risks move. First, purpose: data collected for billing may not be allowed for training a model, so a new AI use needs the same purpose check as any new use, and contracts with customers might restrict it. Second, sensitivity and access. If a document search assistant indexes everything on the file shares, it can show people content they could never open directly, so it must respect the same permissions, and restricted data should be kept out or masked. Third, provenance: we need lineage from source data to each training set, model or search index, so we can answer what it learned from, and remove data when a deletion request or retention rule applies. Quality matters more too, because a model repeats bad data confidently. I'd add AI uses to the existing review and catalog, not build a separate process from scratch.”
Treating AI as out of scope for governance, or proposing a blanket ban with no path to approved use.
Access Control
Column masking: sensitive columns show a masked or partial value unless the user is allowed.
Row filters: users only see rows that match their attributes, like their own region.
Why: one governed table instead of many copies with different content.
“Both let many people use the same table while each sees only what they should. Column masking changes what a sensitive column returns based on who's asking. An HR group sees full ID numbers, while everyone else sees a masked value, or maybe just the last four digits. Row-level security filters rows: a regional sales manager querying the orders table only gets rows for their region, without needing a separate table. The big win is fewer copies. Without these, teams make a trimmed-down extract for each audience, and every extract is another place personal data can leak and another thing to delete later. Many platforms support both natively. In Databricks Unity Catalog, for example, you attach a masking function to a column and it applies to every query, whatever tool the user connects with.”
-- Unity Catalog style: mask a column for everyone outside HR
CREATE FUNCTION hr.mask_national_id(id STRING)
RETURN CASE WHEN is_account_group_member('hr_team') THEN id
ELSE '*****' END;
ALTER TABLE hr.employees
ALTER COLUMN national_id SET MASK hr.mask_national_id;
Solving it by creating a separate copy of the table for every audience.
The control: what it was and why it made sense at the time.
The signal: how you learned it was hurting, with evidence.
The change: how you kept the protection but removed the friction.
“After a small leak of customer data at my last company, I set up a rule that every access request to customer tables needed the data owner's personal approval. It felt safe. A few months later I noticed analysts were exporting old extracts instead of asking for access, and requests were taking over a week because the owner was travelling. So my control had made things less safe, not more. I looked at the requests and saw most were for the same few data sets for standard reporting. We changed it: a pre-approved group for the common reporting data, with personal columns masked by default, and owner approval only for unmasked access. The owner delegated routine approvals to a steward. Requests dropped to a day, the stray extracts stopped, and the sensitive columns were better protected than before. I learned to check how a control actually changes behaviour, not just what it says.”
Never having changed a control, or telling a story where people are blamed for working around it.
Tools
Start from needs: what problems, which platforms, who the users are, business or technical.
Know their angle: business-led catalog and workflows, a governance suite tied to one cloud and office ecosystem, or governance built into one data platform.
Test: a short proof of concept on real, messy sources and real users.
“I'd start with the problem and the estate, not the tool. Are we mainly trying to get business definitions and stewardship working, or to control access and lineage inside our data platform? And where does the data live? Broadly, Collibra is strong as a business-led catalog with a glossary, stewardship roles and approval workflows across many sources. Microsoft Purview fits well when a company is mostly on Azure and Microsoft 365, since it scans sources into a data map, classifies sensitive data and ties in with sensitivity labels. Unity Catalog is built into Databricks, so it governs access, lineage and auditing for data and AI assets on that platform very tightly, but it isn't meant to catalog everything outside it. Many companies combine a platform-level tool with an enterprise catalog. I'd shortlist two, run a proof of concept on real sources with real stewards, and check connectors, cost and effort to run.”
Picking a tool by brand before defining the problem, or claiming one tool does everything equally well.
Hierarchy: a metastore holds catalogs; catalogs hold schemas; schemas hold tables, views, volumes, functions and models.
Permissions: granted in SQL to groups; privileges on a parent are inherited by its children.
Beyond grants: lineage and audit logs captured by the platform, tags, row filters and column masks.
“Unity Catalog gives you a hierarchy with a three-level name for data: catalog, schema, then table, so something like sales.orders.daily_totals. Above catalogs sits the metastore, which is the top container for a region. Schemas can hold tables and views, and also volumes for files, functions and machine learning models, so data and AI assets are governed in one place. Permissions are granted with SQL, like GRANT SELECT on a schema to a group, and they're inherited downwards, so a grant on a catalog applies to everything in it. That makes it sensible to organise catalogs by domain or environment. It also captures lineage for queries and jobs that run on the platform, keeps audit logs of who accessed what, and supports tags, row filters and column masks for finer control. I'd grant to groups, never individuals, and keep production catalogs locked down.”
GRANT USE CATALOG ON CATALOG sales TO `sales_analysts`;
GRANT USE SCHEMA ON SCHEMA sales.orders TO `sales_analysts`;
GRANT SELECT ON SCHEMA sales.orders TO `sales_analysts`;
Granting privileges to individual users, or not knowing that grants on a parent flow down to its children.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.