Who should approve access requests to a domain's sensitive dataset, and how do you keep approvals fast without making them meaningless?
answer
- accountability sits with the data owner
- tier the data by sensitivity
- policy-based auto-approval for low tiers
- human review where harm is high
- measure lead time and revocations
basics
~20 sThe data owner should be accountable for access to sensitive data, but not click every request. Tier data by sensitivity, auto-approve low tiers by policy, require owner review with a stated purpose for high tiers, make grants expire, and track lead time.
solid answer
~50 sAccountability belongs with the **data owner** — the team that understands the data and answers for its misuse — not with a central platform team that cannot judge business need. But routing every request to owners makes them either a bottleneck or a rubber stamp. I would **tier data by sensitivity**. For low tiers, access is **policy-approved automatically** when the requester's attributes match (same department, completed training), and logged. For high tiers, the owner reviews a request that states a **purpose** and a duration, with an option to grant **masked or aggregated** access instead of raw. All grants **expire**. The platform team runs the workflow and measures **lead time**, approval and revocation rates; a high-sensitivity tier with 100% approval is a sign reviews are meaningless, and a long lead time means people will find workarounds.
go deeper
Know that access to sensitive data should be approved by someone accountable for that data.
Explain why data owners, rather than a central team, judge need, and why grants should expire.
Design a tiered workflow with automatic policy approval, purpose statements and least-privileged alternatives.
Own the trade-off between speed and meaningful review, set tier definitions and lead-time targets, and use approval and revocation data to adjust them.
## The tension Access decisions need two things that pull apart: - **Judgement**: does this person have a legitimate need for this data? Only the owner of the data can answer. - **Speed**: analysts blocked for two weeks export data, reuse colleagues' credentials or build shadow copies — which is worse for governance than a slightly generous grant. A central security or platform team approving everything is fast to set up but cannot judge need. Owners approving everything by hand have the judgement but not the time. There is no single right answer; the design is about **where to spend human judgement**. ## A tiered model | Sensitivity tier | Example data | Approval | Grant form | |---|---|---|---| | Internal | product usage aggregates | automatic for employees, logged | standing, reviewed yearly | | Confidential | customer-level behaviour, masked identifiers | automatic when attributes match policy (department, training), otherwise owner | expires in months | | Restricted | raw personal data, payroll, health | owner review with stated purpose and duration | expires in weeks; masked alternative offered first | ## Making approvals meaningful 1. **Require a purpose** for restricted access, recorded with the grant; it makes the decision reviewable later. 2. **Offer the least-privileged alternative**: masked, aggregated or synthetic data often satisfies the need. 3. **Expire grants** so approval is not permanent by default. 4. **Separate duties**: the requester's manager confirms the need; the data owner confirms the data is appropriate. 5. **Watch the numbers**: an approval rate near 100% for restricted data, or approvals decided in seconds, suggest a rubber stamp. ## Making approvals fast - Put the request **inside the catalog**, on the dataset's page, pre-filled with the options available. - Automate everything policy can decide, so owners see only the requests that need judgement. - Set a **target lead time** per tier and publish it; escalate stale requests. - Give owners **delegates** so absence does not block the queue. ## Signals that the model works - Lead time for confidential access measured in hours, for restricted in days. - Falling volume of unmanaged exports and shared credentials. - Revocations at expiry not immediately re-requested, showing grants matched real need. ## Why interviewers ask it It is a genuine organisational trade-off. Principal candidates are expected to place **accountability** correctly, spend review effort **where harm is highest**, and measure both **speed** and **meaningfulness** rather than optimising one.
- Why not let the central platform team approve all access?It can check that a request follows the process, but it cannot judge whether the requester needs this data for a legitimate purpose. Approvals then become procedural, and the accountable owner learns about access to their data only after something goes wrong.
- What would you do if restricted-tier requests are approved 99% of the time?Sample recent approvals and check the stated purposes and actual usage. Often the tier is too broad, so the data is split into a restricted core and a masked version that can be granted by policy; otherwise owners need better context, such as usage and alternatives, at decision time.
saying these in an interview costs you the question
- Routing every access request to one central team regardless of data
- Granting permanent access to restricted data with no stated purpose
- Measuring only approval speed and ignoring whether reviews mean anything
- Offering raw access when masked or aggregated data would do