skip to content

In a data lake, table-level access rules are enforced by the query engine; why can a user sometimes still read the protected data, and how do platforms close that path?

level: seniorimportance: should knowfreq 40%

answer

  1. the files sit in object storage
  2. storage permissions bypass table rules
  3. one enforcement point per path
  4. engines get scoped, short-lived access
  5. users never touch raw paths

basics

~20 s

Table rules run in the query engine, but the files sit in object storage with separate permissions, so a direct storage reader skips filters and masks. Deny users direct storage access and give only governed engines scoped, short-lived credentials.

solid answer

~40 s

In a lake, a table is **files in object storage** plus metadata. Row filters, column masks and table grants are enforced by the **query engine** when it plans the read. If a user, notebook or job also holds **read permission on the storage location**, it can open the files directly and get every row and column — the engine's policy never runs. This happens easily: broad storage roles for data scientists, a bucket policy granted to "all analysts", a job role reused interactively. Platforms close the path by making the governed engine the **only** reader: users and ad hoc compute get **no direct storage permissions** on governed paths; engines receive **scoped, short-lived credentials** for the specific table after the access decision; and storage-level audit logs are watched for direct reads that bypass the engine.

go deeper

for a junior

Know that lake tables are files in storage, and that storage permissions are separate from table permissions.

for a middle

Explain how direct storage access skips the engine's row filters and column masks.

for a senior

Design the controls that close the bypass, including removed direct access, scoped short-lived credentials and storage log monitoring.

for a principal

Set the platform rule for which engines may read governed data and how new engines are admitted without opening a second door.

## Two layers that both grant access A lakehouse table has two layers: 1. **Storage**: data files (for example columnar files) in an object store, protected by the store's own permissions — bucket policies, identity roles, access-control lists. 2. **Table and policy**: the catalog's view of the table, plus grants, row filters and column masks, enforced by the **query engine** that plans the read. Access is only governed if **every path to the data goes through the second layer**. ## The bypass Row filters and masks are applied while the engine executes a query. A reader that fetches the **files directly** from storage never passes through that step: - a data scientist's notebook role has read on the whole bucket; - a batch job's credentials, granted broad storage access, are reused for interactive exploration; - a legacy bucket policy grants read to an entire analyst group; - a second engine without the policy integration reads the same files. Any of these returns **every row and every column**, regardless of the table rules. The platform looks governed while the data is effectively open. ## Closing the path | Control | Effect | |---|---| | Remove direct storage permissions for users and ad hoc compute on governed locations | the file path is closed for people | | Engines obtain **scoped, short-lived credentials** for the table's files after the access check | the engine can read only what was authorised, only for a while | | One policy integration for every engine that may read governed tables | no second engine becomes an unguarded door | | Separate governed and ungoverned storage locations | clear boundary for permissions | | Monitor storage-level access logs for reads not made by governed engines | detects remaining bypasses | The credential step is often called **credential vending**: the catalog or policy service issues temporary storage credentials scoped to the table after deciding the request is allowed. ## Fine-grained rules need a trusted engine A row filter or column mask can only be enforced by something that reads the rows. If an engine receives raw file access and is **trusted** to apply the policy, it must be an engine under the platform's control. Handing file-level credentials to an arbitrary client for a table with row filters defeats the filter, so such tables are usually served only through governed engines or pre-filtered views. ## Why interviewers ask it It is the most common way lake access control fails in practice. A senior answer identifies the **two layers**, explains why storage permissions **bypass** engine-level rules, and closes the path with removed direct access, scoped credentials and monitoring.

  • Why can a governed engine receive file-level credentials while users cannot?
    Because the engine is part of the platform and applies the row filters and masks before returning results. A user with the same credentials would read the raw files with no policy in between, so file access is limited to components trusted to enforce it.
  • How would you find existing bypasses in an established lake?
    Inventory who holds read permissions on governed storage locations, and compare storage access logs against the identities of governed engines. Any direct read by a person, notebook or unknown role is a bypass to remove or justify.

It is like locking the front door of a records office while the loading dock at the back stays open to anyone holding a building pass.

saying these in an interview costs you the question

  • Believing table-level grants protect files that users can read directly
  • Giving data scientists broad read access to the whole data bucket
  • Letting several engines read governed tables with only one enforcing policy
  • Issuing long-lived storage credentials to clients of filtered tables