In Tableau, what's the difference between publishing a workbook and publishing a data source?
answer
- one upload carries views, one carries data
- ask who else needs these numbers
- how many refresh schedules do you now own
- only one of the two can be certified
basics
~20 sPublishing a workbook uploads the sheets and dashboards, and by default a private copy of their connection and extract. Publishing a data source uploads the connection, model and extract on its own, so many workbooks share one governed, separately refreshed copy.
solid answer
~50 sPublishing to Tableau Server or Tableau Cloud moves two different kinds of content. A **workbook** is the sheets, dashboards, formatting and workbook-level calculations — and unless you change the publish option, it carries its own embedded copy of the connection and, if you built one, its own extract. A **published data source** is only the data layer: the connection details, joins or relationships, data-source filters, field naming, hierarchies and the extract. Once that data source exists on the site, any author can connect to it from Desktop or the web, it refreshes on one schedule, it can be certified, and permissions on it are managed independently of the workbooks that use it. The practical rule: if more than one workbook needs the same numbers, publish the data source separately and point workbooks at it; embedded data is for genuinely single-use analysis.
code
text · 9 lines-- Embedded (default)
Workbook A -> its own extract of orders -> refresh schedule A
Workbook B -> its own extract of orders -> refresh schedule B
Workbook C -> its own extract of orders -> refresh schedule C
-- Published separately
Published data source "Orders (certified)" -> one extract -> one schedule
^ ^ ^
Workbook A Workbook B Workbook Cgo deeper
Be able to state the two content types you can publish and name one consequence of each — one extract shared by many workbooks versus a private copy per workbook.
Explain what actually lives in a published data source: connection, joins or relationships, data-source filters, field metadata, calculations and the extract, plus its own permissions and refresh schedule.
Show the operational judgment: you are trading a dependency for consistency. Talk about how you promote a prototype to a certified source, and how you manage breaking changes to a data source other teams consume.
Own the policy: which projects require certified published data sources, who is allowed to publish them, and how you keep the count small enough to govern rather than ending up with four hundred near-duplicates.
## What publishing actually moves When you publish from Tableau Desktop to Tableau Server or Tableau Cloud, you are creating one of two kinds of content on the site. A **workbook** is everything you see and interact with: worksheets, dashboards, story points, mark cards, formatting, parameters, and the calculated fields defined inside the workbook. It also carries a *data connection* — and here is the fork in the road. In the publish dialog, each data source the workbook uses is listed with a publish type: **embedded in workbook** or **published separately**. Embedded means a private copy of that connection (and its extract, if it has one) is uploaded inside the workbook and belongs to that workbook alone. A **published data source** is the data layer on its own: the connection to the database or file, the joins/relationships that form the model, data-source filters, field renames and folders, default aggregations and formats, hierarchies, calculations defined at the data-source level, and the `.hyper` extract if you chose one. It appears on the site as its own content object, with its own owner, its own permissions and its own refresh schedule. ## Why the difference matters **Refresh.** An embedded extract refreshes on a schedule attached to that workbook. Ten workbooks with embedded extracts of the same table mean ten extracts, ten schedules, and ten chances for one of them to fail or drift. One published data source means one extract, refreshed once, and every workbook connected to it sees the same numbers at the same moment. **Consistency.** Business logic tends to live in the data layer: how you filter out test orders, whether returns are negative rows, what "active customer" means. Put it in a published data source and every author inherits it. Leave it embedded and each author re-implements it slightly differently — the classic "three dashboards, three revenue numbers" complaint. **Governance.** A published data source can be **certified**, which surfaces a badge and promotes it in search so authors find the blessed one first. It has its own permission rules, so you can let a wide group *connect* to a data source without letting them download the underlying rows. None of this is possible for data buried inside a workbook. **Credentials.** Publishing the data source separately also gives you a single place to embed the database credentials used by scheduled refresh, rather than scattering them across workbooks. ## The cost of publishing separately It is not free. The workbook now has a dependency: if someone deletes or renames the published data source, or revokes your access to it, the workbook's views break. A viewer downloading the workbook needs permission on the data source too. Changing the model — dropping a field, changing a data type — can break workbooks you do not own, so a shared data source needs the same change discipline as any other shared interface. And for a genuinely one-off analysis, the ceremony is overhead with no payoff. ## How you choose Ask one question: *will anyone else need these numbers?* If the answer is yes — or if the workbook is going into production where a stale or wrong number will be noticed — publish the data source separately. If it is a personal exploration, a one-time board slide, or a tiny static lookup that will never be reused, embedding is fine and faster. A useful maturity pattern is a two-tier site: authors prototype with embedded connections in a sandbox project, and anything promoted to a production project must connect to a certified published data source. That converts an individual habit into something you can actually audit. ## Publishing mechanics worth knowing At publish time you also choose the **project** the content lands in (which determines its default permissions), the content **name** and description, whether to include external files such as a local Excel workbook, and whether **Show sheets as tabs** is on. Permissions default to the project's rules unless you override them, and on a well-governed site you should not be overriding them by hand. One more asymmetry to remember: publishing a workbook does not publish its data source, but *unchecking* embedded and choosing "published separately" creates both objects in one action — the workbook and the data source it now points at.
- What happens to a published workbook if the data source it connects to is deleted?The workbook's views fail to load — the connection resolves to a server-side object that no longer exists, and there is no silent fallback to a local copy. That is why shared data sources need change discipline: announce breaking changes, and check downstream usage before deleting or renaming. Lineage tooling helps, but the safe habit is to treat a published data source as a public interface.
- When is embedding the data inside the workbook the right call?When the analysis is genuinely single-use: a personal exploration, a one-off deck, or a small static lookup nobody else will consume. Embedding avoids creating another governed object and another refresh schedule to babysit. The moment a second author asks for the same numbers, or the workbook moves into a production project, promote it to a published data source.
- Can a single workbook mix embedded and published data sources?Yes. The publish dialog sets the publish type per data source, so a workbook can connect to a certified published data source for the fact data and keep a small embedded lookup — a manually maintained target list, say — inside the workbook. It is a reasonable pattern, but the embedded piece is invisible to governance, so keep it small and obviously disposable.
An embedded connection is a photocopy stapled inside one report; a published data source is the original in the filing cabinet that every report cites.
saying these in an interview costs you the question
- Thinks a published data source is just a saved connection string
- Assumes every workbook must own its own extract and schedule
- Says publishing a workbook automatically shares its data with other authors
- Believes certification applies to workbooks the same way it applies to data sources
- Forgets viewers need permission on the data source, not just the workbook