skip to content

With the Kubernetes Multi-Cluster Services API, how would you expose a loyalty-points accrual service running in two clusters under one name to a third cluster?

level: seniorimportance: nice to knowfreq 30%

answer

  1. ClusterSet and namespace sameness
  2. export in each serving cluster
  3. implementation creates the import
  4. svc.clusterset.local, not cluster.local
  5. network and DNS not included

basics

~10 s

Create a ServiceExport with the Service's name and namespace in both clusters that run it. An MCS implementation then creates a matching ServiceImport across the ClusterSet, and callers use name.namespace.svc.clusterset.local.

solid answer

~50 s

The Multi-Cluster Services API is a set of CRDs in the `multicluster.x-k8s.io` group, defined by SIG Multicluster. It is not built into kube-apiserver. It works over a **ClusterSet**: a group of clusters that trust each other and treat a namespace name as meaning the same thing in every cluster (**namespace sameness**). In both clusters that run the service, you create a `ServiceExport` with the same name and namespace as the `Service`, for example `loyalty-accrual` in `loyalty`. The implementation's controller combines the exports and creates a `ServiceImport` with that name and namespace in the ClusterSet's clusters. Its type is `ClusterSetIP` (one virtual IP for the whole set) or `Headless`. It also creates EndpointSlices labelled `multicluster.kubernetes.io/service-name` and `multicluster.kubernetes.io/source-cluster`. Callers in the third cluster use `loyalty-accrual.loyalty.svc.clusterset.local`. The `cluster.local` name still resolves only within its own cluster. You still need an MCS implementation, DNS for `clusterset.local`, and pod-to-pod network reachability between the clusters. If the exports disagree, for example on ports, the ServiceExport reports a `Conflict` condition.

code

yaml · 18 lines
yaml
apiVersion: v1
kind: Service
metadata:
  name: loyalty-accrual
  namespace: loyalty
spec:
  selector:
    app: loyalty-accrual
  ports:
  - name: grpc
    port: 9443
    targetPort: 9443
---
apiVersion: multicluster.x-k8s.io/v1alpha1
kind: ServiceExport
metadata:
  name: loyalty-accrual
  namespace: loyalty

go deeper

for a junior

Remember the pair: ServiceExport marks a Service as shared, and ServiceImport is how other clusters see it under the clusterset.local name.

for a middle

Explain the flow: a same-named export, an implementation-created import, the ClusterSetIP or Headless type, and the separate clusterset.local DNS name.

for a senior

Show that you know what MCS leaves to you: an installed implementation, DNS, pod-network reachability, and handling Conflict conditions when exports disagree.

for a principal

Decide whether namespace sameness can hold across your ClusterSet, and whether MCS or a mesh should own cross-cluster discovery for the platform.

## The problem cluster.local cannot solve Inside one Kubernetes cluster, a Service `loyalty-accrual` in namespace `loyalty` resolves as `loyalty-accrual.loyalty.svc.cluster.local`. That name only ever covers the cluster's own endpoints. When the accrual service runs in two regional clusters and a checkout service in a third cluster needs it, the options are: - a hand-maintained external load balancer per cluster; - hard-coded per-cluster hostnames in the caller; - a service mesh's multi-cluster mode (a separate system, not covered here); - the **Multi-Cluster Services (MCS) API**, the Kubernetes-native API for this. ## ClusterSet and namespace sameness MCS defines a **ClusterSet**: a group of clusters under one authority that trust each other. Its key assumption is **namespace sameness**. Namespace `loyalty` in cluster A and `loyalty` in cluster B belong to the same owner and mean the same thing. A Service exported from `loyalty` merges with same-named exports from other clusters instead of colliding with them. This assumption is why MCS fits clusters run by one platform team and fits poorly when unrelated tenants pick namespace names independently. ## The export/import flow 1. The service owner has a normal `Service` in each cluster that runs the workload. 2. In each of those clusters, the owner creates a `ServiceExport` with the **same name and namespace** as the Service. The ServiceExport carries almost no configuration, because the Service already defines it. 3. The **MCS implementation** (a controller you install; the API alone does nothing) checks each export and records conditions: `Valid` (for example `NoService` if the Service is missing), `Ready`, and `Conflict`. 4. The implementation creates a **`ServiceImport`** with the same name and namespace in the ClusterSet's clusters, and fills in EndpointSlices labelled `multicluster.kubernetes.io/service-name` and `multicluster.kubernetes.io/source-cluster`. 5. A DNS component answers **`loyalty-accrual.loyalty.svc.clusterset.local`**, and the node data plane routes traffic to endpoints in any exporting cluster. ```yaml apiVersion: v1 kind: Service metadata: name: loyalty-accrual namespace: loyalty spec: selector: app: loyalty-accrual ports: - name: grpc port: 9443 targetPort: 9443 --- apiVersion: multicluster.x-k8s.io/v1alpha1 kind: ServiceExport metadata: name: loyalty-accrual namespace: loyalty ``` The SIG Multicluster repository serves both `v1alpha1` and `v1beta1` of these CRDs. Which versions a cluster accepts depends on the implementation installed. ## ServiceImport types | `ServiceImport` type | Behaviour | Comparable single-cluster Service | |---|---|---| | `ClusterSetIP` | One virtual IP (in the `ips` field) load-balances across endpoints in all exporting clusters | A ClusterIP Service | | `Headless` | DNS returns the individual backend addresses; no virtual IP | A headless Service | ## What MCS does not provide - **Network connectivity.** Pods in the third cluster must be able to reach pod IPs in the exporting clusters, either through a flat routable network or through gateways the implementation adds. - **The controller and DNS.** Without an installed implementation and a DNS plugin serving `clusterset.local`, the CRDs are just stored objects. - **Traffic policy such as retries, mTLS or failover weighting.** That belongs to a service mesh or to the implementation's own extensions. - **Exposure outside the ClusterSet.** Exporting makes nothing reachable from the internet. ## Conflicts and failure modes - If the two exports disagree on **ports**, **type** (headless or not) or **session affinity**, the implementation sets a `Conflict` condition with a reason such as `PortConflict`, `TypeConflict` or `SessionAffinityConflict`. The KEP resolves the conflicting property by giving precedence to the oldest export. - An export whose Service is missing reports `Valid=False` with reason `NoService`, and it contributes no endpoints. - An `ExternalName` Service cannot be exported; this is reported as `InvalidServiceType`. - If callers keep using `cluster.local`, they never reach remote endpoints, a common reason teams conclude that failover did nothing.

  • After exporting, does loyalty-accrual.loyalty.svc.cluster.local start reaching the other cluster's pods?
    No. `cluster.local` names keep resolving to the local Service and its local endpoints only. To use backends in other clusters, callers must switch to `loyalty-accrual.loyalty.svc.clusterset.local`, which the ServiceImport backs with endpoints from every exporting cluster, including the local one if it exports too. The name change is deliberate, so no caller starts crossing clusters without choosing to.
  • Why does namespace sameness make MCS a poor fit for a ClusterSet shared by unrelated tenants?
    MCS assumes a namespace with a given name has the same owner in every cluster of the set. If two unrelated teams each create `loyalty` in different clusters and export a Service with the same name, their endpoints merge behind one ServiceImport. One team could then receive the other's traffic. A ClusterSet therefore needs central control of namespace names.

saying these in an interview costs you the question

  • ServiceExport and ServiceImport are built into kube-apiserver and work without an add-on.
  • Creating a ServiceExport makes the Service reachable from the internet.
  • You hand-write a ServiceImport in every consuming cluster.
  • The cluster.local name starts resolving to other clusters' pods after export.
  • MCS itself tunnels traffic, so pod networks need no cross-cluster reachability.
  • The same namespace name can belong to different owners across one ClusterSet.