skip to content

A revoked Apple distribution certificate breaks every iOS build; how do you recover the match repository?

level: seniorimportance: should knowfreq 47%

answer

  1. CI is readonly; it cannot self-heal
  2. regenerate from a write-capable machine
  3. nuke revokes and clears stored assets
  4. force regenerates profiles, not certificates
  5. coordinate: revocation hits the whole team

basics

~20 s

Regenerate from a machine with write access, never from CI. Run sync_code_signing without readonly and with force to recreate the profiles, or nuke the unusable material first, then let developers and readonly runners sync the new encrypted assets down.

solid answer

~40 s

The stored assets are now scrap: the certificate is revoked and every provisioning profile that referenced it is invalid, so the falconry weight-log iOS build fails on every machine at once — including CI, which runs `readonly` and *cannot* fix itself. Recovery happens on one machine holding portal credentials and write access to the storage backend. Clear the dead material — `nuke` revokes and deletes the stored certificates and profiles — then re-run `sync_code_signing` without `readonly` for each `app_identifier` and `type` you ship, letting it create a fresh certificate and profiles and write them back encrypted. `force` regenerates profiles that were not deleted. Everyone else, and the runners, simply sync again. Announce it first: revocation hits the whole team.

code

ruby · 10 lines
ruby
lane :regenerate_appstore_signing do
  sync_code_signing(
    type: "appstore",
    app_identifier: "com.mews.falconry-weight-log",
    storage_mode: "git",
    git_url: "[email protected]:mews/falconry-signing.git",
    readonly: false,
    force: true
  )
end

go deeper

for a junior

Know that a revoked certificate invalidates the profiles built on it, and that the fix is regenerating shared assets centrally rather than signing locally on your own machine.

for a middle

Explain why the readonly CI lane fails and cannot repair itself, and which option does what: nuke removes and revokes, force rebuilds a profile, renew handles expiry.

for a senior

Run the incident: coordinate the window, regenerate from one write-capable machine, understand that outside-store builds already installed can stop working, and verify with a real release build.

for a principal

Own the prevention and the policy — who may revoke, how many certificates the estate keeps, and how a signing outage is rehearsed rather than improvised.

## What a revocation actually breaks An Apple distribution certificate is the anchor: every provisioning profile the team uses names a certificate, and a profile whose certificate has been revoked is not repairable — it is invalid. So the visible symptom is disproportionate to the act. One person clicked *Revoke* in the portal to tidy up, and now the falconry weight-log app fails to sign on every laptop and every CI runner simultaneously, with errors that mention profiles rather than the certificate that actually died. The encrypted storage backend is fine. The bytes inside it are the problem: `sync_code_signing` (the canonical action behind `match`) will happily fetch and install a revoked certificate, because *fetch and install* is all a readonly run does. It has no way to know the portal disowned it. ## Why CI cannot heal itself, and should not CI runs with `readonly: true`, which forbids creating, renewing or revoking anything. That is working as designed, and it is the reason the outage is loud rather than silent. The wrong reflex under release pressure is to flip `readonly` off in the pipeline so the runner regenerates the identity — that turns one incident into a permanent policy in which unattended jobs mint certificates against a small team quota and push private keys into the shared vault. The right move is to fix the vault once, from a machine with a person on it. ## The recovery, in order 1. **Tell the team before you touch anything.** Regeneration changes what everyone signs with, and anything still using the old material stops working the moment you proceed. 2. **Work on one machine** that holds Apple Developer Portal credentials (an `api_key`, `api_key_path` or `username`) and write access to the backend named by `storage_mode`. 3. **Clear the dead material.** `nuke` revokes the stored certificates and deletes them along with their profiles, which is what you want when what is in storage is unusable or possibly compromised. It is destructive and team-wide: schedule it, do not run it to unstick one build. 4. **Recreate.** Run `sync_code_signing` without `readonly` for each `app_identifier` and `type` you actually ship — typically `appstore` plus `adhoc` for testers. It creates a new certificate and matching profiles and writes them back, encrypted, to storage. 5. **Use `force` where you did not delete.** For profiles that still exist but must be rebuilt against the new certificate, `force` regenerates them; `force_for_new_certificates` targets exactly the 'the certificate underneath changed' case. 6. **Let everyone re-sync.** Developers and CI both run their normal lanes; the readonly runners install the new assets like any other run and need no new permissions. 7. **Re-run one real release build** end to end before declaring it over, because a profile that installs is not proof that a build signs. ## nuke, force and renew are not interchangeable | operation | what it acts on | when it is the right tool | |---|---|---| | `nuke` | revokes at the portal and deletes from storage | the stored certificates are dead, wrong, or possibly leaked | | `force` | regenerates a provisioning profile | the certificate is fine but the profile must be rebuilt | | `force_for_new_devices` | regenerates a profile after the device list changed | a tester's device was registered and ad-hoc builds no longer install | | `renew_expired_certs` | certificates that have passed their expiry | routine lifecycle, not an emergency | | `safe_remove_certs` | removing certificates deliberately | tidying the team's certificate list on purpose | Reading that table is most of the answer. A candidate who reaches for `nuke` first for every problem is telling you they have not lived through one; a candidate who reaches for `force` when the certificate itself is revoked will simply generate more invalid profiles. ## Blast radius and coordination Revoking a distribution certificate is not a local action. Builds distributed outside the App Store and signed with the revoked certificate can stop being trusted on the devices that already have them, so a `nuke` in the middle of a beta cycle takes the testers' installed weight-log builds with it. Coordinate it: pick a window, tell the beta group, and have the replacement build ready to distribute. `include_all_certificates` and `certificate_id` matter here too, because in a team with several certificates you want to be precise about which one you are acting on rather than sweeping the lot. ## Preventing the next one - **Name an owner** for the signing repository, so nobody revokes in the portal on impulse. - **Keep the number of certificates small**, which is exactly what readonly CI protects. - **Watch expiry deliberately** rather than discovering it through a red build. - **Rehearse the recovery** once, while nothing is on fire, so the runbook is not written during an incident. None of this touches Android: its keystore is self-signed with nobody to revoke it, and it is configured in the Gradle build, so an Android release lane keeps working straight through an Apple certificate revocation.

  • What would you put in place so a revoked or expiring certificate is not discovered by a red build?
    Ownership and warning. Name an owner for the signing repository so revocation is never casual, keep the certificate count low so the team is not one click from a rebuild, and run a scheduled read-write check that reports what is stored and when it expires — using options such as `renew_expired_certs` deliberately rather than stumbling into them mid-release.
  • Why not simply let the CI runner regenerate the certificate when it finds none?
    Because every clean runner would enrol another certificate against a small team quota, and those writes would land in the shared vault from an unattended job. Creation changes what the whole team signs with, so it belongs on a machine with a person on it while CI stays `readonly`.

saying these in an interview costs you the question

  • Fixes it by giving the CI runner write access
  • Thinks nuke only clears a local clone
  • Runs nuke mid-beta without telling anyone
  • Believes a readonly run will renew the certificate
  • Deletes the storage repository instead of regenerating
  • Assumes each developer can regenerate independently