Your service can run its own OAuth 2.0 authorization server or rent a managed one — what does each choice leave your team to operate?
answer
- itemise the work, then assign it
- same six jobs, two columns
- correlated failure of the login path
- the consent record stays your obligation
- local rows and join key never move
basics
~20 sRenting moves the login path's uptime, its abuse handling and most of its key custody onto someone else; running your own leaves you all three plus the on-call rotation behind them. Neither choice moves your obligation to show who consented to what.
solid answer
~50 sRefuse the abstract argument and itemise the work, because every item exists either way — only the pager changes. Five jobs: keeping the **login path available**; holding the **private key** that mints tokens your services trust; keeping the **record of what each person consented to** and when; handling **abuse on registration, password reset and code delivery**; and staffing the **queue for someone locked out at the worst hour**. Renting moves the first four decisively and shares the fifth. Running it yourself moves none of them, and adds a correlation you should name out loud: an issuer inside your own platform is down during your own incident, exactly when nobody can sign in to diagnose it. What renting never moves is your local account rows, your integration code, and the regulator's question, which is still addressed to you.
code
json · 22 lines{
"rented": {
"login_path_uptime": "provider",
"registration_and_reset_abuse": "provider",
"token_signing_key_custody": "provider",
"consent_and_signin_evidence": "provider stores, you must be able to export",
"locked_out_user_queue": "shared: their mechanism, your queue",
"local_account_rows_and_join_key": "you",
"callback_and_session_code": "you",
"regulator_asks": "you"
},
"self_run": {
"login_path_uptime": "you, and it fails inside your own incident",
"registration_and_reset_abuse": "you, continuously",
"token_signing_key_custody": "you",
"consent_and_signin_evidence": "you",
"locked_out_user_queue": "you",
"local_account_rows_and_join_key": "you",
"callback_and_session_code": "you",
"regulator_asks": "you"
}
}go deeper
Remember the two options exist and that an authorization server is a running service, not a library you add. Be able to say that renting one still leaves your application with its own account records and its own callback code.
Itemise the jobs before assigning them: login-path uptime, key custody, the consent record, abuse handling, the lockout queue, and who builds a new authentication method. Then say which column each one falls into and why.
Bring the operational detail that shows you have run one: the correlated-failure argument, what an export clause has to say about evidence, and the fact that abuse handling is a standing job rather than a feature. Give a verdict, not a balanced essay.
Frame it as a reversibility decision. Say what would make you change your mind later, what you would be locked into, and how you would keep the option open — then insist that the exit is priced at the same time as the entry.
## The question behind the question An **authorization server** is the component that authenticates a person, records what they agreed to, and issues the tokens your own services then accept. Deciding whether to run one or rent a managed one reads like a procurement question, and it is really an operations question: the software is the cheap half of either answer. What you are choosing is **whose pager rings** for a fixed list of jobs, every one of which exists in both worlds. So the strong answer refuses the abstract argument and itemises. Work it against something concrete — a filing assistant that reads a payroll system on the filer's behalf. Its population is tens of thousands of people for ten months of the year and millions for the fortnight around a statutory deadline; a regulator can ask what each filer agreed to; and a failed sign-in at 23:40 on deadline night is a missed filing, not a support ticket. ## The inventory 1. **Availability of the login path.** Somebody keeps it up, patched and reachable. The subtlety is **correlation**: an issuer you run on your own platform goes down *inside your own incident* — the moment your staff most need to sign in to the tooling that diagnoses it, and the moment the people you serve most need to be let in. A rented issuer fails on its own schedule, uncorrelated with yours. That is a real and unglamorous benefit, and its price is that you cannot influence that schedule at all. 2. **Custody of the key material.** Whoever runs the issuer holds the private key that mints tokens every one of your services trusts, and owns the consequences if it leaks — a compromise there is indistinguishable from a compromise of every account. Rent, and that custody sits behind someone else's controls and someone else's audit report. Run it, and it is now a secret your team must store, restrict and account for. (How that key is generated, published and rolled over is a separate subject from whether you own it.) 3. **The consent and authentication record.** Who agreed to what, when, and from where — plus the sign-in events a dispute is later reconstructed from. Renting moves the *storage*; it moves the *obligation* only to the extent that you can pull the evidence out on demand and keep it after the contract ends. Check that at signing, not at subpoena. 4. **Abuse on the open surfaces.** Registration, password reset and out-of-band code delivery face the internet unauthenticated. They attract automated sign-ups, credential stuffing and delivery-cost abuse, and they need throttles, monitoring and a human who watches the graphs. This is a standing adversarial job, not a feature you finish. 5. **The recovery queue.** A person locked out the night before a deadline is a support conversation with a real deadline behind it. A provider owns the reset *mechanism*; you own the *person*, because they contacted you. 6. **Change over time.** A new second factor, a new jurisdiction's requirement, an accessibility obligation. Rented, these arrive on the provider's roadmap and you wait. Self-run, they arrive on yours and you build them. ## What renting does not move - **Your account rows.** Local records still exist and are still joined to whatever identity the provider returns, so the join key is your design problem either way. - **Your integration code.** The callback, the session you issue afterwards and the sign-out path are yours in both worlds. - **Your obligation.** The regulator, the auditor and the affected person all address you, whoever stores the record. - **Your incident, sometimes.** A provider outage is your outage, with the added misery that you cannot shorten it. Your status page still carries your name. ## Scoring the two columns | Job | Rented | Self-run | |---|---|---| | Login-path uptime | Provider's, and uncorrelated with your incidents | Yours, and correlated with your incidents | | Abuse on registration and reset | Provider's, priced into the contract | Yours, continuously | | Private key that mints tokens | Provider's custody | Yours to store and account for | | Consent and sign-in evidence | Provider's store, yours to export | Yours end to end | | Locked-out person at 23:40 | Shared: their mechanism, your queue | Yours entirely | | Local account rows and the join key | Yours | Yours | | Roadmap for new authentication methods | Provider's, you wait | Yours, you build | ## How to answer it Give a verdict in one sentence, then defend it with the inventory rather than with adjectives. For most teams the honest default is to rent, because four of the six jobs above are continuous adversarial operations work with no product value, and a small team cannot staff them. Name the conditions that reverse it: a residency or control requirement that forbids the identities leaving your estate; a login path that must survive an outage at a supplier you did not choose; authentication behaviour no provider will express for you; or a population large enough that per-identity pricing dominates the engineering cost. Then say the word **exit** before you are asked — a decision that is cheap to make and expensive to undo is not finished until you have priced undoing it.
- Why does it matter that a self-run authorization server fails at the same time as the application it serves?Because the failures are correlated. During a platform-wide incident the login path is down too, so your own responders cannot sign in to the tools that diagnose it and your users cannot reach anything, including whatever self-service path would have reduced the support load. A rented issuer fails independently, which does not make it more reliable — only differently timed.
- If you rent, what must the contract say about the consent and sign-in record?That you can export it on demand, in a documented format, including after termination, and for the full retention period you are obliged to keep. Renting moves where the evidence lives, not who has to produce it. A provider that shows the record only in its own console leaves you unable to answer a regulator without their cooperation.
- Does renting remove abuse handling on registration and password reset?It removes the mechanism, not the accountability. The provider runs the throttles, the bot detection and the delivery limits, and generally does it better than a small team would. You still own the policy decisions those controls express, the bill when delivery abuse spikes, and the conversation with people the controls block by mistake.
saying these in an interview costs you the question
- Treats it as a build cost only, ignoring who operates it afterwards
- Says renting removes the consent and audit obligation
- Misses that a self-run login path dies inside your own incident
- Assumes local account rows disappear once identity is outsourced
- Argues purely on control, with no inventory of the work involved
- Believes a provider outage is not your outage