Cloud Platform Concepts
What every provider sells under a different name: regions, accounts, identities, managed tiers, a private network and a bill — from the largest platforms to a one-command host.
on this pageshowhide
explore
- Provider Responsibility Model22 questions
- Managed Service Tiers4 questions
- Shared Responsibility Line4 questions
- Managed Tier Trade-offs5 questions
- Outgrowing a Managed Tier4 questions
- Key Custody & Defaults5 questions
- Global Footprint & Placement25 questions
- Regions & Availability Zones5 questions
- Choosing a Region5 questions
- Regional Service Parity5 questions
- Edge Presence5 questions
- Data Residency & Sovereignty5 questions
- Tenancy Layout21 questions
- The Account Boundary4 questions
- Organization Hierarchy4 questions
- Environment Separation4 questions
- Organization Guardrails5 questions
- Account Vending & Baselines4 questions
- Platform Access Model32 questions
- Principals & Policy Attachment5 questions
- Short-Lived Credentials4 questions
- External Identity Trust4 questions
- Least Privilege & Scoping5 questions
- Privilege Escalation Paths5 questions
- Human Access & Break-Glass4 questions
- Credentials from the Machine5 questions
- Renting Compute22 questions
- Workload Hosting Tiers4 questions
- Machine Profiles & Sizing4 questions
- Purchase & Commitment Models4 questions
- Designing for Reclaimable Capacity5 questions
- When Capacity Runs Out5 questions
- Platform Storage Services23 questions
- Object, Block & File5 questions
- Durability vs Availability4 questions
- Tiers & Lifecycle Rules5 questions
- Versioning & Deletion Safety4 questions
- Storage Access Control5 questions
- Provider Network Model33 questions
- Private Network & Subnets5 questions
- Routing & Outbound Paths5 questions
- Private Service Access5 questions
- Joining Private Networks4 questions
- Hybrid Connectivity4 questions
- Network Boundary Filtering5 questions
- Managed Entry Points5 questions
- Reliability & Failure Domains25 questions
- Provider SLA Semantics5 questions
- Zone & Region Outages6 questions
- Control Plane Dependency4 questions
- Shared Platform Dependencies5 questions
- Recovery Objectives5 questions
- Billing & Charge Model29 questions
- Pricing Dimensions5 questions
- Commitments & Discounts5 questions
- Data Transfer Pricing4 questions
- Attribution & Tagging5 questions
- Idle & Orphaned Spend5 questions
- Budgets & Spend Signals5 questions
- Platform Operations22 questions
- Management API Surface4 questions
- Quotas & Service Limits4 questions
- Control Plane Throttling4 questions
- Maintenance & Version Lifecycle5 questions
- Platform Audit Trail5 questions
- Exit & Switching Costs18 questions
- Dimensions of Lock-In4 questions
- Multi-Cloud Postures5 questions
- Portability Through Abstraction5 questions
- The Cost of Leaving4 questions
- AI & Data Scientistroleanchors this topic
- AI Engineerroleanchors this topic
- Cyber Security Expertroleanchors this topic
- Data Engineerroleanchors this topic
- DevOps / SRE Engineerroleanchors this topic
- MLOps Engineerroleanchors this topic
- Machine Learning Engineerroleanchors this topic
- Server-Side Game Developerroleanchors this topic
- Software Architectroleanchors this topic
- Forward Deployed Engineerrole
questions
272 · 11 sectionsAn internal platform offers one service on a rented machine, a managed runtime, or a whole managed capability — what does each rung take over?
basics
~20 sEach rung hands the provider more of the stack. A rented machine leaves everything above the virtualization layer to you; a managed runtime adds the operating system and runtime; a whole managed capability adds patching, scaling, backups and failover.
A managed store reports that encryption at rest is enabled by default - which risk does that remove, and which does it leave?
basics
~20 sDefault encryption at rest protects bytes on the provider's media against drive loss, disposal or raw block access. It restricts nobody who calls the service: an authorized caller, and the provider holding the key, still get plaintext.
A service moves from a rented machine to a managed runtime — why does that rung remove your access along with the work?
basics
~20 sA provider can only promise an outcome it fully controls. Each duty a rung takes over becomes a guarantee made across many tenants at once, and any change you could still make by hand is state the fleet operator cannot assume.
Why does a managed database tier disable certain engine extensions and tuning knobs, and what does that force on your design?
basics
~20 sA managed tier exposes only what it can operate, persist and support for every tenant at once, so settings that need host access, load foreign code into the engine, or weaken the tier's own promises are withheld, and the work they would have done moves into your application.
When choosing the region for a new service's first deployment, which input removes candidates outright and which ones are traded?
basics
~20 sA residency obligation filters the candidate list before anything is compared: regions outside the required jurisdiction simply leave. Distance to users, the rate the region charges for the same service, and whether it offers what you need are then traded against each other.
An edge location sits in front of your single region for far-away users - what can it absorb, and what must travel back?
basics
~20 sAn edge location can terminate the connection and TLS, serve a cached response, issue a redirect, and make a cheap decision from the request itself. The authoritative dataset, durable writes and heavy computation still travel back to the region.
What separates a region from an availability zone inside it, and what does each boundary isolate?
basics
~20 sA region is an independent geographic deployment unit; the availability zones inside it are separate facilities with their own power, cooling and network paths, a metro hop apart. The zone boundary isolates a facility failure, the region boundary isolates everything larger.
A mobile backend serves users a continent away and feels slow — what latency floor does that distance set, and what cannot fix it?
basics
~20 sDistance fixes a round-trip floor: light in fibre travels about 200,000 km per second and real paths are longer than the map, so a far continent costs roughly a tenth of a second per round trip. Bigger machines and more instances cannot shorten it.
Why does terminating the client connection at a nearby edge location cut time to first byte for an uncacheable response?
basics
~20 sConnection setup costs several round trips before any request bytes move. Terminating nearby pays those against the short hop, while the forwarded request crosses the long distance once over a connection the edge already holds warm.
A team runs everything on the single provider account they opened - what does that one account bundle together?
basics
~20 sA provider account is one container bundling four things at once: one isolation boundary, one bill, one pooled set of usage ceilings, and one set of administrators. Every resource you create lives inside exactly one account.
Your staging ledger shares production's account, kept apart only by a name prefix and an environment label — why is that not an isolation boundary?
basics
~20 sA name prefix and an environment label are data every caller in the account can read or ignore; nothing enforces them. Isolation, quota and billing attach to the account, so only a separate account makes the split real.
A side project's single account now carries three products - which symptoms say that one account has been outgrown?
basics
~20 sThree symptoms matter: products contending for ceilings counted once for the whole account, a bill nobody can split by product, and a change or mistake in one product reaching the other two. Each follows from the account being one shared container.
An internal platform issues each team a new cloud account from a template — what does that account land with before any workload runs?
basics
~20 sA vended account arrives with its baseline already applied: audit and platform logs delivered to a central collection account, a private address range allocated from the organisation's plan, guardrail policies in force above it, and a named owner recorded.
A policy set above the account refuses your action although you hold full administrative permissions inside it - how does that work?
basics
~20 sA permission ceiling above the account subtracts from everything inside it. The effective permission is the intersection of the ceiling and the account's own grants, so an administrator can only grant within what the ceiling still leaves.
A workload on a rented machine stores no platform key anywhere, yet its API calls are authorized — where does its credential come from?
basics
~20 sThe platform issues it on demand. An identity is attached to the machine, and a credential endpoint reachable only from that machine hands the workload a short-lived credential set, which is replaced before it expires.
In a cloud platform's access model, what is a principal, and how do human, group and workload principals differ?
basics
~20 sA principal is the identity a platform authenticates and names when it decides whether a call is allowed: a person's login, a workload such as a batch job or service, or a group that holds grants for the humans inside it.
A long-lived platform access key is found in a public repository — why is issuing a replacement key not the first move?
basics
~20 sCreating a new key does not disable the old one. Revoke the exposed credential first so it stops authenticating immediately, then issue replacements, then read back what the leaked credential actually called before it was stopped.
Why do engineers sign in to a production cloud account through the company directory instead of each holding a local platform login?
basics
~20 sFederated workforce logins keep one copy of each person, in the company directory: cloud access follows group membership, and disabling the directory account cuts every federated sign-in at once. Local platform logins are extra credentials nobody remembers to delete.
A team can create workloads and attach any existing identity to them, but holds no administrator permission — why is that grant administrator access anyway?
basics
~20 sAttaching an identity to a workload means running your own code as that identity. If any stronger identity can be attached, the team can launch code that borrows its permissions, so the attach grant is worth the strongest identity it reaches.
A nightly cleanup job runs about twenty minutes - why does an event-driven runtime rule itself out, and which tier fits?
basics
~20 sAn event-driven runtime caps how long one invocation may run, so a twenty-minute job is stopped part-way. Put it on a tier that keeps a process alive: a managed container platform running it on a schedule, or a machine that stays up.
You are moving a measured search service onto rented machines — which numbers set the machine size, and which do you ignore?
basics
~20 sSize from the service's own observed processor, memory, disk and network use over a full demand cycle, taken at a high percentile with headroom added. The specification of the server being replaced records a purchase decision, not the workload, and is ignored.
A launch request for a new machine is refused in one zone — how do you tell an empty capacity pool from an administrative ceiling?
basics
~20 sA refused launch means one of two shortages: the provider has no machine of that shape free in that zone, or your account is not permitted more. A ceiling is a number you can look up; pool depth is never published.
Machine families are weighted toward processor, memory, local disk or accelerators — which measurements pick one?
basics
~20 sThe ratio in a measured profile picks the family — which resource saturates first while the others idle. Memory-heavy work wants a memory-weighted family, compute-bound work a processor-weighted one, heavy local input and output a disk-weighted one, dense parallel numeric work an accelerator.
A reporting service has a flat weekday base and a sharp month-end peak - which purchase posture suits each part of that demand shape?
basics
~20 sBuy the always-on base with a term commitment, serve the month-end peak with metered on-demand capacity, and put only interruption-tolerant extra work on reclaimable capacity. Price each layer of the shape separately rather than picking one posture for the fleet.
An external scanner reports that your receipts store is readable by anyone on the internet — which grant produces that, and what was not leaked?
basics
~20 sA grant written on the store side — at store or object level — naming any caller as allowed to read produces public access. Nothing was leaked: the store is doing exactly what its policy says, so no credential was stolen and no defect was exploited.
An object store advertises durability with many nines, but ingest requests fail for an hour - which promise did that figure never make?
basics
~20 sDurability is about bytes surviving; availability is about bytes being reachable now. A many-nines durability figure estimates how unlikely it is that the store loses an object, and says nothing about an hour of failed requests.
Which storage shape fits a transcoder's scratch space while rendering, and which fits the finished videos many clients fetch?
basics
~20 sScratch belongs on a block volume: it attaches to one machine and behaves like a local disk, so seeks and in-place writes are cheap. Finished videos belong in an object store, fetched by key over HTTP by any number of readers.
Weekly-read application logs are moved to the coldest archive storage tier to save money — what goes wrong?
basics
~20 sA colder tier discounts rent by charging separately for reads and by making them slow. Data read every week pays a retrieval charge every week and waits for a restore each time, so the bill usually goes up, not down.
A job overwrote every document in a versioned object store with a corrupt render — how do you get the originals back?
basics
~20 sVersioning kept each pre-overwrite copy as a previous version under the same key, so recovery is promoting that version back to current, key by key. Nothing is restored from a backup, and every retained version keeps being stored and charged.
A stateless subnet filter fronts a workload whose own rule set is stateful — what does each need written for return traffic?
basics
~20 sThe stateful rule set needs only the request direction written; it matches the reply to the flow it already accepted. The stateless subnet filter judges each packet alone, so the reply needs its own rule allowing the client's ephemeral port range.
A managed entry point keeps its public address when the machines behind it are replaced, so what makes that address a separate rented resource?
basics
~20 sThe address is allocated from the platform's pool and held by the entry point resource, not by any machine. It outlives instance replacement, is usually charged while you hold it, and goes back to the pool only when you release it.
What actually makes one subnet in your private cloud network public and another one private?
basics
~20 sRouting, not the name. A subnet is public when the route table it uses sends traffic for destinations outside the network to a target that reaches the internet. Private means no such route exists. The label itself configures nothing.
Why would a team reach a managed store through an endpoint inside its own address range instead of the store's public address?
basics
~20 sAn endpoint inside your own range keeps the call on the provider's internal network: the workload connects to an address in one of your subnets, so the request does not take the public path and the subnet needs no route to the internet.
A batch job in a subnet with no internet route must call partner APIs without ever being reachable — what do you add?
basics
~20 sTwo things: a default route on that subnet naming an address-translating gateway, and the gateway itself on the routed side. Outbound flows leave translated; unsolicited inbound packets match no translation entry, so nothing outside can open a connection.
What does a recovery point objective promise about a managed service, and what does a recovery time objective promise instead?
basics
~20 sA recovery point objective caps how much recent work the business accepts losing, measured backwards from the failure. A recovery time objective caps how long the service may stay unavailable, measured forwards from the same moment.
A provider publishes an availability commitment for a managed service — is that a guarantee, and what does it pay you when missed?
basics
~20 sAn availability commitment is a contract term, not a promise the service stays up. If the provider's own measurement falls short, the remedy is a service credit against your bill for that service, which you normally have to claim.
A checkout service runs in three availability zones and one zone goes dark at peak — what keeps serving, and what had to be true beforehand?
basics
~20 sInstances in the two healthy zones keep serving, but only if the traffic entry point health-checks the dead ones out, the writable data copy is not stranded in the lost zone, and the survivors already had spare capacity.
During a provider incident, running workloads keep serving but no new instance launches - which half of the platform is degraded, and what stops?
basics
~20 sThe control plane - the platform's management API that creates, changes and deletes resources - is degraded, while the data plane that carries request traffic keeps serving. Launching, scaling, replacing and failing over stop; already-running capacity does not.
Why does a standby replica of a managed store give a near-zero recovery point yet fail to protect against a mistaken bulk delete?
basics
~20 sA replica copies committed writes, so it reproduces the mistaken delete as faithfully as any other write. Only a point-in-time restore rewinds the data to a moment before the mistake, at the cost of a much longer recovery time.
What does a label attached to a cloud resource actually change about the bill, and what does it not change?
basics
~20 sA label copies a key-value pair onto the charges a resource generates, so the provider's detailed cost report can be grouped by team, environment or service. It changes reporting only: the rate, the metered quantity and the total are unaffected.
A monthly cloud budget threshold is crossed at noon — what does crossing it actually do, and what does it not do?
basics
~20 sA budget threshold is a notification rule, not a spending cap. Crossing it emits a message about money already spent. It does not pause running resources, block new ones, or cancel anything, and the figure it fired on lags the usage.
On a cloud transfer bill, a service uploads large video masters and ships finished renditions out to viewers — why is only one direction charged?
basics
~20 sTransfer is metered by direction. Bytes arriving from outside are normally unmetered; bytes leaving toward the internet are priced per gigabyte after a small monthly allowance. So delivery to viewers is the charge, not the upload of masters.
A managed service with no hourly rate is the largest line on your cloud bill — which pricing dimensions does a provider meter?
basics
~20 sProviders meter usage along several independent dimensions: running time, requests served, gigabyte-months stored, gigabytes moved, and provisioned capacity units. One service can charge on all of them at once, so a large bill line exists with no hourly rate anywhere.
After you delete a virtual machine, the bill barely moves - which charges outlive the machine, and how do you find them?
basics
~20 sStorage and reservations outlive the machine. A block volume that was not marked to be deleted with it, the snapshots and images built from it, and a reserved public address all keep billing. Find them by listing resources attached to nothing, not by reading the machine's record.
Your create call returns an identifier and reports success in under a second, yet the new resource refuses connections — why?
basics
~20 sThe call was accepted, not completed. A management API create is asynchronous: it validates and records the request, hands back an identifier immediately, and only then works the resource through build states before anything can serve traffic.
A resource created in a provider's web console is identical to one made from the command line — why?
basics
~20 sBoth are clients of the same management API. The console is a hosted application that turns a form into the same request the command line tool sends; neither has a private path into the platform.
Your managed database has a weekly maintenance window - what may the provider do inside it, and what does your application see?
basics
~20 sA maintenance window is a recurring slot you nominate in which the provider may patch and restart your managed instance. The application normally sees dropped connections and a short interruption, or a failover to the standby replica, not a seamless change.
What separates a soft quota you can ask to have raised from a hard limit, and what does hitting each cost?
basics
~20 sA soft quota is a provider ceiling an increase request can raise, so hitting it costs lead time. A hard limit is fixed by the platform's design, and the only way past it is an architecture change.
Your deployment tool is rate-limited while the running service it deploys serves user traffic normally - which request rate is being metered?
basics
~20 sThe management API meters requests per account, separately from the traffic your workload serves. Deployment tools, dashboards and scripts all spend that management budget; user requests do not consume it unless the service itself calls the management API.
Why would a team run a component itself on rented machines instead of using the provider's managed version?
basics
~20 sRunning the component yourself keeps the software identical wherever the machines are rented, so a move becomes a reinstall and a restore rather than a rewrite. You pay for that with the patching, backups and on-call the managed tier was doing.
A team says it is locked in to its cloud platform "because of the code" — what does lock-in actually mean, and which kinds are not code?
basics
~20 sLock-in is the cost of leaving, not an inability to leave. Code is one bill of four: data that is priced to move, operating knowledge the team has only here, and an unexpired term commitment are the other three.
A migration off one platform is budgeted purely as engineer-months to rewrite code — which cost lines does that estimate miss?
basics
~20 sA rewrite estimate covers one of four exit lines. The others are the outbound data charge plus the weeks the copy takes, both platforms billed through the overlap window, and the months still running on a term commitment.
A managed service is wire-compatible with an open interface - what does that compatibility usually not cover?
basics
~20 sCompatibility with an open interface normally covers the data path your application speaks, and stops at the edges: provisioning, sizing, authentication, backup and restore, telemetry, quotas and error behaviour all stay the provider's own design.
Your board asks for multi-cloud after another company's outage, so which three distinct postures can that one word mean?
basics
~20 sMulti-cloud covers three different architectures: the same workload live on two platforms at once, different workloads split across platforms with one home each, and a single live platform plus a documented exit plan. Each has a different bill and a different failure behaviour.