skip to content

Serverless Architecture

Fully managed, event-triggered compute where you ship functions rather than servers. This area covers the execution model, triggers, cold starts, forced statelessness, managed backing services, and the very different cost curve.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

A two-person startup wants their app to have user login, a database, and push notifications, but nobody on the team wants to run or patch servers. Using Backend-as-a-Service (BaaS) products like Firebase Authentication, Firestore, and Firebase Cloud Messaging, how does building the app this way differ from writing and hosting a traditional custom backend for those same three features?

level: juniorimportance: must knowfreq 70%

answer

  1. rent don't build
  2. client talks to vendor API directly
  3. security rules replace controller code
  4. pay-per-request pricing
  5. vendor owns ops

basics

~10 s

BaaS means renting ready-made backend pieces (login, database, notifications) from a vendor's cloud instead of writing and running your own server code and infrastructure for each of them.

solid answer

~40 s

With BaaS, the client app talks almost directly to managed services over SDKs/APIs — Firebase Auth issues and verifies identity tokens, Firestore is a managed document database with declarative security rules instead of a hand-written API layer, and FCM handles push delivery. There's no server fleet to provision, patch, or scale; the vendor owns availability and elasticity. You still write business logic, usually as small serverless functions triggered by events, but the bulk of 'plumbing' code (auth flows, connection pooling, notification fan-out) disappears. The cost is giving up control: you're bound to the vendor's data model, query limits, pricing curve, and outage schedule.

go deeper

for a junior

Should describe BaaS in plain terms — using vendor products for login/db/notifications instead of writing that code — and name at least one concrete example service.

for a middle

Should explain the mechanism (client SDKs talking to managed APIs, security rules replacing controller logic) and name at least one real trade-off, like schema/query limits or vendor lock-in.

for a senior

Should discuss failure modes concretely — misconfigured rules causing data exposure, hot-partition throttling, blast radius of vendor outages — and how they differ from a self-hosted backend's failure surface.

for a principal

Should reason about when BaaS is and isn't the right call for an organization — team size, expected traffic growth curve, compliance/data-residency needs, and the long-term cost/control trade-off of committing to a vendor's primitives.

## What BaaS actually replaces Backend-as-a-Service replaces the traditional 'write a server, deploy it, keep it patched' model with a set of **managed, API-addressable primitives** that the client application (or a thin layer of serverless functions) calls directly. Concretely, for the three features in the question: - **Firebase Authentication** runs a hosted identity service — it stores credentials, issues signed JSON Web Tokens after login, handles password reset emails, and validates multi-factor codes, all reachable through a client SDK. - **Firestore** is a managed NoSQL document database that clients can read and write to directly from mobile or web code, with access control expressed as declarative 'security rules' (e.g., 'a document under `/users/{uid}` is writable only when `request.auth.uid == uid`') rather than as hand-rolled controller/service/repository code sitting in front of a database driver. - **Firebase Cloud Messaging** is a managed push-delivery pipeline: you hand it a device token and a payload, and it handles the fan-out to Apple's and Google's push gateways, retry, and delivery receipts. In each case the vendor owns the operational surface — provisioning capacity, patching the database engine, replicating data for durability, absorbing traffic spikes — and exposes only a narrow API/SDK contract. ## Why it exists This exists because most of what a backend engineer writes for CRUD-plus-auth-plus-notifications apps is **undifferentiated heavy lifting**: the same login flow, the same pagination logic, the same retry-with-backoff for push delivery, rebuilt project after project. BaaS vendors amortize that work across thousands of customers and sell it as a metered service. For a two-person startup this is decisive: there is no backend engineer needed to stand up a server fleet, configure a load balancer, run database backups, or write an auth service from scratch. Time-to-first-working-app drops from weeks to days, and the pay-per-use pricing model (per read/write/auth-verification) matches an early-stage product's near-zero traffic, so there's no idle server cost either. ## What you gain and what you give up The trade-offs run in both directions. On the velocity side you gain: - no server ops, automatic scaling, built-in high availability; - security primitives (rules engines, managed MFA) that are hard to get right by hand. On the cost side you give up: - **control over the data model** — Firestore is a document store with limited multi-record transactions and no arbitrary joins, so schemas that need complex relational queries (ad hoc reporting, multi-table joins) are awkward or require denormalizing data at write time; - **control over the request path** — you can't add custom middleware, rate-limiting logic, or a caching layer the vendor doesn't already offer; - **predictability at scale** — per-request billing that looked cheap at 100 users can become expensive at 100,000 users in ways that are hard to forecast or cap. Local development and testing are also harder: emulators for Firestore/Auth exist but don't perfectly replicate production quotas, latency, or edge-case error codes, so integration bugs surface for the first time in staging or production. ## Failure modes Failure modes cluster around three areas. 1. **First, throttling and quota limits:** Firestore and DynamoDB both throttle 'hot' partitions or documents under heavy concurrent writes (e.g., a single counter document incremented by thousands of users at once), which shows up in production as sudden write latency spikes or rejected requests that never appeared in testing with low traffic. 2. **Second, misconfigured access rules:** because clients talk to the database directly, a security-rules bug (an overly permissive rule, or one that forgets to check `request.auth != null`) can expose an entire collection publicly — this is a recurring class of real-world Firebase data exposure, distinct from a traditional backend where a bug in a controller only exposes what that one endpoint touches. 3. **Third, blast radius from vendor incidents:** because auth, database, and messaging are all the same vendor's infrastructure, a regional outage in that vendor can simultaneously break login, reads/writes, and notifications app-wide, whereas a self-hosted backend spreads that risk across whatever infrastructure choices the team made independently. ## Where it shows up A concrete real-world pattern: many production mobile apps are built entirely on Firebase — Firebase Auth for sign-in, Firestore as the sole data store, Cloud Functions for the handful of operations that need server-side logic (e.g., charging a payment, aggregating a leaderboard), and FCM for push — with zero custom server code deployed anywhere, scaling from a prototype demoed to five people to an app serving hundreds of thousands of users without the team ever provisioning a virtual machine.

  • If Firestore's security rules are the only thing standing between the internet and your data, what's the main risk compared to a traditional backend with a server-side controller layer?
    There's no application code in the middle to catch mistakes — a single overly permissive rule (or a missing request.auth check) exposes the raw collection to any client, whereas in a traditional backend a bug typically only affects the one endpoint it lives in. This is a well-documented class of real exposures in Firebase apps that shipped with default or overly broad rules. It pushes the burden of authorization correctness from a small number of server engineers onto a declarative rules file that's easy to under-test.
  • Why might per-request BaaS pricing become a problem specifically at scale, even though it looks cheap for a prototype?
    Metered pricing (per read/write/auth check) scales linearly or worse with traffic, and chatty client patterns — like re-fetching a whole document on every UI re-render — that are invisible at 100 users can generate millions of billed operations at 100,000 users. Unlike a fixed server budget, there's no natural cap forcing efficiency, so costs can surprise a team that never had to reason about request-level billing before.
  • What happens to local development and CI testing when the backend is entirely managed services accessed via client SDKs?
    Teams rely on vendor-provided emulators (e.g., the Firebase emulator suite) that approximate but don't perfectly reproduce production quotas, latency characteristics, and error responses, so some bugs — especially throttling and rules edge cases — only surface in staging or production. This is a real gap compared to a self-hosted backend where the same database engine runs in dev, CI, and prod.

BaaS is like moving into a furnished serviced apartment instead of building a house: you get working plumbing, electricity, and security on day one for a monthly fee, but you can't rewire the walls or choose your own water heater.

saying these in an interview costs you the question

  • Says BaaS means 'no backend code at all' with no mention of serverless functions still needed for business logic
  • Assumes the client-side security rules are equivalent to server-side authorization and doesn't flag them as a bug-prone attack surface
  • Can't name a single trade-off / says it's strictly better than a custom backend
  • Confuses BaaS with generic PaaS (like a container-hosting platform) and describes deploying your own server code
  • No awareness that pricing is usage-metered and can spike unpredictably

context

open as a page

In a serverless platform like AWS Lambda, what is a 'cold start' and why does it add latency to the first request?

level: juniorimportance: must knowfreq 75%

basics

~20 s

A cold start is the extra delay when a serverless function runs for the first time (or after being idle) because the cloud provider has to set up a fresh container and boot the code before it can process the request, instead of reusing one that's already running.

open as a page

In serverless computing, what does 'pay-per-invocation' pricing mean, and what are the two components (requests and GB-seconds) that make up the bill for a function like AWS Lambda?

level: juniorimportance: must knowfreq 70%

basics

~20 s

You pay for two things: how many times your function ran (per request), and how much compute it used each time (memory allocated multiplied by how long it ran, in GB-seconds). No traffic means no bill.

open as a page

A serverless platform can invoke a function two different ways depending on the event source: some sources 'push' an event straight to the function, while others are 'polled' by the platform on the function's behalf. Using an HTTP API Gateway trigger and a message-queue trigger as examples, what is the basic difference between these two invocation models?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Push: the source (like an API Gateway) calls the function immediately for each request, so it runs right away. Poll: the platform itself checks a source (like a queue) on a timer, gathers waiting messages, then invokes the function with a batch.

open as a page

In a Functions-as-a-Service (FaaS) platform like AWS Lambda, what is the difference between a 'cold start' and a 'warm start' when a function is invoked, and why does it happen?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A cold start is when the cloud has to build a fresh environment to run your function, which takes extra time. A warm start reuses an environment left over from an earlier call, so it responds almost instantly.

open as a page

In a serverless function-as-a-service platform like AWS Lambda, why is it unsafe to cache a logged-in user's session data in a plain in-memory global variable inside the function's code?

level: juniorimportance: must knowfreq 78%

basics

~10 s

Because each request might land on a different, freshly-started copy of your function that never saw that variable, so the data would just be missing sometimes.

open as a page

You're designing an order-processing feature using managed backend services: AWS Cognito for auth, DynamoDB to store orders, and EventBridge plus SNS to notify a shipping partner when an order is placed. Concretely, how do these three managed services get wired together into a working flow, and what glue code (if any) is still required?

level: middleimportance: must knowfreq 65%

basics

~20 s

Cognito checks who the user is, DynamoDB stores the order data, and when a new order is saved, a small event rule automatically tells EventBridge/SNS to notify the shipping partner — no server is running in between.

open as a page

Break down where the time actually goes during an AWS Lambda cold start — what are the distinct phases, and which one typically dominates for an interpreted language like Python versus a JVM-based language like Java?

level: middleimportance: must knowfreq 70%

basics

~20 s

A cold start has three parts: finding a machine and starting an isolated sandbox, booting the language engine (like the JVM or Python interpreter), and running your own startup code (imports, DB connections). For Python it's usually your own code and the interpreter that dominate; for Java it's usually the JVM and framework startup.

open as a page

What is a 'concurrency limit' in a serverless compute platform, and what happens to incoming requests when a function's executions hit that limit at the same moment?

level: middleimportance: must knowfreq 65%

basics

~10 s

It's a cap on how many copies of your function can run at once. If more requests arrive than the cap allows, the extra ones get rejected or queued/retried instead of running immediately.

open as a page

When an API Gateway is configured as an HTTP trigger for a serverless function, the invocation is synchronous end-to-end: the gateway waits for the function's response before replying to the client. What practical constraints does this synchronous push model impose on the function, and what happens if the function takes too long or errors out?

level: middleimportance: must knowfreq 60%

basics

~20 s

The client is waiting live, so the function has to answer fast — there's a hard time limit shorter than the function's own max runtime. If it's too slow or crashes, the gateway just returns an error straight to the client; nothing retries automatically.

open as a page

A function is triggered by an SQS queue with a batch size of 10. If 7 of the 10 messages in a batch process successfully but 3 throw errors, what happens to each message by default, and how does 'partial batch response' (reporting individual item failures) change that behavior?

level: middleimportance: must knowfreq 65%

basics

~20 s

Normally, if any message in the batch fails, the whole batch of 10 gets retried later — including the 7 that already succeeded, so they run twice. Partial batch response lets the function tell the queue exactly which 3 failed, so only those get retried and the other 7 aren't repeated.

open as a page

In a FaaS handler like AWS Lambda's `handler(event, context)`, why is it considered a best practice to initialize things like a database connection or an AWS SDK client outside the handler function rather than inside it, and what state actually persists between invocations on a warm environment?

level: middleimportance: must knowfreq 85%

basics

~20 s

Code written outside the handler function runs once per environment and gets reused across many warm calls, while code inside the handler runs every single time. Putting slow setup (like opening a connection) outside means later calls skip that work and run faster.

open as a page

AWS Lambda gives each function instance a writable /tmp directory (up to a configurable size, e.g. up to 10 GB). Why is writing files to this directory not a substitute for durable storage, even though multiple invocations sometimes reuse the same warm instance and can still see files written by an earlier invocation?

level: middleimportance: must knowfreq 70%

basics

~10 s

/tmp is just a scratch disk attached to one function instance; it disappears when that instance is recycled, and other instances never see it, so it can't reliably hold data you need to keep.

open as a page

Your team has built a production app on Firebase (Firebase Auth + Firestore + Cloud Functions). Leadership asks: 'How locked in are we, really, and what would it take to leave?' As the senior engineer, how do you answer — what specifically creates lock-in with managed backend services, what's actually portable, and what design choices reduce switching cost?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Lock-in comes from writing your app's rules and data shape in the vendor's specific style. Some things (like standard login tokens) transfer easily; your database's exact structure and vendor-specific security rules usually don't, so leaving means real rewrite work.

open as a page

How does AWS Lambda's provisioned concurrency mitigate cold starts, and what are the cost and operational trade-offs of relying on it versus letting the platform auto-scale on demand?

level: seniorimportance: must knowfreq 80%

basics

~20 s

Provisioned concurrency pays to keep a set number of function instances pre-warmed and ready at all times, so requests hit already-initialized code instead of waiting for a fresh one to boot — but you pay for that idle readiness even when there's no traffic.

open as a page

When does a serverless (pay-per-invocation) compute model come out cheaper than running an always-on server or container for the same workload, and when does it flip to being more expensive?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Serverless wins when traffic is spiky or low-volume, since you pay nothing while idle. An always-on server wins when traffic is steady and high, since its flat rate beats serverless's per-request premium at high utilization.

open as a page

A serverless function is triggered by a Kinesis (or Kafka) stream with several shards/partitions. One shard has a record that always throws an exception when processed. What happens to that shard's throughput while the bad record isn't skipped or fixed, and why doesn't it affect the stream's other shards?

level: seniorimportance: must knowfreq 55%

basics

~20 s

That one shard gets stuck: every batch keeps re-trying starting from the same bad record, so new records behind it never get processed — like a traffic jam on one lane. Other shards each have their own separate reader, so they keep moving fine; the jam doesn't spread.

open as a page

What specifically causes a FaaS cold start's latency (walk through the init phase), and what mechanisms do AWS Lambda, Google Cloud Functions, and Azure Functions each offer to reduce or avoid it - and what do those mechanisms cost you?

level: seniorimportance: must knowfreq 75%

basics

~20 s

A cold start is slow because the platform has to boot a brand-new environment: download your code, start the language runtime, and run any setup code before your function can even begin. All major clouds let you pay extra to keep a pool of environments pre-warmed and ready, so real users never hit that slow first call.

open as a page

A team implements a multi-step order-fulfillment process (validate payment, reserve inventory, ship, notify) as a single long-running Lambda function that keeps intermediate results in local variables across the steps, retrying the whole function from the top on any failure. What's wrong with this design, and how would a workflow orchestrator like AWS Step Functions change it?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Keeping progress in a running function's local variables means a crash or timeout loses everything and forces starting over, possibly repeating things like a payment charge; a workflow orchestrator instead saves each step's result externally so the process can resume exactly where it stopped.

open as a page

A team composing managed backend services needs to fan out an 'order placed' event to a handful of consumers today, but expects new consumer teams to keep adding themselves over the next year. They're choosing between publishing to an Amazon SNS topic versus an Amazon EventBridge event bus. What's the actual difference in how each routes events to consumers, and which fits the team's stated growth pattern better?

level: middleimportance: should knowfreq 55%

basics

~20 s

SNS pushes a message to every subscriber, full stop. EventBridge lets you write rules that pick which events go where based on their content, so it's easier to plug in new, unrelated consumers later without touching the publisher.

open as a page

Why is pinging a serverless function on a schedule (a 'keep-alive' or 'warming' ping) considered a weak, sometimes actively misleading, mitigation for cold starts?

level: middleimportance: should knowfreq 55%

basics

~20 s

A scheduled ping keeps one instance warm, but real traffic often needs many instances at once, and pings can't predict or prevent the cold starts that happen when demand suddenly scales up beyond what's already warm.

open as a page

In a serverless platform's concurrency model, what is the difference between 'reserved concurrency' (a dedicated slice of an account's shared concurrency pool set aside for one function) and 'unreserved concurrency' (the shared pool available to all other functions), and what problem does each solve?

level: middleimportance: should knowfreq 55%

basics

~10 s

Reserved concurrency carves out a guaranteed, capped slice of the shared pool for one function only. Unreserved concurrency is the leftover shared pool every other function draws from.

open as a page

On AWS Lambda, configuring more memory for a function also increases the CPU and network bandwidth it gets, and every function has a maximum execution timeout (15 minutes on Lambda). How do these two limits shape the way you design a FaaS-based system, and what happens when a function exceeds either one?

level: middleimportance: should knowfreq 65%

basics

~30 s

FaaS functions have two hard caps: how much memory (and therefore CPU) you give them, and how long they're allowed to run before the platform kills them. If a function tries to use more memory than allowed it crashes, and if it runs past its time limit it gets forcibly stopped - so long or heavy jobs must be split into smaller pieces or moved to a different kind of compute.

open as a page

When a stateless function needs to remember something between separate invocations — e.g., a shopping cart or a rate-limit counter — what are the trade-offs between externalizing that state to a key-value store like DynamoDB or Redis versus an object store like S3?

level: middleimportance: should knowfreq 60%

basics

~10 s

DynamoDB/Redis are built for fast lookups of small structured records, while S3 is built for storing large files cheaply; picking the wrong one costs you either speed, correctness, or money.

open as a page

How does AWS Lambda SnapStart reduce cold starts for Java functions, and what correctness pitfalls does its snapshot-and-restore approach introduce that provisioned concurrency doesn't have?

level: seniorimportance: should knowfreq 45%

basics

~20 s

SnapStart boots and initializes your function once, then saves a memory snapshot of that fully-warmed state; every future cold start just restores from that snapshot instead of re-running startup, making it much faster. The catch: cached state from before the snapshot (like random numbers or timestamps) can leak into later invocations unless you handle it explicitly.

open as a page

What does 'scaling to zero' mean in a serverless platform, and what latency and cost trade-off does it introduce via cold starts?

level: seniorimportance: should knowfreq 60%

basics

~20 s

When idle, the platform fully shuts down your function's instances, so you pay nothing. The catch: the next request after idle has to wait for a fresh instance to start up — that delay is a cold start.

open as a page

Object-storage triggers, like S3 event notifications invoking a function on every object upload, are push-based and deliver events with at-least-once semantics. What does 'at-least-once' mean operationally for a handler, and what other delivery quirks (ordering, timing) does this trigger type have that a handler needs to account for?

level: seniorimportance: should knowfreq 45%

basics

~20 s

At-least-once means the same upload event might trigger your function more than once, so it can't assume it only runs exactly one time per file. Events also might not arrive in the exact order files were uploaded, and there can be a short delay before the event fires.

open as a page

Two concurrent invocations of the same stateless Lambda function both read a counter value of 5 from a DynamoDB item, increment it locally to 6, and write 6 back. What went wrong, and what two DynamoDB mechanisms would prevent it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Both invocations read the same starting number before either wrote back, so one increment silently overwrites the other and the counter ends up one short. DynamoDB fixes this with atomic updates or conditional writes that check the value hasn't changed.

open as a page

You're the principal engineer advising on architecture for a new B2B SaaS product that will need complex multi-table transactions, strict data-residency guarantees for enterprise customers in the EU, and predictable infrastructure costs at a contracted scale of millions of requests/day. A junior architect proposes building the entire backend on Firebase/Firestore because 'it worked great for our last consumer app.' What's your reasoning for where BaaS composability breaks down here, and what would you recommend instead?

level: principalimportance: should knowfreq 45%

basics

~20 s

BaaS tools are great for simple apps, but this project needs complex multi-step transactions, guaranteed EU-only data storage, and predictable costs at huge scale — things Firestore-style services aren't built to guarantee, so a traditional or hybrid backend fits better.

open as a page

For a serverless API with a strict p99 latency SLA, why can average or even p50 latency numbers hide a cold-start problem entirely, and what architectural options exist beyond provisioned concurrency to keep tail latency in check?

level: principalimportance: should knowfreq 50%

basics

~20 s

Most requests hit already-warm functions and look fast, so the average looks great, but the rare slow ones (cold starts) show up only in the tail (p99, the slowest 1%). Fixing that tail needs more than just averages — options include keeping capacity pre-warmed, timing out and retrying to a warm instance, or picking faster-starting runtimes.

open as a page

showing 1–30 of 34