What breaks when an ADK agent that worked under adk web is deployed to Cloud Run?
answer
- nothing about the agent changed
- many instances, none of them yours
- in-memory services die with the process
- your gcloud identity does not travel
- UI and tracing are opt-in flags
basics
~20 sConversations vanish and tool calls start getting denied. In-memory session, artifact and memory services are per-process, and Cloud Run runs many replaceable instances; local user credentials are replaced by the service account, and the dev UI and tracing are opt-in flags at deploy time.
solid answer
~50 s`adk deploy cloud_run` packages the agent into a container serving the same FastAPI app you had locally, so the code path is identical — what changes is the environment around it. Three things bite. **Statelessness**: locally one process held every session in memory; on Cloud Run requests spread across instances that scale to zero, so a follow-up turn can land somewhere that never saw the session. Point the session, artifact and memory services at real backends via the deploy command's service-URI options. **Identity**: locally your tools ran under your `gcloud` application-default credentials; in the container they run as the service's service account, so anything not explicitly granted is denied. **Surface and visibility**: the dev UI is not deployed unless you pass `--with_ui`, and traces do not reach Cloud Trace unless you pass `--trace_to_cloud`. Add cold starts and the request timeout against long agent turns.
code
bash · 7 linesadk deploy cloud_run \
--project=my-project \
--region=us-central1 \
--service_name=weather-agent \
--with_ui \
--trace_to_cloud \
./agents/weathergo deeper
Know that the deployed agent runs in a container with no memory of your laptop: sessions are not kept unless configured, and your personal Google credentials are not there.
Explain that adk deploy cloud_run serves the same FastAPI app, so the breakage is environmental — in-memory services across multiple instances, service-account identity, and opt-in flags for the UI and tracing.
Walk the failure from the user's symptom back to the cause: intermittent amnesia means multi-instance plus in-memory sessions. Cover timeouts against long turns, streaming, cold starts and ingress control.
Own the deployment contract: which state must be durable, which identity each tool runs under, what the public surface is, and how you keep production observability equivalent to the dev UI's trace view.
## What the deploy actually produces `adk deploy cloud_run` generates a container around your agent directory and deploys it as a Cloud Run service. Inside, it serves the same FastAPI application you exercised locally with `adk api_server` — the same `POST /run`, `POST /run_sse` and session endpoints. If you want to control the container yourself, you build your own image and construct the app with `get_fast_api_app()` from `google.adk.cli.fast_api`, which is what you do when the agent must live inside an existing service. Either way the agent code is unchanged, which is exactly why the failures that follow are surprising: nothing about the agent broke, everything about its environment did. ## Statelessness is the headline failure Locally, `adk web` ran one process. The default in-memory session service, artifact service and memory service all lived in that process's heap, and every request hit the same heap. Cloud Run is the opposite by design: instances start on demand, scale out under load, scale to zero when idle, and are replaced freely. Consequences, in the order users notice them: 1. A user sends turn two and the agent has no idea who they are, because turn one landed on a different instance. 2. Traffic dips, the service scales to zero, and every conversation in flight evaporates. 3. Artifacts a tool wrote are unreadable on the next request because they were held in another instance's memory. The fix is configuration, not code: at deploy time point the session service at a database URI (or a Vertex-managed backend) and the artifact service at durable storage, using the deploy command's corresponding URI options. This is the number-one thing to say in an interview, because it is the failure that gets shipped. ## Identity changes underneath your tools On your laptop, a tool calling a Google API picked up your application-default credentials — an account that probably has generous project access. In the container, calls run as the Cloud Run service's service account. Anything you never granted is denied, and the denial surfaces as a tool error the model then narrates apologetically to the user. Deploy checklists should include: which APIs do my tools touch, what roles does the runtime service account hold, and does the model endpoint itself (Vertex AI or an API-key provider) work under that identity. Non-Google API keys must move out of the local `.env` and into the service configuration or Secret Manager — the local `.env` is a development convenience, not a deployment mechanism. ## The surface you get is narrower than the dev UI The browser dev UI is not part of a deployment by default; `--with_ui` includes it, and you should think hard before enabling it on anything public, because it is an unauthenticated-by-design developer console over your agent. Cloud Run's own access control (authenticated invocations, IAM, or an identity-aware proxy in front) is what stands between the internet and `POST /run`. An ADK service deployed with unauthenticated access is a bill and a data-exfiltration path with a chat box on it. ## Observability has to be turned on The Trace tab you leaned on locally does not follow the agent to production. `--trace_to_cloud` sends spans to Cloud Trace so you can still see which LLM or tool call consumed a turn's latency. Without it, a slow agent in production is a black box, and you will be reduced to reading logs for tool names. ## Timing and shape mismatches Agent turns are long by web standards — several model round-trips plus tool latency. Cloud Run applies a request timeout, and a multi-step turn can bump it; streaming through `/run_sse` both improves perceived latency and keeps the connection productive. Cold starts add seconds to the first request after a scale-to-zero, which users experience as the agent being slow specifically when they have not used it recently. Concurrency settings matter too: an agent instance spends most of its time waiting on model and tool I/O, so a very low concurrency setting wastes instances while a very high one can starve the event loop. ## The checklist worth reciting Durable session and artifact backends; service-account roles for every tool's API; secrets out of `.env` and into the platform; ingress restricted and `--with_ui` deliberate; `--trace_to_cloud` on; streaming endpoint used; timeout and concurrency reviewed against real turn length. Everything on that list is environment, not agent code — which is the point of the question.
- Users report the agent forgets them intermittently, not always. What does intermittent tell you?That some requests still land on the instance holding the session. Intermittent amnesia is the signature of in-memory session storage behind a multi-instance service: sticky-by-luck routing makes it look like a flaky bug rather than a configuration gap. Confirm by checking instance count during the failures, then move the session service to a durable backend rather than chasing it as a race.
- How do you deploy an ADK agent inside an existing FastAPI service instead of as its own Cloud Run service?Build the ADK app with `get_fast_api_app()` from `google.adk.cli.fast_api`, mount it into your own application, and ship your own container. You then own the container image, dependencies and routing, and you inherit the host service's auth and networking — which is usually the reason to do it in the first place.
- Would you deploy with --with_ui on a production service?Only behind real access control. The dev UI is a developer console with a chat box, event inspector and state view over your agent; exposed publicly it is both a cost and a data-exposure path. If the service is internal and already gated by IAM or an identity-aware proxy, it can be a useful support tool; on a public endpoint, leave it off.
- Why does --trace_to_cloud matter more in production than locally?Locally you have the dev UI's Trace tab showing every LLM and tool span for the turn you just ran. In production that tab is gone and the turns you care about are ones you did not run. Sending spans to Cloud Trace preserves the ability to attribute a slow or failed turn to a specific model call or tool, which logs alone will not give you.
saying these in an interview costs you the question
- Assuming the in-memory session service is fine because it worked locally
- Expecting local gcloud credentials to be available in the container
- Shipping the .env file as the production secret mechanism
- Deploying with the dev UI enabled on a public endpoint
- Never enabling cloud tracing and debugging production from logs