Field-provisioned solar inverters accept firmware and config over the air — how do you model that update path as an entry point?
answer
- Which way does the arrow point?
- The device is the one that must verify
- Availability and safety outrank confidentiality
- One credential, the whole installer base
- Configuration is as dangerous as firmware
basics
~20 sModel the update channel as an inbound flow crossing into the device, so the question is what the device verifies before accepting, not just how you authenticate it. A shared installer credential and a fleet-wide push make one compromise reach every unit at once.
solid answer
~50 sMost models draw the device-to-cloud flow and stop. The update path runs the other way: it is an inbound data flow crossing a trust boundary into equipment that controls a physical process, so the security question is what the inverter demands before it accepts an image or a setpoint change. Spoofing of the update source and Tampering with what is accepted dominate, and because the asset is safety and grid availability rather than customer records, Denial of Service outranks Information Disclosure — the reverse of a typical web model. Two structural properties make it worse: a shared installer credential means one leaked value provisions or reconfigures any unit in the field and cannot be revoked per unit, and a single fleet-wide push is a lever that turns one compromise into simultaneous mass impact. I would push for per-unit enrolment credentials, device-side verification including a refusal to downgrade, staged rollout with a halt, and a recovery path that does not require a truck.
go deeper
Be able to say that an over-the-air update is a way into the device, not just a maintenance convenience, and that the device needs some way to tell a real update from a fake one.
Explain the flow direction and the trust zones involved — back end, transport, installer, premises network — and why a shared commissioning credential cannot be revoked or attributed.
Show judgment by ranking threats against a safety and availability asset rather than by habit, demanding version monotonicity and equal treatment of configuration, and designing staged rollout so one push cannot reach the whole fleet.
Own the fleet-level strategy: how much update reach any single action is allowed, how units already deployed under a weaker policy get migrated, and how you keep field commissioning workable for installers while removing the shared secret.
## The system A manufacturer runs a fleet of grid-tied solar inverters installed on homes and small commercial roofs. Each unit accepts firmware images and configuration changes over the air. Installers commission units in the field using a credential that is the same across the installer base, and units sit on customer premises networks the manufacturer has no visibility into. ## The modeling move: the flow points inward Teams instinctively model telemetry — the device reporting outward to the cloud — and reason about who can read it. The update and provisioning path is the flow that matters, and it points the other way: data crossing a trust boundary *into* equipment that controls a physical process. Once you draw the arrow correctly, the question stops being "how does the cloud know it is talking to the inverter" and becomes "what does the inverter demand before it accepts an instruction that changes its behaviour". Both directions need answering, but only the inbound one can hurt someone. Drawing it also exposes the third participant. There are not two zones here but at least four: the manufacturer's back end, the update transport, the installer at commissioning time, and the premises network the unit lives on for the next fifteen years. ## Which threats dominate, and why that ordering is unusual The asset is the safe operation of a physical process and the availability of generation, not customer records. That reorders STRIDE relative to a typical web model: - **Tampering** (integrity) — an altered image or a changed setpoint makes the unit behave outside its safe envelope, or makes it stop exporting. - **Denial of Service** (availability) — a bad push that bricks units removes generation and needs a physical visit per unit to recover, which is the expensive kind of outage. - **Spoofing** (authentication) — something that is not the manufacturer persuades the unit to accept an update, or something that is not an authorised installer commissions a unit. - **Elevation of Privilege** (authorization) — a maintenance interface reachable from the premises network grants more than the local network deserves. - **Information Disclosure** ranks *low* here. Generation data is not the prize. Saying so out loud is part of a good answer, because a model that spends its effort on the least valuable asset is a failed model. ## The two structural properties worth calling out **A shared installer credential.** One value, used by every installer, embedded in commissioning tooling and passed between contractors. It never rotates because rotating it means retraining and re-tooling the field force. It cannot be revoked for one leak without stopping all commissioning. And it gives no attribution: when a unit is misconfigured you cannot tell which installer touched it, which is a Repudiation problem with real commercial consequence. The fix is per-installer or per-unit enrolment identity established at commissioning, with the shared value demoted to at most an identifier and never an authenticator. **Fleet-wide reach.** The update channel is the one mechanism that touches every unit simultaneously. That is exactly what makes it valuable to an attacker and dangerous when you are merely wrong: the blast radius of a compromised or mistaken push is the entire fleet, at once, on physical equipment. Model the channel itself as a high-value asset and design so that no single compromise or single mistake reaches everything — staged rollout by cohort, a halt condition based on health telemetry, and rate limits on how much of the fleet can change in a window. ## Device-side requirements the model produces The model does not need to specify mechanisms to state the requirements, and the requirements are what an interviewer wants: - The unit must be able to verify that an image or configuration originated from the manufacturer, using something it holds rather than something the sender asserts. - It must refuse to install an older version than it runs, because otherwise every fixed defect remains reachable by anyone who can offer the old image. - Configuration must be treated with the same suspicion as firmware. Setpoints, export limits and grid-protection parameters change behaviour just as effectively as code, and they are routinely protected far less. - Acceptance must be atomic with a fallback: a unit that fails mid-update should come back on its previous image, because the alternative is a physical visit. - Local maintenance interfaces must authenticate as strictly as remote ones. The premises network is not a trusted zone just because it is behind a home router. ## Attacker positions to enumerate Write them down explicitly, because each reaches a different part of the diagram: someone on the site's local network probing the unit's maintenance interface; a compromised update channel or the credentials that drive it; a departed installer still holding the shared commissioning secret; and the owner of the equipment, who may want to raise an export limit for their own benefit. That last one matters — as with any device in someone else's hands, the legitimate owner is a valid attacker position and a model that only imagines outsiders will miss half the misconfiguration it eventually sees. ## Residual The honest model ends with what remains: units already in the field with an older acceptance policy cannot be fixed by design changes to new units, so you need a migration plan and an inventory of which units are on which policy. The manufacturer's own signing and release process becomes a concentrated risk, and how that pipeline is protected is a separate discussion from this one — but the model should name it as a dependency rather than pretend it does not exist.
- Why does refusing to install an older version matter as much as verifying the source?Because an attacker who can offer any legitimately produced image can offer the vulnerable one. If the unit accepts a valid but older image, every defect you have ever fixed remains permanently reachable, and your patch programme buys nothing durable. Version monotonicity — the device tracks what it runs and declines anything below it — is what turns a fix into a lasting change, and it needs a plan for the rare legitimate rollback.
- The team argues the premises network is a trusted zone because it sits behind a home router. How do you answer?A home network holds the customer's own devices, guest devices, and anything already compromised on it, none of which the manufacturer controls or can see. It is a distinct trust zone from the manufacturer's back end and deserves its own boundary on the diagram. The practical conclusion is that a local maintenance interface must authenticate and authorise as strictly as a remote one, and that network position must never substitute for identity.
- Why does Information Disclosure rank low in this model when it usually leads?Because the assets are the safe operation of a physical process and grid-side availability, not personal or financial records. Generation telemetry has modest value if read. Ranking threats by the asset rather than by habit is the point of the exercise — a model that spends its mitigation budget encrypting telemetry while the update path accepts anything has optimised the wrong property.
- How would you bound the damage from a bad push that is entirely your own fault?By treating the update channel as a fleet-wide lever and refusing to pull it all at once: cohort the fleet, roll out in stages, define a halt condition on health telemetry, and rate-limit how much of the fleet can change in a window. Pair that with an atomic install that falls back to the previous image, because on physical equipment the difference between an automatic fallback and a truck roll per unit is the entire cost of the incident.
A hospital's drug cabinet is refilled by a supplier. The lock on the ward door matters less than whether the cabinet checks who is doing the refilling.
saying these in an interview costs you the question
- Models only outbound telemetry and never the inbound update flow
- Treats the premises network as a trusted zone
- Protects firmware but leaves configuration changes unverified
- Accepts any validly produced image regardless of version
- Assumes a shared installer credential can be rotated in practice
- Ranks confidentiality of generation data above availability and safety