What makes text-message and push checks unsuitable for a suite that runs on every change?
answer
- Every send is billed and real
- The recipient belongs to somebody
- Volume from one sender looks abusive
- Throttling degrades the product, not the suite
- Substitute the transport, schedule the real one
basics
~20 sEach send costs real money and reaches a live endpoint belonging to somebody. A suite firing them on every change burns budget, needs consented recipients, and produces sender volume that anti-abuse limits throttle, degrading the product's own deliveries.
solid answer
~40 sThree constraints, and none is about test quality. **Cost** is charged per send, and a check that runs on every change across every branch multiplies a fraction of a cent into a line item nobody budgeted. **Consent**: the recipient is a real number or a real device belonging to a person, so invented numbers, live customer accounts and colleagues' handsets are all the wrong pool. **Volume**: repeated sends from one sending identity are exactly what receiving networks throttle, and the throttle degrades the product's real deliveries, not just your suite. The design answer is to substitute the transport for routine runs — hand the outbound message to a recorder inside the run and assert on it — and keep a small scheduled set of real-channel checks against a dedicated, consented pool, capped and non-gating.
code
yaml · 13 linesnotifications:
transport: recording-stand-in # routine runs: nothing leaves the system
assert_at: handoff
real_channel_checks:
schedule: nightly
transport: real
recipients:
- consented-number-pool
- team-owned-device-pool
max_sends_per_run: 12
gates_change: false
alert_on: delivery-rate-trendgo deeper
Be ready to say that a text message or a push in a check is a real send: it is billed and it reaches a live endpoint, so a suite does not fire them on every change.
Expect to explain the substitution: routine runs assert the outbound message where the product hands it off, while the real channel is exercised on a schedule against recipients that exist for that purpose.
Show judgment about blast radius — that volume from one sending identity is throttled by the receiving side, and the throttle degrades the product's real deliveries rather than only your run.
Own the trade explicitly: how much real-channel coverage the organisation buys, who holds and retires the consented recipient pool, and what evidence justifies that spend while the substituted checks stay green.
## Three constraints, none of them about test quality A check that sends a real text message or a real push notification is a good check. The reason it does not belong in the suite that runs on every change has nothing to do with what it verifies and everything to do with what it does outside the run. - **It costs money per send.** Not per suite — per send. - **It reaches a live endpoint** that belongs to somebody, which makes the recipient a consent question rather than a configuration question. - **It produces volume from a single sender**, which is precisely the signature that receiving networks and delivery services throttle. Only the third is commonly a surprise, and it is the one that hurts, because the damage lands on the product rather than on the suite. ## Cost, worked Prices move, so do the arithmetic with numbers you state rather than numbers you half-remember. Suppose a text send costs a tenth of a cent and the suite holds twelve cases that each send one. A team merging ten times a day, plus a nightly run, is about 330 suite runs a month — roughly 4,000 sends, or a few dollars. Trivial. Now scale one axis at a time: building on every push rather than on every merge multiplies the runs several times over; a policy that re-runs failed cases doubles the worst days; two more suites copy the pattern; and the price per message is not the same everywhere you send. The honest conclusion is not "it is expensive". It is that the cost is **unbounded by anything in the check's own design** — it grows with pipeline decisions made by people who are not thinking about your channel at all. ## Consent is not a detail you can defer The recipient of a text message is a number on a live network, and the recipient of a push notification is a device somebody carries. Both belong to a person, and that creates obligations no amount of "it is only a test" removes. - **Numbers you did not allocate may be in service.** A plausible-looking number invented by a case can be somebody's real one, and what they get is an uninvited message from your brand. - **Live customer accounts are never test recipients.** Matching production conditions does not justify messaging people who never agreed to it. - **Colleagues' personal handsets are a poor pool.** Their agreement is informal and revocable, the numbers churn with staffing, and a nightly run wakes a real person at three in the morning. The workable answer is a **dedicated recipient pool** — numbers rented for this purpose and devices the team owns — documented, consented, and disposable. ## Volume, throttling, and whose blast radius it is Where a receiving network or a delivery service applies anti-abuse limits, those limits usually attach to the **sender**, not to the individual message. A suite firing hundreds of sends a day from the same sending identity is, from the outside, indistinguishable from a sender behaving badly. The consequence is not that your check fails. It is that the sending identity is throttled or downgraded, and the next messages to be delayed or dropped are the ones real users are waiting for. That is the argument that settles the discussion: routine real-channel checks put a production capability at risk in order to verify something you can verify more cheaply somewhere else. ## The split that works | | Routine runs (every change) | Periodic real-channel runs | |---|---|---| | Transport | A recording stand-in the product hands the message to | The real channel | | What it asserts | Recipient, content, and that exactly one message was produced | That a message addressed to a controlled endpoint actually arrived | | Frequency | Every change | Scheduled, a handful of sends per run | | Recipients | None — nothing leaves the system | The dedicated, consented pool | | Gates the change | Yes | No | The stand-in is the load-bearing piece. The product is configured so the outbound message is handed to a receiver inside the run instead of to the real channel, and the case asserts on what it was handed. That check is free, instant, deterministic, and it catches the great majority of real defects: wrong recipient, wrong content, sent twice, not sent at all. ## Sizing and governing the small real set 1. **Cap the sends per run explicitly**, in configuration, so nobody grows it by accident. 2. **Run it on a schedule, not on a change**, and keep it non-gating; a channel you do not own must never block a merge. 3. **Name an owner for the recipient pool** — who rents the numbers, who holds the devices, who retires them. 4. **Alert on the trend, not on a single miss**, because one absent arrival is expected and a falling delivery rate is not. What you give up is early warning on the last hop: a misconfigured sending identity, an expired credential, a message that reads correctly in the outbound record and badly on a device. That is exactly what the small scheduled set exists to find, and finding it a few hours late is a trade almost every team should take.
- If routine runs never touch the real channel, what can only the scheduled set catch?Everything between the handoff and the recipient: a misconfigured sending identity, an expired credential for the delivery service, a message that reads correctly in the outbound record and badly on a device, a recipient format the network rejects. None of that is visible where the product hands the message off, which is the whole reason a small real-channel set exists.
- Someone proposes sending to colleagues' own phones instead of a dedicated pool. Why refuse?Their agreement is informal and revocable, their numbers churn with staffing, and a nightly run wakes real people. A dedicated pool rented for the purpose is auditable, disposable and replaceable, and it keeps an on-call engineer's personal handset from becoming an undocumented dependency of the suite.
Checking the alarm by setting off the real siren works, but do it hourly and the neighbours stop listening and the council takes the siren away.
saying these in an interview costs you the question
- Assuming test sends are free because production sends are budgeted
- Sending verification texts to live customer numbers from a run
- Treating throttling as the suite's problem rather than the product's
- Running real-channel checks on every commit to catch more
- Using colleagues' personal numbers as the suite's recipient pool