A merchant-risk model's most predictive columns are all self-declared - what do you report?
answer
- importance says nothing about who sets the value
- the historical applicants had no motive
- put price beside gain, column by column
- the cheapest edit set, in currency
- the load-bearing columns are the free ones
basics
~20 sAccuracy was measured on applicants with no reason to misstate those fields. Report the share of score sitting on columns the subject retypes for free, and what the cheapest application that flips a decline costs.
solid answer
~50 sReport the importance ranking next to two columns nobody put in the model card: who sets each field, and what changing it costs that person. If most of the score rests on fields the applicant asserts, the model is accurate on a population that had no motive to lie, and that is all the offline metric establishes. Then price the cheapest edit set that turns a decline into an approve — a retype costs nothing, a domain registration costs a fee, tenure costs months, a funded balance costs cash — and state that figure. Finally give options rather than a verdict: corroborate free fields against priced ones and treat disagreement as a signal, cap how much of the score a free column can move, route disagreement to review, or drop the column. Each costs accuracy on honest traffic, which is why it is a decision, not a bug fix.
code
text · 9 linescolumn importance value set by cost to change
--------------------------- ---------- --------------- ----------------------
stated_monthly_volume 0.31 applicant retype
stated_business_category 0.18 applicant retype
website_age_days 0.12 applicant one registration fee
verified_avg_balance 0.09 bank connection hold real funds
months_settled_history 0.07 card network months of real trading
...
self-declared columns account for 0.61 of total importancego deeper
Know that a feature-importance ranking answers how much a column moves the score and nothing about who fills that column in or what changing it would cost them.
Explain why a self-declared column can be genuinely predictive historically and worthless under a motivated author, and why the most predictive column is the most attractive to move.
Show the deliverable: importance beside authorship beside price, plus the currency cost of the cheapest application that flips a decision, plus options with their accuracy costs attached.
Own the escalation: each remedy spends accuracy, integration budget or reviewer hours, and choosing which one to fund is your call to frame with numbers rather than to hand over as a warning.
## The finding You rank the model's columns by importance and the top of the list is entirely fields the applicant types in. This is not a coincidence and it is not a modelling mistake. Self-declared fields are often genuinely informative, because in the historical data the people filling them in had no reason to misstate them, and because they capture things nothing else in the record captures. The model found what was there. The problem is what the ranking is silent about. Feature importance says how much a column moves the score. It says nothing about who sets the column's value, or what setting it differently costs that person. Put those two facts beside the importance and the picture changes: the columns carrying the decision are exactly the ones an unfunded, model-blind adversary controls completely. ## Why the offline evaluation looked fine A held-out split is drawn from the same historical population. Those applicants were not optimising against this model — most had never thought about it. Predictiveness measured on a non-adversarial population is a statement about that population, and it does not survive an author who knows the column exists and wants a particular outcome. This is the same direction-of-claim error that runs through the whole area: a high number proves the attack that was run failed, and no attack was run here at all. Worse, the incentive runs the wrong way. Whatever column the model leans on hardest is the one worth most to move, so predictiveness and attractiveness to an adversary are correlated by construction. A column that is both load-bearing and free is the worst combination available. ## What the report should contain **One table.** Column, importance share, who sets the value, and what a change costs the subject. That table is the deliverable; it is usually the first time anyone has seen those facts together. **One number.** The cost of the cheapest set of edits that changes a decline to an approve. Sum the per-field prices of the fields you would have to move: a retype is free, a domain registration is a fee, tenure is months of waiting or the price of an aged entity, a funded balance is cash held for as long as the check looks back. Report it in currency and calendar time. What that figure is worth against the value of an approved account is somebody else's exercise; your deliverable is the price your feature set charges. **One reproducibility statement.** Say whether the cheapest edit set flips the decision reliably or only for applications already close to the line, because those are different findings with different urgency. ## The options, and why none of them is free - **Corroborate.** Cross-check a free field against a priced one and treat disagreement as its own signal. This is often the strongest move, because the adversary now has to move a priced field to keep the story consistent. It costs an integration and it produces false disagreements on honest applicants with unusual profiles. - **Cap the influence.** Bound how much of the score any free column may contribute. This costs accuracy on the honest majority, and the loss is real and measurable. - **Route to review.** Send disagreement or high-declared-value cases to a human. This costs reviewer hours at a rate that scales with volume. - **Drop the column.** Cleanest, most expensive in accuracy, and sometimes the right answer when a column is both dominant and costless to set. - **Add attested columns.** Moves score mass onto fields that cost money, which is the right direction, and costs integration, consent, latency and applicants who drop out or genuinely cannot provide them. ## What not to do Do not report the importance ranking on its own, and do not report a robustness figure produced by perturbing all columns slightly — no applicant does that. Do not present the fix as obvious. And do not accept 'the fraud rules cover it' without asking which rule reads which of these columns, because a rules layer that reads the same self-declared fields inherits the same price list. ## The chair you are sitting in This is the red-teamer's report, not the modeller's. The deliverable is a priced inventory of the input surface and the cost of the cheapest successful application. Which trade to make against that price is a decision somebody else funds — but they cannot make it until the price is on paper.
- How do you put a money figure on this?Cost the cheapest edit set that flips the outcome, field by field: a retype is zero, a domain is a registration fee, tenure is months or the price of an aged entity, a funded balance is cash held over the lookback window. Sum them and report a currency figure and a calendar delay per fraudulently approved account. Comparing that against what such an account is worth to an abuser is a separate exercise; your deliverable is the price your columns charge.
- Why did the offline evaluation look fine?Because the historical rows were written by applicants with no reason to misstate them, so a self-declared column genuinely correlated with outcome. Predictiveness measured on a non-adversarial population is a fact about that population. It says nothing about behaviour under an author who knows the column exists, knows it is weighted heavily, and wants a particular decision.
- Does adding more attested columns solve it?It shifts score mass onto fields that cost money, which is the right direction, but every attested column costs an integration, consent, latency at signup and applicants who drop out or legitimately cannot provide it. You are trading funnel for adversarial cost. Make that trade deliberately against the priced inventory, rather than by adding whichever integration happens to be easiest to buy.
- The fraud team says their rules already handle this. What do you check?Which columns those rules read. A rules layer that keys on the same self-declared fields inherits exactly the same price list, so it raises the applicant's cost by nothing. Rules that read a priced or attested field, or that compare a declared value against one, are doing real work. Ask for that mapping before treating the rules layer as a compensating control.
saying these in an interview costs you the question
- Reports importance without saying who sets each field
- Calls high offline accuracy evidence of resistance to gaming
- Proposes dropping every self-declared column with no accuracy estimate
- Assumes a downstream rules layer is automatically a compensating control
- Prices the attack in effort but never in money or calendar time