Tuning a surveillance alert queue
A surveillance rule fires 2,823 alerts to catch 193 real ones. What raising the threshold saves, what it costs, and how to price every step.
The deliverable is a tuning memo for one surveillance rule. It answers three questions before anyone touches a setting: what the rule costs to review today, what a higher threshold would save, and how many real findings the firm gives up to get there. Supervision teams raise thresholds all the time. The memo is what makes it a decision instead of a guess.
The rule here is excessive trading, measured as an annualized turnover ratio. The working version is the supervision and trade surveillance demo, where the threshold is a slider and the tradeoff redraws as you move it. This post is the memo behind the slider.
Synthetic data throughout. Seeded generator, no real firm, branch, or representative.
The queue the rule sits in
- Latest month, firm-wide: 4,310 alerts across eight rule types and four regions. 77.3 percent of them were false positives.
- 2,872 alerts routed to a reviewer. The other 1,438 auto-closed on enrichment before a person saw them, and 2,872 plus 1,438 is 4,310, so nothing falls off the edge of the report.
- 2,580 were manually reviewed, 255 became cases, 80 were substantiated, and 37 escalated. Each stage is the subset of the one before it that survived, so the counts only fall.
- The queue takes in 20.5 alerts per reviewer per day, with a median of 5.6 days to disposition. That reads healthy until the next number: 18.0 percent of open alerts are already past the 15 day review SLA.
- The case-open rate is 5.9 percent: 255 cases out of 4,310 alerts. Substantiated findings are 1.9 percent, and escalations 0.9 percent.
A rule with that hit rate is not a control. It is a tax on the review team, and it buries the alerts that do matter.
One rule, thirteen settings
- At its current 3.0 turnover threshold, the excessive trading rule fires 2,823 alerts. 193 of them are real. The other 2,630 are noise.
- The tuning panel draws one population of 6,000 accounts once. Each account carries a turnover ratio and a hidden flag for whether the underlying activity was genuinely problematic. Every threshold is a recount over that same population.
- The invariant that proves it: true positives plus missed findings equals exactly 196 at all thirteen settings. The rule cannot invent a finding or lose one, only move findings between the caught column and the missed column. That is what makes this a structural recount rather than a shape somebody drew to argue for tuning.
- Watch the two colors invert. At a 3.0 threshold the true positives are 7 percent of the bar. At 8.0 they are 64 percent. The rule gets more precise the whole way up, so the question is never whether tightening helps precision. It is what precision costs.
What the recommended move costs
Moving from the current 3.0 setting to the recommended 4.5:
- False positives fall 67 percent, from 2,630 to 878.
- Reviewer hours fall 62 percent, from 988.0 to 372.8. That is 615.2 hours returned.
- Missed findings rise from 3 to 9, so the move costs six findings.
- 615.2 hours divided by 6 findings is about 103 reviewer hours per finding given up.
That last number is the one to put in front of a supervisory principal, because it is the only version of the tradeoff a person can argue with. "Cuts false positives by 67 percent" sells itself. "Costs one finding per 103 hours saved" invites a real conversation about risk appetite.
Price every step, not just the endpoints
The endpoints hide the shape. Step the threshold one notch at a time and price each step in reviewer hours per additional missed finding:
| Step | Hours saved | Findings given up | Hours per finding |
|---|---|---|---|
| 2.0 to 2.5 | 346.5 | 1 | 346.5 |
| 2.5 to 3.0 | 332.2 | 2 | 166.1 |
| 3.0 to 3.5 | 271.9 | 0 | free |
| 3.5 to 4.0 | 203.0 | 1 | 203.0 |
| 4.0 to 4.5 | 140.3 | 5 | 28.1 |
| 4.5 to 5.0 | 97.7 | 4 | 24.4 |
| 5.0 to 5.5 | 67.6 | 4 | 16.9 |
| 5.5 to 6.0 | 43.0 | 6 | 7.2 |
| 6.0 to 6.5 | 37.5 | 10 | 3.8 |
| 6.5 to 7.0 | 22.7 | 8 | 2.8 |
| 7.0 to 7.5 | 19.6 | 13 | 1.5 |
| 7.5 to 8.0 | 11.9 | 9 | 1.3 |
- The step from today's 3.0 to 3.5 is free on this data: 271.9 reviewer hours back for zero additional missed findings.
- Past 4.0 the price of a finding collapses, from 203.0 hours at one step to 28.1 two steps later and 1.3 at the far end. Findings get cheap to give up, which is the point where tuning stops being tuning and becomes a coverage cut.
- The flat stretch between 3.0 and 4.0 is where the honest gain lives. Everything past it is a policy decision about risk appetite, not an efficiency win, and it should be written up that way.
Be careful with that free step. There are only 196 problem accounts in this population and the slider moves in half point increments, so a step landing on exactly zero is partly an artifact of how wide the bins are. The transferable part is the method, not this row: price every step, find the flat stretch, then check the flat stretch is not just a thin bin before you act on it. A memo that reports a free step without testing it is how a rule quietly loses coverage.
What the numbers say
- Excessive trading is the firm's loudest rule, at 1,218 of the latest month's 4,310 alerts. Tuned separately against its own account population, 93 percent of what it fires is noise at the current setting.
- The firm-wide 77.3 percent false-positive rate and the 18.0 percent past-SLA backlog are the same problem measured twice. Reviewers are not slow. They are reviewing the wrong alerts.
- The same supervision team owes an inspection on every one of its 120 branches under the FINRA Rule 3110 cycle. It has completed 104, at 86.7 percent coverage, with 16 branches overdue. Freed review hours are the most obvious place to fund that backlog.
- None of this required a new system. It required knowing which alerts became findings, which is data the firm already holds in its closed cases.
How I would run this on the job
From 2015 to 2017 I was a branch office examiner at Advisor Group, now Osaic, conducting roughly 100 to 120 branch office examinations a year under FINRA Rule 3110. Those were the firm's own internal inspections of its own branches, not examinations conducted by FINRA. I later worked in regional supervision as a manager and then a director, and after that on the supervisory controls team. The alert queue in this demo is the artifact I spent those years inside.
The process I would run:
- Start from the disposition funnel, not the alert count. What matters is how many of a rule's alerts survive to a case and to a substantiated finding. A rule whose alerts almost never reach a case is a tuning candidate. A rule whose alerts often do is one to leave alone, however loud it is.
- Fix the population before touching the threshold. Draw the accounts once and recount at every candidate setting. If each threshold gets its own sample, the curve moves for reasons that have nothing to do with the rule, and nobody can tell which reason is which.
- Interrogate the label. The hidden problem flag in this demo stands in for a real dispositioned case history. On the job that label comes from closed cases, and it is only as honest as the reviewers' notes. Tuning against a bad label is how a firm automates its own blind spot, so I would sample closed cases and re-read the dispositions before trusting the curve.
- Take the flat stretch first, then stop. The nearly free steps are an efficiency win. The steps past the flat stretch change what the firm is willing to miss, and that decision belongs to the supervisory principal, not to the analyst holding the slider.
- Deliver it as a memo, not a screen. Current setting, proposed setting, alerts before and after, findings given up, reviewer hours returned, and the curve attached. One page. A slider in a meeting is a demo. A memo is a record.
- Write down what you gave up. Six findings is not a rounding error, it is the price of the decision, and an examiner will ask what that price was. A tuning record that reports only the false-positive reduction is not a record.
- Re-run it on a schedule. A threshold set two years ago is tuned to a population that no longer exists. Quarterly on the high-volume rules, annually on the rest, with the prior memo attached so the drift is visible.
- Spend the freed hours where coverage is short. 615 hours is worth nothing if it evaporates back into the same queue. Tie it to something specific, and 16 overdue branch examinations is something specific.
Tools I would use
- The surveillance platform is the source, not the reporting layer. The systems I have actually worked supervision in are FIS Protegent Surveillance for the alert queue, NetX360, Envestnet WMP, and Wealthscape for account and trading detail, and Smarsh for communications review.
- The tuning curve computed upstream in SQL or Python against the dispositioned case history, on a schedule, so every point traces back to specific accounts and specific closed cases. Never computed in the browser, never inside the report.
- Power BI for the memo view and the funnel, refreshed on the same schedule, so the supervisory principal and the analyst read the same numbers on the same day.
- A threshold table the compliance team can version and review, so a setting change leaves a dated record with a name on it instead of living in someone's saved filter.
Key takeaways
- A rule that fires 2,823 alerts to surface 193 real ones is not a control, it is a backlog. Judge a rule by how many of its alerts reach a substantiated finding, not by how many it fires.
- Tune against one fixed population recounted at every threshold. The invariant to check is that true positives plus missed findings stays constant, here 196 at all thirteen settings. Without it you are reading noise.
- Endpoint comparisons hide the shape. Price each step in reviewer hours per additional missed finding and the flat stretch, where hours come back nearly free, becomes obvious.
- Here the next step up returns 271.9 hours for zero measured findings, but with 196 problem accounts and half point increments a zero is partly a bin artifact. Test a free step before you bank it.
- The number that belongs in the memo is hours per finding given up, because it is the only form of the tradeoff a supervisory principal can weigh against the firm's risk appetite.
See it live
- The supervision and trade surveillance demo has the tuning panel as a live slider, plus the funnel, the aging buckets against the review SLA, and branch exam coverage, all filterable by region.