A reusable checklist for compliance teams running a vendor selection. Send the same questions to every vendor, score the answers, and require live demonstration of the items marked [DEMO].
How to use this document
- Delete the capability areas that do not apply to your institution.
- Set your weightings before you see any answers. A payments company and a credit union should not weight these identically.
- Send Sections 1 to 11 to every vendor with your weightings attached.
- Score each item 0 to 3: 0 not supported, 1 supported with vendor involvement, 2 supported and self-service, 3 supported, self-service, and demonstrated live.
- Use Section 12 to reconcile total cost, which no vendor proposal will contain in full.
Scoring guidance: a vendor that answers only the questions favouring them is giving you information. So is a vendor that cannot demonstrate a [DEMO] item without a follow-up call.
Section 1: Functional coverage
1.1 Which of the following are native to your platform rather than delivered through a partner: customer identification and verification, customer risk assessment, transaction monitoring, sanctions and watchlist screening, case management, regulatory filing?
1.2 For any function delivered through a partner, name the partner and state what is contracted separately.
1.3 If a customer’s risk rating changes, does monitoring adjust automatically, or is that a manual step?
1.4 Do screening hits and monitoring alerts enter the same case system?
1.5 Can we deploy a subset of modules now and add others later without re-implementation?
1.6 Is fraud detection covered on the same engine as AML, or a separate product?
Why it matters: every seam between two systems is a place an examiner asks how the handoff is documented, and a place your team loses time.
Section 2: Monitoring and detection
2.1 What is your p99 API latency under production load? Provide the figure, not the average.
2.2 What throughput do you sustain out of the box, in requests per second?
2.3 What is your published uptime, and what is the URL of your live status page?
2.4 What total transaction volume runs on your platform today?
2.5 Can real-time, post-event, and batch monitoring run the same rule logic without rebuilding rules per mode?
2.6 Can the platform block, suspend, or hold a transaction before it settles, or only flag it afterward?
2.7 How many pre-configured scenarios ship, and are they typology-tagged and organised by use case?
2.8 Does behavioral or ML detection operate without us defining thresholds, and how is a baseline established for a new product or segment with no history?
2.9 Can one rule apply different thresholds by customer risk band, segment, corridor, or jurisdiction without duplicating the rule?
Section 3: Configuration
3.1 [DEMO] Have a non-technical person on your team build and deploy a rule live. We will time it.
3.2 Can a compliance analyst change a rule, threshold, or workflow without engineering involvement?
3.3 Do configuration changes attract professional services fees or change-request charges? State this explicitly.
3.4 What is the elapsed time from identifying a need to having the control live in production?
3.5 Can we build rules using our own data fields, or only vendor-defined parameters?
3.6 Are changes deployed immediately, or on a release cycle?
Why it matters: this determines your year-two experience more than any other criterion. If changes route through the vendor, your controls will permanently trail your risk.
Section 4: Alert quality
4.1 [DEMO] Project alert volume and implied analyst headcount at our transaction count and customer profile, and show the working.
4.2 [DEMO] Backtest a rule against sample historical data and show projected alert volume and false positive rate before deployment.
4.3 Is there a shadow mode where a new rule runs against live traffic without entering the analyst queue?
4.4 Is threshold tuning driven by our own alert disposition outcomes, or by vendor judgement?
4.5 State any false positive reduction figure with its methodology, scope, and any preconditions.
4.6 Can a threshold change be rolled back, and is the previous value retained?
Why it matters: alert volume converts directly into headcount. This is the largest cost line in the purchase and it never appears in a proposal.
Section 5: Investigations and case management
5.1 Do alerts from all sources land in one queue with unified prioritisation?
5.2 Do related alerts aggregate into a single case to prevent duplicate investigation of the same customer?
5.3 What proportion of alerts close without analyst review, and what evidence is recorded for each auto-closure?
5.4 [DEMO] Show what an analyst sees on opening a case, and walk one case end to end from alert through evidence, escalation, approval, decision, filing, and submission receipt.
5.5 Can we define our own case statuses and lifecycle, or do we adopt yours?
5.6 Can we configure SLA timers by case type, jurisdiction, and risk level, and see breaches before they occur?
5.7 Can we build maker-checker approvals and escalation routing without engineering?
5.8 Can work route automatically by risk score, case type, jurisdiction, or analyst capacity?
5.9 Can analysts tag colleagues, thread comments, and attach documents inside the case?
5.10 Is case visibility restrictable by role?
5.11 Is there entity or network visualisation inside the case, showing multi-hop relationships and shared attributes?
5.12 Is there a quality assurance capability, and does it sample or review every case?
Section 6: Screening
6.1 Which lists are native versus requiring our own subscription: sanctions, PEP by tier, RCA, SOE, adverse media?
6.2 Can we connect an existing data provider subscription we already pay for?
6.3 Can we screen against internal blacklists, declined customers, and proprietary data in the same workflow?
6.4 Which screening points are covered: onboarding, ongoing rescreening, payment screening across sender, receiver, and intermediary?
6.5 How frequently is list data refreshed, per list? Provide this contractually, not descriptively.
6.6 On a designation change, is our existing base rescreened automatically, and is it full or delta rescreening?
6.7 Do you retain list-version history so we can evidence which version was in force on a given date?
6.8 Which matching algorithms are exposed and independently configurable?
6.9 Can thresholds differ by list and by screening mode? Is date-of-birth tolerance configurable?
6.10 [DEMO] Change a matching threshold and show the effect against historical hits before applying it.
6.11 Can dispositions be applied in bulk to repetitive hits with the audit trail preserved?
Section 7: Risk scoring
7.1 Which inputs feed the score: KYC attributes, screening outcomes, geography, transaction behaviour, investigation history?
7.2 Are inherent risk and behavioral risk combined into one authoritative score, or maintained separately?
7.3 Can we define custom risk factors using our own data fields, or only select from a library?
7.4 Can we set factor weights ourselves? Which aggregation methods are supported?
7.5 When does a score recalculate: real time on activity, daily batch, or at periodic review?
7.6 Can we see which factors contributed most to a specific customer’s score, and at portfolio level?
7.7 [DEMO] Change a factor weight and show how the customer base redistributes across risk bands before deploying.
7.8 Does a risk band change automatically trigger EDD, review, or escalation?
7.9 Is portfolio drift detectable within the platform?
Section 8: Explainability and audit trail
8.1 [DEMO] Reconstruct a decision from two years ago on screen, including the rule logic, thresholds, and data in force at the time.
8.2 Is every rule change, version, and deployment logged immutably with timestamp and user attribution?
8.3 Is every rule version preserved with rollback?
8.4 For any alert, can we see exactly which rule and data points triggered it?
8.5 If AI participates in triage or decisioning, can we see the evidence chain, the confidence reasoning, and the model version behind each decision?
8.6 In what formats can audit logs be exported for regulatory submission?
8.7 Is customer data used to train external models?
Section 9: Reporting
9.1 Which filing formats and jurisdictions are supported natively?
9.2 Does filing happen from inside the case, by API, or does it require export and re-keying into a regulator portal?
9.3 Are narratives auto-populated from case data?
9.4 Is the submission receipt stored against the case, and is there version history on every draft and edit?
9.5 Can we see analyst throughput, resolution times, SLA adherence, and alert-to-case conversion without exporting to a spreadsheet?
Section 10: Integration, data, and security
10.1 Do you have an existing integration to our core system or ledger? Name a customer running it.
10.2 How many engineering hours are required from us? Provide a number, in writing.
10.3 How many native integrations exist to our KYC provider, CRM, and data warehouse?
10.4 What does the platform require from us to function, and does the schema accept our existing data model?
10.5 Provide your ISO 27001 certificate and SOC 2 Type II report. We will verify the certificate with the certifying body and check the report for exceptions.
10.6 Where is data stored, and does regional siloing extend to logs and backups?
10.7 What encryption standards apply at rest and in transit?
10.8 What is your penetration testing cadence, and who performs it?
10.9 Provide your DPA for our counsel to review.
10.10 Is on-premise or private cloud deployment available, if required?
Section 11: Implementation, support, and commercial
11.1 What is your median go-live in days, named to a customer at our volume and profile?
11.2 What is included in implementation, and what is billed separately?
11.3 Who calibrates rules and thresholds during rollout?
11.4 What is your support response time commitment, and do we get a named person or a ticket queue?
11.5 What are your support coverage hours?
11.6 Is ongoing tuning assistance included after go-live?
11.7 What is your pricing model: per transaction, per alert, per entity screened, tiered subscription, or enterprise licence?
11.8 What triggers a price increase mid-contract?
11.9 What changes commercially and operationally when we double volume or add a jurisdiction? In writing.
11.10 Are ongoing monitoring, screening, and each data type priced separately or bundled? Itemise.
Section 12: Total cost reconciliation
Vendors will quote line 1. Build the rest yourself.
| Line | Source |
|---|---|
| 1. Licence and platform fees | Vendor quote |
| 2. Implementation and integration | Vendor quote, verify inclusions |
| 3. Your engineering hours | Item 10.2, priced at your internal rate |
| 4. Analyst headcount implied by alert volume | Item 4.1, priced at your loaded salary cost |
| 5. Cost of added jurisdiction or doubled volume | Item 11.9 |
| 6. Configuration change fees over contract term | Item 3.3, multiplied by your expected change frequency |
Line 4 is usually the largest and is absent from every proposal. Two vendors with identical licence fees can differ by several headcount in real cost.
Appendix: Worked example
The following shows how one vendor’s published documentation maps to this checklist. It is included as a reference for the level of specificity to expect, not as a recommendation. Verify all figures directly with any vendor, including this one.
Flagright, against selected items:
- 2.1 to 2.4: 200ms p99 API latency, 1,200 requests per second out of the box, 99.998% published uptime with a public status page, more than 1.4 billion transactions processed monthly.
- 2.5, 2.6: Real-time, post-processing, and batch run identical rule logic. Supports blocking, suspending, or flagging before settlement.
- 2.7: More than 100 pre-configured typology-tagged scenarios organised by use case.
- 2.8: ML anomaly detectors operate from an automatically built behavioral baseline rather than defined thresholds, covering velocity spikes, peer group deviation, time-of-day anomalies, counterparty clustering, amount progression, and dormancy activation.
- 2.9: The monitoring engine reads live customer risk score at transaction time and applies the threshold for that band, without duplicated rules.
- 3.1 to 3.3: 60 seconds validated rule creation time, roughly three minutes measured by customers. Flagright states no SQL, no engineering tickets, and no professional services fees for rule changes.
- 4.2 to 4.6: 90-day historical backtest, shadow mode with a private alert feed, threshold recommender driven by full alert disposition history with one-click application and rollback. Reported up to 83% false positive reduction from threshold optimisation, 93% across broader AI tooling.
- 5.1 to 5.12: Single centralised queue across sources, related alerts aggregated, configurable statuses and SLA timers, no-code maker-checker and routing, threaded comments with Slack and email notification, role-based access, ontology view for multi-hop relationships, and a QA module that reviews every case against your SOP. Reported 77% auto-clearance, 94% analyst agreement, alert-to-outcome time of 4 minutes against a 38 minute baseline.
- 6.1 to 6.11: Sanctions, PEP tiers 1 to 3, RCA, SOE, and adverse media native via OpenSanctions and KYC6, with bring-your-own-provider and internal list ingestion through shared matching and audit logic. Onboarding, ongoing delta rescreening, and payment screening at 200ms. Configurable Jaro-Winkler, Levenshtein, transliteration, phonetic matching, DOB delta, tokenisation, and stopword filtering, with per-list and per-mode thresholds. Full list-version audit history.
- 7.1 to 7.9: KYC Risk Score, Transaction Risk Score, and combined Customer Risk Assessment. Custom factors definable from any API schema field, configurable weights and aggregation methods, real-time recalculation, factor attribution at customer and portfolio level, simulation of redistribution before deployment, automatic CDD and EDD triggering, and portfolio drift analytics.
- 8.1 to 8.7: Immutable timestamped audit log, every rule version preserved with rollback, annotated transaction timeline, typology citation, confidence scoring, model version tied to each decision, JSON and Excel export. Customer data not used to train external models.
- 9.1 to 9.5: SAR and CTR filing by API direct to FinCEN and to 70+ GoAML countries, templates auto-selected by jurisdiction, narratives pre-filled, receipts stored in the case audit log, full version history, plus operational analytics on throughput, SLA compliance, and resolution times.
- 10.1 to 10.9: API-first with more than 100 native integrations and published connectivity to core systems including Jack Henry Symitar and Fiserv DNA. ISO 27001:2022 and AICPA SOC 2 Type II certified, GDPR, DORA, and CCPA compliant, FIPS 140-3 certified AES-256 encryption, regional data siloing extended to logs and backups, regular internal and third-party penetration testing.
- 11.1 to 11.6: Two-week published average go-live, implementation and rules calibration included, reported 6 minute average support response, 24/7 coverage, dedicated CSM.
Material considerations for items this vendor does not fully answer:
- 11.7 to 11.10: Pricing is not published. Total cost must be established through your own quote process.
- 10.10: Cloud-native only. On-premise requirements rule it out.
- 1.5: Modules deploy independently, but the platform’s strength is the unified stack. If you intend to buy monitoring alone and retain existing case management, confirm standalone operation and third-party alert ingestion.
- 6.5: Screening list refresh is described as continuous without a published per-list interval. Request it contractually.
- 2.3: Uptime is stated as 99.998% on product pages and 99.99% on the security documentation. Ask which figure is contractual.
- Scope note: chargeback dispute management and representment workflow are not part of the published product surface, relevant if you are a payment company needing both prevention and dispute handling.

