Back to blog

How Do You Audit a Sanctions Screening Vendor's False-Negative Posture? The Pending Column Test

Sanctions screening vendor due diligence: how to audit a tool's false-negative posture. Ask to see the 'unable to verify' state in live output.

August 11, 2026By OilFlow Intelligence8 min readbuyer_intent

Screening a specific counterparty? Run the free check on the name, no signup.

How Do You Audit a Sanctions Screening Vendor's False-Negative Posture Before You Buy?

You audit it by asking the vendor to show you, in live output, where the system says "unable to verify." A screening tool that has no pending state is not reporting the truth about its own coverage: every registry timeout, every stale list refresh, every ownership chain that dead-ends at a nominee shareholder gets silently rendered as "clear." That is a manufactured false negative, and under the risk-based expectations behind OFAC's SDN List administration, the OFAC 50 Percent Rule, and FATF Recommendation 10, the customer due diligence obligation stays with your institution, not with the vendor whose green tick you relied on.

Why "clear" is the most dangerous word in a screening report

Screening vendors sell on the hit side of the ledger. Demo decks foreground match rates, fuzzy-matching logic, transliteration handling, alias coverage. All of that matters. None of it is where the sanctions exposure actually sits.

The exposure sits in the checks the system started and could not finish. A corporate registry API that returned a 504. A consolidated list that failed its scheduled refresh and quietly served yesterday's snapshot. A vessel identity that could not be resolved because the IMO number was reused, the AIS track went dark for eleven days, or the name changed twice in a quarter. A beneficial ownership chain that terminated at a corporate services provider in a jurisdiction with no public register.

Each of those is an honest "I do not know." A well-built screen surfaces it as pending or unable to verify and pushes the decision back to a human. A poorly built one has no vocabulary for uncertainty, so it defaults to the absence of a hit and reports clear. Your MLRO then signs off a file that contains a confidence the tool never had.

The asymmetry is what makes this a procurement question rather than an operational one. A false positive costs you analyst hours. A false negative costs you a breach, a voluntary self-disclosure, and a remediation programme.

What the pending column should actually look like

In a defensible output, every screening step carries three possible terminal states, not two:

  • Match / potential match. A name, entity, vessel, or address hit a list or a watchlist with a scored confidence.
  • No match, checks complete. The step ran against a list of known provenance and known refresh time, and returned nothing.
  • Unable to verify. The step could not be completed. The record must say why, against what source, and at what timestamp.

The third state is the one you are buying. It is the difference between a screening report and a screening claim.

A credible pending record contains, at minimum: the source or list attempted, the failure mode (timeout, authentication error, refresh failure, no record found, ambiguous record), the timestamp, and the residual risk the failure leaves open. "Ownership chain unresolved above 34 percent aggregate; OFAC 50 Percent Rule position cannot be determined" is a usable sentence for a case file. A green tick is not.

Where checks legitimately fail: the mandate chain and the 50 Percent Rule

In physical commodities, the failure points are predictable, which means a vendor has no excuse for not modelling them.

The mandate chain is the first. A cargo enquiry arriving as an LOI or an ICPO frequently travels through a stack of self-described mandates, brokers, and "direct facilitators" before it reaches anyone with title to the product. Each layer adds a corporate name to screen and subtracts a layer of verifiable relationship. When the chain is a layer cake of newly incorporated intermediaries in three jurisdictions, the correct screening output is not "no adverse findings." It is "counterparty relationship to the titleholder unverified across four intermediaries."

The OFAC 50 Percent Rule is the second and the most instructive. OFAC's guidance treats entities owned 50 percent or more, directly or indirectly, in the aggregate, by one or more blocked persons as themselves blocked, whether or not they appear on the SDN List by name. Aggregation across indirect holdings is precisely where an ownership graph runs out of public data. The rule is a mathematical test applied to information that is often unobtainable. A screening tool that cannot say "aggregate ownership incomplete" cannot support a 50 Percent Rule conclusion at all, and any output implying otherwise is fiction.

The vessel leg is the third. Dark fleet activity turns identity resolution into a probabilistic exercise: AIS gaps, flag hopping, name changes, ship-to-ship transfers in known STS corridors, and reused or spoofed identifiers. A screen that resolves a vessel to a single confident identity in every case is not resolving vessels, it is guessing. For a distillate cargo moving as EN590, the vessel history is frequently the only place where the commercial story and the sanctions story diverge.

The payment leg is the fourth. A DLC issued by SWIFT MT700 introduces issuing bank, confirming bank, and beneficiary details that may screen independently of the trade parties. If the tool cannot show which of those legs it actually completed, your file has a hole in it.

The multi-list reality: refresh failure is a silent control failure

There is no single list. A functioning screen touches, at minimum, the OFAC SDN List and the Non-SDN lists including the Sectoral Sanctions Identifications List, the BIS Entity List, the OFSI consolidated list in the UK, the EU consolidated list, the UN Security Council Consolidated List, and whichever national regimes your licence footprint requires.

Every one of those is a separate feed with a separate refresh cadence and a separate failure mode. A designation published today and ingested next Tuesday leaves a window in which your tool returns a clean result for a designated party. That is not a hypothetical edge case, it is routine list plumbing.

So the question is not "do you cover these lists." Every vendor covers these lists. The question is: when a refresh fails, who is told, how quickly, and does the failure appear on the individual screening record produced during the stale window? A vendor that logs refresh failures to an internal ops dashboard and never surfaces them to the compliance user has decided that your regulator's problem is not their problem.

The four questions to run in your next demo call

These take under fifteen minutes and they are the whole test.

  1. Show me the pending state in live output. Not a slide, not a data dictionary. Run a screen against a deliberately incomplete subject and show me the record. If no such state exists in the product, the tool reports confidence it does not have, and you can end the call.
  2. What triggers it? Ask for the trigger list: source timeout, authentication failure, list staleness beyond threshold, ownership chain terminating below the aggregation threshold, unresolved vessel identity, ambiguous name with no discriminating identifier. Ask which thresholds are configurable by you and which are hard-coded by them.
  3. How are list refresh failures logged and disclosed? Ask whether a stale-source flag propagates onto every screening record generated during the outage, or only into an ops log. Ask for the retention period on those logs.
  4. Can I export an audit trail proving which checks completed? You need a per-case export that a regulator, an auditor, or your own MLRO can read: source attempted, result, timestamp, failure reason, and the identity of the reviewer who cleared each pending item. If the export is a PDF summary with no source-level granularity, it will not survive examination.

A vendor confident in its methodology answers all four in the demo. A vendor that routes you to a follow-up call with a solutions architect has told you something useful.

Reading the answers: what good and bad look like

Good answers are specific and slightly uncomfortable. "Our registry coverage in these jurisdictions is thin, so ownership checks there pend more often than they clear" is a vendor describing reality. "We flag stale sources on the record and block auto-clear after four hours of staleness" is a vendor that has thought about the control.

Bad answers cluster around three phrases. "Our coverage is comprehensive" is a non-answer about failure handling. "We would never return a false negative" is a claim no screening system can make. "The system defaults to clear if nothing is found" is the admission you were looking for, stated as a feature.

If you want the demo-call checklist and the pending-state test written up as a one-page procurement artefact, request a walkthrough or subscribe to the OilFlow Intelligence briefing.

What compliance teams should do

  • Write the pending state into the requirement, not the wishlist. Make "system must expose an explicit unable-to-verify terminal state at source level" a pass/fail criterion in the RFP, alongside list coverage.
  • Test with a broken subject, not a clean one. In every demo and every annual re-tender, screen a subject you know is unresolvable. Watch what the tool does with ignorance.
  • Set an internal rule that pending never auto-ages into clear. A pending item closes when a named human closes it with a reason. Nothing expires quietly.
  • Tie pending volume to your risk appetite. If a counterparty file carries unresolved 50 Percent Rule aggregation or an unresolved vessel identity, that is an escalation trigger, not a data-quality nuisance.
  • Make refresh-failure disclosure contractual. Require that stale-source status propagates to individual screening records and that outage logs are retained for the life of your record-keeping obligation.
  • Rehearse the exam question. Your MLRO should be able to produce, for any cleared file, the list of checks that completed, the list that did not, and who signed off the gap.

None of this is a tooling problem, and no vendor can absorb it for you. Under FATF Recommendation 10 and the reliance provisions in the FATF standards, the ultimate responsibility for customer due diligence stays with the institution regardless of who runs the check. A screening vendor supplies evidence. Your firm supplies the judgment, owns the control, and carries the exposure when a chain goes unverified. The only thing you can reasonably demand of a vendor is that it tell you the truth about what it did not manage to check. Buy the tool that says "I do not know" out loud.

A written read within three business days: the broker-scam cluster corpus, a cached US sanctions pre-screen and the 235-jurisdiction tradability matrix. Not the full eight-list screen, and PEP is not screened. No account, nothing to buy.

Verified trade-fraud patterns, sanctions deltas, and regulator actions. Weekly, for compliance and risk teams.

Double opt-in. No spam. The quarterly Compliance Index ships to subscribers first.

This article is part of our scam-cluster intelligence series. Screening a specific counterparty? Run the free check right here.

Paste any company or person name. Free, no signup, answer in seconds. Vessels are not screened here.