Back to blog

How Do You Audit a Screening Vendor's False-Negative Posture Before You Buy?

Screening vendor false negatives: four written questions that reveal whether a tool reports unverifiable checks as pending or silently clears them.

August 16, 2026By OilFlow Intelligence8 min readbuyer_intent

Screening a specific counterparty? Full 7-step dossier — $25, no account, report by email within the hour.

How Do You Audit a Screening Vendor's False-Negative Posture Before You Buy?

You audit a screening vendor's false-negative posture by asking it to show you a result where the check failed to complete, not a result where it found a hit. An honest screening system returns three distinct states, cleared, flagged, and not established, and it never silently promotes the third into the first. This matters because FATF Recommendation 10 requires that where customer due diligence cannot be completed, the institution should not establish the relationship and should consider filing a suspicious transaction report, and because OFAC's 2019 A Framework for OFAC Compliance Commitments names "sanctions screening software or filter faults" among the root causes of apparent sanctions violations it has observed in enforcement matters.

Most vendor demos are engineered around the hit. The salesperson types in a name that sits on the OFAC Specially Designated Nationals and Blocked Persons List, the screen turns red, and the room nods. That demonstration tells you almost nothing about the risk you are actually buying. The clearance gap lives at the other end of the workflow: what the tool does when a list endpoint times out, when coverage does not extend to the jurisdiction of the counterparty, when a name transliterates three ways from Arabic or Cyrillic, or when the ultimate beneficial owner behind a trading intermediary cannot be resolved at all.

The failure mode: a green result that is actually an incomplete result

A screening result is a claim about evidence. "Cleared" is a claim that the named lists in scope were queried, the query returned, and no match above threshold was found. "Flagged" is a claim that a match was returned. "Not established" is a claim that the check did not finish, for a reason the system should be able to name.

The third state is the one that carries the operational risk, and it is the one most tools collapse into the first. When a source is unreachable and the interface still renders green, your file now contains a positive assertion that no evidence supports. Your MLRO will sign off on it. Your first line will onboard against it. And when an examiner reconstructs the decision two years later, the record will say the counterparty was screened clean, when in fact the counterparty was never screened at all against that regime.

That is not a rare edge case in commodities. Heightened Middle East supply risk widens counterparty churn, and churn raises the cost of any screen that quietly reports unknown as clean.

Question one: show me a screening record where a source was unavailable

Put this in writing, in the RFP, and require a screenshot or an exported record rather than a verbal answer. You are asking for an artefact, not a policy statement.

What you are looking for in the artefact:

  • A named state that is visibly distinct from cleared. "Pending", "not established", "incomplete", any label will do provided it does not read as a pass.
  • A reason code. Source timeout, coverage gap, jurisdiction out of scope, insufficient identifiers, ambiguous transliteration. "Error" is not a reason code.
  • A timestamp for the attempt, and the specific list or source that failed.
  • Whether the overall case status inherits the incomplete state. A record with nine green checks and one pending check is a pending record, not a ninety percent clean record.

If the vendor cannot produce such an artefact from a live system, you have your answer. A system that has never needed to render the incomplete state has probably never been built to hold it.

Question two: which lists are in scope, named, with refresh cadence per list

Ask for a table, one row per list, with source, jurisdiction, refresh cadence, and last successful refresh timestamp exposed in the user interface. At minimum you should expect explicit naming of:

  • OFAC SDN List and the OFAC Consolidated Sanctions List, including the Sectoral Sanctions Identifications List and the Non-SDN Menu-Based Sanctions List where relevant to your book.
  • The EU Consolidated List of persons, groups and entities subject to EU financial sanctions.
  • The UK OFSI Consolidated List of Financial Sanctions Targets.
  • The UN Security Council Consolidated List.
  • Applicable national and sectoral regimes for the jurisdictions you actually trade into, plus export control lists such as the US BIS Entity List where dual-use exposure exists.

A count of lists is a weaker signal than a set of named lists with cadence. "Over 1,400 watchlists" tells you nothing about whether the SDN List refreshed this morning or last quarter. Ask which list in the stack has the longest refresh interval, and ask what the system displays if that refresh fails. A vendor that answers the second question without hesitating has thought about the problem.

Ask separately how the tool handles ownership. OFAC's 50 Percent Rule treats entities owned fifty percent or more, directly or indirectly, in the aggregate, by one or more blocked persons as themselves blocked, even where the entity is not itself listed. A screening tool that matches only against literal list strings is not screening for that rule. It should tell you so, in the record, rather than returning green.

Question three: what is returned on transliteration and alias ambiguity

This is where physical oil trade fraud does its most reliable work. A mandate chain for an EN590 gasoil parcel can run through five or six intermediaries, each presenting an ICPO, an LOI, and a draft DLC MT700, and each layer introduces a new corporate name in a new jurisdiction with a new transliteration of the same principal. The layer cake is not incidental. It is the product. Its function is to ensure that no single screening query resolves the party who actually controls the cargo, and dark fleet tonnage supplies the same ambiguity on the vessel side.

So ask the vendor directly: when a name admits multiple defensible transliterations, or when the identifiers supplied are insufficient to disambiguate between a listed party and a similarly named unlisted party, what does the system return?

There are only two acceptable answers. Either it returns a flag for analyst adjudication, or it returns not established with a reason code. "It returns clean" is a false negative dressed as efficiency. Note also that the fuzzy-matching threshold is a policy setting, not a technical one, so ask who can change it, whether the change is logged, and whether results screened at a prior threshold are marked as such.

Question four: does the audit log preserve the incomplete state after a re-run?

Re-runs are where records get laundered. A check fails on Monday, the analyst re-runs it on Wednesday, the source is up, the result is green, and the Monday state disappears from the file. The record now shows a clean screen with no trace that clearance was ever in doubt.

Demand append-only behaviour. The audit log should retain the failed attempt, its reason code, its timestamp, and the identity of whoever triggered the re-run, alongside the later successful result. Ask whether the incomplete state can be overridden manually, and if so whether the override captures a user, a timestamp, and a written justification. Then ask the question most vendors have not rehearsed: can you export the full state history for a single counterparty as evidence, in a format an examiner can read without access to the platform?

OilFlow's screening logic is built to that standard by design. Unverifiable steps render as pending and stay pending until the underlying source returns, and the record retains the pending state rather than overwriting it. That is the posture we think buyers should require from any vendor, including us. You can request a walkthrough of the clearance record and see the incomplete state rendered rather than described.

Why the documentation is the actual deliverable

You are not buying detection. You are buying a defensible record of the enquiries you made, the enquiries that returned, and the enquiries that did not. Under FATF Recommendation 10, the obligation is to identify and verify using reliable, independent source documents, data or information, and to know when verification has failed. A tool that cannot distinguish failure from success has removed your ability to discharge that obligation, while making you feel you have discharged it.

The test of a screening record is not whether it looks clean. It is whether, when read cold by a regulator or by counsel in a dispute over a cargo that turned out to be sanctioned, it accurately describes what was and was not known at the moment of the decision.

What compliance teams should do

  1. Rewrite the demo script. Ask every shortlisted vendor to produce a live result in the not-established state before you look at a single hit.
  2. Put the four questions in the RFP in writing. Failed-source artefact, named lists with per-list refresh cadence, transliteration and alias behaviour, and audit log persistence after re-run. Score them.
  3. Reject list counts as a coverage answer. Require named lists, named jurisdictions, named cadence, and a visible last-refresh timestamp.
  4. Test the ownership logic separately. Confirm whether the tool addresses aggregated indirect ownership under the OFAC 50 Percent Rule, or whether it matches list strings only, and confirm it says which.
  5. Set the internal rule before procurement, not after. Unverifiable renders as pending. Pending blocks onboarding and blocks payment release until an analyst adjudicates it. Write it into the policy so the tool has to meet the policy rather than the policy drifting to meet the tool.
  6. Brief the MLRO on the three states. Sign-off authority should distinguish clearing a screened file from clearing a file the system could not screen.

A vendor that converts unreachable sources into green results is selling comfort. What survives examiner review is clearance, and clearance requires that the system be willing to tell you when it does not know.

For weekly typology briefings on mandate chain, layer cake and dark fleet structures in physical oil trade, subscribe to the OilFlow Intelligence newsletter.

Verified trade-fraud patterns, sanctions deltas, and regulator actions. Weekly, for compliance and risk teams.

Double opt-in. No spam. The quarterly Compliance Index ships to subscribers first.

This article is part of our scam-cluster intelligence series. Screening a specific counterparty? Run the free check right here, or order the full 7-step dossier.

Paste any company, person, or vessel name. Free, no signup, answer in seconds.