Scope and limitations

Where this stops.

Written before anyone asked for it. If you are evaluating this for a regulated collections or disputes operation, this page will tell you more than the rest of the site put together, and it is the one worth reading first.

01What it decides

Fifteen facts can override the score outright.

These are the cases where the arithmetic is not allowed a say, because a weighted model can average away a disqualifying fact. Each one is a decision you can point at afterwards and explain.

GateApplies toDecides
G-REFUNDEDDisputesAlready refunded
G-DUPLICATEDisputesDuplicate charge confirmed
G-WINDOW-CLOSEDDisputesRepresentment window closed
G-NET-VALUEDisputesBelow the net value floor
G-DEADLINE-BUFFERDisputesInside the filing buffer
G-HIGH-VALUEDisputesAbove the unattended filing ceiling
G-DOCUMENTED-DEFECTDisputesDefect corroborated by the cardholder
G-NO-EVIDENCEDisputesNothing to file
G-INSOLVENCYReceivablesInsolvency proceedings
G-UNAPPLIED-CREDITReceivablesCredit note covers the balance
G-DISPUTEDReceivablesOpen deduction on the invoice
G-STATUTEReceivablesBeyond the limitation period
G-CONTACT-CAPReceivablesContact frequency cap reached
G-BELOW-FLOORReceivablesBelow the pursuit floor
G-PROMISE-OPENReceivablesUnexpired promise to pay

02Where it fails

Known failure modes, including the ones still open.

It cannot see what is not in a system of record

The whole design rests on evidence being retrievable and machine-readable. A merchant whose proof of delivery lives in a shared mailbox, or whose terms acceptance was never captured with a timestamp, gets a correct decision on incomplete information — which is to say, a concession. The engine will tell you that is why, but it cannot conjure the exhibit.

Corroboration weighting is reasoned, not fitted

Evidence within a family counts the strongest signal in full and the rest at 40%, because three separate proofs that the buyer was the real cardholder are not three times as convincing as one. That ratio is argued from how these cases are actually assessed, not fitted to outcome data. It is the kind of number a pilot should move.

Win probability is a ranking, not a forecast

The probability shown on a represented case is a monotonic function of its evidence score, anchored so that a case at the floor is a little better than even. Until representments have been filed and adjudicated on your book, use it to order the queue rather than to forecast a recovery line.

Propensity has no macroeconomic input

It reads the payer's own history and nothing else. A sector-wide liquidity shock, a currency control, or a customer's largest client failing will all show up in the score only after that payer has already started paying late.

Thresholds do not adapt on their own

The policy is static until someone edits it, and deliberately so: a threshold that moves by itself is one nobody can be accountable for when a regulator asks. The cost is real — a book whose mix shifts needs its floors revisited by hand rather than drifting to fit.

Reason-code coverage is deliberate, not exhaustive

Extractors are built for Visa 10.4, 13.1, 13.3, 13.6 and 12.6.1, and Mastercard 4853, 4837, 4855 and 4841 — the codes that carry the volume. American Express and Discover are modelled in the type system but have no extractor yet. An unrecognised code yields few signals and therefore a concession: graceful, but not coverage.

One representment, and no pre-arbitration path

The product assembles and decides the first representment. It does not model pre-arbitration, arbitration, or the earlier collaboration flows some networks now offer. A case that loses is out of scope from that point on.

Generated text is English only

Composing a WhatsApp nudge in English to a debtor who works in Portuguese or Swahili is worse than sending nothing. Localised templates are the obvious next piece of work, and until they exist the receivables side is honest only in English-speaking books.


03How to evaluate it

Run it behind your own team first.

The honest way to assess this is in shadow mode: the engine decides, your analysts decide, and nothing is filed or sent on its say-so. After a few hundred cases the disagreements are the whole story — where it conceded something you would have fought, where it fought something you would have dropped, and which threshold was responsible each time.

Every disagreement is traceable to a specific threshold on the policy page, which is the point of holding them in one place. Tuning is a conversation about numbers you can see, not about a model nobody can inspect.

What to measure

  • Agreement rate against your analysts, split by reason code and by gate.
  • Recovery on the cases it chose to fight, against your current baseline on comparable cases.
  • Cases it declined to fight that you would have won — the expensive miss, and the one the 58-point evidence floor is set to trade against.
  • Time to assemble a packet, against the 20–25 minutes an analyst currently spends.
  • On receivables: relationship outcomes on accounts it left alone, which is the number conventional dunning never captures.

Check the working now

Every claim on this page is inspectable in the product. Open any case and you can see the records it read, the facts drawn from each, the weight applied and why, the gate that fired, the artifact produced, and the self-check that had to pass before release.