Why accuracy claims built on training data don’t survive contact with reality

By Ian Venner, CTO, Hurricane Commerce

“If a vendor tests its model on the data it trained that model with, the accuracy score tells you how good it’s managed to train its model, nothing about its actual performance in the real world.“

The number that’s too good

Trade compliance software has an accuracy inflation problem, and unbounded marketing hype is where it starts. Headline accuracy figures for HS classification and landed cost get quoted with confidence and very little explanation of where they came from.

A lot of these figures are based on testing against a vendor’s own curated dataset, some even admit it, the same data they used to build and tune the model. The model isn’t determining new responses, it’s just regurgitating what it’s been trained on.

The security industry fell foul of this thirty years ago. It didn’t end well for the buyers who believed the hype then either.

Scan the zoo, detect the zoo, quote the percentage – a history lesson

In the early days of antivirus, vendors built their products from their own libraries of virus samples. Then they tested those products against the same libraries. Researchers called these “zoo” collections: samples that existed in a lab, many of which had never infected a real user’s machine.

The result was a run of detection claims in the high nineties, 99.99% included for some of the more unscrupulous vendors. A product tuned to a vendor’s own library will score brilliantly on that library. It says very little about the virus arriving on a floppy disk in the accounts department on Monday morning.

The correction came from outside the vendors. In America, researcher Joe Wells started the WildList, a monthly record of viruses confirmed spreading on real users’ machines.  In the UK we had the Virus Bulletin giving the same sort of information.  Testers finally began measuring products against what was loose in the world rather than what sat in a vendor’s vault, that they had been supplied with.

Sarah Gordon and Fraser Howard, writing in the proceedings of NIST’s National Information Systems Security Conference in 2000, called these real world criteria “a reality check to the antivirus industry”. They were blunt about the marketing too: “if the advertising claims are to be believed, all products are not merely created equal, they are all created superlative!” They also warned that any test relying on vendor-supplied material meant trusting the one party with the most to gain from the result.

By 2008 the problem was serious enough that vendors and testers formed the Anti-Malware Testing Standards Organization, set up “to address a perceived need for improvement in the quality, relevance and objectivity of anti-malware testing methodologies.”

Security buyers learned to ask where the test samples came from.

The time has come for Trade Intelligence buyers to be asking the same questions.

Same trick, new industry

Machine learning makes the zoo problem easier to hide. Build a classification model on a dataset of products and their HS codes. Hold back a slice of that same dataset. Test on the slice. Publish the score.

That score measures how well the model reproduces patterns in its own training data. Your catalogue isn’t in that data. Your supplier’s vague product descriptions, the “misc. parts” entries, the tariff changes since the dataset was frozen: the test doesn’t contain them, so the score can’t account for them.

Data scientists call this leakage, where information from the training side seeps into evaluation and inflates the result. Sayash Kapoor and Arvind Narayanan at Princeton documented leakage problems across 30 scientific fields, affecting 648 papers, published in Patterns in 2023. In one field, once the leakage was fixed, sophisticated models showed “no substantive improvement over decades-old Logistic Regression models.”

That’s peer-reviewed research. Vendor accuracy claims get no peer review. When one company controls the training data, the test data, the definition of “correct” and the press release, the number means whatever that company needs it to mean.

A curated test set is clean, labelled and consistent. Live traffic is typos, missing fields, translated descriptions and products that didn’t exist last quarter. A model can be correct on the first but fail on the second, and a self-graded score will hide the gap, and build false confidence in a product that really doesn’t deserve it.

How Hurricane measures accuracy

Hurricane’s accuracy is above 99%, measured across more than one billion real-world transactions processed through our service.

Those transactions are live calls from merchants, marketplaces and logistics partners shipping real goods across real borders. They include the messy descriptions, the incomplete data and the edge cases that a curated test set filters out.

No team of humans can check a billion transactions by hand, so we audit them statistically. We analyse outcomes across the full population of live calls, which gives a measure grounded in production traffic rather than in a lab sample we cherry picked ourselves.

The difference is the one the antivirus industry learned the hard way. A zoo score tells you how a product performs on the vendor’s terms.

A wild score like Hurricane’s tells you how it performs on yours.

Five questions to put to any vendor

Before you trust an accuracy figure, get straight answers to these:

  1. Was the test data drawn from the same source as the training data?
  2. Is the figure measured on live production traffic, or on a curated sample?
  3. How many transactions sit behind the number, and over what period?
  4. Who defines a “correct” result, and who checks it?
  5. Can we run our own shipments through your service ourselves to test these claims?

A note on the last point, if the vendor does this test for you, it’s not a real-world test, as training a model on your data before they actually run the test, is back to where we started this blog.

A vendor confident in its numbers will answer all five without flinching.  Hurricane is confident in its numbers, this is real world results, using real world data complete with all its flaws.

A vendor that hesitates on the first question, or wants you to send them your data to test themselves, has told you everything you need to know.

The “fake it ’til you make it” hyperbole which appears to be omnipresent in some quarters of our industry, is no basis for rational business decisions.