Why we rebuilt a model that was already ahead
Our previous models were tuned hard for ecommerce, and that tuning worked.
What kept showing up, on both the ecommerce and the postal side, is how little we often have to work with. An ecommerce listing can come with a descriptive title, a category, and something about what the item is made of. Plenty do. Plenty more are a brand name and a product name, and a listing called Summit Ridge Mid does not tell you whether it is a hiking boot, a backpack, or a jacket, let alone what it is made of.
Postal is where this gets hardest. Zonos handles the large majority of all postal shipments entering the United States, and a parcel there is often described in two or three words, typed onto a customs form by the person who packed the box, and there is nothing else behind it.
That is what we rebuilt for: the items where the description does not tell you much, whichever channel they come through.
"T-shirt" can mean any of 16 HS codes - but which one?
Say a declaration reads t-shirt and nothing else. That could be any of 16 HS codes. The right one depends on what the shirt is made of, who it is cut for, and how it was knit, and the duty owed changes with each of those answers. Cotton and man-made fiber alone sit in different subheadings at materially different rates.
Multiply that by a catalog. Kid-Silk Mohair 25g is the kind of title we see all day, and it still has to come out as one line of the tariff schedule.
Those trickier ones are where we focused this round. The well-described items get the same attention, they are just less likely to trip up a classifier.
How we measured it
We measured the new model the way we would want a vendor to measure for us: on live traffic, not on a test set of our own choosing. We used two weeks of classification requests that came through our platform, which means the titles and descriptions our customers actually send, including the ones that are barely a description at all.
Across that traffic the new model makes 29.4% fewer errors than the model it replaces.
This is a new model, not a retuned version of the old one. We rebuilt the corpus it learns from, training on official customs rulings alongside millions of the partially described listings our classifier sees every day.
Against a competitor, on the same shipments
Comparing a model to your own older model only tells you so much. We took 1,000 real shipments and ran them through three systems: a leading competitor's classifier, the model Classify Zonos AI uses today, and the new one. Every shipment was scored the same way.

The model Classify Zonos AI runs today already made 33% fewer errors than the competitor. The new one makes 70% fewer.
Some mistakes cost more than others
Not every classification error carries the same bill. If a pair of socks is classified one subheading away from where it belongs, the duty owed is usually close, and nobody notices. If those same socks end up in a different chapter of the tariff schedule entirely, booked as something they are not, the duty is not close at all. Those are the ones that turn into an unexpected bill later, and they are a large part of why we rebuilt the model.
What this means for you
- The right code more often. 29.4% fewer errors than the model it replaces, measured on live cross-border traffic.
- Built for thin product data. We aimed the rebuild at short, incomplete descriptions, which is most of what a postal shipment provides and more of an ecommerce catalog than most merchants expect.
- Fewer surprises at the border. Fewer wrong codes means fewer shipments that arrive with a duty bill nobody expected.
- Nothing to change on your side. The new model is already serving classifications. If you call Classify Zonos AI today, your request and the response you get back stay the same, so there is no integration work to do and no setting to switch on.
Measured on two weeks of live cross-border traffic and a 1,000-shipment comparison against a leading competitor. Classification accuracy varies by product category and by how much detail the item description carries.
Zonos' new Classify model makes 29.4% fewer classification errors
Every parcel that crosses a border needs an HS code, and that code decides what the shipment owes and how cleanly it moves. Zonos has spent more than 16 years on that problem, and Zonos Classify Zonos AI solves it with proprietary AI models we build ourselves.
We train our models on cross-border data we see firsthand and we test them ourselves, rather than running a general-purpose model with a customs prompt bolted on.
That work has kept us in front. Against a leading competitor's classifier, on the same 1,000 shipments, the model Classify Zonos AI runs today makes 33% fewer errors.
We have now rebuilt it to be even better. Across two weeks of live cross-border traffic, the new model makes 29.4% fewer classification errors than the model it replaces. It is live now.