Skip to content

Which carrier actually sails on time?

The Navo24 desk SchedulesMCP

No carrier is ‘reliable’ in the abstract: on-time performance is a lane fact. The only honest answer compares each sailing’s scheduled arrival to the one that actually happened, per lane, over enough sailings to mean something. Read the p50 and p90 spread, not just the average, and check the sample size.

Ask a carrier’s sales rep whether they’re reliable and you’ll get a yes. Ask their marketing deck and you’ll get a number that’s been kind to itself. The only version of “reliable” you can book on is the one built from arrivals that actually happened: scheduled ETA versus the day the box really came off the vessel, counted over enough sailings to mean something. Everything else is a promise, and promises don’t clear customs.

Here’s how to read on-time reliability like someone who’s been burned by an average before.

Reliability is a lane fact, not a carrier fact

The first mistake is asking “is Carrier X reliable?” as if it’s one number. It isn’t. A carrier can run a tight, punctual service on a mature headhaul lane and be all over the place on a thin feeder route with a wobbly transshipment. On-time performance lives at the lane level, origin to destination, on that service, because that’s where the physical reality is: this port pair, this rotation, this transshipment hub.

So the useful question is never “which carrier is best,” it’s “on my lane, for the window I’m booking, who’s actually been arriving when they said.” A carrier that’s a disaster somewhere else may be the safe pick on yours.

The average lies; the distribution tells the truth

Even at the lane level, a single on-time percentage hides the thing that hurts you. Two services can both be “80% on time” and be completely different risks. One is reliably a day or two off. The other is usually early but occasionally two weeks late when a transshipment falls over; it’s that long tail that blows up your demurrage, your production line, your customer promise.

That’s why the number to look at isn’t the average transit: it’s the spread. The p50 (median) tells you the typical crossing. The p90 tells you the bad-but-plausible case: nine sailings in ten arrive by this day. The gap between them is the risk. A tight p50-to-p90 spread is a service you can plan around; a wide one is a service that will surprise you when you can least afford it.

SAME LANE, SAME AVERAGE: DIFFERENT RISK ILLUSTRATIVE 28d32d36d40d transit days (illustrative shape) Carrier A · p50 32d p90 34d: tight Carrier B · p50 32d p90 40d: fat tail same median · one you can plan around, one that bites
Illustrative distributions: not observed data for any named carrier. Identical p50; the fat p90 tail on the right is the demurrage-and-missed-promise risk an average hides.

Where does an honest reliability number come from?

You can’t buy your way to a trustworthy score: you have to observe it. The only defensible on-time number is built by matching each sailing’s scheduled arrival against the arrival that actually occurred, drawn from real vessel tracking, and aggregating enough of them per lane to be more than noise. Two things make or break it:

Sample size. A “95% on-time” score off three sailings is a coin-flip dressed as a fact. A lane with dozens of observed arrivals earns a number you can lean on; a thin lane should be labelled as early-data, not quietly averaged into false confidence. If a tool won’t show you the sample size, don’t trust the percentage.

Provenance. The arrivals have to come from what was actually seen, the vessel discharging and tracked, not from the carrier grading its own homework. A reliability score is only as honest as the arrivals feeding it.

HOW THE SCORE IS EARNED SCHEDULED carrier's promised ETA vs ACTUAL observed via tracking per-lane reliability + sample size shown thin lane → early-data flag Aggregate enough sailings and the delay you keep seeing becomes a number you can book on. No observed arrivals, no score: a reliability figure is never invented to fill a gap.
Scheduled vs actual, per lane, with the sample size in the open: the only reliability number worth booking on.

How to actually use it when booking

Put together, the workflow is short:

  1. Pick the lane and window you’re actually shipping, not the carrier in the abstract.
  2. Read the p50 and the p90, not just the on-time percentage. Budget your buffer against the p90, because that’s the case that costs you.
  3. Check the sample size. A confident number off a handful of sailings is worth less than a modest number off many.
  4. Compare the options on that lane side by side, and weigh reliability against transit and cut-off: sometimes the slightly slower service that always shows up is the cheaper one once you price in the demurrage the fast-but-flaky one causes.

The point

“Which carrier sails on time?” has no honest answer in the abstract: only “on this lane, over this many observed sailings, here’s the median, here’s the p90, and here’s how big the sample is.” Anything more confident than that is selling you something.

That’s the bet behind SchedulesMCP: reliability scored from arrivals we actually observed through the tracking layer, per lane, with the sample size in the open and thin lanes honestly flagged as early-data rather than faked. Compare sailings the way a desk books them, by transit, cut-off and the reliability you can defend, at the compare view, or call compare_sailings and get_lane_reliability from your own system. The distributions above are illustrative; for the real per-carrier, per-lane reliability we’ve observed, see the carrier reliability view. If you want the case for grouping those sailings by the physical vessel so slot-partners don’t clutter the picture, we made it here: book the hull that shows up.

Common questions about carrier reliability

Which carrier is the most reliable?

There’s no honest answer in the abstract. On-time performance lives at the lane level: a carrier can be punctual on a mature headhaul and all over the place on a thin feeder route. Ask who’s been arriving on time on your lane, for the window you’re booking, not which carrier is best overall.

What is a good on-time reliability score?

A single percentage hides the risk. Two services can both be 80% on time and be completely different bets: one reliably a day or two off, the other usually early but occasionally two weeks late. Look at the p50-to-p90 spread; a tight spread is a service you can plan around.

What’s the difference between p50 and p90 transit time?

The p50, or median, is the typical crossing: half of sailings arrive by then. The p90 is the bad-but-plausible case: nine sailings in ten arrive by this day. The gap between them is the risk, and you should budget your buffer against the p90.

How many sailings make a reliable score?

Enough that it’s more than noise. A 95% score off three sailings is a coin-flip dressed as a fact; a lane with dozens of observed arrivals earns a number you can lean on. If a tool won’t show you the sample size, don’t trust the percentage.

Built by people who move boxes for a living.

Tracking, schedules and load planning, as components you can adopt one at a time.