The math vendors do not multiply. The previous generation hid its failure in capture. The new one hides its failure in the decision. Two different forms of unmanaged complexity; the same result: nobody knows, in time, what is true.

USD/year per capture role · Parseur
$28.5k
inventory errors from manual capture
55%
declared confidence vs real accuracy
90% a 75%
0.75³ chain reliability
~42%

In 30 seconds

Declared confidence is not accuracy. Individual accuracy is not chain reliability. And no database, however luxurious, fixes what was captured wrong or never captured. Bring a calculator: how many people does the system need to stay current, and who reviews the case where the machine doubts.

What the previous generation cost

The previous generation of transport software solved its data problem with a method nobody boasts about in the demo. People. Dozens of clerks, coordinators, and analysts typing statuses, rates, and invoices to keep the luxury database alive. The cost of that decision is already measured. A Parseur survey with QuestionPro in 2025 calculated that manual capture costs about $28,500 USD per year per employee involved. And quality is not saved either. Industry digests attribute about 55% of inventory errors to manual capture.

The hidden cost of manual capture

The keyer is not cheap once you add error, turnover, and rework.

$28,500

USD / year · fully loaded capture cost

Parseur / manual entry synthesis 2025

55%

inventory errors tied to manual capture

A portal nobody feeds well produces data nobody can defend

Source · Parseur / QuestionPro · manual entry cost

Declared confidence is not accuracy

The new generation of software promises the opposite remedy. No people. AI agents that decide alone, each with its confidence number on screen. The vendor shows it, the buyer writes it down, and both assume it means what it appears to mean. The problem is that the number lies systematically, and the lie multiplies when agents work in a chain.

Start with the original defect. Models trained with human feedback tend toward overconfidence as a structural trait. Training rewards answers that sound sure, so the model learns to sound sure even when it should not. Digital Applied's escalation-design analysis, which collects production calibration data, documents the magnitude. A declared confidence of 90% often corresponds to real accuracy near 75%.

Stated confidence vs real accuracy

A self-assured model is not a reliable system.

The rule

Anything that touches money always escalates to a human, no matter how sure the agent claims to be.

Source · Digital Applied · model calibration

The chain multiplies the error

Fifteen points of difference look manageable on an isolated agent. But almost nobody deploys an isolated agent. Serious deployments use a chain. One agent extracts data from the document, another matches it against sources, another drafts the decision. If each link is miscalibrated by those fifteen points, the probability that all three steps are correct is not 90%. It is 0.75 cubed. About 42%.

0.75³ ≈ 0.42
Three links at 75% real accuracy: chain reliability near a coin flip.

Error chain across agents

Reliability multiplies. It does not average.

Agent 1

75%

cumulative reliability

Agent 2

56%

cumulative reliability

Agent 3

42%

cumulative reliability

Source · Digital Applied · 0.75³ ≈ 42%

And here is the detail that turns the technical problem into a business problem. The system does not warn anyone. The chain presents its result as high confidence because each individual link reported being sure. Nobody multiplied. The buyer who evaluated the system on a controlled demo is operating something whose real reliability is that of a loaded coin flip.

The escalation gate

Notice the symmetry with the previous generation, because that is the full argument. Traditional TMS hid its failure in capture. Data arrived late, typed, and expensive. The loose agent hides its failure in the decision. Data arrives alone, but the certainty with which it is processed is fictitious. Two different forms of unmanaged complexity, the same result. Nobody knows, in time, what is true.

The way out is not to choose between both failures but to correct both at once. On the capture side, the system goes get the data where the action happens, without clerks. On the decision side, the escalation gate. An agent with 75% real accuracy, with a human reviewing the doubtful cases with a ready file, is a usable and improvable system. A chain at 42% presented as certainty is a liability waiting for a date. The gate does not exist because the model is dumb. It exists because the model's own confidence signal does not know when it is wrong.

The best production designs no longer depend on a single numeric threshold. Pickaxe's guide describes how good handoffs combine signals. Model confidence as a weak input, not a deciding vote. Conversation sentiment. Loop detection. And hard rules by action type that admit no exception. Everything that touches money always escalates. Galileo's technical guide adds the operational nuance almost nobody implements. Thresholds are not copied from an external benchmark. They are calibrated against your own production data.

Escalation threshold by cost of error

Calibrating confidence is not enough: tie the threshold to the money at stake.

Low impact

Status · messages · routine

The agent can proceed with a moderate threshold. The cost of being wrong is reversible.

High impact

Payment · rate · claim

Always escalate to a human. Stated model confidence does not authorize releasing money.

Operating rule: if the case touches money, open a ticket. Calibration prioritizes the queue — it does not skip it.
Source · HITL synthesis · thresholds by cost of error
Auditor reviewing a freight file before releasing payment
Everything that touches money always escalates. The threshold is calibrated on your data.

The arithmetic on a thousand shipments

Translated to transport, the arithmetic becomes tangible. An operator with a thousand monthly shipments and five documents per shipment generates five thousand pieces to match each month. With the old architecture, an entire team types and reconciles, at $28,500 USD per year per head, errors included. With the fashionable architecture, a chain of agents approves at 42% real dressed as 90. With the correct architecture, the system captures alone, automates 85% with complete evidence, and escalates 750 cases a month as tickets to an auditor who decides with a ready file. A volume one person with a tool can operate, not a department.

Correct

Capture · automate · escalate

  1. Capture

    At source

  2. Match

    ~85% with evidence

  3. Tickets

    ~750 / month

  4. Human

    Ready file

Two questions before you buy

Declared confidence is not accuracy. Individual accuracy is not chain reliability. And no database, however luxurious, fixes what was captured wrong or never captured. Anyone buying operations software this year would do well to bring a calculator and ask two questions. How many people does the system need to stay current, and who reviews the case where the machine doubts.

At OCL Cargo those answers are operational: agents on the voyage file; tickets to Finance when money or certainty is at stake; each correction retrains. It can stamp invoice and Carta Porte.

Key takeaways5 points
  1. Previous generation hid the failure in capture: ~$28,500 USD/year per clerk; ~55% of inventory errors start in manual entry.
  2. 90% declared confidence often equals ~75% calibrated real accuracy.
  3. Three agents at 75% real: 0.75³ ≈ 42% chain reliability. Nobody multiplies in the demo.
  4. Gate: model confidence is a weak input; everything that touches money always escalates.
  5. Correct architecture: capture at source + ~85% automated with evidence + tickets to a human.

What is your real chain reliability?

We review capture, calibration, and the escalation gate on your flow.

Sources

Related reading

FAQ

By Gibrán Ramírez, CEO of OCL Cargo.