The math vendors do not multiply. The previous generation hid its failure in capture. The new one hides its failure in the decision. Two different forms of unmanaged complexity; the same result: nobody knows, in time, what is true.
- USD/year per capture role · Parseur
- $28.5k
- inventory errors from manual capture
- 55%
- declared confidence vs real accuracy
- 90% a 75%
- 0.75³ chain reliability
- ~42%
In 30 seconds
Declared confidence is not accuracy. Individual accuracy is not chain reliability. And no database, however luxurious, fixes what was captured wrong or never captured. Bring a calculator: how many people does the system need to stay current, and who reviews the case where the machine doubts.
What the previous generation cost
The previous generation of transport software solved its data problem with a method nobody boasts about in the demo. People. Dozens of clerks, coordinators, and analysts typing statuses, rates, and invoices to keep the luxury database alive. The cost of that decision is already measured. A Parseur survey with QuestionPro in 2025 calculated that manual capture costs about $28,500 USD per year per employee involved. And quality is not saved either. Industry digests attribute about 55% of inventory errors to manual capture.
The hidden cost of manual capture
The keyer is not cheap once you add error, turnover, and rework.
$28,500
USD / year · fully loaded capture cost
Parseur / manual entry synthesis 2025
55%
inventory errors tied to manual capture
A portal nobody feeds well produces data nobody can defend
Declared confidence is not accuracy
The new generation of software promises the opposite remedy. No people. AI agents that decide alone, each with its confidence number on screen. The vendor shows it, the buyer writes it down, and both assume it means what it appears to mean. The problem is that the number lies systematically, and the lie multiplies when agents work in a chain.
Start with the original defect. Models trained with human feedback tend toward overconfidence as a structural trait. Training rewards answers that sound sure, so the model learns to sound sure even when it should not. Digital Applied's escalation-design analysis, which collects production calibration data, documents the magnitude. A declared confidence of 90% often corresponds to real accuracy near 75%.
Stated confidence vs real accuracy
A self-assured model is not a reliable system.
Stated confidence
by the agent
Real accuracy
calibrated single model
3-agent chain
errors compounded
The rule
Anything that touches money always escalates to a human, no matter how sure the agent claims to be.
The chain multiplies the error
Fifteen points of difference look manageable on an isolated agent. But almost nobody deploys an isolated agent. Serious deployments use a chain. One agent extracts data from the document, another matches it against sources, another drafts the decision. If each link is miscalibrated by those fifteen points, the probability that all three steps are correct is not 90%. It is 0.75 cubed. About 42%.
Error chain across agents
Reliability multiplies. It does not average.
Agent 1
75%
cumulative reliability
Agent 2
56%
cumulative reliability
Agent 3
42%
cumulative reliability
And here is the detail that turns the technical problem into a business problem. The system does not warn anyone. The chain presents its result as high confidence because each individual link reported being sure. Nobody multiplied. The buyer who evaluated the system on a controlled demo is operating something whose real reliability is that of a loaded coin flip.
The escalation gate
Notice the symmetry with the previous generation, because that is the full argument. Traditional TMS hid its failure in capture. Data arrived late, typed, and expensive. The loose agent hides its failure in the decision. Data arrives alone, but the certainty with which it is processed is fictitious. Two different forms of unmanaged complexity, the same result. Nobody knows, in time, what is true.
The way out is not to choose between both failures but to correct both at once. On the capture side, the system goes get the data where the action happens, without clerks. On the decision side, the escalation gate. An agent with 75% real accuracy, with a human reviewing the doubtful cases with a ready file, is a usable and improvable system. A chain at 42% presented as certainty is a liability waiting for a date. The gate does not exist because the model is dumb. It exists because the model's own confidence signal does not know when it is wrong.
The best production designs no longer depend on a single numeric threshold. Pickaxe's guide describes how good handoffs combine signals. Model confidence as a weak input, not a deciding vote. Conversation sentiment. Loop detection. And hard rules by action type that admit no exception. Everything that touches money always escalates. Galileo's technical guide adds the operational nuance almost nobody implements. Thresholds are not copied from an external benchmark. They are calibrated against your own production data.
Escalation threshold by cost of error
Calibrating confidence is not enough: tie the threshold to the money at stake.
Low impact
Status · messages · routine
The agent can proceed with a moderate threshold. The cost of being wrong is reversible.
High impact
Payment · rate · claim
Always escalate to a human. Stated model confidence does not authorize releasing money.

The arithmetic on a thousand shipments
Translated to transport, the arithmetic becomes tangible. An operator with a thousand monthly shipments and five documents per shipment generates five thousand pieces to match each month. With the old architecture, an entire team types and reconciles, at $28,500 USD per year per head, errors included. With the fashionable architecture, a chain of agents approves at 42% real dressed as 90. With the correct architecture, the system captures alone, automates 85% with complete evidence, and escalates 750 cases a month as tickets to an auditor who decides with a ready file. A volume one person with a tool can operate, not a department.
Correct
Capture · automate · escalate
Capture
At source
Match
~85% with evidence
Tickets
~750 / month
Human
Ready file
Two questions before you buy
Declared confidence is not accuracy. Individual accuracy is not chain reliability. And no database, however luxurious, fixes what was captured wrong or never captured. Anyone buying operations software this year would do well to bring a calculator and ask two questions. How many people does the system need to stay current, and who reviews the case where the machine doubts.
At OCL Cargo those answers are operational: agents on the voyage file; tickets to Finance when money or certainty is at stake; each correction retrains. It can stamp invoice and Carta Porte.
Key takeaways5 points
- Previous generation hid the failure in capture: ~$28,500 USD/year per clerk; ~55% of inventory errors start in manual entry.
- 90% declared confidence often equals ~75% calibrated real accuracy.
- Three agents at 75% real: 0.75³ ≈ 42% chain reliability. Nobody multiplies in the demo.
- Gate: model confidence is a weak input; everything that touches money always escalates.
- Correct architecture: capture at source + ~85% automated with evidence + tickets to a human.
What is your real chain reliability?
Sources
- Parseur / QuestionPro: cost of manual capture
- Digital Applied: production calibration / escalation design.
- Pickaxe: human-in-the-loop handoff guide.
- Galileo: operational thresholds and patterns.
Related reading
- The war on complexity
- What Decagon understood before the market
- A 1983 irony that transport software perfected
FAQ
A Parseur / QuestionPro survey (2025) estimated about $28,500 USD per year per employee dedicated to capture. Industry digests attribute about 55% of inventory errors to manual capture.
Models trained with human feedback tend toward overconfidence. Production calibration analyses document that a declared 90% confidence often corresponds to real accuracy near 75%.
If each link has real accuracy of 0.75, the probability that all three are correct is 0.75³ ≈ 42%. The system can still present “high confidence” because nobody multiplied.
Everything that touches money. Also: high-amount disputed deductions, permit changes, and data deletion. Thresholds are calibrated on your own production data, not an external benchmark.
By Gibrán Ramírez, CEO of OCL Cargo.
