01 · Context

A thirty-year-old error meets a system that never hesitates

On 30 June 2026, Gartner named agentic AI and physical AI the leading supply chain technology trends of the year. Two months earlier, on 29 April 2026, the same firm published a survey of 140 senior supply chain leaders in which 56% named integration with legacy systems a major challenge. Director Analyst Snigdha Dewal put the finding plainly: the greatest friction in scaling AI is not the technology but the environments into which it is deployed. At the Barcelona symposium on 20 May 2026, low data quality was named the top barrier to scaling AI.

Those are the right constraints, stated at the wrong altitude. "Data quality" is a category, not a decision. Underneath it sits one specific number that operations has been quietly failing at since long before anyone shipped an agent: whether the stock position in the system matches the stock position on the floor.

The evidence is old, consistent and uncomfortable. DeHoratius and Raman, publishing in Management Science in 2008, examined nearly 370,000 inventory records across 37 stores of a single retailer and found 65% inaccurate. More usefully, they decomposed the variance: 26.4% of it sat between product categories and only 2.7% between stores. ECR Retail Loss, working with seven retailers, reports roughly 60% of records wrong and a 4–8% sales recovery when they are corrected. A 2026 study by Rekik, Oliva, Glock and Syntetos covering some 24,000 SKUs across 11 stores found inaccuracy rising with inventory level, restocking frequency and perishability, and a field quasi-experiment in which an audit produced an 11% store-wide sales lift — with the effect concentrated on particular items rather than spread evenly.

Latin America has no reason to assume it sits at the better end of that distribution. The Datup Supply Chain Trends and Digitalization Study published on 28 January 2026, surveying 155 regional leaders, found 58.1% reporting stockouts and 54.2% reporting oversupply at the same time, and 63.5% reporting critical integration problems. Simultaneous stockout and oversupply is not a forecasting signature. It is what a wrong stock position looks like from two directions at once.

The record was always wrong. What is new is that something now acts on it thousands of times a day without pausing to look at the shelf.

02 · Framework

The three states of a stock position

Socradata separates a stock position into three states, because operations has historically managed the first two and is now being asked to govern a third that did not exist in its current form two years ago. Treating all three as one number is what allows an accuracy programme to run for years, report improvement, and still deliver an agent that commits the wrong order every Tuesday.

State 1 — Physical

What is actually in the bin, on the pallet, in the trailer. It is the only state that is true, and it is unobserved between counts. Every technology that claims to improve inventory accuracy — cycle counting, RFID, computer vision, drone imaging — is a method for sampling this state more often at lower cost. Corvus Robotics reports imaging a cosmetics distribution centre 52 times a year, a genuine advance that still leaves the position unobserved for most of the week.

State 2 — Record

What the ERP stock ledger and the WMS bin-level on-hand believe. It is observed continuously, reported monthly, and wrong for roughly six items in ten. Crucially, it is wrong non-randomly: the variance concentrates in categories, in fast-restocked items, in perishables and in high-handling locations. That non-randomness is the whole opportunity, because it means a uniform remedy is the most expensive possible remedy and a targeted one is affordable.

State 3 — Decision

What an agent commits on the basis of State 2: an order raised, a store allocation written, a wave released, a transfer triggered. This state is new at scale, and it changes the economics of the error. A planner reviewing forty proposals a day applies suspicion to the ones that look odd, which is an unglamorous but real error filter. An agent clearing four thousand decisions a day applies none. The same record error that used to produce one bad order now produces a hundred, faster, and with a clean audit trail showing that policy was followed exactly.

So what: an agent does not read your shelf. It reads your record — and then it acts on that record more often, more confidently, and with less hesitation than the planner it replaced.

The spanning metric that follows is decision-weighted record accuracy. Conventional accuracy asks what share of SKUs match on count day, treating a slow-moving spare part and a promoted perishable as one unit each. Decision-weighted accuracy asks instead: of the decisions an agent committed this period, what share were taken against a record verified inside that class's tolerance window, weighted by the value of the decision. The first number is a hygiene report. The second states how much committed spend rests on an unverified belief. KPIs before APIs applies literally: rebuild the measurement before granting the write authority, because afterwards the agent generates the very volume that makes the old measurement misleading.

03 · Use Cases

Three patterns, three systems, one precondition

The patterns below are anonymised composites drawn from operating and advisory work in Argentina and the Southern Cone. Figures are illustrative targets rather than audited client results. Each names the system implicated, the decision loop affected, the human override path, and the measurable outcome.

01

CABA consumer-goods distributor — the decision-exposure map. System: SAP S/4HANA stock ledger and the MRP run that raises purchase requisitions. Decision loop: automated replenishment proposals committed without planner review below a value threshold. The count plan was a textbook ABC schedule ranked by inventory value, so expensive slow movers were counted monthly and cheap fast movers quarterly — the precise inversion of where the empirical literature says error concentrates. The remedy was not more counting but re-ranking: each record scored by decisions committed against it per quarter multiplied by the reversibility of those decisions, then existing count hours reallocated onto the top decile. Override path: the planner may suspend automatic commitment for a material group, with a logged reason and a mandatory expiry at the end of the planning cycle. Illustrative target: decision-weighted accuracy above 90% on the top decile within two quarters, at flat count-hours.

02

Buenos Aires province regional grocery DC — counting where the error lives. System: WMS bin-level on-hand and the cycle-count module, feeding a store allocation engine. Decision loop: daily allocation of short-shelf-life categories across roughly seventy stores. Headline accuracy had improved for three years while out-of-stock complaints on fresh categories held flat — the classic symptom of an average concealing a concentration. Following the Rekik and colleagues finding that inaccuracy rises with restocking frequency and perishability, the count plan was re-cut by handling intensity rather than value, and blind verification was inserted at commitment for the top fresh categories rather than on a calendar. Override path: the shift supervisor may release against a stale record, time-boxed to one shift, reason-coded, visible to the category team next morning. Illustrative targets: record age at commit below 7 days on fresh categories, and variance — not the mean — brought inside a declared band before any further automation is granted.

03

Multi-client 3PL, AMBA — teaching the agent to refuse. System: multi-tenant WMS wave release and slotting, writing against client-owned inventory the operator does not own and cannot write off. Decision loop: ambient wave release and replenishment to pick face. The operator's exposure is contractual rather than balance-sheet: a wave released against a phantom position produces a short ship, a client claim and an argument about whose record was wrong. The control introduced was a refusal condition rather than an accuracy target — the agent may not release a wave containing a line whose bin record exceeds the client's agreed verification window, and must instead escalate. Override path: the operations duty manager may authorise release, which writes an exception record visible in the client's monthly service review. Illustrative target: a decision-blocked rate that is non-zero and falling, since a blocked rate of zero after implementation means the threshold was set to permit whatever the operation was already doing.

All three patterns share a structure worth naming, because it is what separates a production control from POC theater. None of them buys a new accuracy technology first. Each one changes what the existing verification capacity is pointed at, and each one makes the agent's write authority conditional on the freshness of what it is reading. Technology selection — drones, RFID, vision-assisted picking — becomes a legitimate investment question only once the exposure map exists, because until then nobody can say which bins are worth instrumenting.

04 · Implementation

Make freshness a precondition, not a report

The sequence is short and does not begin with procurement. First, build the decision-exposure map from data already held: the agent's commit log joined to the stock ledger gives decisions per record per period, and the requisition or allocation value gives consequence. Second, reallocate existing verification capacity onto the top decile of that map, holding count hours flat so the change survives the budget conversation. Third, write the freshness precondition into the agent's write credential rather than into a policy document.

The third step is where most programmes stop short, and it is the one that converts a metric into governance. A record-age threshold in a standard operating procedure is a recommendation. The same threshold enforced at the credential means an agent that meets a stale record blocks and escalates, producing a visible queue that shows exactly where verification capacity is short. From pilot to policy means that queue is designed and staffed before go-live, because an unstaffed refusal is indistinguishable from an outage and will be switched off within a fortnight.

So what: inventory accuracy was a reporting metric for thirty years because nothing acted on it automatically. The moment an agent holds write authority, accuracy stops being a number in a monthly pack and becomes a precondition for the decision.

Governance

One control: the verification covenant, a written, per-decision-class record-freshness precondition enforced at the agent's write credential and owned by the inventory control lead reporting to the COO, not by IT. Each entry carries the decision class, the maximum permitted record age, the blind-count sampling rule applied at commitment, the named human who may override and for how long, the escalation path when an agent blocks, and a quarterly review against realised decision-weighted accuracy. Where a third party holds the record — a 3PL, a consignment partner, a client-owned pool — the covenant becomes a contract clause with a verification window and an exception-reporting obligation, because a precondition only one party can observe is not a control. Overrides are logged as decision events carrying the same fields as agent commitments.

KPIs

Decision-weighted record accuracy: share of agent-committed decisions taken against a record inside its class window, weighted by decision value. Baseline effectively unmeasured everywhere; empirical proxy is the roughly 60% of records found wrong in the ECR work; target above 90% on the decision-dense decile within two quarters. Record age at commit: median and p95. Baseline is usually the count cycle itself, 30–90 days; target under 7 days on high-exposure classes. Decision-blocked rate: share of commits refused on a stale record. Baseline zero because no precondition exists; zero after implementation is a finding, not an achievement. Blind-count variance at commit: reported as spread rather than mean, because the mean is what has been improving while the decisions kept failing. Count-plan yield: service level or margin movement per verification hour, by class — the number that defends the reallocation.

90D 180D 360D

12-month roadmap

0–90: join the commit log to the stock ledger and publish the decision-exposure map for the top two decision classes, however unflattering; compute decision-weighted accuracy retrospectively over four quarters so the baseline is historical rather than aspirational; name the covenant owner. 90–180: reallocate existing verification capacity onto the top decile at flat count-hours and measure count-plan yield against the prior ABC schedule; write the freshness precondition for one decision class and enforce it at the credential; staff the blocked queue before go-live. 180–360: extend the covenant to the remaining consequential classes and into third-party contracts at renewal; make a verification window a standing clause in 3PL and consignment agreements; put decision-weighted accuracy and record age at commit on the operations scorecard beside service level and working capital, and only then evaluate sensing technology against the bins the map has identified.

Socradata Perspective

Accuracy was a report. Now it is an interlock.

There is a reason inventory record accuracy has survived three decades of documented failure without becoming a board-level constraint. Its consequences were absorbed by people. A planner who had learned not to trust the record for a particular material group quietly adjusted; a supervisor walked the aisle before committing a wave. That absorption was invisible, unpaid and load-bearing. Automation removes the absorber and leaves the defect, which is why the first year of an agentic replenishment programme so often produces excellent adherence metrics and worse service.

The correction is not a better model and it is not, in the first instance, a sensing investment. It is a change in what the number is for. An accuracy figure in a monthly pack is information. The same figure expressed as a freshness threshold on a write credential is an interlock — a condition the system checks before it may act. That is a small engineering change and a large governance change. Interoperability or it doesn't scale has a quieter corollary here: a precondition the ERP, the WMS and the partner platform do not all recognise is not a precondition, because the decision will simply be taken in whichever system is not checking.

Socradata transforms ERP, WMS and supply-chain data into predictive intelligence and governed operational decision systems — and the first thing a governed decision system has to know is whether it is allowed to believe what it is reading.

Find out what your agent is believing

Every Wednesday, The Operational AI Dispatch takes one consequential AI signal and translates it into an operating model, a KPI set and a practical action plan for leaders running enterprise operations, ERP, WMS, supply chains and public systems. Published weekly by Socradata from Buenos Aires and New York.