01 · Context

The industry has agreed on a number that cannot exist

Three things happened within ten days this month, and together they describe a governance problem rather than a maturity curve. On 24 September 2026, Gartner predicted that by 2030 only 5% of organisations implementing supply chain planning automation will make at least 10% of planning decisions autonomously — against a survey of 243 senior leaders at companies above USD 500 million in revenue, fielded 11 November to 18 December 2025, in which 83% had already spent USD 3 million or more on planning automation. Buse Aras, Director Analyst in the firm's supply chain practice, framed the gap plainly: investment alone does not create readiness, and leaders risk funding technology that does not improve the decisions that matter.

On 16 September, the same firm named suggestive and semiautonomous agents as one of four AI trends reshaping warehousing, defined as the band between manual operation and full autonomy. And a World Economic Forum study published in September, summarised by Infobae on 24 September, mapped 612 startups holding close to USD 7 billion in funding and projected logistics autonomy rising from 1.5 to 3.5 on a four-point scale. It also found only 7% of organisations able to execute decisions in real time.

Notice what the forecast assumes: autonomy as a scalar, one level for one network travelling one way. Two experiments published this year say the scalar is a modelling error.

The first is the Agentic AI Autonomy Assessment framework of Lennart Trumpler, Rodrigo Furlan de Assis, Elias Ribeiro da Silva, Luis Antonio de Santa-Eulalia and Christian Hendriksen, posted 28 July 2026. They define a bounded autonomy score from an initiative rate and a consultation rate, then run 240 experiments across the four tiers of a beer distribution game. The cost response reverses by position: −195.64 euros per unit of autonomy score at the factory and −211.14 at the distributor, turning positive at +60.04 for the wholesaler and +68.45 for the retailer, with a pooled gradient of +109.6 euros per tier step at p=0.015. Their own reading is that this "positions autonomy as a governance variable rather than a performance variable", and that it is not a one-size-fits-all design parameter.

The second explains the mechanism. Carol Xuan Long, David Simchi-Levi, Feng Zhu, Huangyuan Su, Andre Calmon and Flavio Calmon, in work posted in May 2026, replaced every player in that game with a language-model agent and ran thirty identical repetitions. They name the result agent bullwhip: decision instability amplified by the chain itself. With demand held fixed, order variability across runs is barely visible at the retailer and most pronounced at the distributor and the factory. Coefficients of variation on total cost reach 0.46 in one configuration and fall to 0.13 only after task-specific post-training. Sampling more did not help — majority voting over one hundred samples did not reduce run-to-run variability — which makes this a policy problem, not a compute problem.

Two independent 2026 experiments agree on direction: the further a node sits from the customer signal, the more autonomy earns and the more variance it exports upstream.

02 · Framework

The autonomy gradient

A practitioner forecast reached us this month through private circulation; its publication could not be independently verified, so it is engaged here as a position rather than cited as authority. It frames the next twelve months of enterprise AI as a rotating bottleneck — capability, context, authority and proof, attention, economics — and argues that cost per accepted outcome should replace token price as the unit of account. The sequence is sound and the unit is right. What a general enterprise vantage point cannot see is that in a multi-echelon system the rotation is not only temporal: all five constraints bind at once, at different nodes. A single moving bottleneck produces exactly the scalar the evidence refuses. Socradata therefore governs autonomy on three layers rather than one level.

Layer 1 — Position

Where the node sits between the customer signal and the constrained asset. Position is not an org chart fact but a measurable distance: information lead time plus intermediaries between the node and the demand it serves. It determines which constraint binds and the direction in which autonomy pays. Upstream nodes face amplified, delayed, uncertain signals, and proactive decomposition helps them. A node facing the customer faces the signal itself, where autonomy amplifies noise the rest of the chain inherits.

Layer 2 — Band

A declared minimum and maximum for that node, not a target, expressed in the two rates the research measures: how often the node originates its own tasks, and how often it consults a human before acting. It needs a floor, because an agent that escalates everything makes the reviewer the runtime, and a ceiling, because drift above it is where incidents live. Gartner reached the same structure on 26 May 2026: agents "operate at different autonomy levels and across different trust boundaries", and by 2027 40% of enterprises will demote or decommission autonomous agents over governance gaps found only after a production incident.

Layer 3 — Budget

The node's cost per accepted decision: inference, tool calls, review minutes and rework, over decisions the process owner accepted. This converts the band from a risk preference into an economic choice, and it must be computed per node, because review minutes are cheap upstream and expensive at the customer-facing node where every exception carries a service consequence. A programme-level figure averages two opposite economies and conceals both.

DECLARED BAND AND OBSERVED AUTONOMY, BY NODE 0.0 1.0 AUTONOMY SCORE STORE ORDER DC ALLOCATION DRIFT WAVE RELEASE PLANT SCHEDULE FACING THE CUSTOMER FACING THE CONSTRAINT ILLUSTRATIVE BAND SHAPE. THE VALUES ARE YOURS TO CALIBRATE.

The band widens with distance from the customer signal. Drift above the ceiling, shown here at the DC allocation node, precedes most agent incidents — and is invisible unless someone declared the ceiling.

So what: an enterprise-wide autonomy policy is not a conservative choice. It is a guarantee of being wrong at one end of your own chain, and the end it is wrong about is the one facing your customer.

Two numbers span the framework. Autonomy drift is the observed score minus the declared band midpoint, per node per week — the only way to catch the behavioural change the AAAA authors warn cannot be guaranteed at runtime once a system is calibrated in development. Its companion is the amplification ratio: run-to-run dispersion of a node's committed quantities under identical inputs, over the same measure at the node immediately downstream. Above one, the node is manufacturing variance for everyone upstream. KPIs before APIs is not a slogan here: neither number needs new software, and neither is on any operations scorecard today.

03 · Use Cases

Four nodes, four bands, one chain

The patterns below are anonymised composites from operating and advisory work in Argentina and the Southern Cone. Figures are illustrative targets, not audited client results. Each names the system, the decision loop, the position, the band and the human override path.

01

CABA grocery chain — the store order, narrow by design. Systems: ERP replenishment, WMS pick-face minimums at the serving DC. Decision loop: a daily order proposal per store and article, committed to the DC. Position: adjacent to the customer signal — where the evidence says autonomy raises cost, and where the failure is familiar to every category manager: an agent reading yesterday's promotion lift as a level shift. The band is deliberately narrow, with the agent committing freely inside a category-specific quantity envelope and consulting above it. Override path: the category planner, two-hour window before DC cut-off, logged reason code. Illustrative outcome: order-line amendments down a quarter, promotional overstock down 10–15%.

02

Southern Cone consumer goods — the allocation node, wide but fenced. Systems: APS distribution requirements planning across three DCs, TMS tendering, a customs filing platform on a Mercosur corridor. Decision loop: allocate constrained stock across DCs and release the replenishment and freight orders behind it. Position: mid-chain, the transition point the AAAA experiment locates at the wholesaler, where the slope on cost is near zero and the sign is ambiguous. The right response to ambiguity is not a middling setting but a split one: a wide band on reversible actions such as allocation between owned DCs, and a hard ceiling on anything running on an external clock — a carrier tender, a customs filing — where consultation is mandatory regardless of confidence. Override path: the supply planning manager for allocation, the trade compliance lead for anything filed. The prize is publicly estimated: the World Economic Forum mapping puts disruption response 62% lower and recovery 60% lower with autonomous capability.

03

AMBA multi-client 3PL — the warehouse labour node, wide and instrumented. Systems: WMS wave release, engineered labour standards, workforce management, dock scheduling. Decision loop: release picking waves and allocate labour across the shift. Position: upstream of the customer signal by a full order-to-ship cycle, working against forecast volume rather than live demand — where proactive decomposition earns. Gartner's warehousing analysis of 16 September names labour forecasting and slotting as the proven entry points. The band is wide and the discipline is the instrument rather than the approval: amplification is computed nightly on wave sizes under repeated identical inputs, and a ratio above one suspends widening until explained. Override path: the shift manager, with a release-anyway authority that is logged rather than assumed. Illustrative outcome: overtime down 10–15% and carrier cut-off misses roughly halved on peak days.

All three share what separates a production control from POC theater. None changes the model, the vendor or the integration. Each takes an undeclared setting, gives it an owner in the operations line, and puts a number against it that reveals when it stops being true.

04 · Implementation

Declare the band before you widen it

Three actions, in order. First, map position for every node where an agent holds write authority: information lead time, intermediaries to the demand signal, and whether the commitment runs on your clock or a counterparty's. It is two weeks of work, and it produces the first honest inventory of where agents act. Second, declare a band per node and name the process owner who holds it. The floor matters as much as the ceiling: an agent that escalates everything has moved the bottleneck onto the reviewer, and that failure arrives at renewal. Third, instrument drift and amplification before widening anything.

State the boundary as plainly as the finding. Neither experiment is a field deployment. The AAAA study reports coefficients rather than correlations and autonomy clustered in a narrow high range; the agent bullwhip paper defines amplification ratios it never computes and reports no significance tests. They establish direction and mechanism, not a calibrated value for your chain. The falsifier is clean: if bands are instrumented for two quarters, drift shows no relationship to incident rate, and cost per accepted decision converges across positions, the positional argument is wrong and a single enterprise band is cheaper. From pilot to policy means running that test on purpose rather than learning the answer through an incident.

So what: you cannot govern a gradient with a level. Declare four bands, measure the drift, and let the economics of each position tell you which one to widen next.

Governance

One control: the autonomy register, a per-node record enforced at the agent's write credential and owned by the operations director, not the AI team and not the vendor. Each entry names the node and its position, the declared floor and ceiling, the accountable process owner, observed drift and amplification with measurement dates, the human who may act outside the band, and an expiry after which the band is re-declared rather than inherited. Two rules make it real. An expired band collapses to its floor automatically, so neglect fails closed. And where the commitment runs on a counterparty's clock, the ceiling is a contract term rather than a configuration value, because a setting only one party can change cannot govern a shared decision.

KPIs

Autonomy drift: observed score minus declared band midpoint, weekly; unmeasurable today because no band exists; target absolute drift under 0.05. Amplification ratio: run-to-run dispersion of committed quantities under identical inputs, relative to the next node downstream; target at or below 1.0 at every internal node, since above one the node exports variance. Cost per accepted decision, per node: inference, tool calls, review minutes and rework over accepted decisions; baseline in sixty days, then falling quarter over quarter while volume grows. Review minutes per accepted decision: the attention constraint, per position. Band coverage: share of write-authorised nodes with an owner, a band, a drift alarm and a baseline; typically zero at the start, and the only one of the five that should reach 100%.

Four-quarter roadmap

Q4 2026 — declare. Map position for every write-authorised node, declare a band and an owner for each, and baseline cost per accepted decision. A node without an owner, a band and a baseline is an experiment on production credentials. Q1 2027 — instrument. Weekly drift reporting and nightly amplification measurement on the two highest-value nodes, and make the expiry real by letting one band lapse to its floor on purpose. Q2 2027 — widen selectively. Raise the ceiling only where amplification sits at or below one and review minutes per accepted decision are falling; hold or narrow at the customer-facing node whatever the vendor roadmap proposes. Q3 2027 — renew on evidence. Take cost per accepted decision by position into renewal, and put drift and amplification on the operations scorecard beside service level.

Socradata Perspective

The setting nobody owns is the one that decides

Operational decision intelligence sits at an awkward place in this landscape, and the awkwardness is the point. Vendors sell autonomy as a product capability, analysts measure it as organisational maturity, and boards approve it as an investment level, because those are the forms in which it ships, benchmarks and budgets. None of them can hold a gradient, and a gradient is what a multi-echelon operational system is. The variable that decides the economics — how much initiative this node may take before it consults — has no owner, no declared value, no measurement and no expiry in most enterprises running agents today.

This lands harder in Buenos Aires than in the markets writing the roadmaps. The same World Economic Forum mapping reports North America holding 49% of the funding and hosting 48% of the funded agentic companies, with Latin America at materially lower venture capital — so the defaults arriving in regional deployments were calibrated against other chains, other lead times, other labour agreements. Position is local by construction: it encodes this corridor's border dwell, this DC's cut-off, this workforce's notice period. No vendor default supplies it, and interoperability or it doesn't scale applies to the setting as much as the interface — a band the ERP, the WMS and the partner platform cannot all read governs nothing.

Socradata transforms ERP, WMS and supply-chain data into predictive intelligence and governed operational decision systems — and a governed decision system is one where every node that can act has a declared range, a named owner and a number that shows when it drifted out. If you want that mapped against your own chain before you widen anything, request an operational diagnostic.

Find out which of your nodes is drifting

Every Wednesday, The Operational AI Dispatch translates one consequential AI signal into an operating model, a KPI set and an action plan for leaders running enterprise operations, ERP, WMS, supply chains and public systems. Published weekly by Socradata.