01 · Context

The oversight everyone built governs the wrong moment

On 14 May 2026, SAP used Sapphire to announce Autonomous Supply Chain Management: six Joule Assistants and more than 60 purpose-built agents phased into general availability through the year. The language is careful and worth reading closely. People define goals, assistants orchestrate, and agents execute the work — validating inbound receipts, aligning labour with real workload, resolving inside defined guardrails. This is not a copilot drafting a recommendation. It is software that writes to the system of record and, through it, to a counterparty.

The market context is no longer speculative. Gartner forecasts that 40% of enterprise applications will carry task-specific AI agents by the end of 2026, from under 5% in 2025. McKinsey's State of AI in 2026 reports 40% of organisations above USD 1 billion in revenue scaling agents, against 27% a year earlier. And since 2 August 2026 the EU AI Act's high-risk obligations have applied, including the Article 14 requirement that a designated person be able to interrupt the system or bring it to a halt in a safe state.

Three announcements, one architectural omission. The vendor guardrail, the analyst checklist and the statutory oversight interface all govern the moment before the agent acts: approve or decline, monitor, stop. None specifies what happens to the commitment the agent already made. A stop button halts the next action. It does nothing to the production order released eleven minutes ago, the load tendered to a carrier who has already accepted it, or the shift roster published to a workforce under a notice period. The operational question is not whether a human can intervene, but whether intervention, when it arrives, still has anything to act on.

Governance that ends at the approval gate is not governance of an autonomous system. It is governance of a recommendation engine that was quietly promoted.

02 · Framework

The reversibility window

The analytical lens arrives with a proof. In Revisable by Design, Zhai, Li and Wang formalise a reversibility taxonomy classifying every action an agent can take as idempotent (it does not change the world — a query, a read), reversible (an exact inverse exists), compensable (no inverse exists, but a further action restores an acceptable state at a cost) or irreversible (no compensation exists, or its cost is prohibitive). Their claim is stated plainly: an agent's flexibility is bounded by its reversibility, and the cost of conflicting irreversible actions is a property of the action space rather than of the algorithm. No amount of model improvement removes it.

The authors note that this mirrors the saga pattern from distributed database theory, described by Garcia-Molina and Salem in 1987. Enterprise architecture solved this problem forty years ago for money movement and is now rediscovering it for physical operations; the recovery side is being formalised again in work such as SagaLLM and Robust Agent Compensation. What has not happened is the translation into an operating model a COO can own.

Socradata's extension is one sentence and it is the whole argument: a reversibility class is not a property of an action, it is a property of an action at a point in time. A purchase order is reversible until the supplier confirms it, compensable until material is cut, irreversible once perishable input is issued against it. Nothing about the action changed. The clock moved. Three layers follow.

Layer 1 — Class

What kind of action is this at the instant it executes? The classification is made per action and per counterparty, never per agent. A single replenishment agent raises requisitions (reversible), confirms orders to a business network (compensable) and triggers material issue (irreversible) — three classes inside one workflow, under one credential. Granting that agent a single autonomy level is the design error, and almost every deployment makes it.

Layer 2 — Window

How long does the class hold before it degrades? Nobody measures this, because it does not live in your systems. It lives in the supplier's acknowledgement rule, the carrier's tender acceptance terms, the collective agreement's notice period, the customs filing deadline. The window is set by the counterparty and runs on the counterparty's clock — which makes it an interoperability problem, not a configuration problem. Interoperability or it doesn't scale applies as much to contract terms as to APIs.

Layer 3 — Compensation

What named, tested action restores an acceptable state, what does it cost, and who may run it? A compensating action is not a rollback: it is a new transaction leaving its own trace, and it must appear in the decision log beside the commitment it reverses. If the answer is "we would call the supplier," that is not a compensation path. It is an improvisation, discovered at the worst moment by the least senior person available.

So what: an approval gate governs the moment before a decision. A compensation path governs everything after it. Enterprises have built the first and assumed the second.

The spanning metric that falls out of the three layers is window margin: median time from commitment to irreversibility, minus median time to detect the commitment was wrong. A positive margin means errors are correctable by design. A negative margin means the organisation has automated decisions whose mistakes it can only absorb, never reverse — and it will discover this in the variance account, not the control room. Most operations running agents cannot compute this number, which is itself the finding. KPIs before APIs is not a slogan here: window margin is what should have been instrumented before the first credential was issued.

03 · Use Cases

Three decision loops, three clocks

The patterns below are anonymised composites from operating and advisory work in Argentina and the Southern Cone; figures are illustrative targets rather than audited client results. Each names the system implicated, the decision loop, where the window closes, the human override path and the measurable outcome.

01

CABA food and beverage plant — the release that cannot be recalled. System: ERP production planning and execution, integrated to an MES on the packaging lines. Decision loop: an agent validates material availability and work-centre capacity, then releases the production order without planner intervention. Mapped against the taxonomy, the release is reversible for roughly 40 minutes until components are staged, compensable through changeover for perhaps 90 minutes more, and irreversible once short-shelf-life input is issued to the order. Detection latency for a bad demand signal was four hours. Window margin was structurally negative and nobody had written the number down. The correction cost nothing in software: agent autonomy was scoped to the pre-staging window, with a named shift supervisor holding release authority past that boundary and a two-hour escalation clock. Illustrative outcome: scrap from late plan changes down 15–20% in two quarters, with no reduction in agent throughput.

02

Southern Cone cross-border lane — the tender that becomes a contract. System: TMS load building and tendering, integrated to carrier portals on the Argentina–Brazil Mercosur corridor. Decision loop: an agent builds loads from confirmed orders and tenders to the routing guide. The window is short and external: the tender is reversible until the carrier accepts, which on this lane averaged eleven minutes, and compensable thereafter only at a cancellation and repositioning cost the operator had never separated from general accessorial spend. KPMG's 2026 South American supply chain outlook notes that automation lowers transaction costs precisely where inflationary pressure makes them most punishing — which is also where an unpriced cancellation hurts most. The fix was a configurable pre-tender hold sized to detection latency, plus a compensation cost code so the price of reversal became visible. Illustrative outcome: cancellation and repositioning spend identified at 1.8% of tendered freight value; hold window set at 25 minutes; service impact under 0.5 percentage points.

03

Regional e-commerce fulfilment — the roster that is already payroll. System: WMS predictive labour planning, of the kind SAP is now shipping in Extended Warehouse Management, feeding a workforce management module. Decision loop: a demand forecast produces a shift roster that becomes a published commitment. The scale context is public: Mercado Libre's first-quarter 2026 results report more than 50 fulfilment facilities handling 55% of shipments, with unit shipping costs in Brazil down 17% year on year in local currency. Networks of that density commit labour days ahead. Here the reversibility window is not a system parameter at all: it is the notice period in the applicable collective agreement, after which the roster is paid regardless of whether the volume arrives. Override sits with the site operations manager and the escalation path must reach an industrial relations owner, not an IT ticket queue. Illustrative outcome: forecast revision cut-offs aligned to the notice period, cutting paid-but-unworked hours by 8–12% on volatile weeks.

The three share a structure that separates a production control from POC theater. None begins by buying anything. Each rewrites an autonomy grant written once for an agent into grants written per action class and per window. And each makes visible a cost already being paid — in scrap, accessorials, idle paid hours — but never attributed to the decision that caused it.

04 · Implementation

Price the window before you widen the autonomy

The sequence begins with data the organisation already holds. First, build the commitment ledger: take the agent's execution log, join it to the transaction records it wrote, and classify each action type by reversibility class and by the elapsed time at which that class degrades. Most of this is discoverable from timestamps — requisition to PO confirmation, tender to acceptance, roster publication to notice cut-off. Where the window sits in a contract rather than a system, read the contract; that work is legal and commercial as much as technical.

Second, compute window margin retrospectively over four quarters so the baseline is historical, and price compensation by opening a reversal cost code on each class. Third — where most programmes stop short — enforce the result at the agent's write credential, not in a standard operating procedure. A rule in an SOP is a recommendation. The same rule as a credential condition means an agent reaching an irreversible class without a named human blocks and escalates, producing a queue that shows where human capacity is short. Staff that queue before go-live. From pilot to policy is this transition: the pilot proves the agent can act; the policy decides what it may not finish alone.

So what: autonomy is not a property of the agent. It is a property of the window in which its action can still be undone — and that window closes on your counterparty's clock, not on yours.

Governance

One control: the compensation covenant, a per-action-class register enforced at the agent's write credential and owned by the operations director in the COO line, not by IT. Each entry carries the action class, its reversibility class at execution, the window length and the fact that sets it, the named compensating transaction with its tested cost, the human authorised to run it, and the escalation clock. The binding rule is short: no agent executes an action whose class at execution is irreversible unless a named human holds the decision. Where a third party owns the clock — carrier, 3PL, supplier network, works council — the covenant becomes a contract clause specifying a reversal window and an exception-reporting obligation, because a window only one party can observe is not a control.

KPIs

Window margin: median hours from commitment to irreversibility minus median detection latency, per decision class. Baseline unmeasured almost everywhere and frequently negative once computed; target positive on 100% of agent-executable classes within two quarters. Compensation cost ratio: reversal spend over the value of agent-committed transactions; baseline hidden inside accessorials, scrap and premium labour; target measured within 90 days, then held under 1.5% of committed value. Irreversible-class autonomy share: proportion of irreversible-class actions executed without a named human; target zero, and zero is a control rather than an achievement. Compensation success rate: share of attempted reversals that restore an acceptable state without escalation; target above 90%, and an untested reversal is reported as unproven, not available. Decision-blocked rate: volume and ageing of the escalation queue, read as human capacity rather than agent failure.

12-month roadmap

0–90: build the commitment ledger for the two highest-value decision classes, compute window margin retrospectively over four quarters, open the reversal cost code, name the covenant owner in the operations line. 90–180: enforce the covenant at the write credential for one class, staff the escalation queue before go-live, test every declared compensating transaction at least once in production conditions, put reversal windows into the next carrier and 3PL renewals. 180–360: extend the covenant to remaining consequential classes, carry the window question into supplier onboarding and collective bargaining, put window margin and compensation cost ratio on the operations scorecard beside service level and working capital, and only then widen autonomy — class by class, never agent by agent.

Socradata Perspective

The enterprise already solved this. For money, not for materials.

There is something instructive in the fact that the reversibility taxonomy now being formalised for AI agents restates a 1987 database paper. Finance systems never assumed a committed transaction could simply be undone. They were built on compensating entries, reversal journals and audit trails precisely because settlement is irreversible and everyone knew it. The discipline is old, mature and boring, which is the highest compliment an operational control can receive.

Physical operations were spared that discipline for a structural reason: a human sat between the plan and the commitment and absorbed reversibility informally. Planners held orders back when a number looked wrong. Dispatchers waited a beat before tendering. Supervisors knew which rosters could still be pulled. That absorption was unpaid, invisible and load-bearing, and it was the first thing removed when the agent received the credential. What replaced it governs a moment already past by the time anything goes wrong.

The correction is not a better model, a stricter gate or a richer dashboard. It is the recognition that an autonomous operational system is a transactional system, and transactional systems are governed by what can be compensated, at what price, inside what window. That is an architecture question, a contract question and a decision-rights question at once — which is why it keeps falling between the vendor, the integrator and compliance.

Socradata transforms ERP, WMS and supply-chain data into predictive intelligence and governed operational decision systems — and a governed decision system is one that knows, before it acts, what it would take to take the action back.

Find out what your agent cannot take back

Every Wednesday, The Operational AI Dispatch takes one consequential AI signal and translates it into an operating model, a KPI set and an action plan for leaders running enterprise operations, ERP, WMS, supply chains and public systems. Published weekly by Socradata.