01 · Context

The agents arrived before the sign-off did

Three things happened this year that an operations executive should read together. On 20 May 2026, at Momentum in Las Vegas, Manhattan Associates moved agents out of the demo track and into the product line, launching Solution Design Studio and Manhattan Marketplace on top of the AgentFoundry platform released a year earlier. Giant Eagle's director of supply-chain technology described the Wave Coordinator Agent running inside grocery distribution operations where, in his words, three quarters of product is shipped and received the same day. Manhattan's own framing is instructive: every agent runs on a platform offering unified data, governance, APIs and deterministic execution. The vendor has already conceded that the scaffold beneath the agent must be deterministic. What it cannot supply is the customer's rule for when the agent is allowed to use it.

Second, agent creation stopped being an engineering activity. Kinaxis released Maestro Agent Studio, a no-code environment in which supply-chain teams compose agents grounded in their own operating context. That is a genuine advance in accessibility and a governance event at once. When a planner can assemble a decision-making agent in an afternoon, the constraint on how many agents reach your ERP is no longer the IT backlog.

Third, the acceptance discipline did not arrive with them. The Observatorio Agentic AI 2026, produced by NTT DATA with CIONET from 130 executives of large organisations in Spain and Latin America, found 59% operating dispersed pilots and 3.8% reaching industrial deployment. Underneath that headline sit the mechanics: 40% of executives cannot construct a solid economic case, only 10% of initiatives carry direct executive sponsorship, and while 80% report medium or high data quality, only 26% consider themselves prepared to support autonomous agents at scale. The same study reports that a large majority of organisations lack explicit operational mechanisms to govern agent autonomy at all — no kill switch, no containment procedure for a failing or hallucinating agent. Gartner reached the destination from the other direction, forecasting that over 40% of agentic AI projects will be cancelled by the end of 2027, naming escalating cost, unclear business value and inadequate risk controls as the causes.

Read the two datasets together and the diagnosis sharpens. Organisations are not failing to build agents. They are failing to decide, in writing, what would make one acceptable. Without that decision the pilot cannot end, because there is no finish line, and the project eventually exhausts its sponsor. Cancellation and permanent piloting are the same failure wearing different clothes.

A pilot with no acceptance criterion cannot succeed. It can only continue until somebody stops paying for it.

02 · Framework

The Acceptance Protocol

Socradata uses a three-gate protocol to decide whether an operational agent may act inside an ERP, WMS or planning engine. Each gate is an artefact, not a meeting, and nothing proceeds until the previous artefact exists and is signed by a named person on the operations side.

Gate 1 — Envelope

A written statement of the decision class the agent may take, the tolerance band inside which it may act unassisted, and the conditions on which it must refuse and escalate. The refusal conditions are the load-bearing part, and they are written before any evaluation runs, so that the envelope is a specification rather than a rationalisation of whatever the agent turned out to do. An envelope is specific: this agent may release replenishment waves for ambient SKUs within a labour-hour band; it must refuse on cold-chain, controlled-substance and any wave exceeding a stated share of shift capacity. A vendor can propose an envelope. Only the operator can sign one.

P95
Gate 2 — Distribution

Acceptance evidence produced by replaying the agent against a fixed set of historical operating periods many times, and scoring the resulting distribution rather than a demonstration. Three numbers matter: median decision quality, the p95 worst case, and the variance ceiling. The p95 is the one that governs, because operational risk lives in the tail. Beneath it, the deterministic scaffold is gated separately. Zhang and colleagues, publishing on 10 June 2026, decompose a deployed ordering agent into layers — ontology, intent, routing, decomposition, escalation, safety, memory — and gate each with a no-LLM regression harness of 238 cases across 23 slices running against locked baselines, with a coverage-honesty rule that refuses to score an untested layer. Their observation is the practical one: end-to-end success tells you the agent regressed, not where.

Gate 3 — Revocation

Authority that expires, with a named owner who can withdraw it and has proved it. Acceptance is issued as a certificate against one decision class, with an expiry date and a scheduled re-evaluation, because the model behind the agent will be updated, the operating context will drift and the certificate you issued in March describes a system that no longer exists in September. The test is not whether a kill switch is documented. It is whether somebody has pulled it, under observation, and timed the result. Deployers with European exposure will recognise the same requirement in Article 26 of the EU AI Act, which obliges human oversight and the retention of automated logs; the substantive obligation for Annex III systems now runs to December 2027, which is time to build the control rather than a reason to postpone it.

The spanning metric is revocation latency: the elapsed time between a detected out-of-envelope decision and the agent losing write authority for that decision class, measured by unannounced drill rather than asserted by policy. It is deliberately unflattering. Most organisations discover on the first drill that revocation requires a vendor support ticket, that nobody knows which credential the agent uses, or that removing its access also removes an integration three other processes depend on. KPIs before APIs: an agent whose authority you cannot withdraw inside an operating shift is not accepted, whatever the pilot deck says.

So what: the question is not whether the agent works. It is whether you have written down what working means, measured it on a distribution, and proved you can take the authority back before Monday's shift ends.

03 · Use Cases

Three operating patterns

The patterns below are anonymised composites from operating work in the Southern Cone and the Andean region. Ranges are illustrative targets, not audited results. Each names the system, the decision loop, the human override path and the artefact that constitutes acceptance.

01

E-commerce fulfilment operator, Buenos Aires — the written envelope. System: WMS wave release and labour management, integrated to the order management system. Decision loop: which orders enter which wave, when the wave releases, and how many pickers are assigned. Before evaluation begins, the operations director signs a one-page envelope: the agent may release ambient waves inside a labour-hour band of ±8% against the shift plan and no later than ninety minutes before carrier cut-off; it must refuse and escalate on any wave containing cold-chain or controlled items, any wave exceeding 70% of remaining shift capacity, and any release inside the final thirty minutes before cut-off. Override path: the shift supervisor may release outside the band with a reason code, and the override is logged as an agent event rather than a human one. Acceptance target for the first quarter: cut-off adherence at or above the human baseline at p95, not at the mean, across 180 replayed shift-days.

UNTESTED
02

Mining-services group, Santiago de Chile — layer isolation on the requisition. System: ERP purchase requisition, supplier master and contract-price records. Decision loop: approval of maintenance, repair and operations requisitions against framework agreements. The architectural rule is that the checks with a legal or financial consequence never run on a language model. Supplier validity, tax treatment, contract-price conformity and budget availability stay deterministic and are gated by a no-LLM assertion suite that runs on every change against a locked baseline; the model is confined to drafting justifications, ranking substitutes and flagging anomalies for a buyer. Coverage honesty is enforced: if a layer has no assertion slice, the release is blocked rather than scored as a pass. Override path: any requisition the agent escalates goes to a named category buyer with the agent's reasoning attached. Acceptance artefact: a per-layer baseline report, not a demonstration.

T-MINUS
03

Omnichannel retailer, Bogotá — the revocation drill. System: allocation and replenishment planning engine writing store-level allocation decisions. Decision loop: how much of a constrained article each store receives. This organisation issues acceptance as a certificate with a two-quarter expiry, and treats revocation as an operational drill on the model of a fire test. Once per quarter, unannounced, the planning governance lead declares a simulated out-of-envelope allocation and the clock starts; it stops when the agent has lost write authority for that decision class and the affected allocations have reverted to the prior rule. The first drill is customarily embarrassing and that is its value. Target: p95 revocation latency under 15 minutes by the third drill, with authority scoped by decision class so that revoking allocation does not sever the demand feed. Override path: revocation is a one-person action, and reinstatement requires a fresh certificate.

04 · Implementation

From demonstration to certificate

Start with an inventory, not an evaluation. List every agent currently touching an operational system, whether procured, composed in a vendor studio or assembled by a team that did not think of it as software. For each one, record the decision class it affects, the credential it writes with, the person accountable for a bad decision and the date its behaviour was last assessed. In most enterprises this takes a fortnight and produces two uncomfortable findings: several agents have no accountable owner outside IT, and at least one is writing with a shared service credential that nobody can trace back to a decision.

Then build the evaluation set before the evaluation. Take a fixed sample of historical operating periods — shift-days, planning cycles, requisition batches — that includes the difficult ones: the promotion week, the port strike, the month the supplier failed. Replay the agent against them, score the distribution, and hold the deterministic layers to a locked regression baseline. This is the step that separates productionisation from POC theater, because a demonstration is a single favourable sample and everybody in the room knows it. From pilot to policy means the pilot ends with a signed certificate or it ends with a decision not to proceed. Those are the only two acceptable outcomes.

So what: your vendors will keep shipping agents faster than your controls mature. The one artefact that closes the gap is an expiring, per-decision-class acceptance certificate that names an owner and has been tested by revocation.

Governance

One control, owned by operations. Every agent with write access to an operational system holds a dated acceptance certificate for one decision class, carrying the signed envelope and refusal conditions, the evaluation set and date, the p95 result against the human baseline, the named accountable owner, the revocation procedure, the record of the last drill and an expiry no longer than two quarters. No valid certificate, no write authority, enforced at the credential rather than in policy. The owner is the operations or planning governance lead, not IT, because the consequence of a bad wave release lands on the floor. Interoperability or it doesn't scale: the certificate must be recognised by the ERP, the WMS and the planning engine, or you will have three registers disagreeing about which agent is allowed to act.

KPIs

Envelope coverage: share of agent-executable decision classes carrying a signed envelope with explicit refusal conditions. Baseline unmeasured; the regional proxy is the large majority of organisations the Observatorio finds without any explicit mechanism to govern agent autonomy. Target 100% on classes with financial or safety consequence within one quarter. Revocation latency: median and p95 elapsed time from detection to loss of write authority, measured by unannounced drill. Baseline undrilled; target p95 under 15 minutes by the third drill. Scaffold regression coverage: share of deterministic layers holding a locked, no-LLM assertion slice in continuous integration, with untested layers scored as failures. Baseline typically zero; target above 90% within two quarters. Distributional acceptance margin: agent p95 decision quality no worse than the human p95 across at least 150 replayed operating periods. Certificate freshness: share of live agents inside a valid, unexpired certificate. Target 100%.

90D 180D 360D

12-month roadmap

0–90: inventory every agent touching an operational system with its decision class, credential and accountable owner; write and sign envelopes for the two decision classes with the largest financial exposure; assemble the replay set from historical operating periods including the hard weeks; run the first revocation drill and publish the time, however poor. 90–180: stand up the no-LLM regression harness on the deterministic scaffold with locked baselines and coverage honesty; issue the first acceptance certificates against p95 rather than mean performance; enforce certificate validity at the credential so an expired certificate withdraws write access automatically. 180–360: extend certification to composed and studio-built agents at the point of creation rather than at go-live; add revocation latency and certificate freshness to the operations scorecard beside service level; move re-acceptance from calendar-driven to trigger-driven on model updates and material context change.

Socradata Perspective

Acceptance is the new integration.

For twenty years the hard part of enterprise systems work was integration: getting the warehouse to speak to the ledger, the planning engine to speak to both. That problem has not disappeared, but it has been substantially commoditised by platforms and APIs, and the vendors have earned that. The new hard part sits one layer up. When a system stops executing instructions and starts producing decisions, the binding constraint moves from whether the data can flow to whether the organisation can say, in advance and in writing, which decisions it is willing to delegate and on what evidence it would take the delegation back.

Treating that capability as a compliance exercise produces the worst of both outcomes: a policy document and an ungoverned agent. The useful precedent is industrial rather than regulatory. No refinery operator accepts a new control loop because the vendor demonstrated it; they accept it against a written operating envelope, they trip it deliberately during commissioning to see how it fails, and they re-certify it on a schedule. That discipline was not invented for artificial intelligence. It was invented for any system that acts faster than a human can review, which is precisely what an operational agent now is.

This is the layer Socradata works in, and the entry point is unglamorous: write the envelope, build the replay set, run the drill, issue the certificate. Socradata transforms ERP, WMS and supply-chain data into predictive intelligence and governed operational decision systems — and in this domain, governance means an agent whose authority you granted deliberately and can withdraw before the shift ends.

Time how long it takes to switch yours off

Every Wednesday, The Operational AI Dispatch takes one consequential AI signal and translates it into an operating model, a KPI set and a practical action plan for leaders running enterprise operations, ERP, WMS, supply chains and public systems. Published weekly by Socradata from Buenos Aires and New York.