A frontier argument that lands on the shop floor
INSEAD Knowledge published Alexandra Roulet's account of her onstage conversation with Yann LeCun at INSEAD's AI Forum Europe, held at Station F in Paris. LeCun's position is well known and worth stating precisely, because the precision is where the operational value sits. Large language models, he argued, are rote learners whose reasoning holds in mathematics and code, domains where the symbols themselves carry the reasoning. A world model does something different: it is, in his words, what allows you to anticipate what is going to happen as a consequence of your own actions.
The venture behind the claim is real and capitalised. AMI Labs, which LeCun founded in Paris after leaving Meta, raised USD 1.03 billion in March, with LeCun describing the investor base as roughly 40% European, a third American and the remainder largely Asian. He has also joined Project Tapestry, the AI Alliance effort to train open frontier models through federated contribution; its first two proofs of concept completed on 1 September 2026, trained across India, Australia and two cloud regions while exchanging model weights and nothing else.
None of that is an operations story yet. This is: the same week, the empirical case against acting without a forward model got sharper. The SIMMER benchmark, which evaluates language-model planning against a curated symbolic world model, found that frontier models produced error-free plans in at most 17% of cases, that up to 56% of plans carried latent failures which satisfy every stated precondition, and that most of those failures were irreversible. Adding explicit counterfactual simulation — making the model predict the consequence before committing the step — cut latent failures by up to 72%.
Read those two items together and the frontier debate collapses into an operating decision. The question is not whether world models arrive. It is whether the agent now writing to your ERP is required to evaluate a consequence before it commits one.
Operations does not lack a world model. It lacks the discipline of making the thing that acts consult the one it already has.
The consequence stack
Socradata separates an operational decision into three layers, because enterprises fund the first, neglect the second, and leave the third in a planner's head. Collapsing them is what allows an agent programme to pass every integration test and still commit dates the factory cannot meet.
What the systems believe is true right now: the ERP stock ledger, the WMS bin-level on-hand, MES equipment status, TMS in-transit. It is measured continuously, it is imperfect, and — importantly — agents do read it. Almost every deployed operational agent is a State-only actor. It checks whether something is available and commits.
What happens to State if this action is taken. It lives in the finite-capacity model of the planning engine, the routing and changeover matrix, the engineered labour standard, the transit-time distribution, the shelf-life clock. This is the enterprise's world model, and it is generally a decade old, re-estimated by nobody, and absent from the agent's execution path. It is also the only layer that can answer the question LeCun says intelligence requires.
What the action is for, expressed as a scored goal with constraints that may not be traded away. LeCun drew the distinction that matters here: a system can be built so that it is structurally incapable of breaking a rule, rather than merely deterred from breaking it. In operations that distinction is concrete. A constraint written into the objective function of the planning engine is structural. The same constraint written into a policy document is deterrence, and deterrence is what an agent clearing four thousand decisions a day will quietly erode.
So what: a system that cannot predict the consequences of its own action is not making a decision. It is placing a bet and recording it as a plan.
The spanning metric is consequence coverage: the share of agent-executable actions, weighted by decision value, for which a maintained consequence model exists and is evaluated before commitment. Its companion is model-reality divergence, the gap between what the consequence model predicted and what actually happened, tracked per decision class. Coverage without divergence testing is worse than no coverage, because it grants authority to a model nobody has checked. KPIs before APIs applies with unusual force here: these two numbers decide whether an agent deserves a write credential at all.
Three commitments, three missing models
The patterns below are anonymised composites from operating and advisory work in Argentina and the Southern Cone; figures are illustrative targets rather than audited client results. Each names the system, the decision loop, the consequence model that was bypassed, the human override path and the outcome.
CABA consumer-health plant — the promise made from a stock figure. Systems: ERP sales order management, an APS holding the finite-capacity and changeover model, an MES on the packaging lines. Decision loop: an agent returns a delivery date on an incoming order. It was reading available-to-promise from the stock ledger — pure Layer 1 — while the question the customer was actually asking was Layer 2: what does accepting this order do to the changeover sequence and to the three orders already committed on that line? The APS could answer it. It simply was not in the agent's path. The fix required no new software: a capable-to-promise evaluation became a mandatory pre-commitment call, the agent was permitted to promise only dates the constraint model returned feasible, and a named planner retained override authority with a logged reason code. Illustrative outcome: promise-date adherence improving 12–18 points and expedite freight falling by roughly a fifth, with no change in quoted lead time.
Multi-client 3PL, AMBA — the wave released against no model of the shift. Systems: multi-tenant WMS wave release, engineered labour standards, dock and yard scheduling. Decision loop: an agent releases picking waves through the day. The consequence model here is a travel-and-labour simulation that predicts when the shift actually finishes; the operator held the standards but had not re-estimated them since a slotting change two years earlier. That this is solvable is no longer theoretical: at NVIDIA's GTC in March, KION, working with Accenture and Siemens, showed large-scale warehouse digital twins used to train and test autonomous forklift fleets for GXO before deployment. World models are entering the warehouse through simulation, not through chat. The intervention was narrower: re-estimate standards from observed pick data quarterly, gate wave release on predicted shift completion, and give the shift manager an explicit release-anyway authority that is logged rather than assumed. Illustrative outcome: overtime hours down 10–15% and carrier cut-off misses roughly halved on peak days.
Southern Cone export lane — the mean that hid the tail. Systems: TMS transit planning integrated to a customs filing platform on a Mercosur cross-border corridor. Decision loop: an agent commits a delivery window to the consignee at order confirmation. Its consequence model was a single mean transit time per lane. The real consequence model is a distribution with a long right tail at the border crossing, where dwell is driven by documentation completeness and inspection selection rather than by distance. Committing against a mean guarantees that roughly half of commitments are late by construction, which is a modelling error rather than an execution failure. The fix was to commit against a service percentile agreed with the commercial function, to make the percentile itself a named decision right rather than a buried parameter, and to re-estimate the distribution monthly from observed crossings. Illustrative outcome: delivery-window adherence rising from the high fifties to above 85% with no change in physical routing.
All three share the structure that separates a production control from POC theater. None begins with a model purchase. Each takes a consequence model the enterprise already owns, restores it to a maintained state, and makes its evaluation a precondition for the agent's commitment rather than an optional reference the agent may skip.
Fund the second layer before widening the first
The sequence is short. First, inventory the consequence models you already hold: the finite-capacity model, the routing and changeover matrix, the labour standards, the transit distributions, the shelf-life and stability rules. For each, record when it was last re-estimated from observed data rather than last edited. Most organisations find that the answer is measured in years and that nobody owns the number.
Second, join the agent execution log to that inventory to compute consequence coverage by decision value. The result is usually uncomfortable and always actionable, because it names exactly which commitments are being made blind. Third, run divergence testing: for each class, compare what the model predicted against what happened, and publish the p95 error alongside the median, since the tail is what destroys a promise. Fourth, put the evaluation in the execution path at the credential, not in a runbook — an agent that meets a stale or missing consequence model should block and escalate, and the queue that produces must be staffed before go-live. From pilot to policy is precisely this move: the pilot proves the agent can act; the policy decides what it may not decide alone.
So what: you do not need to wait for world models. You need to decide whether your agent is allowed to act without consulting the one you already own.
Governance
One control: the consequence mandate, a per-action-class register enforced at the agent's write credential and owned by the planning director in the COO line, not by IT. Each entry names the action class, the consequence model that must be evaluated before commitment, the owner accountable for re-estimating that model, its last validation date and observed divergence, the named human who may commit anyway, and what the agent must do when the model is stale or unavailable. The binding rule is structural rather than advisory: constraints declared non-tradeable are expressed in the objective function of the planning engine, so that infeasible commitments cannot be produced, rather than in a policy that asks an agent not to produce them. Where a partner owns the model — a 3PL's labour standards, a carrier's transit performance — the mandate becomes a contract clause obliging periodic re-estimation and disclosure, because a model only one party can inspect cannot govern a shared decision.
KPIs
Consequence coverage: value-weighted share of agent-executable actions with a maintained consequence model evaluated before commitment. Baseline typically under 20%, because only availability is checked; target above 80% on the top decision-value decile within two quarters. Model-reality divergence: median and p95 error between predicted and realised outcome per class; an untested model is reported as unvalidated, not as available. Consequence model age: days since last re-estimation from observed data, against each class's own drift period; baseline frequently measured in years. Hard-constraint violation rate: commitments breaching a non-tradeable constraint; target zero, achieved by structural infeasibility rather than by review. Feasibility-blocked rate: volume and ageing of commitments the constraint model refused, read as a measure of model quality and human capacity rather than of agent failure.
12-month roadmap
0–90: inventory the consequence models and their true re-estimation dates, compute consequence coverage for the two highest-value decision classes, and name an owner for each model in the operations line. 90–180: re-estimate the two models that carry the most decision value, stand up divergence testing with published p95 error, and enforce pre-commitment evaluation at the credential for one class, staffing the blocked queue first. 180–360: move non-tradeable constraints into the objective function so infeasibility becomes structural, write re-estimation and disclosure obligations into 3PL and carrier renewals, put consequence coverage and divergence on the operations scorecard beside service level and working capital, and only then widen agent autonomy — one decision class at a time.
The world model arrives from below, not from above
The frontier framing invites a passive posture: wait for the architecture to mature, then buy it. That posture misreads where the asset sits. Operations research has been building consequence models since the first finite-capacity scheduler, and every enterprise running an APS, a labour standard or a transit-time distribution owns one. What it usually does not own is a maintenance budget, a named owner, a divergence test, or any obligation on the software that acts to consult the thing before committing.
There is a second reason this matters more in Buenos Aires than in Palo Alto. A consequence model is local by construction. It encodes this plant's changeover times, this corridor's border dwell, this workforce's agreement, this regulator's stability rules. A general-purpose model trained elsewhere cannot supply it, and the federated, sovereignty-minded direction that Project Tapestry represents only makes the point sharper: the part that cannot be imported is the part that describes your operation. Interoperability or it doesn't scale applies to models as much as to interfaces — a consequence model that cannot be read by the ERP, the WMS and the partner platform alike governs nothing.
Socradata transforms ERP, WMS and supply-chain data into predictive intelligence and governed operational decision systems — and a governed decision system is one that has to say what it expects to happen before it is allowed to make it happen.
Find out what your agent is predicting
Every Wednesday, The Operational AI Dispatch takes one consequential AI signal and translates it into an operating model, a KPI set and an action plan for leaders running enterprise operations, ERP, WMS, supply chains and public systems. Published weekly by Socradata.