01 · Context

The week the model stopped being the product

On July 15, Mira Murati's Thinking Machines released Inkling: a 975B-parameter mixture-of-experts model with 41B active parameters, a 1M-token context window and weights on Hugging Face. Artificial Analysis scored it the leading US open-weights release. The positioning is without precedent for a debut at this scale: the lab states outright that Inkling is not the strongest model available and ships it as a base to be finished — through its Tinker post-training platform, where enterprises supply the corpus, the reward specification and the evaluations. Two days later, on July 17, Tinker raised prices roughly 50% on sampling and 10% on training. The base model is free. The loop that finishes it just repriced upward. That is what a market sounds like when it decides where value lives.

Twenty-four hours after Inkling, Moonshot AI announced Kimi K3: 2.8T parameters, 16 of 896 experts active per token, native vision, weights promised under a Modified MIT license by July 27. Its published results — 93.5% on GPQA Diamond and 91.2% on BrowseComp, both the strongest open-weight scores on record — trail only the models the US government has gated behind vetted-partner programs. The certified top of the frontier is restricted; the open tier now presses against everything below it. Production behavior confirms the shift — open-weight models carried 29% of tokens routed through enterprise gateways in June, up from roughly one-ninth in April, while absorbing under 4% of spend.

The third event frames the other two. On July 14, Google DeepMind's Demis Hassabis proposed an independent, FINRA-style standards body for frontier AI — industry-funded, voluntary submission up to 30 days before release, designed as an on-ramp to a mandatory regime, operational before year-end. The following day the EU's AI Office published its frontier-AI expert findings and the Commission opened a call to build third-party model evaluation capacity by 2027. Assurance is institutionalizing at the exact moment the artifact to be assured stops being the lab's finished model and starts being your fine-tune of it.

Capability is commoditizing from below while the certified top stays gated. What remains scarce is the loop that turns a base model into your model.

02 · Framework

The Unbundled Frontier Stack

For three years the enterprise question was which model to buy. This week made that question obsolete in its old form, because the thing being sold split into three separable layers. The Unbundled Frontier Stack names them, and the metric that spans all three is the portable-behavior ratio: the share of your production model behavior encoded in artifacts you own and can move across substrates — adapters, evaluation suites, training corpora, reward specifications — rather than locked inside a vendor's product.

Layer 1 — Base weights

Pretrained capability, now approaching commodity status at frontier scale. Inkling and K3 land within one week of each other, both open, both at 1M context, both explicitly designed to be modified. The discipline here is hygiene, not selection: a license register (Modified MIT is not Apache), provenance tracking, and a substitution path.

Layer 2 — The post-training loop

The layer that finishes the model: your corpus, your adapters, your reward specification, your evaluation harness. This is where differentiation now lives — Inkling's entire commercial design assumes it — and it is the layer no vendor can ship, because its raw material is your operational data and your decision standards. Every adapter cycle deepens an asset no competitor can buy. The loop, not the model, is the thing worth owning.

Layer 3 — Assurance

Who certifies the finished artifact. The Hassabis proposal, the EU's evaluation-capacity call and China's standardized safety benchmark all target the lab's model — but the artifact in production is your fine-tune, which no external body has seen. Under EU AI Act Article 25, a downstream modifier can inherit provider obligations outright, with substantial modification presumed above one-third of the systemic-risk compute threshold. Customize the behavior and you own its assurance. The evaluation harness stops being optional tooling and becomes your regulatory posture.

So what: the frontier stopped shipping finished goods. Base weights are becoming a commodity; the loop that turns them into your model — your corpus, your evals, your adapters — is where value and liability now concentrate. Own that loop, or rent your own behavior back from a vendor.

03 · Use Cases

Three LATAM operators, three tuning postures

The patterns below are composites drawn from Socradata engagement work in the region. They share one design move: the fine-tune is treated as a governed, versioned, portable asset — not a notebook experiment.

01

CABA Tier-1 bank — a credit-language layer the vendor never sees. A Buenos Aires universal bank fine-tunes an Inkling-class open base with LoRA adapters on its Río de la Plata Spanish credit-explanation corpus, running in-jurisdiction under Ley 25.326 with Tier-2 advisory human-in-the-loop per EU AI Act Article 14. Post-training compute is deliberately held below the Article 25 substantial-modification presumption and documented per release. The behavior is the bank's: portable-behavior ratio reaches 92%, cost per decision falls 47% against the frontier-API baseline, override rate holds at 5.4% — and in a failover drill, the full behavior re-based onto a second substrate in 19 days.

02

São Paulo industrial logistics — adapters on a release train. A Brazilian operator runs dock scheduling and customs documentation on a K3-class agentic base, retrained monthly on its own operational corpus. No adapter ships without 100% evaluation regression coverage and a pass^k gate of ≥0.9 at k=5; every version is signed into an adapter registry with model provenance carried under LGPD Article 20. Forecast error falls 24%, on-time-in-full improves 7.9pp, and the monthly cycle — corpus refresh, tune, evaluate, sign, release — runs in 11 days end to end.

03

Multi-country grain exporter — an assurance board before the regulator asks for one. A fourteen-port exporter across Argentina, Paraguay and Uruguay mirrors the FINRA-style design internally: no fine-tune reaches production without review by an internal standards board that tests against the strictest regime it touches. Regulated customs and sanctions screening runs on a sovereign fine-tune anchored at Latam-GPT/CENIA. Sovereign coverage on regulated workloads reaches 44%; identity-attested action holds at 100%; substrate concentration stays under 60%; portfolio inference cost falls 39%.

04 · Implementation

Implementation: owning the loop

The gap is the opportunity. Across enterprise production stacks, retrieval-augmented generation runs at 51% adoption while fine-tuning sits near 9% — even as 41% of enterprises tell a16z they will expand open-model use, and another 41% would switch the moment performance matches. Most organisations rent behavior through prompts and retrieval on someone else's model — differentiation parked in the least defensible layer of the stack.

This is also where POC theater hides in 2026. A fine-tune in a notebook is a pilot. Production is a governed loop: a versioned corpus, a signed adapter registry, an evaluation harness that runs before every release, a rollback path, and a named owner. The distinction is not technical sophistication — it is whether the behavior you shipped last month can be reproduced, audited and moved.

So what: when every enterprise ships its own fine-tune, the lab's safety card no longer covers the deployed artifact. Assurance follows customization downstream — the evaluation harness is now a deployer asset, not a vendor courtesy.

Governance

Map EU AI Act Article 25 exposure per workload before the first tune: document post-training compute, purpose and capability deltas so you know when you become the provider. Keep a license register: Modified MIT, Apache and research licenses carry different rights. Stand up an internal assurance board mirroring the external pattern: pre-release review of every adapter against the strictest regime you touch, with Ley 25.326, LGPD Article 20 and Article 14 human-in-the-loop checkpoints wired into the release gate, and a named accountable owner per fine-tune.

KPIs

Portable-behavior ratio at 90% or above. Tune-to-production latency at 30 days or less per adapter cycle. Evaluation regression coverage at 100% before every release. Adapter provenance — signed versions — at 100%. Cost-per-decision delta of at least 35% versus the frontier-API baseline. Substrate concentration under 60%. Sovereign coverage above 30% on regulated flows. Override rate under 8%; decision auditability at 100%.

90D 180D 360D

12-month roadmap

0–90: build the tuning corpus and the evaluation baseline, stand up the license register, map Article 25 exposure per workload, and name a model owner per planned fine-tune. 90–180: ship the first LoRA adapter on a non-regulated workload through a gateway, launch the signed adapter registry, and convene the internal assurance board. 180–360: move one regulated workload onto a sovereign-substrate fine-tune, reach third-party evaluation readiness ahead of the EU's 2027 capacity, and report portable-behavior ratio and tune-to-production latency to the board quarterly.

Socradata Perspective

The weights are free. The behavior is yours to build.

Two weeks ago the operating question was which frontier models an enterprise would be permitted to run; last week, who answers once it runs them. This week resolves the strategy underneath both: the model itself is no longer where advantage lives. When the leading US open-weights release ships with a disclaimer that it is deliberately unfinished, the lab is telling you — in its pricing, not its marketing — that the valuable asset is the loop that finishes it. That loop is made of things only you have: your corpus, your decision standards, your evaluation discipline.

For Latin American operators this is the most favorable reframing of the frontier in three years. The region will not win a pretraining race, and it sits outside the vetted-partner perimeter of the gated frontier. But open weights do not ask Washington for clearance, and post-training is a discipline, not a capital program. Latam-GPT is a post-training play; the substrate to run it on is being poured now, from Rio AI City's planned 1.5 GW to Argentina's 500 MW nuclear-baseload positioning. A CABA bank that owns a governed tuning loop holds frontier-class behavior, in Spanish, in jurisdiction, on hardware it can name. KPIs before APIs — and from now on, evals before adapters. From pilot to policy, the loop is the asset.

Measure your portable-behavior ratio

Socradata runs Post-Training Readiness Diagnostics for LATAM operators in financial services, logistics, agribusiness and the public sector. The output is a workload-by-workload map of tuning opportunity and Article 25 exposure, a corpus and evaluation baseline, a license register, and a board-ready scorecard anchored on portable-behavior ratio.

Request an Operational Diagnostic