Every phase your governance knows survives. Discovery to Live maps onto Frame to Operate: each phase is re-pointed, none is skipped.
Blackstone& field note · article two · for the people who run governance
Who holds the bar?
Governing an agentic build does not ask you to give up control; it moves the control from the plan to the gate.
This note is for delivery managers, PMO, finance, investment and assurance: the people asked to report a probabilistic reality in a deterministic vocabulary. GDS is the worked example; the mapping holds for any stage-gated enterprise.
picture is true
Report the uncertainty truthfully and it reads as failure. Smooth it over and you have manufactured this: the green that collapses the day it meets production. That is a system failure, not a personal one, and the fix is a better vocabulary.
Fixed scope, phase gates, a red-amber-green status, a burndown, and a budget approved once against a business case. For deterministic work, every one of those is the rational choice.
"Done" is a measured band, not a fixed state. The number of experiments to reach the bar is unknown before you start, so a date and a cost are guesses. Requirements move when the model moves, not only when the customer changes their mind.
Running gates, articulating progress, managing expectations, holding a room's confidence through uncertainty: these are the skills an agentic delivery needs most. The instrument set changes; the profession does not.
This note is written in two voices. › The dark panels are the engineer speaking. The white cards are the same fact as the room hears it. The delivery manager is the person who carries a fact from one voice to the other without losing the truth on the way.
Move the control from the plan to the gate.
The assessments and business-case stages you already run keep their authority. They stop receiving narratives and demos, and start receiving eval scores, drift records and cost per outcome.
One up-front budget becomes a chain of capped tranches, each behind a gate that can stop. The downside between any two gates is bounded.
Assurance, funding discipline and clear reporting are exactly what separate the surviving agentic projects from the more than 40 percent Gartner expects to be cancelled by the end of 2027. Governance is not the enemy. The vocabulary is.
The phase mapping
The probabilistic method does not skip the lifecycle. It keeps every phase your governance already knows and re-points what each one produces. GDS already gates the major boundaries with a service assessment; the eval gate does not compete with that panel. It feeds it.
◇ the service assessment keeps its place and its authority. Diamonds sit where GDS puts them: at Alpha, before public beta, and at Live. There is no Discovery assessment; nothing external gates the exit from Discovery, which is why the kill finding at Frame has to be the team’s own output.
The eval bar defined: a measurable definition of "good", or the finding that no agentic build is warranted. The kill option starts here.
Golden datasets, a first end-to-end slice, the first eval gate passing on real data.
Iterate to threshold under flow, card by card, with limited live traffic behind guardrails.
Integration, human-in-the-loop checkpoints, rollback, cost per outcome inside budget.
Hold the eval band, monitor for drift, re-verify evals, transfer the capability.
better inputs, same authority. point-in-time and narrative-led becomes point-in-time and instrumented.
The Service Standard checks far more than output quality: user needs, team health, accessibility, security, technology choices. The eval gate is the agentic-strength version of the quality-and-metrics strand, and it feeds the panel rather than replacing it.
Not on GDS? Read "service assessment" as your stage-gate review or gateway. Green Book SOBC / OBC / FBC is already tranche-gated approval; gateway reviews (IPA, now NISTA) are advisory assurance, while approval control sits with finance. The mapping holds.
The funding upgrade
Most organisations that changed how they build never changed how they pay. Delivery went agile; the money stayed on a waterfall. But the staged machinery already exists: SOBC to OBC to FBC is tranche-gated approval, and "no Alpha pass, no Beta money" is already the GDS instinct. The gates are not missing. They release on documents and dates: a business case updated, a milestone reached, a financial year survived.
For deterministic work that wiring is survivable. For agentic work it is fatal, because you cannot know the scope or cost of research up front. A budget fixed against a pretended scope resolves one of two ways: the overspend keeps climbing, or the project is cancelled outright. Both feed Gartner's cancellation causes at once: escalating cost, unclear value, weak risk controls.
not exotic finance. cooper’s stage-gate R&D model, pharmaceutical portfolio management and venture tranching are the same discipline. the claim is not "adopt a new funding model." it is "your funding model already stages; make the gates release on measured quality and value."
"More control, not less" is earned, not asserted. More decision points only mean more control if the gates have teeth. Three conditions, stated up front:
Written at Frame, before anyone is invested: the eval floor, the cost ceiling, the escalation-rate limit beyond which the business case fails. A gate with no pre-agreed failure condition is a ceremony.
Gate decisions taken by a small standing forum on evidence already produced: the eval run, the cost line, the drift record. Not a board paper commissioned per gate. The documented failure mode of gateway regimes is continuation bias, reviews that never stop anything; naming it is how you avoid it.
Guardrails and budget caps mean the maximum loss between any two gates is the tranche, not the programme.
The whole point of a kill point is that it sometimes kills. The spend to date was capped at the released tranches. The golden dataset, the platform components and the eval harness survive the cancellation and are reusable by the next use case, which materially softens the write-off. The options at a failed gate are stop, hold, or pivot, and each is a legitimate outcome of the machinery running as designed.
Public-sector budgets are voted annually and in-year controls are rigid. Tranches do not fight that; they live inside the annual envelope, with gates governing release within it. Year-end carry-over risk is real and belongs on the risk register, not in the small print.
Where the build runs under a customer contract, tranche release is a contract-shape question. Two shapes work: milestone payments tied to gate outcomes, or time-and-materials with continued authorisation gated on the evidence pack attached to each period's delivery report and invoice. Both give the paying customer a real decision point on evidence at every gate.
The business case stops being a one-time hurdle and becomes a living document. The plan is stable: the pipeline, the stages, the criteria. The estimates inside it update as evals reveal the real economics.
The translation layer
A two-way dictionary: engineer reality on the left, the same fact rendered in governance language on the right. Two rules govern it. Every technical fact has a governance meaning, and the translation must neither soften it into false comfort nor dump raw uncertainty on the room. And every rendering lands as a decision: release, hold, contain, escalate, re-baseline, or stop. A report a governance function cannot act on is noise. The bracketed values in the cards below are slots: fill each with the real figure from the incident in front of you.
One row matters more than all the others. A vocabulary that can only carry good news is the fake-green machine rebuilt with better words. This one carries a stop recommendation as the proof the vocabulary works.
Quality gate at 88 of a required 92 and converging. The team forecasts two iterations to pass; if the bar is not cleared within this tranche, the decision returns to this forum at the gate.
This stage is research, so it is funded a tranche at a time against the gate, not committed to a fixed date. The maximum exposure this period is the tranche.
The gate has not passed and the evidence does not support a path to the standard at acceptable cost. Recommendation: stop. Total exposure was capped at the released tranches; the golden dataset, platform components and eval harness are retained and reusable by the next use case.
Live quality fell below the band and the monitoring control caught it at 09:40. Affected traffic has routed to human handling since detection, so exposure is bounded to roughly [X] cases over [Y] hours. Remediation is underway; return above the bar expected by [date]. This is the control doing its job, and this is what it caught.
The cost per outcome has breached its envelope. Mitigation levers are model routing and caching, expected to recover [X] percent; if the revised curve still breaches at the gate, the business case comes back to this forum.
A guardrail fired and prevented an out-of-policy action, with a full log of the attempt. For assurance purposes this is direct evidence the control layer works under real conditions.
Not yet at the standard to move to public beta. Here is the gap, and here is the plan and cost to close it.
A dependency is blocking the evidence the next gate needs. Every week it slips moves the gate decision by a week. Escalating for a date.
Alpha did its job: the evidence changed the plan. Scope re-baselined at the Alpha gate. Here is what came out of scope, and what that does to the benefits case.
Read the right-hand column as a set. It carries a stop recommendation, a cost breach, a live incident and a blocked dependency without flinching, and it still sounds like a team in control. That is the trick, and it is not a trick: the control is real, so nothing needs softening.
Who owns the bar?
A governance reader will, and should, ask the assurance question: if the delivery team builds the eval, runs the eval and reports the eval, the gate is marking its own homework. A control must be independent of what it controls. Three assignments fix it.
owns the golden dataset
Domain experts on the business side curate and sign the examples that define "good". The delivery team can propose additions; it cannot quietly change the exam.
owns the bar
The threshold, and the pre-agreed kill criteria around it, move only with sign-off at the forum. A bar that drifts to meet the score is Goodhart's law in action, and it is the eval-world version of moving a milestone.
makes every claim checkable
Every eval run is versioned: dataset version, model version, score, date. Durable evidence that survives an internal audit or an NAO visit, which is more than a RAG history can say.
run 042 · ds v1.3 · model 06-02 · 90.1 · 28 Jun
run 043 · ds v1.4 · model 06-02 · 89.4 · 12 Jul
With those in place, the eval band becomes a status far harder to fake than a RAG colour, precisely because the person reporting it does not control the exam, the bar, or the record.
Reporting a probabilistic world to a deterministic board
If the room expects a deterministic, 100 percent-consistent result, trust dies the first time reality varies. Say early: quality is a measured band, visible uncertainty is a property of the method, and the controls are built for it. Then report like this.
88 of 92 and converging. Forecast: two iterations to the bar; if not cleared within the tranche, the decision returns to the forum at the gate.
replaces ● green · the gap to the bar is visible, and visible is the point
Prove gate passed 14 June on real data. Build gate next; beta assessment follows it.
progress = which gate has passed and which is next · not percent-of-plan
Tranche 3 of 5 released and capped. Spend reports against the released tranche.
the next tranche is a decision, not an entitlement
Resolves nine in ten cases inside cost; the tenth goes to a person, and the rate is tracked weekly.
"88 of 92, converging" is a status you can defend in front of anyone, because the dataset, the bar and the run log are all owned outside the team reporting it.
Progress is which gate has passed and which is next, not percent-of-plan.
Spend reports against the released tranche, and releasing the next one is a fresh decision each time.
An early amber is worth more than a green that collapses at go-live. The cure for the fake-green status is a vocabulary in which amber is a normal, actionable state, not a career event.
The bridge, zoomed out.
A delivery manager who speaks both languages becomes the trusted bridge in the room. The engineers get a delivery wrapper that does not force them to pretend the work is deterministic. The governance function keeps its authority, better fed: the same gates, releasing on evidence instead of documents; the same assessments, receiving instruments instead of narratives; more decision points with the downside bounded in between. And the business gets what it actually wanted all along, which was never a fixed date. It was confidence that the money follows evidence, and that someone can tell them, in their own language, exactly where things stand, including when the answer is stop.
Next in this series: the literacy floor. The shared language only works if the people being reported to understand the concepts underneath it: an eval bar, a golden dataset, a tranche gate. Building that floor across a governance community is its own piece of work.