Blackstone& field note · article two · for the people who run governance

Who holds the bar?

Governing an agentic build does not ask you to give up control; it moves the control from the plan to the gate.

This note is for delivery managers, PMO, finance, investment and assurance: the people asked to report a probabilistic reality in a deterministic vocabulary. GDS is the worked example; the mapping holds for any stage-gated enterprise.

The bind · your instruments on one side, the work on the other, you in the middle
What your instruments assume
plan status burndown
fixed scope · phase gates · percent complete · pass or fail
you
the delivery manager
asked to say which
picture is true
What the work actually is
the bar · 92
a measured band, not a fixed state · iterations to the bar unknown up front
Status · Green · on plan Status · Amber · at go-live
go-live

Report the uncertainty truthfully and it reads as failure. Smooth it over and you have manufactured this: the green that collapses the day it meets production. That is a system failure, not a personal one, and the fix is a better vocabulary.

The playbook you were trained on

Fixed scope, phase gates, a red-amber-green status, a burndown, and a budget approved once against a business case. For deterministic work, every one of those is the rational choice.

Where agentic work breaks it

"Done" is a measured band, not a fixed state. The number of experiments to reach the bar is unknown before you start, so a date and a cost are guesses. Requirements move when the model moves, not only when the customer changes their mind.

Your craft, upgraded

Running gates, articulating progress, managing expectations, holding a room's confidence through uncertainty: these are the skills an agentic delivery needs most. The instrument set changes; the profession does not.

This note is written in two voices. › The dark panels are the engineer speaking. The white cards are the same fact as the room hears it. The delivery manager is the person who carries a fact from one voice to the other without losing the truth on the way.

Move the control from the plan to the gate.

Same lifecycle

Every phase your governance knows survives. Discovery to Live maps onto Frame to Operate: each phase is re-pointed, none is skipped.

Same gates, better fed

The assessments and business-case stages you already run keep their authority. They stop receiving narratives and demos, and start receiving eval scores, drift records and cost per outcome.

More decision points

One up-front budget becomes a chain of capped tranches, each behind a gate that can stop. The downside between any two gates is bounded.

Assurance, funding discipline and clear reporting are exactly what separate the surviving agentic projects from the more than 40 percent Gartner expects to be cancelled by the end of 2027. Governance is not the enemy. The vocabulary is.

01

The phase mapping

GDS as the worked example · generalises to any stage-gated enterprise

The probabilistic method does not skip the lifecycle. It keeps every phase your governance already knows and re-points what each one produces. GDS already gates the major boundaries with a service assessment; the eval gate does not compete with that panel. It feeds it.

GDS SERVICE LIFECYCLE Discovery Alpha Private beta Public beta Live no discovery assessment alpha assessment beta assessment live assessment Frame Prove Build Harden Operate or: the finding that no agentic build is warranted · the kill option starts at Frame the probabilistic path · article one’s five gates

◇ the service assessment keeps its place and its authority. Diamonds sit where GDS puts them: at Alpha, before public beta, and at Live. There is no Discovery assessment; nothing external gates the exit from Discovery, which is why the kill finding at Frame has to be the team’s own output.

DiscoveryFrame

The eval bar defined: a measurable definition of "good", or the finding that no agentic build is warranted. The kill option starts here.

AlphaProve

Golden datasets, a first end-to-end slice, the first eval gate passing on real data.

Private betaBuild

Iterate to threshold under flow, card by card, with limited live traffic behind guardrails.

Public betaHarden

Integration, human-in-the-loop checkpoints, rollback, cost per outcome inside budget.

LiveOperate

Hold the eval band, monitor for drift, re-verify evals, transfer the capability.

The assessment, re-fed · the node does not change
WHAT THE PANEL RECEIVED a narrative a demo WHAT IT RECEIVES NOW eval 88 / bar 92 · converging drift record · versioned runs cost per outcome · in envelope service assessment proceed · or not same size · same seat · same authority

better inputs, same authority. point-in-time and narrative-led becomes point-in-time and instrumented.

The Service Standard checks far more than output quality: user needs, team health, accessibility, security, technology choices. The eval gate is the agentic-strength version of the quality-and-metrics strand, and it feeds the panel rather than replacing it.

Not on GDS? Read "service assessment" as your stage-gate review or gateway. Green Book SOBC / OBC / FBC is already tranche-gated approval; gateway reviews (IPA, now NISTA) are advisory assurance, while approval control sits with finance. The mapping holds.

02

The funding upgrade

the centre of gravity · your machinery already stages; rewire what releases it

Most organisations that changed how they build never changed how they pay. Delivery went agile; the money stayed on a waterfall. But the staged machinery already exists: SOBC to OBC to FBC is tranche-gated approval, and "no Alpha pass, no Beta money" is already the GDS instinct. The gates are not missing. They release on documents and dates: a business case updated, a milestone reached, a financial year survived.

For deterministic work that wiring is survivable. For agentic work it is fatal, because you cannot know the scope or cost of research up front. A budget fixed against a pretended scope resolves one of two ways: the overspend keeps climbing, or the project is cancelled outright. Both feed Gartner's cancellation causes at once: escalating cost, unclear value, weak risk controls.

How the money moves today
One block, then variance
the plan · time runs left to right ↑ spend over plan one budget approved up front actual spend the variance report looks bad overspend keeps climbing cancel · the line just ends
decision points · 1 · at approval, then variance reporting
Eval-gated tranche funding
Capital follows evidence
the tranche = max loss between gates T1 T2 T3 T4 T5 any gate can stop · exposure capped at the released tranches dataset · platform · eval harness survive any stop · reusable next time
decision points · 5 · one at every gate · more control, shown

not exotic finance. cooper’s stage-gate R&D model, pharmaceutical portfolio management and venture tranching are the same discipline. the claim is not "adopt a new funding model." it is "your funding model already stages; make the gates release on measured quality and value."

"More control, not less" is earned, not asserted. More decision points only mean more control if the gates have teeth. Three conditions, stated up front:

1 · Pre-agreed kill criteria

Written at Frame, before anyone is invested: the eval floor, the cost ceiling, the escalation-rate limit beyond which the business case fails. A gate with no pre-agreed failure condition is a ceremony.

2 · A forum that can actually stop

Gate decisions taken by a small standing forum on evidence already produced: the eval run, the cost line, the drift record. Not a board paper commissioned per gate. The documented failure mode of gateway regimes is continuation bias, reviews that never stop anything; naming it is how you avoid it.

3 · Bounded downside between gates

Guardrails and budget caps mean the maximum loss between any two gates is the tranche, not the programme.

When a gate fails

The whole point of a kill point is that it sometimes kills. The spend to date was capped at the released tranches. The golden dataset, the platform components and the eval harness survive the cancellation and are reusable by the next use case, which materially softens the write-off. The options at a failed gate are stop, hold, or pivot, and each is a legitimate outcome of the machinery running as designed.

Friction 1 · Annuality

Public-sector budgets are voted annually and in-year controls are rigid. Tranches do not fight that; they live inside the annual envelope, with gates governing release within it. Year-end carry-over risk is real and belongs on the risk register, not in the small print.

Friction 2 · Whose money it is

Where the build runs under a customer contract, tranche release is a contract-shape question. Two shapes work: milestone payments tied to gate outcomes, or time-and-materials with continued authorisation gated on the evidence pack attached to each period's delivery report and invoice. Both give the paying customer a real decision point on evidence at every gate.

The business case stops being a one-time hurdle and becomes a living document. The plan is stable: the pipeline, the stages, the criteria. The estimates inside it update as evals reveal the real economics.

03

The translation layer

the delivery manager’s tool · dark panel: the engineer says it · white card: the room hears it

A two-way dictionary: engineer reality on the left, the same fact rendered in governance language on the right. Two rules govern it. Every technical fact has a governance meaning, and the translation must neither soften it into false comfort nor dump raw uncertainty on the room. And every rendering lands as a decision: release, hold, contain, escalate, re-baseline, or stop. A report a governance function cannot act on is noise. The bracketed values in the cards below are slots: fill each with the real figure from the incident in front of you.

One row matters more than all the others. A vocabulary that can only carry good news is the fake-green machine rebuilt with better words. This one carries a stop recommendation as the proof the vocabulary works.

Eval score is 88 against a 92 bar, converging. Forecast is two more iterations.
Holdtranche continues · gate decision pending

Quality gate at 88 of a required 92 and converging. The team forecasts two iterations to pass; if the bar is not cleared within this tranche, the decision returns to this forum at the gate.

We do not know how many iterations it will take to hit the bar.
Releasebounded

This stage is research, so it is funded a tranche at a time against the gate, not committed to a fixed date. The maximum exposure this period is the tranche.

Eval has plateaued at 81 against 92 across two tranches. We do not see a path to the bar at acceptable cost.
Stopthe machinery working, not a scandal

The gate has not passed and the evidence does not support a path to the standard at acceptable cost. Recommendation: stop. Total exposure was capped at the released tranches; the golden dataset, platform components and eval harness are retained and reusable by the next use case.

Drift alert fired on live traffic at 09:40. We are re-grounding the retrieval layer and adjusting routing.
Containthen hold until back in band

Live quality fell below the band and the monitoring control caught it at 09:40. Affected traffic has routed to human handling since detection, so exposure is bounded to roughly [X] cases over [Y] hours. Remediation is underway; return above the bar expected by [date]. This is the control doing its job, and this is what it caught.

Token spend is running three times forecast because turns per case have doubled.
Escalateenvelope breached · levers named

The cost per outcome has breached its envelope. Mitigation levers are model routing and caching, expected to recover [X] percent; if the revised curve still breaches at the gate, the business case comes back to this forum.

The agent attempted an out-of-policy action and the guardrail blocked it.
Containcontrol evidenced

A guardrail fired and prevented an out-of-policy action, with a full log of the attempt. For assurance purposes this is direct evidence the control layer works under real conditions.

The demo works, but production accuracy is 84, below the bar.
Holdgap stated · plan attached

Not yet at the standard to move to public beta. Here is the gap, and here is the plan and cost to close it.

We cannot run evals against real data until the integration environment exists.
Escalatea dependency, not the delivery

A dependency is blocking the evidence the next gate needs. Every week it slips moves the gate decision by a week. Escalating for a date.

Alpha evidence changed the approach; the original intent was too broad.
Re-baselinebenefit impact stated

Alpha did its job: the evidence changed the plan. Scope re-baselined at the Alpha gate. Here is what came out of scope, and what that does to the benefits case.

Read the right-hand column as a set. It carries a stop recommendation, a cost breach, a live incident and a blocked dependency without flinching, and it still sounds like a team in control. That is the trick, and it is not a trick: the control is real, so nothing needs softening.

04

Who owns the bar?

the independence clause · what makes the whole argument assurance-proof

A governance reader will, and should, ask the assurance question: if the delivery team builds the eval, runs the eval and reports the eval, the gate is marking its own homework. A control must be independent of what it controls. Three assignments fix it.

the business

owns the golden dataset

Domain experts on the business side curate and sign the examples that define "good". The delivery team can propose additions; it cannot quietly change the exam.

governance

owns the bar

The threshold, and the pre-agreed kill criteria around it, move only with sign-off at the forum. A bar that drifts to meet the score is Goodhart's law in action, and it is the eval-world version of moving a milestone.

the record

makes every claim checkable

Every eval run is versioned: dataset version, model version, score, date. Durable evidence that survives an internal audit or an NAO visit, which is more than a RAG history can say.

With those in place, the eval band becomes a status far harder to fake than a RAG colour, precisely because the person reporting it does not control the exam, the bar, or the record.

05

Reporting a probabilistic world to a deterministic board

the status report, done right · set the probabilistic frame on day one

If the room expects a deterministic, 100 percent-consistent result, trust dies the first time reality varies. Say early: quality is a measured band, visible uncertainty is a property of the method, and the controls are built for it. Then report like this.

delivery report · case-handling agentweek 14 · build phase
Quality · the eval band
88 bar 92 0 100

88 of 92 and converging. Forecast: two iterations to the bar; if not cleared within the tranche, the decision returns to the forum at the gate.

replaces green · the gap to the bar is visible, and visible is the point

Milestone · the gate

Prove gate passed 14 June on real data. Build gate next; beta assessment follows it.

progress = which gate has passed and which is next · not percent-of-plan

Budget · the tranche

Tranche 3 of 5 released and capped. Spend reports against the released tranche.

the next tranche is a decision, not an entitlement

In business terms

Resolves nine in ten cases inside cost; the tenth goes to a person, and the rate is tracked weekly.

The eval band is the real status

"88 of 92, converging" is a status you can defend in front of anyone, because the dataset, the bar and the run log are all owned outside the team reporting it.

The gate is the milestone

Progress is which gate has passed and which is next, not percent-of-plan.

The tranche is the budget line

Spend reports against the released tranche, and releasing the next one is a fresh decision each time.

Visible uncertainty beats false precision

An early amber is worth more than a green that collapses at go-live. The cure for the fake-green status is a vocabulary in which amber is a normal, actionable state, not a career event.

The bridge, zoomed out.

build eval log the engineers · probabilistic reality, unpretended DM the delivery manager · speaks both languages forum gate audit GOVERNANCE · PLANS, GATES, ASSURANCE

A delivery manager who speaks both languages becomes the trusted bridge in the room. The engineers get a delivery wrapper that does not force them to pretend the work is deterministic. The governance function keeps its authority, better fed: the same gates, releasing on evidence instead of documents; the same assessments, receiving instruments instead of narratives; more decision points with the downside bounded in between. And the business gets what it actually wanted all along, which was never a fixed date. It was confidence that the money follows evidence, and that someone can tell them, in their own language, exactly where things stand, including when the answer is stop.

Next in this series: the literacy floor. The shared language only works if the people being reported to understand the concepts underneath it: an eval bar, a golden dataset, a tranche gate. Building that floor across a governance community is its own piece of work.