Trials
Shadowing a candidate model against your incumbent on real traffic -- setup, targeting, sampling, stop conditions, lifecycle and promotion.
Trials
A trial is a shadow-arm comparison of candidate models against an incumbent, on your team's own live traffic. One or more candidates answer the same inputs the model in production actually served, and each pair is judged to say whether a candidate could replace it.
It is a first-class, long-lived object with a lifecycle -- not a request-scoped operation. It is owned by a team, it records who started it, and its verdict is a decision record that outlives the run.
A trial never mutates routing. Promoting a winning candidate is always a separate, explicit human action. Nothing in the platform calls promotion on a verdict or on a stop condition.
The arms
| Arm | Dispatched? | Who sees it |
|---|---|---|
| Incumbent | No. Its response is what your users were already served. | Your users. |
| Candidate | Yes, in the background, off the request path. | Nobody. It is judged and discarded. |
A slate -- more than one candidate -- is judged against the same served response, so it costs linearly more and dilutes nothing. A slate caps at 8 candidates.
Setting one up
The wizard is four questions with the price beside them the whole way. It is deliberately not a gated sequence: there is no next or back, no step is disabled by an earlier one, and nothing hides a field.
Above the four questions is one choice:
| Grouping | What it means |
|---|---|
| A single call type | One model swap, judged per request. Zero exposure -- production keeps serving the incumbent. |
| A whole workflow | A multi-service pipeline, judged one step at a time. Steps whose output your own code reshapes stay on the incumbent, and a mistake that only causes harm further down the pipeline is invisible in this mode. |
On which traffic?
A trial names a scope: the routing profile whose traffic it observes. That is the same object a promotion later writes to, so the thing being measured and the thing being changed are one thing.
Inside the scope you narrow further with targeting conditions, composed with AND in order. An empty set means every request in the scope.
| Field | Condition |
|---|---|
| Workload | is / is not |
| Namespace | is / is not |
| Endpoint | is / is not |
| Requested model | is / is not |
| Issued key | is / is not |
| Invoked tool | is / is not |
| Conversation depth | is / is not |
| Billing regime | is / is not |
That grammar is the whole grammar. Anything richer is refused, not ignored -- and it is a single implementation shared by the setup funnel that estimates your reach and the engine that decides per request, so the preview and the run cannot disagree.
Then a session sample rate, 1 to 100, applied after the conditions. It is deterministic rather than random: two comparisons at the same rate see the same sessions, and re-estimating a retained window cannot move the answer.
Sending zero means "not set" and resolves to 100 -- a trial observing 0% of traffic is a misconfiguration, not a rate.
The setup funnel's pre-targeting figure is a ceiling, not a forecast. It tells you how much traffic is in scope before conditions narrow it.
Which model are you running today?
The incumbent, in the author/slug vocabulary. Its answers are what your
users receive, and candidates are judged against them.
Which models do you want to try?
At least one candidate, up to eight.
When should it stop?
A trial ends on whichever bound binds first. A bound left at zero is not configured.
| Bound | Behaviour on reaching it |
|---|---|
| Spend cap | The trial auto-pauses and its verdict is marked provisional. |
| Max sessions | The trial completes, counted on the same judged-session denominator the verdict uses. |
| Max duration | The trial completes at the next arm poll, so it disarms within one poll interval. |
Every stop condition completes or pauses the trial. None of them promotes.
The cost estimate is deliberately the last thing you read before starting, and the primary button carries the figure. You can save a draft first; reopening it edits the draft rather than duplicating it.
What it will cost
The estimate projects what the configuration would add: shadow inference plus judging. Shadow inference is real money spent on answers nobody reads, which is why the number belongs on the commitment control rather than in a bill afterwards.
| Figure | Meaning |
|---|---|
| Shadow spend per day | The candidate arms. |
| Judge spend per day | One comparative judge call per judged pair. |
| Expected days, expected total | Projected only when a bound actually binds. With no bound configured you get per-day figures and no fabricated horizon. |
| Unpriced models | Surfaced by name rather than priced at $0, which would read as "free". |
| Served spend unchanged | Always true, and stated on the wire on purpose: a shadow trial only ever adds spend. |
The estimate is billing-grade where your contracted rates resolved end to end, and the response says so per model -- including naming the models it had to fall back to list prices for. See AI spend.
What it has cost so far
The running figure is accrued spend: shadow inference plus judging, summed as each judged pair lands.
Accrued spend is list-priced, not contracted. For a team holding a negotiated discount it is therefore systematically high -- in the direction that flatters the vendor. Read it as evidence-grade, never as a charge. It is what the spend cap is enforced against.
The spend cap is absent until you configure one, and a trial with no cap is spending unattended on your own account. When accrued spend reaches the cap the trial auto-pauses and the verdict reads provisional everywhere it appears. Resuming past a cap pause requires raising the cap in the same action -- the remedy is part of the act, not a separate step.
Lifecycle
| State | Meaning |
|---|---|
draft | Created and editable, not collecting. |
running | Collecting and judging. |
paused | Collecting nothing. Reached manually, or automatically at the spend cap. |
completed | Ended. Everything collected is kept. |
The machine is draft to running to completed, with paused reachable from running and returning to running on resume. Inadmissible transitions are refused, never coerced.
A pause records its reason: user or spend_cap. Only the second marks the
verdict provisional.
Progress while it runs
| Figure | Meaning |
|---|---|
| Observed sessions | Distinct judged sessions. This is the denominator the verdict uses. |
| Paired turns | Distinct judged turns. |
| Judged pairs | One per turn per candidate. |
| Shadow spend, judge spend | The two halves of accrued spend. |
| Dispatched, completed, failed | The engine's own funnel. |
| Dropped: breaker, busy, non-dispatchable, no text | Four distinct drop causes, kept apart because they need different fixes. |
| Shadow failure rate, shadow drop rate | Derived from the above. |
| Served latency | Your 95th-percentile serving latency during the trial against an equal-length window before it -- both capped at three days, so a long trial compares its most recent three days. |
If no engine has ever reported counters, the interface says so rather than rendering zeros. A zero that means "we have no measurement" and a zero that means "nothing was dispatched" are different claims, and the product will not conflate them.
The served-latency comparison exists to support one specific claim: that observing the trial did not move your serving latency. It is an observational delta over the same targeting predicate, not an A/B test.
Replay: running a trial over a past window
A replay trial is the same object over traffic that already happened. Both live and replay are trials -- in each, the incumbent's answer already reached a user, so only the candidate is dispatched and the verdict carries full weight.
- You give a window, and optionally a list of sessions -- which is how you replay the conversations that actually went wrong. A session list is always combined with the window, never used instead of it.
- What you ask for is stored verbatim and not clamped to your retention bound, because that bound moves daily. It is narrowed at read time instead, and the estimate reports the reachable bound, the effective window, and a reason for every zero.
- Traffic from before retention was enabled was never stored. Without stating that, an empty date range is indistinguishable from a bug.
Replay requires prompt retention to be on. A live trial works either way.
A replay trial and the live observer are mutually exclusive for one team's traffic: scoring one window under both would inflate the arm populations, which is a silent corruption of the denominator rather than a visible failure.
Workflow trials
A workflow trial reconstructs a multi-service pipeline from correlation data the gateway collects -- never from a description you supply. Each step is classified swappable or locked from one observable: does the next call's prompt carry this step's answer verbatim?
A locked step's output is transformed by your own code, which DevZero does not run, so it stays on the incumbent and nothing is dispatched for it. The lock reason is recorded and named: the caller transformed it, the handoff was inconsistent, the handoff was unreadable, the step's own answers are too short to match on, or there was no evidence either way.
Promotion
Promotion writes the candidate as the scope profile's default model, through the ordinary routing-profile machinery -- so propagation and session pinning apply unchanged.
| Rule | Detail |
|---|---|
| It is the only path by which a trial affects production. | Nothing else can reach it. |
| It is always an explicit human action. | No verdict and no stop condition triggers it. |
| It is refused below the parity floor. | A degraded candidate is never adoptable, whatever its cost advantage. |
| It is refused from an experiment's verdict. | A corpus comparison never touched production. |
| It leaves a receipt on the trial. | Which candidate, which actor, when. |
You may opt in to arming a watch at the moment you promote. If you do, a monthly budget is required and must be positive -- a zero budget is a refusal, not "unlimited". See Watches.
Arming the watch is the one partial outcome in this flow: if the routing change succeeds and the watch fails to arm, the routing change is live and is not undone by the returned error. The response says so explicitly.
Reading the result
See Judging and verdicts for the composite score, the five dimensions, the session floor, the tie band, and what "adoptable", "provisional" and "on the frontier" each mean.
Evals Overview
The family of judged model comparisons -- trials, experiments, compare and watches -- and the one rule that separates traffic evidence from corpus evidence.
Judging and Verdicts
The composite score, its five dimensions, the session floor, the tie band, and exactly what each verdict entitles you to do.