# Cost-per-Verified-Outcome Index: methodology 1.0

Published protocol: 2026-10-08. No measured index or model ranking has been released.
This protocol measures native worker benchmark outcomes, separately from customer
prices and TeamShift's internal per-outcome billing.

## Cohort and verification

Before running anything, freeze a versioned manifest naming ten outcomes, six
distinct stacks, identical task instances for every stack, immutable source and
grader hashes, model/provider IDs, settings, tools, time/step limits, retry policy,
existing aggregate budget and evaluator-owned held-out fixture hashes. Select
outcomes from the actual SMB-Bench templates; do not present mock review handoffs
as completed purchases or warranty actions. Publish exclusions and sample sizes.

An outcome is successful only when its admitted end-state grader passes and its
required evidence is independently verified. Self-reported completion, reference
solutions and instruction installation are not worker successes. Count each
predeclared task instance once; attach all attempts, failures, cancellations and
unknown outcomes to it. Do not replace difficult cases or stop a stack early
because another stack performed better. Unknown outcomes prevent a qualified
comparison; they are reported separately rather than silently discarded.

## Spend, rates and latency

For each outcome and stack, report attempts, task instances, verified successes,
failures, unknowns, settled total spend, currency, success rate and latency.

**Cost per verified outcome = total settled spend / verified successes.**

The numerator includes failed attempts and every retry, model/tool charge,
benchmark compute charge and paid human verification within the frozen boundary.
Disclose allocated shared costs and exclusions. Avoid double-counting provider
usage already represented in a settled ledger. Prices or token estimates do not
replace actual charges. Unknown or unsettled spend prevents a qualified metric.
Do not infer currency or combine currencies without a dated, sourced conversion
rule frozen before the comparison. Retain exact monetary precision until display.

Zero successes gives an undefined cost per verified outcome, never zero. Report
the spend and failed cohort without a numeric cost or a favorable rank. Success
rate is verified successes / all predeclared task instances, not successes / only
the attempts with usable traces. Latency covers initial dispatch through verified
terminal state, including retries and verification; publish median and nearest-rank
p95 for successful cases, plus separate time-to-terminal data for failures and
timeouts. Disclose unresolved/censored cases; do not label those as fast failures.

## 95% uncertainty intervals

For independent task-instance binary outcomes, use the two-sided Wilson 95%
interval (z = 1.959963984540054) for success rate, as described by
[NIST](https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm).
When instances share fixtures or repeated seeds, independence is not assumed:
identify the independent fixture clusters and use the clustered resampling below.

For cost per verified outcome, use a predeclared paired cluster bootstrap: sample
independent fixture clusters with replacement within each outcome, preserve every
attempt/cost/grade in each selected cluster, and use the same sampled clusters
across all six stacks. Recompute total spend / verified successes for each of
10,000 replicates. For clustered success-rate intervals, also recompute verified
successes / all selected task instances in each replicate, including duplicate
clusters with their full multiplicity; a zero-success replicate has rate zero.
Publish the generator/version, seed, cluster counts and exact
nearest-rank 2.5th and 97.5th percentiles. This is our protocol choice using the
[NIST resampling principle](https://itl.nist.gov/div898/handbook/eda/section3/bootplot.htm),
not a claim that NIST prescribes this benchmark design.

Keep zero-success replicates as positive infinity when calculating ratio bounds;
never drop them. Publish an infinite bound as an explicit unbounded value, not a
finite replacement. If there are fewer than two independent clusters, no valid
resampling interval is claimed. Small cohorts, sparse successes and degenerate
resamples remain conspicuous limitations; more replicates cannot create more
independent evidence. Intervals describe sampling uncertainty under this design,
not operational guarantees. Do not declare a universal winner or a causal effect
from overlapping intervals or a synthetic benchmark.
These 95% intervals apply to cost and success rate. Median and p95 latency are
descriptive measurements with the stated successful-case denominator, not latency
confidence intervals.

## Monthly publication

A qualified monthly release requires all 60 outcome/stack cells, complete settled
spend and verification evidence, the frozen cohort and uncertainty protocol, actual
human review of the exact report, and current publication authority. Confirmed
pricing-history rows may explain dated rates; they do not establish actual spend.
Neither JSON reviewer metadata nor a content hash proves a human approved it.

Retain private run/ledger/grader evidence securely. Public tables contain aggregate
metrics, cohort/source hashes, limitations and methodology version; omit customer
data, credentials, raw traces, private fixture contents and tenant/run/task IDs.
Release a versioned report, machine-readable aggregate data and their hashes under
one DOI only after those gates pass. Preserve prior versions, explain corrections,
and disclose cohort or stack changes before presenting month-to-month comparisons.
This protocol grants no paid runs, new access, approval or publication of results.

Current status: the public SMB-Bench scaffold exists; an admitted ten-outcome,
six-stack measured cohort, complete costs and monthly human sign-off are absent.
There is no monthly index value, release DOI or feasibility claim to substitute.
