TeamShift

Benchmark methodology 1.0

Measure the cost of a verified outcome.

A low token price does not tell you what a completed task costs. This protocol compares total settled spend with independently verified outcomes, including the cost of failures and retries. No measured index or model ranking has been released.

Guide

Use this before you automate the work.

A low token price does not tell you what a completed task costs. This protocol compares total settled spend with independently verified outcomes, including the cost of failures and retries. No measured index or model ranking has been released.

Step 1

Use the whole cost and the whole cohort

Cost per verified outcome equals total settled spend divided by verified successes. Include failed attempts, retries and admitted model, tool, compute and verification charges. Unknown spend or outcomes prevent a qualified comparison.

  • Freeze ten outcomes and six distinct stacks before running the same task instances.
  • Use actual end-state grades and independent evidence rather than self-reported completion.
  • Zero successes gives an undefined cost per outcome, never zero.

Step 2

Report uncertainty and time

Report success rate, median and p95 verified latency, sample sizes and 95% uncertainty intervals for cost and success rate. Resample independent fixture clusters with the same samples across stacks; keep zero-success resamples as unbounded cost.

  • Separate failure and timeout timing from successful-task latency.
  • Do not treat repeated seeds as independent evidence.
  • Disclose small cohorts, missing evidence and differences in conditions.

Step 3

Publish only after a real monthly review

A monthly index needs all 60 outcome/stack cells, complete settled costs, verified grades, actual human sign-off on the exact report and publication authority. The versioned report and aggregate data then receive a DOI.

  • The benchmark scaffold is available; an admitted measured matrix and monthly sign-off are still missing.
  • This page publishes the protocol, not a populated index, rankings or a release DOI.
  • Keep customer data, credentials, raw traces and private fixtures out of public reports.

Questions

Before you hand this off

Is this TeamShift pricing?

No. The index is a benchmark measurement of actual spend per verified success. Customer prices and internal per-outcome billing are separate.

Can I compare model prices to get the index?

No. Confirmed pricing history can explain dated rates, but estimates cannot replace complete settled charges, verified outcomes or the actual benchmark cohort.