Blog 2026-09-29

Cheaper tokens ≠ cheaper deployments

A model with cheaper tokens can still produce a more expensive deployment. The missing variables are how many attempts you pay for, how often the workflow finishes successfully, and which costs your comparison includes.

Here is a small, fictional Postgres example: the token rate falls by half, but attributed cost per successful deployment rises from $5 to $9. That is arithmetic from a constructed dataset, not a benchmark or a claim about any model provider.

The useful question is not just “What does a million tokens cost?” It is “What did we spend to produce an acceptable outcome?” The FinOps Foundation’s unit-economics framework connects costs to meaningful units of activity. Its Can You Trust an AI FinOps Agent? discussion also prompted this worked example: let SQL calculate and let the agent explain an inspectable result. The example below is independent; no creator endorsement is implied.

Define one outcome before counting costs

In this example, a deployment is a logical request to ship one change. It retains one ID through retries. An attempt is one execution of that workflow. A successful deployment is one request whose final status passes a fixed acceptance gate—not each successful step, model response, or retry.

  • Cohort: deployment requests started on September 22, 2026, UTC. Include every attempt and attributed charge belonging to those requests, even when an attempt or invoice arrives later.
  • Observation: compare only after the same outcome observation window has closed and billing is reconciled. Pending requests make this report provisional.
  • Numerator: model charges plus attributed CI charges for every attempt, including retries and ultimately failed requests.
  • Denominator: unique deployment requests ending in success. A request succeeding on its second attempt still counts once.
  • Scope: USD, with a fixed allocation rule. These are selected variable costs, not total cost of ownership; engineering time, support, shared overhead, taxes and downstream incident costs are excluded.

A production “success” definition might require no rollback during a fixed observation window. Keep that rule identical across cohorts and record its version. A technically completed deployment is not automatically a valuable deployment.

The cheaper rate loses in this fictional example

MetricStandardBudget
Illustrative model charge per 1M tokens$2$1
Deployment requests33
Attempts, including retries46
Successful deployments21
Model charges$8$6
CI charges ($0.50 per attempt)$2$3
Total attributed cost$10$9
Attributed cost per success$5$9

Each attempt uses exactly one million tokens in this intentionally simplified fixture. Real model billing distinguishes input, output, cached tokens and other charges; use reconciled bill amounts rather than pretending one advertised rate covers all usage.

The budget cohort spends 10% less overall but produces fewer accepted outcomes, so cost per success is 80% higher. This does not establish that cheaper tokens caused the failures. Different task difficulty, prompts, tooling, retry policies or acceptance rules could explain a real comparison. Three requests per cohort are nowhere near enough evidence for a model-selection decision.

Build the fixture outside the read-only connection

Run this setup using an ordinary SQL client in a disposable Postgres database. It creates demonstration tables, so it is not a query to run through Bufflehead. For an existing production dataset, adapt the names and definitions instead of running the setup.

-- Disposable PostgreSQL fixture, not a production migration.
CREATE SCHEMA deployment_cost_demo;
CREATE TABLE deployment_cost_demo.deployments (
  id integer PRIMARY KEY,
  cohort text NOT NULL,
  started_at timestamptz NOT NULL,
  outcome text NOT NULL CHECK (outcome IN ('success','failed','pending'))
);
CREATE TABLE deployment_cost_demo.attempts (
  id integer PRIMARY KEY,
  deployment_id integer NOT NULL REFERENCES deployment_cost_demo.deployments,
  billing_complete boolean NOT NULL
);
CREATE TABLE deployment_cost_demo.cost_lines (
  id integer PRIMARY KEY,
  attempt_id integer NOT NULL REFERENCES deployment_cost_demo.attempts,
  category text NOT NULL,
  usd numeric(12,4) NOT NULL CHECK (usd >= 0)
);
INSERT INTO deployment_cost_demo.deployments VALUES
 (1,'standard','2026-09-22 12:00+00','success'),
 (2,'standard','2026-09-22 12:00+00','success'),
 (3,'standard','2026-09-22 12:00+00','failed'),
 (4,'budget','2026-09-22 12:00+00','success'),
 (5,'budget','2026-09-22 12:00+00','failed'),
 (6,'budget','2026-09-22 12:00+00','failed');
INSERT INTO deployment_cost_demo.attempts VALUES
 (1,1,true),(2,2,true),(3,2,true),(4,3,true),
 (5,4,true),(6,5,true),(7,5,true),(8,5,true),(9,6,true),(10,6,true);
-- Fictional fixed one-million-token workload per attempt; not provider pricing.
INSERT INTO deployment_cost_demo.cost_lines
 SELECT id*2, id, 'model', CASE WHEN id <= 4 THEN 2.00 ELSE 1.00 END
 FROM deployment_cost_demo.attempts;
INSERT INTO deployment_cost_demo.cost_lines
 SELECT id*2+1, id, 'ci', 0.50 FROM deployment_cost_demo.attempts;

Each cost line has a unique ID; the real ingestion pipeline must additionally deduplicate upstream billing identifiers and reconcile currency, credits and allocation. The fixture excludes negative credits for simplicity. A real net-cost calculation needs an explicit credit policy rather than dropping refunds.

billing_complete is supplied by the billing pipeline, not guessed by the AI. It must mean all expected charges for that attempt are loaded. One present line is not proof of completeness. A genuine zero-cost attempt should have an explicit zero row and confirmed completeness; an absent charge is unknown.

Aggregate at the right grain, then divide

Run the following SELECT on the existing tables through a SELECT-restricted Postgres role. Bufflehead provides the read-only query connection. Loading bills, recording deployment outcomes, assigning costs and deploying code remain outside Bufflehead.

WITH billed AS (
  SELECT attempt_id, sum(usd) AS usd, count(*) AS line_count
  FROM deployment_cost_demo.cost_lines
  GROUP BY attempt_id
), per_deployment AS (
  SELECT d.id, d.cohort, d.outcome,
         count(a.id) AS attempts,
         coalesce(sum(b.usd), 0) AS observed_usd,
         count(a.id) > 0 AND bool_and(
           coalesce(a.billing_complete AND b.line_count > 0, false)
         ) AS costs_complete
  FROM deployment_cost_demo.deployments d
  LEFT JOIN deployment_cost_demo.attempts a ON a.deployment_id = d.id
  LEFT JOIN billed b ON b.attempt_id = a.id
  WHERE d.started_at >= timestamptz '2026-09-22 00:00+00'
    AND d.started_at <  timestamptz '2026-09-23 00:00+00'
  GROUP BY d.id, d.cohort, d.outcome
), summary AS (
  SELECT cohort, count(*) AS deployments, sum(attempts) AS attempts,
         count(*) FILTER (WHERE outcome = 'success') AS successes,
         count(*) FILTER (WHERE outcome = 'pending') AS pending,
         count(*) FILTER (WHERE NOT costs_complete) AS incomplete,
         sum(observed_usd) AS observed_usd
  FROM per_deployment
  GROUP BY cohort
)
SELECT cohort, deployments, attempts, successes, pending, incomplete,
       observed_usd,
       CASE WHEN pending = 0 AND incomplete = 0
            THEN round(observed_usd / nullif(successes, 0), 2)
       END AS usd_per_success
FROM summary
ORDER BY cohort;

The first aggregation collapses multiple billing lines to one row per attempt. The second collapses attempts to one row per deployment before counting successes. This avoids counting the same successful request once for every cost line or retry.

cohort   deployments  attempts  successes  pending  incomplete  observed_usd  usd_per_success
budget   3            6         1          0        0           9.0000        9.00
standard 3            4         2          0        0           10.0000       5.00

observed_usd is only the sum of charges currently present. The query suppresses cost per success when any request is pending or has missing/incomplete costs. Zero successful deployments also produces NULL, not zero dollars per success. PostgreSQL’s aggregate documentation is useful here: empty sums and nullable inputs require deliberate handling.

A cohort with no deployment rows will not appear. If you expect both cohorts, check that both are present rather than interpreting absence as a free or inactive cohort. The completeness flags also cannot detect a deployment request missing from the source entirely; reconcile the deployment inventory separately.

Ask the agent to challenge the metric

Using only read-only queries, compare deployment cohorts for the agreed UTC start window. Count each request once, include all its retries and failed-attempt costs, and show missing billing coverage, pending outcomes and zero-success groups. Return the SQL and arithmetic. Do not infer causation, savings or ROI from cost per success.

Before trusting the narrative, test four cases in the disposable fixture: remove an attempt’s billing lines; mark an attempt incomplete; mark a request pending; and give a cohort zero successes. Each should make that cohort’s cost-per-success value NULL. Splitting one bill into two lines with the same total must leave the result unchanged. A deployment without an attempt must be flagged incomplete.

Unit cost is a diagnostic, not ROI

Lower cost per success can still be a poor business result if the accepted changes deliver less value, create more incidents, or require more human review. Compare equivalent work, stable acceptance rules and sufficiently mature cohorts; then examine quality, latency, labor and business outcomes alongside cost.

The practical starting point is a small, reproducible query over data you already own. Use Bufflehead to let your AI inspect that data through a read-only connection. Keep the accounting rules explicit, make the arithmetic reproducible, and do not let an attractive token price substitute for measuring the work that actually gets done.

Try it on your own data.

Native binary, no Docker, no cloud egress. Free for local files and direct database connections — request a demo for the AWS SSM path.

Download free Request a demo