← THE OBSERVATORY

Pricing AI agents on outcomes: the maths we use on ourselves

Seats and hours make no sense for software that works alone at 3am. How we set the metric, fix the baseline and share the upside.

28 JUN 2026 · 11 MIN READ · BY THE AI PRACTICE

Per-seat pricing made sense when software was a tool that people picked up. An agent is not a tool that people pick up. It reads the inbox at 3am, chases the invoice, books the callback, and nobody is sitting at a seat while it does. Charging by the seat for that is charging for the wrong thing.

Why seats and hours both fail

Seats price access, and hours price effort. An agent consumes neither in any way that matters to the client. If it clears a task queue in four minutes that used to take a team all week, an hourly rate rewards us for being slow and a seat licence rewards us for headcount that no longer exists.

The only thing left worth pricing is the result. That sounds like a slogan until you try to write it into an invoice, at which point it becomes three very specific numbers.

$0PER SEAT, PER USER, PER HOUR. THE AGENT LINE IS PRICED ON THE METRIC IT MOVES

The three numbers we agree first

First, the metric: the single figure the agent exists to move. Qualified replies, recovered invoices, orders processed without a human touch. One metric per agent; an agent asked to move everything is accountable for nothing.

Second, the baseline: what that figure was before the agent switched on, measured from the client's own systems over a long enough window to be honest about seasonality. We do not set the baseline; the data does.

Third, the share: what proportion of the improvement comes back to us, over what period, with a floor that covers the running costs and a cap that keeps the arrangement sane when the agent does far better than either side expected.

If we cannot name the number the agent is supposed to move, we have no business charging for moving it.

Running the model on ourselves

Before any client saw this, we ran it internally. Every agent in our own operation has a ledger line: the metric, the baseline, and what it has moved since. Some earned their keep within a month. One did not, and was switched off, which is exactly the discipline the model is supposed to enforce.

The pricing conversation changes shape as a result. It stops being a negotiation about our costs and becomes a joint forecast about the client's outcome, which is a far better place to start an engagement from.

NEXT OBSERVATIONA year of self-hosted n8n: what it costs, what it saves↗︎