Home/Blog/Evaluation & evidence/Twenty-four hours of silence is a billable resolution

Twenty-four hours of silence is a billable resolution

Customer service AI has moved from per-seat to per-resolution pricing. The corpus already established that deflection and resolution count different events, and now that distinction sets an invoice.

TL;DR. The industry has moved to outcome pricing. Intercom's Fin bills $0.99 per outcome, HubSpot dropped to $0.50 per resolved conversation in April 2026, Zendesk sits near $1.50 committed and $2.00 pay-as-you-go, and Salesforce Agentforce launched at $2.00 per conversation regardless of whether anything was resolved. The unit price is half the contract. The definition of the unit is the other half. Per Intercom's own documentation, Fin bills two things as a resolution: a confirmed one, where the customer says it worked, and an assumed one, where the customer goes quiet for 24 hours and does not return. Both cost $0.99. Silence is billable. This corpus already established that deflection and resolution count different events and only one gets reported. That was a measurement dispute. It is now an invoice.

---

Status: prices established, mechanisms documented, comparisons heavily interested. Every pricing comparison cited here is published by a vendor in the market, and each shows its own product favourably. The load-bearing material is different: it is vendors' own documentation of how they define a billable event, which is checkable and which is what this article is actually about.

---

The prices

Intercom Fin: $0.99 per outcome, with a stated minimum around 50 outcomes a month, and no platform fee.

HubSpot Customer Agent: $0.50 per resolved conversation from April 2026, halved from $1.00 per conversation.

Zendesk AI agents: roughly $1.50 per committed automated resolution, about $2.00 on pay-as-you-go overage, falling toward $1.00 at high volume, layered on Suite plans at $55 to $169 per agent per month.

Salesforce Agentforce: $2.00 per conversation, on top of Service Cloud from $175 per user per month.

Quickchat AI: from about $0.50. Decagon, Sierra and Ada use per-outcome models without published rates.

The spread is wide because the market has not settled, and the direction is uniform: away from per-seat, toward billing for something the software did.

The stated logic is sound. Per-seat pricing charges you for humans at precisely the moment you are trying to need fewer of them. A fifteen-person team whose AI handles 75% of inbound pays for fifteen seats whether the system resolved a hundred conversations or ten thousand. Outcome pricing is the obvious correction and the incentives genuinely do line up better.

And then the definitions

Salesforce's unit is the loosest. A conversation, defined as a 24-hour chat window, billed whether or not anything was resolved. A customer who opens a chat, gets nowhere and leaves costs $2.00. One analysis notes the definition proved problematic because a conversation could branch and linger without reflecting business value.

Freshdesk bills per session, so an issue requiring three sessions costs three units.

Intercom's is the one worth reading carefully, because Intercom documents it.

Fin bills two distinct things as a resolution. The first is a confirmed resolution: the customer reads the answer and indicates it worked, by a thumbs-up or a reply. That is the thing a buyer imagines when signing.

The second is an assumed resolution. Per Intercom's own documentation, if the customer goes quiet for 24 hours after Fin's last reply and does not return, the conversation is marked resolved and billed at the same $0.99.

Both appear in the invoice as resolutions. Only one of them is evidence that anything was resolved.

And "outcome" is broader still. Fin's outcome definition covers a resolution, a procedure handoff, or a disqualification, with lead qualification billed separately at $9.99.

Why the corpus was already here

The customer service article established that deflection rates and resolution rates count different events: 90% deflected against 40% resolved, and only the first gets reported.

Deflection means the ticket did not reach a human. That happens when the problem was solved, and it happens when the customer gave up.

An assumed resolution is deflection wearing the word resolution. A customer who read an unhelpful answer and went elsewhere produces exactly the same signal as one whose problem was solved: silence.

Twenty-four hours of it, and the meter fires.

Which is why the pricing shift matters more than a pricing shift usually does. When deflection and resolution were reporting categories, conflating them produced a misleading dashboard. Now it produces a charge, and the party that wrote the definition is the party sending the bill.

The verification move

Zendesk restructured in May 2026 into a three-tier resolution model where only Verified Resolutions are billed, confirmed by LLM evaluation.

That is a real improvement over billing everything that did not escalate, and it is worth crediting as such.

It also means a language model now determines an invoice. A model judges whether a model resolved something, and the judgment is a financial event.

The corpus has been here too. The LLM-as-judge material established that model evaluators carry position bias, verbosity bias and self-preference, and that agreement with human judgement is good rather than perfect.

Applied to a dashboard, an imperfect judge produces a slightly wrong number. Applied to billing, it produces a charge that is sometimes wrong in a direction nobody outside the vendor can audit, because the buyer sees the verdict and not the evaluation.

None of which says Zendesk's implementation is unsound. It says the verification step moved the problem rather than removing it, and the new location has less visibility than the old one.

What a buyer can actually check

Three questions, and they are unusually concrete.

What ends a billable unit? A confirmed positive signal, an inactivity timeout, an escalation, or a clock. The answer determines what fraction of your bill is evidence of anything.

Who adjudicates, and can you see the adjudication? A human review, a model evaluation, or an automatic rule. If a model decides, ask whether its decisions are exportable, because a verdict without a trace is not auditable.

And what happens at the boundary? Zendesk documentation reports that overages bill automatically once a committed tier is exceeded, with no warning. Success at automating and the bill rising are the same event, and the meter does not stop to confirm you meant it.

None of these questions is answerable from a pricing page. All three are answerable from documentation, and the documentation exists, which puts this in an unusually good position relative to most subjects in this corpus.

The layers underneath

The headline unit price is not the cost.

Zendesk stacks resolution charges on Suite plans at $55 to $169 per agent per month plus a reported AI add-on around $50 per agent per month. Salesforce requires Service Cloud from $175 per user per month before Agentforce is available. At twenty agents, platform fees alone run into thousands per month before a single resolution is counted.

Fin charges no platform fee and works with an existing helpdesk, which is the pitch, and which is also why Intercom publishes the comparison.

The structural point survives the sales pitch. A per-unit price is comparable across vendors only if the unit and the surrounding cost structure are comparable, and here neither is. A $0.99 outcome and a $2.00 conversation are not two prices for one thing. They are prices for two different things, and one of them charges for failure.

Where this is going

Salesforce signed a definitive agreement to acquire Fin in June 2026, folding the leading per-outcome product into the platform whose own model bills per conversation. How the definitions reconcile is the thing to watch, and it will be decided commercially rather than analytically.

Zendesk's chief executive framed the shift directly, saying the era of the chatbot, the era of frustration and deflection, is over.

That statement and the assumed-resolution mechanism describe the same industry in the same year. Both can be true: the products are substantially better than 2023 chatbots, and the billing definition still counts a customer walking away as a success.

And the seat model is genuinely eroding. Atlassian reported its first-ever decline in enterprise seat counts in 2026, attributed primarily to agent adoption. Whatever replaces per-seat pricing will be defined by whoever moves first, which is happening now, in documentation most buyers do not read.

What the same unit costs under four definitions

The arithmetic is worth doing once, because it makes the definitional point financial rather than rhetorical.

Take 50,000 conversations a month at a 65% genuine resolution rate, meaning 32,500 conversations where the customer's problem was actually solved.

ModelBillable eventsWhat is being paid for
Per conversation, $2.0050,000Every interaction, resolved or not
Per outcome incl. assumed, $0.99Resolved plus abandoned-and-silentSuccesses and quiet failures alike
Per verified resolution onlyConfirmed subsetSuccesses that produced a signal
Per seatHeadcountHumans, regardless of automation

Under the first, 17,500 failures are billed at full rate. Under the second, an unknown fraction of the 17,500 fall silent for 24 hours and become billable. Under the third, an unknown fraction of genuine successes go unbilled because satisfied customers rarely confirm.

The middle two are the interesting pair, because they err in opposite directions and nobody publishes the size of either error.

Which is the whole finding restated as arithmetic. The rate is public and precise to the cent. The quantity it multiplies is defined in a document the buyer probably has not read, and the number of assumed resolutions in a given invoice is known to exactly one party.

Why the counter-argument is strong here

This corpus usually finds that a defence of a practice is weaker than the practice's critics assume. In this case the reverse holds, and it is worth saying so plainly.

Assumed resolution has a real justification. Most people whose problem was solved do not click a confirmation. They read the answer and get on with their day. A vendor billing only confirmed resolutions would systematically undercount genuine work, and would face immediate pressure to prompt users for confirmation, which makes the product worse for everyone in order to make the invoice cleaner.

Twenty-four hours of silence is genuinely weak evidence, and it is genuinely evidence. The alternative defaults are worse: billing on escalation alone rewards a system that never escalates, and billing on human review does not scale.

And the alignment gain is large. Under per-seat pricing, a vendor's revenue rises when the customer employs more people, which is the opposite of what the customer is buying. Under per-outcome pricing, vendor revenue rises when the product works. That is a structural improvement and it is rare.

So the honest position is not that outcome pricing is a trap. It is that a good pricing model has one load-bearing definition inside it, that definition currently sits in documentation rather than in the contract, and only one party can count the events it governs. Those are three fixable things about an otherwise better arrangement.

Which is a different conclusion from the one the framing of this article invites, and stating it is the point of having a counter-argument section at all.

Three things this establishes

A measurement dispute becomes a financial one the moment the measurement becomes a unit. Deflection against resolution was a reporting problem for years. Pricing per resolution converts it into a line on an invoice, and everything unresolved about the definition is now unresolved about the money.

The party defining the billable event is the party issuing the bill. That is not an accusation; it is the ordinary structure of the market, and it is the reason the definitions belong in the contract rather than in the documentation. A buyer who negotiated a rate and not a definition negotiated half a contract.

And verification relocates the problem. LLM-confirmed resolutions are better than billing everything that did not escalate, and they place an imperfect judge in the billing path. The improvement is real and the auditability is worse, because the buyer sees a verdict rather than an evaluation.

What it does not establish

That any vendor is billing improperly. Every mechanism described here appears in the vendors' own public documentation, which is the opposite of concealment.

That outcome pricing is worse than per-seat. It is better aligned on the central point: per-seat charges you for the humans you are trying to redeploy, and outcome pricing does not.

That the comparisons are neutral. They are not. Every one is published by a competitor, and the totals in each favour whoever published it.

And nothing about which vendor to choose. This corpus does not make purchasing recommendations, and the question a buyer should ask is definitional rather than comparative.

Where it goes when outcomes are harder to define

Customer service is the easy case, and it is worth understanding why before assuming this generalises.

Support has a natural unit. A customer arrived with a problem, and either it went away or it did not. The event is discrete, bounded in time, attributable to one interaction, and observable from both sides. That is why the industry converged here first and why the definitional argument is about a 24-hour timeout rather than about what a resolution fundamentally is.

Most work is not shaped like that.

Legal work is the case already being attempted, with analysts expecting a decisive shift away from per-seat models by the end of 2026. The outcome is much harder to name. A contract reviewed is not a contract improved. A clause flagged is not a risk avoided. The valuable output of legal work is frequently something that did not happen, and non-events cannot be metered.

The same applies across most knowledge work. A financial model that avoided a bad decision, a security agent whose value is an incident that never occurred, an engineering agent whose contribution is a defect caught before merge. In each case the outcome worth paying for is counterfactual, and a counterfactual has no timestamp.

Which predicts what will happen. Where a natural unit exists, outcome pricing arrives quickly and the argument is about edge cases, as here. Where it does not, vendors will define a proxy, and the proxy will be whatever is countable rather than whatever is valuable: documents processed, alerts generated, actions taken, tokens consumed.

And that is the failure this corpus has documented repeatedly, from deflection standing in for resolution to viewability standing in for worth. A countable proxy substituted for an uncountable objective, chosen by the party being paid.

The prediction is falsifiable and worth stating. If outcome pricing spreads into domains without natural units, expect the billable events to be activity measures wearing outcome language. If instead vendors decline to meter where the outcome is counterfactual, and stay on subscription there, that would be evidence of more discipline than this corpus has generally found.

What is unresolved

What fraction of billed resolutions are assumed rather than confirmed. This is the single number that would settle the argument, no vendor publishes it, and every vendor has it.

Whether LLM verification agrees with human review. An agreement rate on a sample of billed resolutions is straightforward to produce and nobody has produced one publicly.

How the Salesforce and Intercom definitions reconcile after the acquisition. Per-conversation and per-outcome are different philosophies, and one will win.

And whether definitions migrate into contracts. At present the rate is negotiated and the definition is documented, which places the two halves of the deal in different documents with different revisability.

What a contract would have to say

The fix is unglamorous and it is short, which is worth demonstrating rather than asserting.

Four clauses would move the load-bearing half of this arrangement out of documentation and into the agreement.

Define the billable event, including its negative cases. Not "a resolution" but the specific triggers: a confirmed positive signal, an inactivity timeout of a stated length, an escalation, or a session close. Naming the timeout is the whole thing, because a timeout is where the disputed cases live.

Require the split. Confirmed against assumed, reported monthly, as a count rather than a percentage. No vendor currently publishes this and every vendor has it, which makes it the cheapest disclosure available in the subject.

Fix the adjudication and its trace. If a model verifies resolutions, the verdicts should be exportable with enough context to sample. A buyer who cannot sample cannot audit, and sampling a hundred billed resolutions a month is a small operational cost against a bill in the tens of thousands.

And handle the boundary explicitly. Automatic overage billing with no warning converts a successful month into an unbudgeted one. A notification threshold is a one-line term and it is absent from the arrangements described here.

None of these is adversarial. A vendor confident in its definition loses nothing by putting it in the contract, and a vendor whose assumed-resolution share is modest gains from publishing it. The reason none of it is standard is that the market is eighteen months old and nobody has asked.

Which is the useful thing this article can offer, since it cannot tell anyone which vendor to choose and can point at the four questions that decide half the money.

The counter-argument

Assumed resolution is a reasonable default and the alternative is worse. Most satisfied customers do not click a thumbs-up; they read the answer and leave. Billing only confirmed resolutions would systematically undercount genuine successes and would push vendors toward nagging users for confirmation, which is a worse product. Twenty-four hours of silence is weak evidence and it is evidence.

The vendor-defined unit problem is universal, not specific. Every metered service defines its own unit: a cloud provider defines a compute-hour, a telecom defines a call minute. Singling out AI pricing for a structure that describes all of metered software imports a standard nobody applies elsewhere.

This article's sourcing is weak in a way it cannot fully escape. The pricing landscape here is assembled almost entirely from vendor comparison pages, each written to win a deal. The documented mechanisms are the reliable part and the market picture is not, and no independent survey of AI customer service pricing exists.

And the alignment argument may be decisive. Under per-outcome billing, a vendor earns more when its product works better, which is the first pricing model in enterprise software where the vendor's revenue tracks the metric the buyer cares about. A definitional imperfection inside a correctly aligned incentive is a much better problem than a cleanly defined unit inside a misaligned one.

The short version

Customer service AI moved to outcome pricing: Fin at $0.99 per outcome, HubSpot at $0.50 per resolved conversation from April 2026, Zendesk near $1.50 committed and $2.00 on overage, Agentforce at $2.00 per conversation whether or not anything resolved.

The rate is half the contract. The unit definition is the other half, and it lives in documentation rather than in the pricing page.

Per Intercom's own documentation, Fin bills a confirmed resolution and an assumed resolution identically. A confirmed one is the customer saying it worked. An assumed one is 24 hours of silence. Both are $0.99, and only one is evidence.

Which is the corpus's own finding with money attached. Deflection and resolution count different events, 90% against 40%, and a customer who gave up produces the same signal as one whose problem was solved. That was a dashboard problem. It is now a charge.

Zendesk's May 2026 restructure bills only Verified Resolutions, confirmed by LLM evaluation, which is a genuine improvement and puts an imperfect judge in the billing path, where the buyer sees a verdict and not the evaluation.

Three questions a buyer can actually ask: what ends a billable unit, who adjudicates and can you see it, and what happens at the tier boundary, where Zendesk documentation reports overages billing automatically with no warning. Success at automating and the bill rising are the same event.

And every comparison in this subject is published by a competitor. The prices are approximately right, the totals favour whoever published them, and the reliable material is the vendors' own documentation of what they count.

Common questions

What is outcome-based pricing for AI customer service? Charging per result rather than per user seat. Current published rates include Intercom Fin at $0.99 per outcome, HubSpot Customer Agent at $0.50 per resolved conversation since April 2026, Zendesk at roughly $1.50 per committed automated resolution and about $2.00 on pay-as-you-go overage, and Salesforce Agentforce at $2.00 per conversation. The logic is sound: per-seat pricing charges for humans at exactly the moment an organisation is trying to need fewer of them.

What is an assumed resolution? A conversation billed as resolved because the customer stopped replying. Per Intercom's own documentation, if a customer goes quiet for 24 hours after Fin's last reply and does not return, the conversation is marked resolved and billed at the same $0.99 as a confirmed resolution where the customer indicated the answer worked. Both appear on the invoice as resolutions and only one is evidence that anything was resolved.

Why does that matter beyond one vendor? Because it is the corpus's existing finding with a price attached. Deflection rates and resolution rates count different events, with one widely reported figure showing 90% deflected against 40% actually resolved. Deflection means the ticket did not reach a human, which happens when the problem was solved and when the customer gave up. Both produce silence. When those were reporting categories, conflating them produced a misleading dashboard; when the unit is billable, it produces a charge.

Is Zendesk's verified resolution model better? Yes, and it relocates the problem. The May 2026 restructure bills only Verified Resolutions, confirmed by LLM evaluation, which is a real improvement over billing everything that did not escalate. It also places a language model in the billing path, judging whether another model resolved something. Model evaluators carry position, verbosity and self-preference biases and agree with human judgement well rather than perfectly. Applied to a dashboard that produces a slightly wrong number; applied to billing it produces a charge the buyer cannot audit, because they see a verdict rather than an evaluation.

What should a buyer ask? Three concrete questions, none answerable from a pricing page and all answerable from documentation. What ends a billable unit: a confirmed positive signal, an inactivity timeout, an escalation, or a clock. Who adjudicates, and are those decisions exportable, since a verdict without a trace is not auditable. And what happens at the tier boundary, given that Zendesk documentation reports overages billing automatically once a committed tier is exceeded, with no warning.

Are these prices comparable across vendors? No, and treating them as comparable is the main error. A $0.99 outcome and a $2.00 conversation are prices for different things, and one of them charges for failure. Platform costs also stack differently: Zendesk layers resolution charges on Suite plans at $55 to $169 per agent per month plus a reported AI add-on, Salesforce requires Service Cloud from $175 per user per month, and Fin charges no platform fee. At twenty agents, platform fees alone can run into thousands per month before a single resolution is counted.

How reliable are the figures in this article? The mechanisms are reliable and the market picture is not. Every pricing comparison cited is published by a vendor competing in this market, and each shows its own product favourably in the totals. What is checkable is what each vendor documents about its own billable events, which is the actual subject here. No independent survey of AI customer service pricing exists.

Is outcome pricing a good development? On the central point, yes, and the strongest argument for it is that a vendor earns more when its product works better, which makes this the first widespread enterprise pricing model where vendor revenue tracks the metric the buyer cares about. The defence of assumed resolution is also reasonable: most satisfied customers never click a confirmation, so billing only confirmed resolutions would undercount real successes and push vendors toward nagging users. Twenty-four hours of silence is weak evidence, and it is evidence. The narrow claim here is that the definition belongs in the contract rather than the documentation, because it decides half the money.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. Cost Per Resolution: How AI Agents Game the KPI BuildMVPFast The documented mechanisms: Intercom's own documentation on assumed resolutions billed after 24 hours of silence, Agentforce's 24-hour conversation window billed regardless of outcome, and Zendesk documentation on automatic overage billing.
  2. AI Agent Pricing Compared: Fin vs Zendesk vs Agentforce Intercom Zendesk's May 2026 three-tier restructure billing only Verified Resolutions confirmed by LLM evaluation. Published by a competitor, and the totals favour the publisher, which is stated in the article.
  3. AI Agent Pricing Models 2026: Per-Resolution vs Per-Seat Compared Quickchat AI The published rate landscape across vendors and the observation that resolution is a vendor-defined term. Also a vendor in this market.
  4. Outcome-Based Pricing: The Real AI-Native Signal guptadeepak.com HubSpot's April 2026 move to $0.50 per resolved conversation, the Atlassian seat-count decline, and the argument that legal AI is the harder case because outcomes there are less definable.

Further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why general benchmarks cannot carry general claims. :: https://arxiv.org/abs/2111.15366 Construct Validity
  • Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in Natural Language Understanding? — what a benchmark must satisfy to support inference. :: https://arxiv.org/abs/2104.02145 Construct Validity

Learn the concepts

← All posts