Insight3 min readBoris, Founder at VendurisPublished

    The Obscurity of Token-Based Pricing: Why AI Costs Are Hard to Forecast

    Seat-based pricing lets you forecast next year's bill from this year's headcount. Token pricing breaks that, and most finance processes haven't adjusted yet.

    Traditional SaaS pricing, whatever its flaws, is at least predictable: a seat costs a fixed amount, and you can forecast next year's bill from this year's headcount. Token-based pricing for AI tools breaks that predictability in a way most finance teams haven't fully adjusted to yet.

    Why token pricing is genuinely hard to reason about

    A token isn't a unit most people have an intuitive feel for. It's a fragment of text, roughly three-quarters of a word on average, but that ratio shifts by language, content type, and model. Usage isn't driven by headcount, it's driven by how people actually use the tool: longer prompts, longer outputs, more back-and-forth conversation, all consume tokens differently, and none of that is easy to estimate in advance from a seat count the way traditional software usage is.

    What you can forecast, and what actually moves the bill

    Headcount
    Known a year out
    Seats licensed
    Known a year out
    Prompt length
    Depends on the workflow
    Output length
    Depends on the workflow
    Calls per integration
    Depends on adoption
    Model tier used
    Changes mid-term

    Two of these come off a headcount plan. The other four decide the invoice.

    What you can forecast, and what actually moves the bill.

    Why this creates real budget risk

    With seat-based SaaS, a surprise invoice is rare, the cost ceiling is roughly known in advance. With token-based pricing, usage can scale in ways nobody explicitly approved. A single team adopting a workflow that generates unusually long outputs, or an integration that calls the API far more often than anticipated, can meaningfully move the bill without any change in headcount or any deliberate purchasing decision.

    What to ask for before signing or renewing

    • A clear breakdown of how the specific model and usage pattern translates to token consumption, in terms you can actually estimate against
    • Whether volume discounts or pricing tiers exist as usage scales, and at what thresholds
    • Whether there's a way to cap or alert on usage before it becomes an unexpected overage
    • Historical usage data, if you're renewing, showing actual token consumption trends over the prior term, not just the contracted volume

    Why this is a newer kind of negotiation

    Most finance and procurement processes were built around predictable, seat-based cost structures. Token-based pricing requires a different kind of scrutiny, one focused on usage pattern monitoring rather than headcount forecasting, and that shift in approach hasn't caught up everywhere yet.

    Common questions

    Let's look at your next renewal together.

    Thirty minutes with the founder. We map your upcoming renewals, flag the notice windows that are about to close, and you decide whether Venduris is worth your time.

    Book a renewal reviewAssess