Definition
Hard Token Cap
A hard token cap is a pricing control used by some AI support vendors that limits how many tokens (units of text processed) an AI can consume per conversation or per billing period. When the cap is hit, the AI stops responding or degrades to a fallback mode — even if the customer's issue is unresolved. Hard token caps create unpredictable support quality and are a major source of hidden AI pricing costs.
How it works
Large language models process text in units called tokens — roughly 0.75 words per token in English. A typical support conversation consumes 500–2,000 tokens depending on length, complexity, and how much of your knowledge base the AI retrieves.
Vendors who charge per-token impose hard caps to control infrastructure costs. When a conversation hits the cap, the AI stops generating output. In a support context, this means a customer mid-conversation suddenly receives no response, or a generic 'please contact us' fallback — at exactly the moment they need help most.
Hard caps are often buried in pricing pages as 'included tokens' or 'conversation credits.' A 1,000-token monthly cap sounds generous until you realize a single BFCM traffic spike can exhaust it in hours. Some vendors charge overage fees per additional 1,000 tokens, adding unpredictable costs during peak periods.
When you'll encounter it
You'll encounter hard token caps when comparing AI support pricing pages, especially from newer AI-native vendors that expose underlying LLM costs directly. Watch for terms like 'conversation credits,' 'AI credits,' 'included tokens,' or 'automation sessions' with monthly limits — these are token caps with different names.
How SupportSyndicate handles this
SupportSyndicate uses flat-rate pricing with no token caps, conversation limits, or per-resolution fees. Every conversation — whether it's 3 messages or 30 — costs the same. During BFCM or a product launch when volume spikes 5x, the AI keeps working without degradation or overage charges.
Related terms
- Per-Seat PricingPer-seat pricing is a billing model where the cost of a software subscription scales with …
- Per-Resolution PricingPer-resolution pricing is an AI support billing model where you pay a fee for each ticket …
- Deflection RateDeflection rate is the percentage of incoming support tickets that are fully resolved by a…