Back to blog
Guides

How to plan your AI support volume: quotas instead of surprises

A practical method for estimating how many AI replies your support actually needs per month — baseline volume, automatable share, seasonality buffer — and turning the estimate into a fixed pool instead of a metered bill.

ReplyPool TeamJuly 2, 20266 min read

Key takeaways

  • Estimate volume from ratios: conversations per 100 customers (SaaS 5–15) or per 100 orders (e-commerce 8–20), adjusted for channel mix and seasonality.
  • Size the pool for automatable volume only — routine, documented questions are 55–70% of inbound at most companies.
  • Pool size = automatable volume × seasonality buffer; size for the p80 month, not the record month.
  • Write the exhaustion policy before you need it: route to team by default, top up only as a deliberate purchase.
  • Upgrade on one trigger — two consecutive months above 80% of the pool — and review usage ten minutes a month.

Most teams meet their AI support bill the way they meet a storm: it happens to them. A launch goes well, ticket volume doubles, the AI answers everything — and the invoice quietly doubles too. The fix isn't to automate less. It's to plan volume the way you plan any other budget line: estimate it, cap it, and review it monthly.

This guide is a practical method for doing exactly that. By the end you'll have a defensible estimate of how many AI replies your support actually needs per month, and a way to turn that estimate into a fixed number on your invoice — a quota you chose, not a meter you discovered.

Why volume planning beats "pay as you go"

Metered AI support sounds fair: pay for what you use. In practice it means your support cost is a function of things you don't control — a viral post, a checkout bug, a seasonal spike. The bill arrives after the volume, so every good month for growth is a bad month for the budget.

Planning inverts that. You estimate volume first, pick a monthly pool of AI replies that covers it with a sensible buffer, and treat anything beyond the pool as a conscious decision rather than an automatic charge. Three things change immediately:

  • Forecasting works. Support becomes a line item you can put in a 12-month plan without an asterisk.
  • Spikes stop being emergencies. When the pool runs out, conversations route to your team — service degrades gracefully instead of the bill escalating silently.
  • Automation decisions get honest. You compare "what does answering this class of question cost us in pool replies" against real alternatives, not against an open-ended meter.

Step 1 — Establish your baseline conversation volume

Start with what already happens. Pull 3 months of history from your current helpdesk or shared mailbox and answer four questions:

  1. Total inbound conversations per month. Not tickets, not emails — conversations. Merge the duplicates.
  2. Conversations per 100 customers (or per 100 orders). This ratio is the engine of every forecast. SaaS products typically see 5–15 conversations per 100 active accounts monthly; e-commerce sees 8–20 per 100 orders, concentrated around delivery and returns.
  3. Channel mix. Chat and messaging channels generate 1.5–2× more conversations than email for the same customer base, because the friction of asking is lower. If you're about to add a widget or WhatsApp, plan for the mix to shift.
  4. Seasonality. Mark your two heaviest months and your two lightest. The ratio between peak and trough is your seasonality factor — for most products it lands between 1.3 and 2.5.

If you're pre-launch and have no history, estimate from the ratios above and your growth plan, then treat the first 60 days as calibration.

Step 2 — Estimate the automatable share

Not every conversation should be answered by AI, so your pool shouldn't be sized for total volume. Sort your conversation history into three buckets:

  • Routine and documented — order status, returns, billing details, plan questions, password-and-settings mechanics. The answer exists (or should exist) in your docs. This is AI territory, and at most companies it's 55–70% of everything inbound.
  • Judgment calls — refund exceptions, angry escalations, custom quotes, anything where the answer depends on a human decision. Route these to people by design.
  • Signal — bug reports, feature requests, churn warnings. AI can acknowledge and tag these, but a person should see them.

Take your monthly baseline, multiply by the routine share, and you have your automatable volume. Example: 1,800 conversations a month × 62% routine ≈ 1,120 conversations the AI should be fielding.

One correction factor: the AI only resolves what your knowledge base covers. If your docs are thin, the realistic share in month one is lower — plan for the AI to hand off questions it can't ground in a source, and use its knowledge-gap report to close the misses. Coverage typically climbs for the first two or three months, then stabilizes.

Step 3 — Convert volume into a pool size

Now the arithmetic that metered pricing never asks you to do:

Pool size = automatable volume × seasonality buffer.

Use your peak-month factor if budget certainty matters more than pool efficiency, or a 1.2× buffer over the average if you're comfortable with occasional handoffs in peak weeks. From the example above: 1,120 × 1.2 ≈ 1,350 replies — so a 2,500-reply plan runs with comfortable headroom, while a 1,000-reply plan would rely on your team absorbing the peaks.

Three sizing rules that save money:

  1. Size for the p80 month, not the record month. The pool's job is to make the bill boring, not to guarantee zero handoffs forever. The graceful-degradation path (AI pauses, humans answer) exists precisely so you don't have to buy insurance twelve months a year.
  2. Don't buy next year's volume today. Quota plans move up in minutes. Upgrade when two consecutive months run above 80% of the pool — that single trigger replaces a quarterly pricing meeting.
  3. Count replies, not conversations, if your vendor meters that way. A conversation can contain several AI answers. Check the definition before comparing plans; a "1,000" on one pricing page is not always a "1,000" on another.

Step 4 — Decide the exhaustion policy before you need it

The most important line in any AI support plan is what happens at zero. Write it down as policy:

  • Route to team (the default): when the pool is spent, new conversations go straight to your inbox with full context. Customers still get answers — from people — and the bill doesn't move.
  • Deliberate top-up: if a launch or a sale makes the extra volume worth it, buying an additional block of replies should be a decision someone makes on purpose, with a price they can see. If your tool converts pool exhaustion into automatic charges, that's not a top-up — that's overage billing wearing a coat.

Put a monthly reminder on the pool-usage report. Ten minutes: usage vs. pool, resolution rate, top unanswered questions. That cadence catches drift long before it becomes either a budget problem or a service problem.

The worked example, end to end

A 7-person e-commerce team, 9,000 orders a month:

  • Baseline: 12 conversations per 100 orders ≈ 1,080 conversations/month
  • Channel shift: adding a chat widget, plan +25% → ≈ 1,350
  • Routine share after doc cleanup: 65% → ≈ 880 automatable
  • Seasonality buffer 1.2× → ≈ 1,050 replies needed
  • Choice: a 1,000-reply plan with peak-week handoffs, or a 2,500-reply plan with full headroom. Either way the number on the invoice is known before the month starts.

That's the whole method. Two hours of history-digging, four multiplications, and your AI support budget stops being a weather forecast.

Where ReplyPool fits

ReplyPool is built around exactly this planning model. Every plan is a fixed monthly price for a fixed pool of AI replies — 1,000, 2,500 or 7,500 — with a hard cap and no overage invoices. When the pool runs out, conversations route to your team inbox with full context until the next cycle, and top-ups are deliberate purchases, never automatic charges. The analytics dashboard shows pool usage and knowledge gaps, so the monthly ten-minute review is one screen.

Want to see your own numbers in it? Start a free trial, connect your channels, and run one calibration month — no credit card required.

Share this article

Frequently asked questions

Pull three months of history from your current helpdesk and compute conversations per 100 customers (SaaS typically sees 5–15 per 100 active accounts) or per 100 orders (e-commerce sees 8–20). Multiply the ratio by your customer or order forecast, adjust for channel mix — chat and messaging generate 1.5–2× more conversations than email — and mark your seasonal peak-to-trough factor, usually between 1.3 and 2.5.

At most companies, routine and documented questions — order status, returns, billing details, plan questions, settings mechanics — make up 55–70% of inbound volume, and that is the share worth automating. Judgment calls and signal (bugs, feature requests, churn warnings) should route to people by design. The realistic share in month one is lower if your knowledge base is thin, and climbs for two or three months as you close documented gaps.

Pool size = automatable volume × seasonality buffer. Use a 1.2× buffer over your average month if occasional handoffs during peaks are acceptable, or your peak-month factor if budget certainty matters more. Size for the p80 month rather than your record month — the graceful-degradation path exists so you do not have to buy insurance twelve months a year.

Decide the exhaustion policy before you need it. The default should be routing: new conversations go straight to your team inbox with full context, customers still get answers, and the bill does not move. Buying an additional block of replies should be a deliberate decision someone makes with a visible price — if pool exhaustion converts into automatic charges, that is overage billing, not a top-up.

Use one trigger: two consecutive months running above 80% of the pool. Quota plans move up in minutes, so there is no reason to buy next year’s volume today. A ten-minute monthly review of pool usage, resolution rate and top unanswered questions catches drift long before it becomes a budget or service problem.

Estimate from industry ratios — conversations per 100 customers or orders — applied to your growth plan, pick the smallest pool that covers the estimate with a 1.2× buffer, and treat your first 60 days as calibration. Watch the pool-usage report weekly at first; one real month of data replaces every assumption.

Ready to put AI support to work?

14 days free. Full platform. We move your data for you.