Forum Discussion
How are people estimating LLM API costs before putting an AI application into production?
It's relatively easy to estimate the number of API requests, but I'm finding the token side more interesting because input and output usage can vary considerably between requests. Things get even more complicated when comparing different models or using long prompts and conversation history.
Do you normally build a spreadsheet for this, use the provider's pricing information directly, or use a dedicated calculator?
I'm especially interested in approaches that estimate monthly cost from expected requests, average input tokens, average output tokens, and the model being used.
4 Replies
- MattBurrSteel Contributor
Briar and jamiecross25 - very suspicious that the two of you are AI bots 🤔
- jamiecross25Brass Contributor
Hello MattBurr
I am not an AI bot. I am human and have been contributing and gaining knowledge on this platform for a while now. I have also stumbled on some of your contributions and always found your comments insightful.
Can I know why you think/suspect otherwise?
- jamiecross25Brass Contributor
That is the million dollar question, is it not? Estimating tokens can be tricky because output length varies so much. A spreadsheet is your best friend here. I would map expected requests against average input and output token usage to get a realistic estimate. The outcome usually gives you a solid ballpark for monthly spend. Just remember to factor in conversation history, because that can grow surprisingly quickly. It is not perfect, but it is certainly better than guessing. I would also use the provider's pricing page as your baseline.
- BriarBrass Contributor
I usually start with a spreadsheet that estimates average input and output tokens per request, then multiply those figures by the expected monthly request volume. I plug in each model’s current pricing to compare costs. I also account for conversation history, retries, and peak usage, since they can significantly affect the final bill. For example, if an AI application uses the Nonprofit Check Plus API for verification, I’d include those API calls in the overall usage estimate. Running a small test with real prompts before launch helps validate the estimates.