1. Enter token usage
Use average uncached input and generated output tokens for one completed task.
2. Add model rates
Enter prices per one million tokens in the same currency.
3. Include retrieval and infrastructure
Allocate vector search, embedding, compute, storage, observability, and related charges on a per-task basis.
4. Set monthly volume
Use completed billable tasks, excluding retries only when their cost is already embedded in another input.
5. Review unit and monthly economics
Compare the model share with non-model costs before choosing an optimization strategy.