1. Count reusable prompt tokens
Measure system instructions and static user-template content.
2. Add variable context
Enter the average retrieved documents, history, or user data appended to each prompt.
3. Set output and executions
Use the expected response size and number of prompt runs for the budget period.
4. Enter token rates
Apply the rates for the model and billing method you intend to use.
5. Find the optimization target
Compare fixed prompt, context, and output costs rather than looking only at the total.