1. Step 1
Define one task by a consistent audio unit and enter its average duration.
2. Step 2
Add the speech-to-text provider charge per audio minute.
3. Step 3
Enter downstream text tokens and their blended price when cleanup or summarization uses a language model.
4. Step 4
Include infrastructure expense not already captured by provider rates.
5. Step 5
Add expected review time and labor cost, then allocate fixed monthly costs across task volume.
6. Step 6
Use the breakdown to identify which cost component changes most between workflow options.