1. Step 1
Enter the audio duration represented by one file.
2. Step 2
Set an expected speaking rate; use a lower value for pauses or interviews and a higher one for dense narration.
3. Step 3
Enter a tokens-per-word conversion suitable for the language and model tokenizer.
4. Step 4
Add prompt, metadata, and expected downstream output tokens per file.
5. Step 5
Enter the number of similar files in the batch and a safety margin for variation.
6. Step 6
Compare planned tokens with context limits, rate limits, and the purchased API budget.