Method
usage = top-k × chunk + prompt + output reserve.
Check how many tokens retrieved chunks consume before the model call.
usage = top-k × chunk + prompt + output reserve.
Use the result as a transparent planning estimate. Change one assumption at a time, compare scenarios and move to Toolkit when you need system-level simulation.
Reranking may send fewer chunks to the model than initially retrieved.
Your inputs are processed in this browser. This utility does not need an account, database or AI API.