Model usage
The Model usage screen is your own spend on model calls: how much you have used in the current period, how much of your budget remains, and which models and providers account for it.
Your budget
Your figure is measured against whichever budget applies to you. That may be a limit set for you specifically, or the organization default. Either way the screen shows the limit alongside the spend, so the number has context.
Two things are worth knowing about how this behaves.
Running out means requests are refused, not throttled or queued. If your spend reaches your limit, further model calls fail until the period rolls over or an administrator raises the limit. The period is calendar-aligned, so a daily budget resets at the start of the next day rather than 24 hours after your first request.
Spend is priced from the provider's reported usage, after each response, so a figure here reflects real token counts rather than an estimate.
Breakdown
The spend is broken down by model and by provider, which is the useful pair when you are trying to reduce it: the model breakdown tells you what to switch away from, and the provider breakdown tells you where the money is going.
If a request is refused
A refused request is usually one of three things: your budget is exhausted, no budget applies to you at all, or the model you asked for has no price configured. This screen distinguishes the first from the others, since an exhausted budget shows as full utilization here.
For the other two, ask an administrator. Both are configuration on their side rather than anything you can change.
Next steps
- Virtual API keys if you need a non-interactive way to authenticate.
- Connect a client to route another tool through the gateway.