Skip to main content

Monitor Usage & Billing

Open Usage & Billing

In the left sidebar, under Dedicated Inference, click My endpoints, then click the Usage & Billing tab.

Usage & Billing tab with KPI cards, filters, and token usage chart

Page title: Usage & Billing Last 14 days · updated every 5 minutes

Filter the View

Use the Filters section to scope all charts and the table simultaneously:

FilterOptions
From / ToDate and time range
ModelAll models or a specific model
API KeyAll keys or a specific key

KPI Cards

CardWhat it shows
TOKENS INInput tokens consumed in the selected period
TOKENS OUTOutput tokens generated in the selected period
GPU-HOURSTotal GPU-hours reserved — this is what drives your bill
SPENDTotal cost in USD for the selected period

Each card shows the percentage change vs the previous 14 days.

note

TOKENS IN and TOKENS OUT are tracked for analytics only. Billing is capacity-based (GPU-HOURS), not per token. Billing is per-minute with a minimum of 60 seconds, rounded up.

Charts

  • Token usage by modelStacked tokens per model / day: shows token consumption per model over time.

Export Data

Click Export CSV to download the filtered data as a CSV file for reporting or finance submissions.

Set a Budget Alert

Click Set budget alert to configure a monthly spend threshold. You receive a notification when spend reaches the configured level.