How Usage-Based Billing Works How Usage-Based Billing Works

How Usage-Based Billing Works

System Admin System Admin

At Acme Tech, some services are billed based on how much you use them. This means your total charges may change from month to month depending on activity, rather than staying fixed every billing cycle.

Usage-based billing is common for AI and cloud services where consumption can vary based on workload, request volume, or model usage.


How Acme Tech pricing works

Acme Tech pricing may include:

  • Fixed platform fees for access to a plan or service
  • Usage-based charges for consumption above or beyond included limits
  • Model-based pricing for different AI models
  • Tiered pricing based on the level of capacity or features included in your plan
  • Custom enterprise pricing for larger deployments or special requirements

Your exact pricing depends on your plan, contract, and the services you have enabled.


AI model pricing

For AI services, pricing may be based on the model you choose and the number of tokens processed.

Tokens are small pieces of text used by AI models when they read or generate content. In simple terms:

  • Input tokens are the text you send to the model
  • Output tokens are the text the model generates in response

Different models may have different pricing depending on size, performance, and workload.

Pricing

ModelInput (per 1M tokens)Output (per 1M tokens)
DeepSeek-R1-Distill-Llama-70B0.70 USD1.40 USD
DeepSeek-V3.1-cb0.15 USD0.75 USD
DeepSeek-V3.13 USD4.50 USD
DeepSeek-V3.23 USD4.50 USD
gemma-3-12b-it0.35 USD0.59 USD
gemma-4-31B-it0.38 USD1.15 USD
gpt-oss-120b0.22 USD0.59 USD
Llama-4-Maverick-17B-128E-Instruct0.63 USD1.80 USD
Meta-Llama-3.3-70B-Instruct0.60 USD1.20 USD
MiniMax-M2.70.60 USD2.40 USD

What counts toward usage?

Depending on the service, usage may be measured by things like:

  • number of AI requests or resolutions
  • input and output tokens processed
  • compute capacity consumed
  • amount of data processed
  • number of users or environments
  • service-specific usage defined in your agreement

For AI workloads, both the text you send and the text the model returns can affect cost.


How charges are calculated

Usage-based billing is usually calculated in three steps:

  1. We track your usage during the billing period.
  2. We apply the pricing terms from your agreement or selected model.
  3. We combine usage charges with any fixed subscription fees.

Simple example

If you use:

  • 1 million input tokens on a model priced at $0.35 per 1M input tokens
  • 1 million output tokens on the same model priced at $0.59 per 1M output tokens

Your AI usage charge would be:

  • Input: $0.35
  • Output: $0.59
  • Total: $0.94

This is only an example. Your actual pricing will depend on the model and plan you choose.

Was this article helpful?

0 out of 0 found this helpful

Articles in this section

Add comment

Please sign in to leave a comment.