Skip to main content
POST
Count tokens
Estimates the number of input tokens for a set of messages without sending the request to an LLM provider. Useful for pre-flight cost estimation, rate-limit planning, and prompt optimization.

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your API key. More info here

Body

application/json
model
string
required

ID of the target model. Format: {author_id}/{model_id}. The gateway uses this to pick the appropriate tokenizer when tokenizer is not provided.

Example:

"openai/gpt-5.2"

messages
object[]

Optional array of message objects to count tokens for. Accepts both OpenAI chat format (with system, user, assistant roles) and Anthropic Messages format; the format is auto-detected from the message structure. Defaults to an empty array.

system

Optional system prompt. Accepts a plain string or an array of Anthropic content blocks. Used when counting tokens for an Anthropic-style request.

tokenizer
enum<string>

Explicit tokenizer override. When omitted, the gateway picks one based on model.

Available options:
cl100k_base,
o200k_base

Response

Token count estimated successfully

input_tokens
integer
required

Estimated number of input tokens for the provided messages. This is an approximation, counts may differ from provider-native tokenizers. Use for estimation and budgeting, not exact billing.

Required range: x >= 0
Example:

42