litellm/litellm-rust/crates/token-counter
devin-ai-integration[bot] 268e8bb735
refactor(rust): share anthropic types, request helpers, and streaming contracts across crates (#43426)
* refactor(rust): standardize Azure Messages module path

* docs(rust): define shared types crate boundaries

* refactor(rust): share request helpers and type Anthropic blocks

* docs(rust): format shared type invariants as bullets

* test(rust): parameterize repeated cases with rstest

* refactor(rust): move Responses transform result into llms

* fix(anthropic): validate chat and batch responses

* docs(rust): clarify API format ownership boundaries

* docs: clarify Rust error message construction

* refactor(auth): keep shared Rust errors provider-neutral

* refactor(rust): separate format contracts from provider policy

* fix(rust): type Anthropic chat response text collection

* fix(rust): pass audio secret sources through hosts

* fix(rust): unblock batch lint and OCR error assertions

* test(rust): assert response failures at the adapter boundary

* refactor(rust): declare error messages with typed context

* wip

* fix(rust): adapt Bedrock error details

* style(rust): cargo fmt bedrock audio transcription

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): adapt tests and dead code to typed error details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep converse error contracts and read env secrets without litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 45 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): tolerate missing usage in converse responses on the transcription route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:53:12 -07:00
..
benches refactor(rust): split token counter backends 2026-09-20 14:10:31 -07:00
src refactor(rust): share anthropic types, request helpers, and streaming contracts across crates (#43426) 2026-09-27 14:53:12 -07:00
tests fix(rust): validate tokenizer ranks and cover backend features 2026-09-20 15:13:59 -07:00
Cargo.toml test(rust): run fast token counter parity tests by default 2026-09-20 21:52:48 +00:00
README.md feat(tokenizer): preserve Python defaults with opt-in Rust dispatch (#42174) 2026-09-22 04:41:11 +00:00

Token counting

Tokenizer is the text-counting interface. TextCodec adds encoding, decoding, and a name. TokenCounter applies LiteLLM request, message, and tool accounting using any Tokenizer

Counts follow the codec: tiktoken treats special-token spellings as ordinary text, while Hugging Face applies its added tokens, post-processing, padding, and truncation. fast=True preserves those semantics and requests acceleration where available. Unsupported configurations use the normal codec, including tiktoken encodings without a scanner and builds without the fast feature. Invalid input and process-guard errors still propagate. Runtime request counting currently uses the normal codec; the custom accelerator is retained for explicit use and testing

FastCounter: TextCodec exposes an optional accelerator over a loaded codec. None means callers should use that codec. The Python bridge caches this selection per immutable tokenizer, shares it with request counters, and initializes it with the GIL released. Hugging Face can also choose the full encoder per input when added tokens require it

The fast feature provides fast::FastTokenizer from litellm-token-counter-fast. TokenCounter::from_json_fast uses this implementation

The huggingface feature provides huggingface::HuggingFaceTokenizer through the upstream tokenizers library. TokenCounter::from_json uses this implementation

The tiktoken feature provides tiktoken::TiktokenTokenizer through tiktoken-rs. Select an encoding with TokenCounter::from_tiktoken. The supported names are cl100k_base, o200k_base, o200k_harmony, p50k_base, p50k_edit, r50k_base, and gpt2

All three backends are enabled by default in this crate and the Python extension. With default-features = false, Rust callers can supply their own Tokenizer to TokenCounter::new without compiling a built-in backend

Python tiktoken and tokenizers remain runtime dependencies and the default implementations. The catalog independently selects the tokenizer and request-counting routes. Enabling Rust changes factory dispatch; existing tokenizer objects keep their backend. Native Hugging Face wrappers provide an immutable encoding and decoding API, while training and mutable configuration remain available through the Python backend

Budget checks, cost calculation, and the max_tokens adjustment policy belong to litellm-core-utils. The counter does not own prices, budgets, or request limits

Run the feature matrix with:

cargo test -p litellm-token-counter
cargo test -p litellm-token-counter --no-default-features
cargo test -p litellm-token-counter --no-default-features --features fast
cargo test -p litellm-token-counter --no-default-features --features huggingface
cargo test -p litellm-token-counter --no-default-features --features tiktoken