litellm/tests/integration/_support/vertex.py
devin-ai-integration[bot] 6d73fa6b49
fix(vertex_ai): forward system and tools to partner model count_tokens (#43900)
* fix(vertex_ai): forward system and tools to partner model count_tokens

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vertex_ai): avoid mutable token request construction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vertex_ai): return partner count_tokens provider errors as values so the proxy falls back locally

* test(integration): cover Vertex AI partner count_tokens forwarding and local fallbacks

Add wire-level cells for /v1/messages/count_tokens, /utils/token_counter,
/v1/responses/input_tokens and the Gemini countTokens route on a Vertex AI
Claude deployment: the system prompt and tools reach the partner
count-tokens endpoint verbatim, null fields stay out of the body, malformed
tools are rejected before any peer call, peer, token-endpoint and connection
failures fall back to the local tokenizer unless disable_token_counter is
set, generation on the same deployment keeps working, and concurrent bursts
survive a peer outage, a slow peer and a worker SIGKILL. The sdk cells cover
litellm.acount_tokens the same way.

The _support/process.py and _support/client.py harness files are brought to
main's content so the self-booting cells read INTEGRATION_PROXY_READY_SECONDS
instead of a fixed 70 s boot budget.

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 16:36:14 -07:00

31 lines
960 B
Python

from __future__ import annotations
import json
from typing import Final
from cryptography.hazmat.primitives import serialization
from cryptography.hazmat.primitives.asymmetric import rsa
def service_account_json(project: str, token_url: str) -> str:
private_key: Final = (
rsa.generate_private_key(public_exponent=65537, key_size=2048)
.private_bytes(
serialization.Encoding.PEM,
serialization.PrivateFormat.PKCS8,
serialization.NoEncryption(),
)
.decode()
)
return json.dumps(
{
"type": "service_account",
"project_id": project,
"private_key_id": "scripted",
"private_key": private_key,
"client_email": f"scripted@{project}.iam.gserviceaccount.com",
"client_id": "0",
"auth_uri": f"{token_url}/_oauth/authorize",
"token_uri": f"{token_url}/_oauth/token",
}
)