test(core-utils): skip token_counter custom-tokenizer assertion when HF hub is unreachable

test_tokenizers downloads Xenova/llama-3-tokenizer from the HuggingFace
Hub via create_pretrained_tokenizer. On the CI runners the Hub keeps
returning 429 Too Many Requests, which propagated into the blanket
except and turned a third-party rate-limit into a hard pytest.fail. The
same test already skips its llama2 differentiation assertion when the
Hub is unreachable; this extends that exact handling to the custom
tokenizer download so a HuggingFace outage/rate-limit no longer fails
the suite while still failing on real assertion or logic errors.
This commit is contained in:
mateo-berri 2026-06-04 17:02:38 +00:00
parent 3f62063314
commit 2fb11f18da
No known key found for this signature in database

View file

@ -200,7 +200,12 @@ def test_tokenizers():
model="meta-llama/llama-3-70b-instruct", text=sample_text
)
llama3_tokenizer = create_pretrained_tokenizer("Xenova/llama-3-tokenizer")
try:
llama3_tokenizer = create_pretrained_tokenizer("Xenova/llama-3-tokenizer")
except Exception as e:
pytest.skip(
f"custom tokenizer download failed (HF hub unreachable): {e}"
)
llama3_tokens_2 = token_counter(
custom_tokenizer=llama3_tokenizer, text=sample_text
)