mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-24 00:52:24 +00:00
* feat(rust): add python-compat crate for Python data formats
Add litellm-python-compat, a PyO3-free crate that reproduces the Python
data formats LiteLLM persists, so Rust readers and writers can interoperate
with state written by the Python proxy:
- literal::literal_eval: a linear recursive-descent port of
ast.literal_eval (prefixes, escapes, implicit concatenation, numeric
underscores and radixes, single unary sign, real +/- complex with 3.14
mixed-mode rules, set(), Python-equality key dedup)
- repr::{repr, to_str}: byte-exact repr()/str(), with a printable table
generated from CPython's str.isprintable (Unicode 16.0.0)
- json::{dumps, from_json, to_json}: json.dumps defaults and the
json.loads mapping
- pickle::{loads, dumps}: plain-data pickles via serde-pickle's serde
interface, which keeps dict insertion order
- truthy::truthy: bool() for plain data
Tests replay fixtures generated by CPython 3.14 (values across every
format and pickle protocol 0-5, plus 154 literal_eval source texts).
Accepted divergences are pinned in a KNOWN table that fails once one
starts matching. A criterion bench covers each format and literal_eval
cost by nesting depth, guarding the linear parse: the py_literal grammar
doubled per nested bracket (105 ms at 16 nested dicts; 19 us at 128 now).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(rust): split python-compat modules and harden the pickle verifier
- Disable class resolution in scripts/verify_rust_pickles.py, and truncate
the export file once instead of removing and appending to it, so the
verifier cannot be pointed at a pre-created file whose rows execute code
through pickle.loads
- Move Error to error.rs and Value to value.rs, leaving lib.rs as the crate
overview, module list and MAX_DEPTH
- Move the generator and verifier to scripts/, beside the Unicode table
generator, leaving tests/ to the Rust tests
- Group the bench by measured surface, give every case a Throughput so
criterion reports bytes per second, and document baseline comparison
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
43 lines
1.4 KiB
Python
43 lines
1.4 KiB
Python
"""Check that CPython unpickles what `pickle::dumps` writes, to the value it was given.
|
|
|
|
PYTHON_COMPAT_RUST_PICKLES=rust.tsv cargo test -p litellm-python-compat --test fixtures
|
|
python scripts/verify_rust_pickles.py rust.tsv
|
|
|
|
The rows are plain data by construction, so this refuses to resolve any class rather than
|
|
handing file-controlled bytes to an unrestricted `pickle.loads`.
|
|
"""
|
|
|
|
import ast
|
|
import io
|
|
import pickle
|
|
import sys
|
|
|
|
|
|
class PlainDataUnpickler(pickle.Unpickler):
|
|
"""An unpickler with `GLOBAL`/`REDUCE` disabled, mirroring `pickle::loads` in Rust."""
|
|
|
|
def find_class(self, module, name):
|
|
raise pickle.UnpicklingError(f"refusing to resolve {module}.{name}")
|
|
|
|
|
|
def loads(data):
|
|
return PlainDataUnpickler(io.BytesIO(data)).load()
|
|
|
|
|
|
def main(path):
|
|
failures = 0
|
|
rows = 0
|
|
with open(path, encoding="utf-8") as lines:
|
|
for line in lines:
|
|
data, expected = line.rstrip("\n").split("\t", 1)
|
|
rows += 1
|
|
actual = repr(loads(bytes.fromhex(data)))
|
|
if actual != repr(ast.literal_eval(expected)):
|
|
failures += 1
|
|
sys.stdout.write(f"mismatch: expected {expected}, got {actual}\n")
|
|
sys.stdout.write(f"{rows} Rust pickles checked, {failures} mismatches\n")
|
|
return 1 if failures or not rows else 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(main(sys.argv[1]))
|