mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-26 01:12:21 +00:00
* feat(anthropic): placement policy for mid-conversation system messages Pure functions over the OpenAI-format message list: split off the leading system run, keep later system messages as role=system at a placement Anthropic accepts on models flagged supports_mid_conversation_system (after a user turn, before an assistant turn or the end, never adjacent), and convert them to user turns in place elsewhere, keeping tool_result first in a merged user turn. * fix(anthropic): keep mid-conversation system out of the chat completions system prompt translate_system_message hoisted every role=system message, at any index, into the top-level system block. On a conversation carrying a mid-session reminder that rewrites the cached prefix, so the provider re-bills the whole history at cache-write pricing on every turn (#36559). #36968 fixed this on /v1/messages; the chat completions path, shared by first-party Anthropic, Vertex, Azure AI and Bedrock Invoke, still hoisted. Only the leading system run becomes the system prompt now. Later system messages go through the placement policy, and anthropic_messages_pt emits a system message instead of rejecting the role. The caller's message list is no longer mutated. Tests pin the two-turn prefix invariant across all four chat configs and both flag states. * refactor(anthropic): single-source the converted system note The /v1/messages pass-through and the chat completions path must prefix a converted system turn with the same operator note. * test(e2e): prove the prompt cache survives a mid-conversation system reminder on chat completions Same priming and assertions as the /v1/messages cases, through /v1/chat/completions with OpenAI-format messages, for first-party Anthropic and Bedrock Invoke on a flagged (Opus 4.8) and an unflagged (Haiku 4.5) model. The reminder sits between the assistant turn and the next user turn, the shape OpenAI-style agent frameworks send, which is the placement the chat path has to translate. * test(anthropic): cover the cache_control rebuild shapes and type the test helpers Codecov flagged the 5m ttl branch and the empty-system path of the wire builder; both now have a test. Greptile asked for full typing on the new test helpers. * refactor(anthropic): read the mid-conversation flag through a public supports_ helper supports_mid_conversation_system joins the other supports_* helpers in litellm.utils, so the chat transformation stops importing the private _supports_factory. * chore(typing): declare the mid-conversation type aliases with TypeAlias The Final sweep tightened LIT010, which exempts TypeAlias declarations but counts a bare alias assignment as an unannotated binding. * fix(anthropic): let add_code_execution_tool take the pass-through message union The translator now emits role=system inside messages for models that accept it, so anthropic_messages_pt returns the pass-through union. add_code_execution_tool still declared the narrower user/assistant union while only ever reading content, so upstream's strip_advisor_blocks_from_messages call in between made the mismatch visible to the type checker. * fix(bedrock): keep mid-conversation system messages in place on converse path * fix: ruff format + multi tool_result order + regression test * fix: satisfy type-discipline gate + update osv ignore for mlflow PYSEC-2026-3865 * fix(bedrock): restore role narrowing in hoisted system loop for basedpyright budget * test(bedrock): cover mid-conversation system conversion branches - non-dict guard in _opens_with_tool_result - in-place conversion without tool context - str/list cache_control preservation in mid-conversation path - drop unreachable non-system guard in hoisted loop * Place type-discipline suppressions on the lines the gate scans * Narrow hoisted loop to system role so basedpyright sees the right TypedDict * fix(anthropic): place mid-conversation system runs by their neighbours only A run after an assistant turn now slides behind the user turn that immediately follows it, and a run that ends the array or precedes an assistant turn becomes a user turn in place. No later message can move an earlier run, so a client that replays the conversation with more turns appended sends a byte-identical prefix and preserved thinking blocks keep their binding * refactor(bedrock): share the converted system note with the anthropic module Converse imports CONVERTED_SYSTEM_NOTE instead of carrying its own copy of the same text, and the reordering helpers lose their comments * test: pin the replayed request prefix across preserved-thinking turns One test per audited feature, through the real entrypoint: the chat transformations for anthropic, bedrock invoke, vertex and converse, the modify_params dummy tool result, dotprompt with unchanged variables, and Presidio masking against an in-process fake. Each serializes system, tools and the earlier messages of turn N and N+1 and asserts they match. The e2e mid-conversation system test imports its content blocks from models.py again and is marked provider_live * fix(anthropic): move mid-conversation system placement into prompt_templates The prompt factory imported the placement helper from the Anthropic provider package, whose common_utils reads a factory constant at import time, so loading the factory first raised ImportError. The module now sits next to anthropic_messages_pt and every consumer imports core utils A user turn with content [] or None puts no block on the wire, so a system run anchored to it landed first in messages or behind an assistant turn. Such a run now converts in place; empty strings and empty text blocks still anchor because the factory fills them with a placeholder * fix(anthropic): anchor system messages only on user turns that reach the wire * fix(bedrock): type the converse system-message helpers over the message TypedDicts * fix(anthropic): read replayed pydantic messages in the Converse helpers and convert a system run whose assistant follower sends nothing A history that replays the previous turn as the litellm.Message object was invisible to the Converse system-message helpers, so a mid-conversation system stayed between a tool call and its result or reached Converse as role: system. The helpers now read fields through the shared message_field and parts_of accessors and drop the local role predicate. Flagged placement anchored a system run on any assistant follower, but anthropic_messages_pt drops an assistant turn that puts no block on the wire (content None, an empty list, an unsigned thinking part), so the system landed directly before the next user turn, which Anthropic rejects. Such a run now converts in place. An empty or whitespace text turn still anchors, since the converter pads it with a placeholder. * fix(anthropic): treat bridged encrypted reasoning as a vanishing assistant turn for system placement An assistant turn whose only blocks carry Responses API encrypted reasoning is dropped by anthropic_messages_pt, so a mid-conversation system run anchored before it landed directly before the next user turn. The unsignable-thinking predicate now lives in common_utils and both the factory and the placement policy consult it. * fix(anthropic): let an inline thinking part hide separate thinking_blocks in system placement anthropic_messages_pt skips an assistant turn's separate thinking_blocks as soon as its content list carries an inline thinking or redacted_thinking part, so a turn whose inline part is unsigned puts nothing on the wire even when the separate block is signed. The placement policy now mirrors that rule. --------- Co-authored-by: Shifat Islam Santo <shifatislamsanto764@gmail.com> Co-authored-by: ege-arhan <egearhany@gmail.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| _support | ||
| agent_tests | ||
| audio_tests | ||
| base_sdk_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| e2e | ||
| guardrails_tests | ||
| image_gen_tests | ||
| integration | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_migration_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| rust-python-harness | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| test_litellm_rust | ||
| unified_google_tests | ||
| unit | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _fake_openai_endpoint_server.py | ||
| _flush_vcr_cache.py | ||
| _live_test_helpers.py | ||
| _openai_record_replay_proxy.py | ||
| _process_helpers.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| _wait_helpers.py | ||
| _ws_vcr.py | ||
| AGENTS.md | ||
| capturing_transport.py | ||
| eval_swe_bench.py | ||
| fake_openai_endpoint.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| pyrightconfig.json | ||
| README.MD | ||
| test_anthropic_compaction_usage.py | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_organizations.py | ||
| test_otel_thread_leak.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_rust_python_harness.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/test_litellm
This folder can only run mock tests.