mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-05 02:41:56 +00:00
Closes LIT-4358. DataDogLLMObsLogger.async_send_batch posted the entire log_queue to /api/intake/llm-obs/v1/trace/spans as one payload with no byte or event-count cap and no failure handling. Worse, the base CustomBatchLogger.flush_queue cleared the queue even when the send failed, so any failed flush window silently dropped every span. Port the _send_with_413_split strategy from the log-intake logger: - Proactively split batches exceeding DD_MAX_BATCH_SIZE (1000) events or DD_MAX_PAYLOAD_SIZE_BYTES (4 MB serialised) before POSTing - On a 413 response, halve and retry; drop single un-splittable spans - On transient errors, re-queue undelivered spans for the next flush - Override flush_queue so the base class cannot wipe re-queued spans - Run the proactive size check inside the send loop's try so a serialization failure re-queues only undelivered chunks Adds pytest cases covering proactive count and byte splits, 413 halving, single-span drop, transient and HTTP 500 re-queue, flush_queue preservation, and the no-duplicate guarantee on size check failure. |
||
|---|---|---|
| .. | ||
| test_datadog_cost_management.py | ||
| test_datadog_llm_obs_agent.py | ||
| test_datadog_llm_observability.py | ||
| test_datadog_logger_batching.py | ||
| test_datadog_metrics.py | ||
| test_datadog_tags_regression.py | ||
| test_datadog_team_handler.py | ||