Stabilises CodSpeed measurements of the LLM-completion benchmarks by
removing GC-induced noise from the per-iteration instruction count.
CPython's cyclic collector fires on its own clock and, because the
multi-turn benchmark only allocates a few KB per iteration, a collection
that lands mid-iteration inflates the per-iteration count by tens of
percent — exactly the magnitude of the flake that caused #32136's
test_completion_multi_turn to be flagged as a -25% regression.
The existing ``inline_logging_executor`` fixture already proved the
pattern works: deferring asynchronous executor work to a per-iteration
inline call removes background-thread scheduling noise. GC is the same
class of artefact — non-deterministic, runs orthogonally to the code
under test — and gets the same treatment.
The deferred collection runs once at session teardown; ``mock_response``
keeps the benchmarks on synthetic allocations so nothing escapes into
real tracing.
Verified locally: the multi-turn benchmark's standard deviation drops
from ~0.37 ms to ~0.001 ms across 20 × 1000-iteration runs, i.e. the
GC-attributable variance is now ~370× smaller.
Guard the per-request CPU cost of the chat completion, MCP tool and A2A
message transforms against regressions on every commit. All benchmarks are
pure in-process work with no network I/O so they stay deterministic under
CodSpeed's simulation mode, and they import under the base dependency set the
benchmark job installs.
Inference covers the full SDK overhead via mock_response (simple, multi-turn,
tools, streaming) plus convert_to_model_response_object as a deterministic
anchor. MCP covers the client-side tool translation and the proxy server-side
tool-name prefix round-trip. A2A covers the client request/response transforms
and the proxy server-ingress message conversion.
Adds the mcp and a2a-sdk packages to the benchmark run since those transform
modules need them, and broadens the workflow triggers to litellm_internal_staging
so the internal branch flow is benchmarked too.