mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-07 02:59:05 +00:00
* add explicit caching to litellm proxy for gemini models via injection
* fix: add missing `supports_function_calling` for deepinfra models
All 55 deepinfra models that had `supports_tool_choice: true` were
missing the `supports_function_calling` flag, causing
`litellm.supports_function_calling()` to incorrectly return False.
Fixes #22619
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* Managed batches - Address PR bot comments from #22464
* feat(togetherai): add support for TogetherAI Qwen3.5-397B-A17B model
* Agent Tracing - support context_id based trace id propogation + nested llm calls (#22626)
* style(ui/): distinguish agent calls from llm calls on ui
* feat: initial grouping working
* feat: set stable contextid for a2a calls - allows for easily passing to downstream llm/mcp calls
* feat(a2a_endpoints.py): fix tracing to avoid recreating logging objects for the same call
allows stable trace id usage
* fix(guardrail_endpoints): handle string ui_type values in _build_field_dict
_build_field_dict unconditionally called .value on ui_type, which crashes
for guardrail configs that use plain strings (e.g. BlockCodeExecutionGuardrailConfigModel
uses "multiselect" and "percentage"). Now checks with hasattr before calling .value.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: propagate trace/session id from headers in MCP server calls
Cherry-picked mcp_server/server.py fixes from 6feb9bab: adds
get_chain_id_from_headers to extract x-litellm-trace-id /
x-litellm-session-id from raw headers, and uses it in call_tool
and list_tools to keep spend logs and tracing consistent with A2A.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* [Feat] UI - Add Open in New Tab on leftnav Bar (#22731)
* Add minimal dev_config.yaml for proxy development
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* feat(ui): wrap left nav items in <a> tags for open-in-new-tab support
Nav items are now rendered as <a> elements with proper href attributes,
enabling right-click → 'Open in new tab', Ctrl/Cmd+click, and
middle-click to open any sidebar page in a new browser tab.
Normal clicks continue to use SPA navigation (no full page reload).
Applied to both leftnav.tsx (query-param routing) and Sidebar2.tsx
(Next.js file-based routing).
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* [Feat] Add Tool Policies for AI Gateway (#22732)
* fix: fix ui render
* fix: fix minor bugs
* refactor: use prisma functions instead of raw sql (safer)
* fix(add-new-tiles-to-tool-policies): allow developer to see what's available
* feat: ensure tool allowlist runs correctly for tool names + mcp's
* refactor: more ui improvements
* feat: working key tool blocking
* feat(tools): show tool logs
* refactor: backend code improvements
* refactor: improve log viewer for tools
* fix: address PR review feedback for tool access control
- Add missing blocked_tools column to root schema.prisma (schema drift)
- Invalidate ToolPolicyRegistry after policy mutations so changes take effect immediately
- Remove dead code: unused get_effective_policies, get_tool_policies_cached, and helpers
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: race condition in permission resolution and remove duplicate allowlist check
- Use atomic update_many with object_permission_id=None to prevent concurrent
requests from creating orphaned permission rows and losing tool blocks
- Remove duplicate allowed_tools enforcement from guardrail (already enforced
in auth layer via check_tools_allowlist)
- Move inline uuid import to module level
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* update to account for userAgent
* UI - Add ToolDetails
* input/output policy
* LiteLLM_PolicyAttachmentTable
* LiteLLM_PolicyAttachmentTable
* fix: add _enqueue_tool_registry_upsert
* fix: tool mgmt endpoints
* tool mgmt endpoints
* Update tests/test_litellm/proxy/db/test_tool_registry_writer.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Update tests/test_litellm/proxy/db/test_tool_registry_writer.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Update tests/test_litellm/proxy/db/test_tool_registry_writer.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix: sync root schema.prisma and fix test_tool_registry_writer for input/output policy
- Migrate root schema.prisma LiteLLM_ToolTable from call_policy to
input_policy/output_policy, add missing user_agent and last_used_at columns
(now consistent with litellm/proxy/schema.prisma and litellm-proxy-extras)
- Fix SpendLogToolIndex comment across all three schema files
- Fix all call_policy references in test_tool_registry_writer.py:
swapped update_tool_policy arguments, wrong get_tools_by_names return type
assertions, _mock_tool_row setting call_policy instead of input_policy
Addresses Greptile review feedback on PR #22732.
Made-with: Cursor
---------
Co-authored-by: Krrish Dholakia <krrishdholakia@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* feat(proxy): add key_alias, key_hash, requested_model DD APM span tags (#22710)
* feat(proxy): add key_alias, key_hash, requested_model tags to DD APM spans
* refactor(proxy): consolidate DD APM tag helpers into DDSpanTagger class
* refactor(proxy): move DDSpanTagger to its own file litellm/proxy/dd_span_tagger.py
---------
Co-authored-by: liweiguang <codingpunk@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Ephrim Stanley <ephrim.stanley@point72.com>
Co-authored-by: Varad Khonde <varadkhonde@gmail.com>
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
117 lines
No EOL
5.4 KiB
Markdown
117 lines
No EOL
5.4 KiB
Markdown
# CLAUDE.md
|
|
|
|
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
|
|
|
## Development Commands
|
|
|
|
### Installation
|
|
- `make install-dev` - Install core development dependencies
|
|
- `make install-proxy-dev` - Install proxy development dependencies with full feature set
|
|
- `make install-test-deps` - Install all test dependencies
|
|
|
|
### Testing
|
|
- `make test` - Run all tests
|
|
- `make test-unit` - Run unit tests (tests/test_litellm) with 4 parallel workers
|
|
- `make test-integration` - Run integration tests (excludes unit tests)
|
|
- `pytest tests/` - Direct pytest execution
|
|
|
|
### Code Quality
|
|
- `make lint` - Run all linting (Ruff, MyPy, Black, circular imports, import safety)
|
|
- `make format` - Apply Black code formatting
|
|
- `make lint-ruff` - Run Ruff linting only
|
|
- `make lint-mypy` - Run MyPy type checking only
|
|
|
|
### Single Test Files
|
|
- `poetry run pytest tests/path/to/test_file.py -v` - Run specific test file
|
|
- `poetry run pytest tests/path/to/test_file.py::test_function -v` - Run specific test
|
|
|
|
### Running Scripts
|
|
- `poetry run python script.py` - Run Python scripts (use for non-test files)
|
|
|
|
### GitHub Issue & PR Templates
|
|
When contributing to the project, use the appropriate templates:
|
|
|
|
**Bug Reports** (`.github/ISSUE_TEMPLATE/bug_report.yml`):
|
|
- Describe what happened vs. what you expected
|
|
- Include relevant log output
|
|
- Specify your LiteLLM version
|
|
|
|
**Feature Requests** (`.github/ISSUE_TEMPLATE/feature_request.yml`):
|
|
- Describe the feature clearly
|
|
- Explain the motivation and use case
|
|
|
|
**Pull Requests** (`.github/pull_request_template.md`):
|
|
- Add at least 1 test in `tests/litellm/`
|
|
- Ensure `make test-unit` passes
|
|
|
|
## Architecture Overview
|
|
|
|
LiteLLM is a unified interface for 100+ LLM providers with two main components:
|
|
|
|
### Core Library (`litellm/`)
|
|
- **Main entry point**: `litellm/main.py` - Contains core completion() function
|
|
- **Provider implementations**: `litellm/llms/` - Each provider has its own subdirectory
|
|
- **Router system**: `litellm/router.py` + `litellm/router_utils/` - Load balancing and fallback logic
|
|
- **Type definitions**: `litellm/types/` - Pydantic models and type hints
|
|
- **Integrations**: `litellm/integrations/` - Third-party observability, caching, logging
|
|
- **Caching**: `litellm/caching/` - Multiple cache backends (Redis, in-memory, S3, etc.)
|
|
|
|
### Proxy Server (`litellm/proxy/`)
|
|
- **Main server**: `proxy_server.py` - FastAPI application
|
|
- **Authentication**: `auth/` - API key management, JWT, OAuth2
|
|
- **Database**: `db/` - Prisma ORM with PostgreSQL/SQLite support
|
|
- **Management endpoints**: `management_endpoints/` - Admin APIs for keys, teams, models
|
|
- **Pass-through endpoints**: `pass_through_endpoints/` - Provider-specific API forwarding
|
|
- **Guardrails**: `guardrails/` - Safety and content filtering hooks
|
|
- **UI Dashboard**: Served from `_experimental/out/` (Next.js build)
|
|
|
|
## Key Patterns
|
|
|
|
### Provider Implementation
|
|
- Providers inherit from base classes in `litellm/llms/base.py`
|
|
- Each provider has transformation functions for input/output formatting
|
|
- Support both sync and async operations
|
|
- Handle streaming responses and function calling
|
|
|
|
### Error Handling
|
|
- Provider-specific exceptions mapped to OpenAI-compatible errors
|
|
- Fallback logic handled by Router system
|
|
- Comprehensive logging through `litellm/_logging.py`
|
|
|
|
### Configuration
|
|
- YAML config files for proxy server (see `proxy/example_config_yaml/`)
|
|
- Environment variables for API keys and settings
|
|
- Database schema managed via Prisma (`proxy/schema.prisma`)
|
|
|
|
## Development Notes
|
|
|
|
### Code Style
|
|
- Uses Black formatter, Ruff linter, MyPy type checker
|
|
- Pydantic v2 for data validation
|
|
- Async/await patterns throughout
|
|
- Type hints required for all public APIs
|
|
- **Avoid imports within methods** — place all imports at the top of the file (module-level). Inline imports inside functions/methods make dependencies harder to trace and hurt readability. The only exception is avoiding circular imports where absolutely necessary.
|
|
|
|
### Testing Strategy
|
|
- Unit tests in `tests/test_litellm/`
|
|
- Integration tests for each provider in `tests/llm_translation/`
|
|
- Proxy tests in `tests/proxy_unit_tests/`
|
|
- Load tests in `tests/load_tests/`
|
|
- **Always add tests when adding new entity types or features** — if the existing test file covers other entity types, add corresponding tests for the new one
|
|
|
|
### UI / Backend Consistency
|
|
- When wiring a new UI entity type to an existing backend endpoint, verify the backend API contract (single value vs. array, required vs. optional params) and ensure the UI controls match — e.g., use a single-select dropdown when the backend accepts a single value, not a multi-select
|
|
|
|
### Database Migrations
|
|
- Prisma handles schema migrations
|
|
- Migration files auto-generated with `prisma migrate dev`
|
|
- Always test migrations against both PostgreSQL and SQLite
|
|
|
|
### Proxy database access
|
|
- **Do not write raw SQL** for proxy DB operations. Use Prisma model methods instead of `execute_raw` / `query_raw`.
|
|
- Use the generated client: `prisma_client.db.<model>` (e.g. `litellm_tooltable`, `litellm_usertable`) with `.upsert()`, `.find_many()`, `.find_unique()`, `.update()`, `.update_many()` as appropriate. This avoids schema/client drift, keeps code testable with simple mocks, and matches patterns used in spend logs and other proxy code.
|
|
|
|
### Enterprise Features
|
|
- Enterprise-specific code in `enterprise/` directory
|
|
- Optional features enabled via environment variables
|
|
- Separate licensing and authentication for enterprise features |