- Add comprehensive sync_models_github.md with API endpoints and examples
- Include Loom video tutorial for Admin UI sync process
- Add cross-references from model_management.md, cost_tracking.md, and ui.md
- Provide both manual and automated sync options
- Include Python SDK usage examples
* fix: test case 1, model hits saturation
* fix: _check_rate_limits test case 2
* fix: _get_priority_allocation
* test_default_priority_shared_pool
* fix: No Rate Limiting when low saturatation
* fix: correctly use model_saturation_check
* fixes priority_descriptors
* fix: tune default PriorityReservationSettings
* Optimize cache performance by avoiding expensive operations when caching is disabled
- Moved cache availability checks before expensive operations to improve performance for non-cached requests
- Updated client code to handle None responses from caching handler
* clean hot path
* Fix TypeError with isinstance check for CustomStreamWrapper in caching
Fixed `TypeError: typing.Any cannot be used with isinstance()` that was
occurring in the caching handler when checking cached streaming responses.
The issue was caused by CustomStreamWrapper being aliased to `typing.Any`
at runtime through the TYPE_CHECKING conditional import pattern. When the
code attempted to use isinstance(cached_result, CustomStreamWrapper) at
lines 222 and 338, it failed because Python's isinstance() cannot be used
with typing.Any.
Solution: Import CustomStreamWrapper at runtime separately from the
TYPE_CHECKING block, while keeping a type alias for static type checking.
This allows isinstance checks to work properly while maintaining type hints.
* fix: remove unnecessary type checking