chore(config): reduce RPS limits for AI model requests

- Lowered max_rps from 20 to 2 to prevent rate limiting issues
- Reduced rps_window from 10 to 1 for stricter request throttling
- Updated default configuration values for better performance stability
This commit is contained in:
jinli.yl 2026-01-14 00:47:28 +08:00
parent de6e3968a3
commit e64667ba77

View file

@ -23,9 +23,9 @@ llm:
backend: openai
model_name: qwen3-max
# temperature: 0.6
max_rps: 20
max_rps: 2
# max_rps: 9
rps_window: 10
rps_window: 1
embedding_model:
default: