ci(test): revert Railway URL swap for tests that depend on Railway-only response shapes

Independent fact-check found the previous follow-up commit (ba1452acb5)
silently shifted test intent on 3 surfaces — revert just those:

* tests/store_model_in_db_tests/test_callbacks_in_db.py + its proxy
  YAML — the test GETs {host}/langfuse/trace/{id} and asserts a
  {"exists": true} body. The Railway endpoint served that route; the
  in-repo mock doesn't reproduce it. Restoring Railway URL keeps the
  pre-existing behavior of proxy_store_model_in_db_tests until we
  decide whether to model that contract in the mock.

* tests/llm_translation/test_triton.py — asserts
  response.data[0]["embedding"] == [0.1, 0.2], a Railway-specific stub.
  The mock returns a 1536-dim vector at /triton/embeddings, so the
  assertion would always fail in llm_translation_testing.

* tests/load_tests/test_vertex_load_tests.py +
  test_vertex_embeddings_load_test.py — both call Vertex paths
  (/v1/projects/.../publishers/google/models/...:generateContent and
  :predict). Railway returns a Vertex-format response; the mock's
  catch-all returns {"status":"ok"}. Load tests are not in CI but are
  broken as standalone scripts after the migration.

The 8 originally failing CI jobs are unaffected — their test files and
the start_mock_openai_server wiring stay as-is.
This commit is contained in:
Yuneng Jiang 2026-05-19 22:32:31 -07:00
parent ba1452acb5
commit 3ad5a48b91
No known key found for this signature in database
5 changed files with 5 additions and 5 deletions

View file

@ -3,7 +3,7 @@ model_list:
litellm_params:
model: openai/my-fake-model
api_key: my-fake-key
api_base: http://host.docker.internal:8090/
api_base: https://exampleopenaiendpoint-production.up.railway.app/
general_settings:
store_model_in_db: true

View file

@ -360,7 +360,7 @@ async def test_triton_embeddings():
litellm.set_verbose = True
response = await litellm.aembedding(
model="triton/my-triton-model",
api_base="http://127.0.0.1:8090/triton/embeddings",
api_base="https://exampleopenaiendpoint-production.up.railway.app/triton/embeddings",
input=["good morning from litellm"],
)
print(f"response: {response}")

View file

@ -59,7 +59,7 @@ def load_vertex_ai_credentials():
async def create_async_vertex_embedding_task():
load_vertex_ai_credentials()
base_url = "http://127.0.0.1:8090/v1/projects/pathrise-convert-1606954137718/locations/us-central1/publishers/google/models/textembedding-gecko@001"
base_url = "https://exampleopenaiendpoint-production.up.railway.app/v1/projects/pathrise-convert-1606954137718/locations/us-central1/publishers/google/models/textembedding-gecko@001"
embedding_args = {
"model": "vertex_ai/textembedding-gecko",
"input": "This is a test sentence for embedding.",

View file

@ -118,7 +118,7 @@ async def make_async_calls(message_type="text"):
def create_async_task(message_type):
base_url = "http://127.0.0.1:8090/v1/projects/pathrise-convert-1606954137718/locations/us-central1/publishers/google/models/gemini-1.0-pro-vision-001"
base_url = "https://exampleopenaiendpoint-production.up.railway.app/v1/projects/pathrise-convert-1606954137718/locations/us-central1/publishers/google/models/gemini-1.0-pro-vision-001"
if message_type == "text":
messages = [{"role": "user", "content": "hi"}]

View file

@ -21,7 +21,7 @@ from openai.types.chat import ChatCompletion
load_dotenv()
# used for testing
LANGFUSE_BASE_URL = "http://127.0.0.1:8090/slow"
LANGFUSE_BASE_URL = "https://exampleopenaiendpoint-production-c715.up.railway.app"
async def config_update(session, routing_strategy=None):