ReMe/tests/unit/test_outbound_proxy.py
jinliyl 1687179f84
Some checks are pending
Pre-commit / run (ubuntu-latest) (push) Waiting to run
Tests ReMe / Unit Tests - py3.13 (push) Waiting to run
Tests ReMe / Unit Tests - py3.11 (push) Waiting to run
Tests ReMe / Unit Tests - py3.12 (push) Waiting to run
Windows Smoke / CLI smoke - py3.11 (push) Waiting to run
feat: add Auto Fin cookbook and managed outbound proxy support (#392)
* feat: add ssh proxy

* feat: add ssh proxy

* feat: add ssh proxy

* feat: add ssh proxy

* feat: add prompt

* feat: add agent wrapper

* feat: add agent wrapper

* feat: add agent wrapper

* feat: add tushare skill

* feat: add tushare skill

* feat: add tushare skill

* feat: add none stream

* chore(deps): update dependency versions in pyproject.toml

- Bump claude-agent-sdk from 0.2.123 to 0.2.126
- Upgrade pre-commit to version 4.6.1 or higher
- Upgrade pytest to version 9.1.1 or higher

* feat(agent_wrapper): add session compaction support and unify session commands

- Introduce compact_session method to BaseAgentWrapper and implement it in AsAgentWrapper, CcAgentWrapper, and CodexAgentWrapper
- Add session_command module with SessionCommandResult dataclass and handle_session_command function for /clear and /compact commands
- Update __init__.py exports to include session_command handlers
- Modify DingTalkWaitStep to handle session commands via handle_session_command function
- Remove streaming mode from DingTalkWaitStep and simplify reply handling to final Markdown replies only
- Add unit tests for session compaction methods and session command handling across wrappers and DingTalk integration
- Clean up and remove obsolete streaming and card rendering code from DingTalk wait step
- Adjust daily_cookbook.yaml to remove stream and card_update_interval config entries for DingTalk wait step

* feat(auto_fin): add Auto Fin simulated portfolio cookbook workflow

- Add comprehensive Auto Fin schema exports for multiple models and enums
- Implement base class and helpers for Auto Fin analysis steps
- Create file, state, and formatting utilities for Auto Fin with atomic file writes and locking
- Define Auto Fin pipeline with four analysis agents: backtest, event, portfolio, and US correlation
- Register Auto Fin package in cookbook workflows and schema initialization
- Add detailed documentation in markdown describing the system design, workflow, and data contracts

* feat(outbound_proxy): add application-scoped outbound HTTP proxy components

- Introduce BaseOutboundProxy and OutboundProxyEndpoint as core contracts
- Implement FixedHttpOutboundProxy for external HTTP proxy integration
- Add SshHttpOutboundProxy providing SSH-backed local HTTP proxy tunnels
- Register outbound proxy components in component registry and enumeration
- Update components package to include outbound_proxy module
- Add dependency on pproxy for SSH HTTP proxy bridging
- Include comprehensive unit tests covering proxy lifecycle, validation,
  environment merging, error handling, readiness, and monitoring mechanisms

* refactor(network): replace SSH proxy with explicit HTTP outbound proxy

- Remove SSH proxy helper implementation and references in codebase
- Add support for explicit HTTP proxy URL in arXiv and HuggingFace clients
- Modify clients to use async context manager for consistent resource handling
- Update daily paper steps to forward outbound proxy configuration explicitly
- Change tests to cover new proxy usage model and remove SSH proxy mocks
- Add outbound proxy component configuration in daily_cookbook.yaml
- Ensure proxy URL usage disables environment trust in HTTP clients
- Fix app context component enum access to be defensive against missing keys

* feat(agent_wrapper): add managed proxy support for command environments

- Introduce BaseOutboundProxy binding in BaseAgentWrapper for outbound proxy management
- Add bash_environment and command_proxy_environment properties to apply proxy settings
- Update WorkspaceBackend instantiation in AsAgentWrapper to use bash_environment
- Inject managed proxy export commands into Claude Code Bash commands via hooks
- Enhance CodexAgentWrapper to include managed proxy in shell environment policy
- Modify daily_cookbook.yaml steps to specify outbound_proxy as default where needed
- Add comprehensive unit tests verifying managed proxy injection and environment isolation
- Ensure subprocess_environment remains unchanged while proxy is applied selectively to commands

* refactor(memory): replace search job_tools with memory in daily cookbook config

- Change workspace_dir default from .reme to reme_workspace
- Replace search job_tools with memory across multiple components and jobs
- Update descriptions to reflect long-term memory retrieval instead of search
- Modify system prompts to instruct using memory for retrieving notes
- Adjust unit tests to verify memory job_tools and job presence instead of search
- Ensure consistency in configuration and tests for memory backend usage

* refactor(config): rename memory to memory_search in daily cookbook config

- Change all occurrences of "memory" to "memory_search" in job_tools and job definitions
- Update related system prompts to reflect the new memory_search terminology
- Modify unit tests to assert the presence of memory_search instead of memory
- Ensure consistency across skills, job tools, and backend configurations in multiple components

* feat(auto_fin): add deterministic quantitative research and ranking fusion

- Introduce new schema models: EtfScore, RankingMetrics, ExtremeAnalysis,
  DimensionRanking, and FusionRanking to represent deterministic research outputs
- Add ranking data to event, backtest, us_correlation, and portfolio analysis outputs
- Implement ranking_section renderer to format Top20 scores and diagnostics in Markdown
- Develop AutoFinQuantStep for deterministic ETF ranking using TuShare data, Polars,
  and a custom extremely randomized tree ensemble
- Integrate quantitative rankings into backtest and portfolio analysis steps and reports
- Extend auto_fin pipeline with new quant_enabled and quant_required config options
- Enforce ranking constraints like unique codes, contiguous ranks, and normalized fusion weights
- Update analysis YAMLs with rules limiting data freshness, universe, and ranking usage
- Incorporate ranking outputs into all major markdown report bodies in Auto Fin pipeline
- Add concurrency-limited asynchronous TuShare client to fetch required market data
- Introduce cross-sectional rank correlation and NDCG metrics for ranking quality evaluation

* feat(auto_fin): implement stage-wise notification and reporting for analysis pipeline

- Refactor notification config in daily_cookbook.yaml to support dispatch steps
- Update AutoFinNotificationStep to deduplicate notifications per run stage
- Add _notify_stage method in pipeline to send notifications for each analysis stage
- Implement persistence and notification for event, backtest, US correlation, and portfolio stages
- Modify pipeline flow to persist reports and notify after each stage completion
- Adjust metadata to track notifications and errors per stage
- Update tests to verify stage-wise notification sending and deduplication
- Remove older combined report persistence in favor of modular stage handling

* feat(auto_fin): add outbound proxy support for Tushare API usage

- Introduce BaseOutboundProxy reference in AutoFinPipelineStep and AutoFinQuantStep
- Update TushareResearchClient and trade calendar fetch to accept and use proxy URL
- Create _ProxiedTushareApi adapter to route Tushare requests via explicit HTTP proxy
- Modify create_tushare_api utility to optionally return proxied API client
- Add unit tests covering proxy forwarding and client behavior with managed proxies
- Ensure proxy usage respects explicit proxy URL over environment fallback
- Integrate outbound proxy into data fetching and quantitative research steps

* feat(auto_fin): enforce checkpoint time validation and add state models

- Introduce AnalysisState base class and specific states for event, backtest, and US correlation analyses
- Replace analysis output types with corresponding state classes in run schemas
- Add require_checkpoint_reached method to validate decision_at/data_cutoff against current time
- Enforce checkpoint time checks before analysis steps in event, backtest, portfolio, and quant analyses
- Refactor quant data loading to include adjustment factors and apply price adjustments without fallback
- Update analysis YAML docs to require real-time checkpoint validation and forbid using future data
- Improve portfolio run serialization by excluding redundant legacy fields and nested proposed actions
- Add helper to extract readable sections from persisted checkpoint documents
- Fix event analysis output validation to reject events and sources with future timestamps

* feat(auto_fin): auto-select latest reached checkpoint if none specified

- Extend checkpoint config to accept empty string for auto selection
- Add static method to compute latest checkpoint reached by current time
- Modify pipeline step to auto-select checkpoint based on trade calendar and time
- Adjust force flag default depending on whether checkpoint is explicit or auto
- Log details when checkpoint is auto-selected to improve observability
- Add comprehensive tests for auto checkpoint selection logic and edge cases
- Remove deprecated default and required constraints from force parameter in config

* refactor(auto_fin): unify datetime comparison with compare_datetimes utility

- Replace direct datetime comparisons with compare_datetimes function calls
- Use cmp_to_key with compare_datetimes for sorting datetime tuples and lists
- Update validation logic in backtest, event, analysis, and ledger modules for consistent datetime handling
- Add unit tests to verify handling of naive and aware datetime comparisons in event and backtest validations
- Ensure marked_at and interval_end timestamps are set and compared consistently using compare_datetimes
- Improve correctness of ordering and conditional checks related to timestamps throughout auto_fin steps and ledger code

* feat(auto_fin): add datetime comparison helper for mixed timezone data

- Implement compare_datetimes function to handle naive and aware datetimes
- Ensure naive datetime is interpreted in the known timezone of the counterpart
- Facilitate comparisons between legacy and timezone-aware Auto Fin data
- Add module docstring explaining purpose of the helpers

* docs(auto_fin): enforce unique ETF representative per sub-theme in analysis rules

- Update backtest.yaml to recommend or highlight only one ETF per sub-theme for ETF analyses
- Modify event.yaml to map only one representative ETF per sub-theme, avoiding duplicate recommendations
- Revise portfolio.yaml to restrict holdings/buys to a single ETF per sub-theme, preventing repeated buys of highly overlapping ETFs
- Adjust us_correlation.yaml to retain only one representative A-share ETF per sub-theme for mapping or recommendation
- Add test to verify presence of new sub-theme uniqueness guidance in step prompts

* feat(auto_fin): separate draft model and include deterministic fusion ranking

- Introduce _PortfolioProposalDraft pydantic model for agent-authored fields before ranking
- Discard any "fusion_ranking" data from draft to prevent conflicts with canonical ranking
- Modify AutoFinPortfolioStep to receive draft, enrich with fusion_ranking, and produce final output
- Update tests to use _PortfolioProposalDraft and validate deterministic fusion ranking propagation
- Add async test verifying fusion ranking is correctly set in portfolio output with no errors

* refactor(auto_fin): rewrite and simplify Auto Fin schema and steps

- Remove legacy Auto Fin analysis step modules and helpers
- Replace complex ranking and portfolio models with simplified current-news models
- Update schema to focus on news-case workflow with new domain models
- Remove A-share decision checkpoints and backtest details from schema
- Simplify recommendation and decision output structures
- Clean up deprecated state and utility functions
- Update Auto Fin steps initialization to new pipeline steps only
- Improve uniqueness validation for themes and ETFs in research plan

* feat(auto_fin): implement full local cache and analysis workflow for Auto Fin

- Add AutoFinDataStep to prepare and cache daily TuShare data with lookback
- Add AutoFinAnalysisStep to analyze cached data and generate Markdown report
- Implement detailed time window, ETF filtering, and historical case validation
- Introduce YAML prompts for planning and decision-making steps
- Update .gitignore to include reme_workspace/
- Clean up config and import structure for auto_fin steps
- Remove old pipeline.py and consolidate functionality into new modules
- Use polars for efficient CSV reading and data processing
- Ensure atomic writes and strict JSON serialization for cache files
- Enforce rules on news timing, ETF universe, and historical case usage

* fix(auto_fin): restrict news data source to '财联社' in analysis and cache

- Update analysis templates to specify current news as from '财联社' only
- Modify news fetching functions to filter by source '财联社'
- Add validation method to check cached news source correctness
- Update news caching logic to exclude non-'财联社' news
- Enhance unit tests with multiple sources to ensure filtering works
- Confirm news API calls include source filter parameter as '财联社'

* refactor(auto_fin): convert I/O methods to asynchronous implementations

- Change _news, _dataset, and _theme_data methods to async for improved concurrency
- Move JSONL and CSV reading operations to asynchronous wrappers using asyncio.to_thread
- Remove synchronous _read_jsonl and _read_csv functions, integrate them as static async class methods
- Update cache validation methods to async, awaiting I/O operations accordingly
- Adjust usage of dataset and news retrieval in analysis step to await asynchronous methods
- Add async unit test to validate JSONL reading with unicode line separators
- Preserve existing functionality while enabling non-blocking file and data access

* fix(nx_file_graph): defer networkx import and improve dependency handling

- Move networkx import inside NxFileGraph constructor for lazy loading
- Raise ImportError with original exception context if networkx is missing
- Remove module-level fallback assignment of nx to None
- Expand test to block loading of multiple optional core dependencies eagerly
- Change exception type in test from ModuleNotFoundError to AssertionError
- Update test comments to reflect broader optional dependency checks

* feat(embedding_store): add quota retry delay mechanism for embedding requests

- Introduce quota_retry_delay parameter to configure wait time before retry on quota exhaustion
- Implement detection of insufficient quota errors in LocalEmbeddingStore without external SDK
- Add retry logic with custom delay when quota is insufficient during embedding requests
- Update configuration to set max_retries and quota_retry_delay defaults for embedding store
- Add unit tests covering quota exhaustion retry behavior with delay and opt-in control
- Ensure existing retry behavior remains unchanged if quota_retry_delay is not set

* feat(auto_fin): add detailed logging to analysis and data fetching steps

- Add _preview static method for bounded diagnostic output in analysis.py
- Log prompt start, completion, errors, and validation details in _reply method
- Add info logs for major processing steps in execute method of analysis.py
- Add debug and info logs for cache validation, data fetching, and pagination in data.py
- Log conditions for skipping reports and cache plans in data.py execute method
- Log download summaries and cache writes for news and ETF data
- Improve error logging with exception details in cache validation functions
- Ensure all logs include context such as record counts, paths, and parameters

* refactor(auto_fin): overhaul Auto Fin workflow and schema contracts

- Replace old Auto Fin schema models with comprehensive new data classes
- Remove legacy Auto Fin analysis step in favor of modular agent-based steps
- Introduce AutoFinAgentStep for validating structured agent replies
- Simplify data cleaning and JSONL writing utilities for news cache
- Remove synchronous and asynchronous dataset methods from analysis step
- Redefine Auto Fin analysis configuration for 360-day news retention and multi-step pipeline
- Remove embedded analysis prompt templates and replace with agent-driven logic
- Update __init__.py exports to match new step implementations and remove deprecated classes
- Improve error handling and validation in agent step reply processing
- Clean up redundant imports and unused code in analysis and data preparation modules

* feat(auto_fin): add detailed logging for analysis and data processing steps

- Add timing logs to measure agent prompt processing duration in analysis.py
- Log news cache hits and news write paths with record counts in data.py
- Include detailed info logs for news download start and completion in data.py
- Add start, progress, and completion logs with topic and event counts in history.py
- Log start and completion of merge step including path and ETF count in merge.py
- Add start and done logs with window and news counts in topic.py

* feat(auto_fin): enhance schema and steps with detailed ETF and event modeling

- Replace and add multiple AutoFin schema classes to support detailed ETF selection,
  historical research, market analysis, forecast models, and report output with validation
- Implement Shanghai timezone normalization and strict validation in schema models
- Remove deprecated AutoFin analysis agent step and consolidate reply handling in base step
- Introduce AutoFinStep base class with shared helpers for prompt handling, data fetching,
  logging, and JSONL file operations
- Add AutoFinDataStep to manage daily news data complete with schedule validation, caching,
  and source validation logic
- Update cookbook configuration to customize auto_fin step parameters and simplify
  outbound proxy settings
- Refactor imports and clean unused code for better maintainability

* feat(auto_fin): introduce detailed historical event resolution and market similarity analysis

- Add AutoFinHistoricalEventReference and AutoFinHistoricalSimilarity models for refined event referencing and similarity judgment
- Implement validation to ensure non-empty critical fields and uniqueness of historical news IDs
- Develop method to resolve Agent-selected historical event references from workspace files with strict path and existence checks
- Enrich historical events with market entry and future returns data after resolution
- Redesign market step to calculate similarity-weighted ETF forecasts based on matched historical event similarities
- Enforce validation on matched historical events for uniqueness and proper weight summation
- Simplify merge step output to final Markdown report without YAML frontmatter and redundant fields
- Update user instructions for history search, market, and merge steps to reflect new data structures and responsibilities
- Adjust test suite to cover new schema and step behavior changes, including enhanced validation and JSON output formats

* feat(auto_fin): add new cron jobs and output analysis jsonl

- Add new cron jobs auto_fin_1145_cron and auto_fin_1800_cron with auto_fin_steps
- Change auto_fin_0930_cron schedule to run Monday to Sunday
- Extend merge step to write analysis data to auto_fin_analysis.jsonl
- Update unit tests to verify new cron jobs and their steps configuration

* fix(auto_fin): improve atomic file write and refresh daily index

- Change temporary file naming to include UUID for uniqueness and hidden prefix
- Replace atomic write method from using Path.replace to os.replace with safe unlink
- Add import and use os.replace for safer file replace operation
- Refresh daily index after writing auto finance markdown and JSONL files
- Import and call refresh_day_index in merge step to update file index asynchronously

* docs(cookbook): add optional SSH proxy configuration in README files

- Introduce optional SSH proxy setup in auto-fin and daily_paper cookbooks
- Provide instructions to enable outbound proxy via `daily_cookbook.yaml` and environment variables
- Add `REME_PROXY_IP` and `REME_PROXY_ACCOUNT` environment variables descriptions in multiple README files
- Update English and Chinese README and README_ZH documents with proxy details
- Maintain consistent formatting of environment variable tables across documents

* fix(file_io): include schema_version in hidden metadata keys

- Added "schema_version" to _INDEX_HIDDEN_METADATA_KEYS in _daily_index.py
- Updated _render_notes_block to always include additional keys regardless of schema_version

fix(deps): move pproxy dependency to later in pyproject.toml

- Removed pproxy from early dependencies list
- Added pproxy back near the end of dependency list for better ordering

fix(outbound_proxy): require pproxy package for ssh_http proxy

- Added importlib.util check for pproxy package presence
- Raise RuntimeError if pproxy is not installed when using SSH HTTP outbound proxy
- Improved error message suggests installing reme-ai with 'core' extra

* docs(readme): update News section with new Cookbook workflows

- Clarify introduction of optional Cookbooks with Daily Paper and Auto Fin workflows
- Update English README to reflect both paper discovery and file-native ETF event research
- Revise Chinese README to include financial news and historical market data research capability
- Maintain announcement of paper acceptance at Findings of ACL 2026

* feat(auto_fin): add calculation results to final Markdown output

- Implement _calculation_results to summarize forecast for each ETF analyzed
- Include program-calculated results in the JSON input for the Markdown report
- Update YAML template to incorporate calculation results and adjust recommendation rules
- Refine recommendation logic to rely on event impact judgments combined with calculation outputs
- Modify tests to verify presence of calculation results and updated report content and format

* up prompt

* fix(keyword_index): ignore non-indexable chunks during keyword sync

- Add is_indexable method to base and BM25 keyword index classes to check text tokenizability
- Update local file store to exclude non-indexable chunks from expected document IDs to prevent rebuild
- Fix JSONL chunker to correctly handle Unicode line separator U+2028 inside JSON strings without splitting
- Add test to ensure non-empty but non-indexable chunk does not trigger keyword index rebuild
- Add test to verify U+2028 character does not cause incorrect JSONL record splitting
2026-07-25 18:09:39 +08:00

341 lines
12 KiB
Python

"""Tests for application-scoped outbound proxy components."""
# pylint: disable=missing-function-docstring,protected-access
import asyncio
import os
import sys
from collections.abc import Callable
import pytest
from reme.application import Application
from reme.components import R
from reme.components.outbound_proxy import FixedHttpOutboundProxy, SshHttpOutboundProxy
from reme.enumeration import ComponentEnum
class FakeProcess:
"""Minimal asyncio subprocess stand-in with observable shutdown."""
def __init__(self, label: str, events: list[str]) -> None:
self.label = label
self.events = events
self.returncode: int | None = None
self.stdout = asyncio.StreamReader()
self.stderr = asyncio.StreamReader()
self._finished = asyncio.Event()
async def wait(self) -> int:
await self._finished.wait()
assert self.returncode is not None
return self.returncode
def terminate(self) -> None:
self.events.append(f"terminate:{self.label}")
self.exit(-15)
def kill(self) -> None:
self.events.append(f"kill:{self.label}")
self.exit(-9)
def exit(self, returncode: int, output: str = "") -> None:
if self.returncode is not None:
return
self.returncode = returncode
encoded = output.encode()
self.stdout.feed_data(encoded)
self.stdout.feed_eof()
self.stderr.feed_data(encoded)
self.stderr.feed_eof()
self._finished.set()
async def _ready_listener(*_args) -> None:
return None
async def _wait_until(predicate: Callable[[], bool], timeout: float = 1.0) -> None:
loop = asyncio.get_running_loop()
deadline = loop.time() + timeout
while not predicate():
if loop.time() >= deadline:
raise AssertionError("condition did not become true")
await asyncio.sleep(0.005)
def test_outbound_proxy_backends_are_registered() -> None:
assert R.get(ComponentEnum.OUTBOUND_PROXY, "fixed_http") is FixedHttpOutboundProxy
assert R.get(ComponentEnum.OUTBOUND_PROXY, "ssh_http") is SshHttpOutboundProxy
@pytest.mark.asyncio
async def test_application_builds_and_manages_fixed_http_proxy(tmp_path) -> None:
app = Application(
workspace_dir=str(tmp_path),
enable_logo=False,
log_to_console=False,
log_to_file=False,
service={"backend": "cli"},
components={
"outbound_proxy": {
"default": {
"backend": "fixed_http",
"url": "http://127.0.0.1:18080",
},
},
},
)
component = app.context.components[ComponentEnum.OUTBOUND_PROXY]["default"]
assert isinstance(component, FixedHttpOutboundProxy)
await app.start()
assert component.http_url == "http://127.0.0.1:18080"
await app.close()
with pytest.raises(RuntimeError, match="start the component"):
_ = component.endpoint
@pytest.mark.asyncio
async def test_fixed_http_publishes_endpoint_and_merges_environment(monkeypatch) -> None:
monkeypatch.setenv("HTTP_PROXY", "http://ambient.example:8080")
component = FixedHttpOutboundProxy(url="http://127.0.0.1:18080")
base = {"CUSTOM": "value", "NO_PROXY": "example.com"}
with pytest.raises(RuntimeError, match="start the component"):
_ = component.http_url
await component.start()
merged = component.merge_environment(base)
assert component.http_url == "http://127.0.0.1:18080"
assert base == {"CUSTOM": "value", "NO_PROXY": "example.com"}
assert os.environ["HTTP_PROXY"] == "http://ambient.example:8080"
assert merged["CUSTOM"] == "value"
for key in ("HTTP_PROXY", "HTTPS_PROXY", "ALL_PROXY", "http_proxy", "https_proxy", "all_proxy"):
assert merged[key] == component.http_url
assert merged["NO_PROXY"] == "127.0.0.1,localhost,::1"
assert merged["no_proxy"] == "127.0.0.1,localhost,::1"
await component.close()
with pytest.raises(RuntimeError, match="start the component"):
_ = component.endpoint
@pytest.mark.parametrize(
("url", "message"),
[
("", "must use http"),
("https://proxy.example:8080", "must use http"),
("http://proxy.example", "include host and port"),
("http://user@proxy.example:8080", "must not contain userinfo"),
("http://proxy.example:8080?mode=x", "must not contain query or fragment"),
("http://proxy.example:8080#fragment", "must not contain query or fragment"),
("http://proxy.example:not-a-port", "malformed"),
],
)
@pytest.mark.asyncio
async def test_fixed_http_rejects_invalid_urls(url: str, message: str) -> None:
component = FixedHttpOutboundProxy(url=url)
with pytest.raises(ValueError, match=message):
await component.start()
assert component.is_started is False
with pytest.raises(RuntimeError, match="start the component"):
_ = component.endpoint
@pytest.mark.parametrize(
("kwargs", "message"),
[
({"host": "", "account": "agent"}, "host is required"),
({"host": "proxy.example", "account": ""}, "account is required"),
({"host": "proxy.example", "account": "agent", "connect_timeout": 0}, "connect_timeout"),
({"host": "proxy.example", "account": "agent", "monitor_interval": -1}, "monitor_interval"),
({"host": "proxy.example", "account": "agent", "restart_initial_delay": "bad"}, "restart_initial_delay"),
({"host": "proxy.example", "account": "agent", "restart_max_delay": float("inf")}, "restart_max_delay"),
],
)
@pytest.mark.asyncio
async def test_ssh_http_validates_configuration(kwargs: dict, message: str) -> None:
component = SshHttpOutboundProxy(**kwargs)
with pytest.raises(ValueError, match=message):
await component.start()
@pytest.mark.asyncio
async def test_ssh_http_requires_ssh_executable(monkeypatch) -> None:
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.shutil.which", lambda _name: None)
component = SshHttpOutboundProxy(host="proxy.example", account="agent")
with pytest.raises(RuntimeError, match="ssh executable was not found"):
await component.start()
@pytest.mark.asyncio
async def test_ssh_http_starts_expected_commands_and_closes_bridge_first(monkeypatch) -> None:
events: list[str] = []
commands: list[tuple[str, ...]] = []
processes: list[FakeProcess] = []
async def fake_spawn(*command, **kwargs):
assert kwargs
commands.append(command)
label = "ssh" if command[0] == "/usr/bin/ssh" else "bridge"
process = FakeProcess(label, events)
processes.append(process)
return process
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.shutil.which", lambda _name: "/usr/bin/ssh")
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.asyncio.create_subprocess_exec", fake_spawn)
component = SshHttpOutboundProxy(host="proxy.example", account="agent")
monkeypatch.setattr(component, "_pick_distinct_ports", lambda: (43123, 43124))
monkeypatch.setattr(component, "_wait_for_listener", _ready_listener)
await component.start()
assert component.http_url == "http://127.0.0.1:43124"
assert commands[0] == (
"/usr/bin/ssh",
"-N",
"-D",
"127.0.0.1:43123",
"-o",
"BatchMode=yes",
"-o",
"ExitOnForwardFailure=yes",
"-o",
"StrictHostKeyChecking=accept-new",
"-o",
"ConnectTimeout=10",
"-o",
"LogLevel=ERROR",
"--",
"agent@proxy.example",
)
assert commands[1] == (
sys.executable,
"-m",
"pproxy",
"-l",
"http://127.0.0.1:43124",
"-r",
"socks5://127.0.0.1:43123",
)
await component.close()
assert events == ["terminate:bridge", "terminate:ssh"]
assert all(process.returncode == -15 for process in processes)
@pytest.mark.asyncio
async def test_ssh_http_cleans_up_when_initial_readiness_fails(monkeypatch) -> None:
events: list[str] = []
process = FakeProcess("ssh", events)
async def fake_spawn(*_command, **_kwargs):
return process
async def fail_readiness(*_args):
raise TimeoutError("not ready")
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.shutil.which", lambda _name: "/usr/bin/ssh")
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.asyncio.create_subprocess_exec", fake_spawn)
component = SshHttpOutboundProxy(host="proxy.example", account="agent")
monkeypatch.setattr(component, "_pick_distinct_ports", lambda: (43123, 43124))
monkeypatch.setattr(component, "_wait_for_listener", fail_readiness)
with pytest.raises(RuntimeError, match="SSH proxy exited before readiness"):
await component.start()
assert events == ["terminate:ssh"]
assert component.is_started is False
with pytest.raises(RuntimeError, match="start the component"):
_ = component.endpoint
@pytest.mark.asyncio
async def test_ssh_http_reselects_ports_only_before_endpoint_is_published(monkeypatch) -> None:
events: list[str] = []
commands: list[tuple[str, ...]] = []
selected_ports = iter(((43123, 43124), (43125, 43126)))
async def fake_spawn(*command, **_kwargs):
commands.append(command)
label = "ssh" if command[0] == "/usr/bin/ssh" else "bridge"
process = FakeProcess(label, events)
if label == "ssh" and len(commands) == 1:
process.exit(255, "bind [127.0.0.1]:43123: Address already in use")
return process
async def fake_readiness(process, *_args):
if process.returncode is not None:
raise RuntimeError("process exited")
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.shutil.which", lambda _name: "/usr/bin/ssh")
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.asyncio.create_subprocess_exec", fake_spawn)
component = SshHttpOutboundProxy(host="proxy.example", account="agent")
monkeypatch.setattr(component, "_pick_distinct_ports", lambda: next(selected_ports))
monkeypatch.setattr(component, "_wait_for_listener", fake_readiness)
await component.start()
assert component.http_url == "http://127.0.0.1:43126"
assert [command[0] for command in commands] == ["/usr/bin/ssh", "/usr/bin/ssh", sys.executable]
await component.close()
@pytest.mark.asyncio
async def test_ssh_http_monitor_restarts_on_original_ports(monkeypatch) -> None:
events: list[str] = []
commands: list[tuple[str, ...]] = []
processes: list[FakeProcess] = []
async def fake_spawn(*command, **_kwargs):
commands.append(command)
label = "ssh" if command[0] == "/usr/bin/ssh" else "bridge"
process = FakeProcess(label, events)
processes.append(process)
return process
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.shutil.which", lambda _name: "/usr/bin/ssh")
monkeypatch.setattr("reme.components.outbound_proxy.ssh_http.asyncio.create_subprocess_exec", fake_spawn)
component = SshHttpOutboundProxy(
host="proxy.example",
account="agent",
monitor_interval=0.005,
restart_initial_delay=0.005,
)
monkeypatch.setattr(component, "_pick_distinct_ports", lambda: (43123, 43124))
monkeypatch.setattr(component, "_wait_for_listener", _ready_listener)
await component.start()
endpoint = component.endpoint
processes[0].exit(7)
await _wait_until(lambda: len(processes) == 3)
assert component.endpoint is endpoint
assert commands[2][0] == "/usr/bin/ssh"
assert "127.0.0.1:43123" in commands[2]
assert sum(command[0] == sys.executable for command in commands) == 1
processes[1].exit(8)
await _wait_until(lambda: len(processes) == 4)
assert component.endpoint is endpoint
assert commands[3] == (
sys.executable,
"-m",
"pproxy",
"-l",
"http://127.0.0.1:43124",
"-r",
"socks5://127.0.0.1:43123",
)
await component.close()