From 67b2f3d037bd84d733cce789103570997492cde3 Mon Sep 17 00:00:00 2001 From: Actual Operator Date: Tue, 12 May 2026 21:56:09 +0000 Subject: [PATCH] Sync context files with ADRs - Update CLAUDE.md (claude) - Update .claude/rules/fcd6a21d-756b-4d3f-98d7-f42d21122afd.md (claude) - Update .claude/rules/da815559-806f-4816-ac62-f2bcba187047.md (claude) - Update .claude/rules/860946d2-f81f-45ed-b3f2-25c8d5f4956f.md (claude) - Update .claude/rules/8743e5ed-c257-4d49-8295-a4481f8523b5.md (claude) - Update .claude/rules/67d628ca-c9de-4922-93cb-1fd54450f492.md (claude) - Update .claude/rules/a4786ef4-01e0-4b65-9c9f-0c3175726c86.md (claude) - Update .claude/rules/b1a53926-9804-45b7-9543-634d35c15734.md (claude) - Update .claude/rules/4dbd1df3-2cdc-4996-a874-d94ee89c9983.md (claude) - Update .claude/rules/c81e5425-fcc1-4177-9246-bd167e4322af.md (claude) - Update .claude/rules/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76.md (claude) - Update .claude/rules/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7.md (claude) - Update .claude/rules/372c20f8-a9ae-4489-b3a3-4825d3343af4.md (claude) - Update .claude/rules/19fbda95-be0a-43ce-88bd-b2cbb1c4075a.md (claude) - Update .claude/rules/36c17245-9a45-4855-affd-3e23712797d8.md (claude) - Update .claude/rules/1267d2fe-4cfa-4c45-9004-ada782af479d.md (claude) - Update .claude/rules/57fb44e9-c959-408b-9e6a-87b058b1ecc7.md (claude) - Update .claude/rules/0ba59d63-b106-4111-a75e-cb18a23f7c56.md (claude) - Update .claude/rules/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498.md (claude) - Update .claude/rules/f5680bc5-1992-44e9-92a3-65214350b7a7.md (claude) - Update .claude/rules/e5731ea0-1ffc-4e87-b5f7-07551952f609.md (claude) - Update .claude/rules/5b2054c7-ba4e-44f2-b4ee-faacb67305d6.md (claude) - Update .claude/rules/f1ba5186-840f-4423-9876-d7cb8ec75614.md (claude) - Update AGENTS.md (agents) - Update docs/adr/fcd6a21d-756b-4d3f-98d7-f42d21122afd-standardize-service-boundary-logging-for-observability-service-implementations-include.md (docs) - Update docs/adr/da815559-806f-4816-ac62-f2bcba187047-standardize-service-boundary-logging-for-observability-service-boundary-logs.md (docs) - Update docs/adr/860946d2-f81f-45ed-b3f2-25c8d5f4956f-standardize-service-boundary-logging-for-observability-log-levels-service.md (docs) - Update docs/adr/8743e5ed-c257-4d49-8295-a4481f8523b5-standardize-service-boundary-logging-for-observability-service-boundary-logging.md (docs) - Update docs/adr/67d628ca-c9de-4922-93cb-1fd54450f492-standardize-service-boundary-logging-for-observability-service-boundary-logs.md (docs) - Update docs/adr/a4786ef4-01e0-4b65-9c9f-0c3175726c86-standardize-service-boundary-logging-for-observability-service-boundary-exit.md (docs) - Update docs/adr/b1a53926-9804-45b7-9543-634d35c15734-standardize-service-boundary-logging-for-observability-service-boundary-entry.md (docs) - Update docs/adr/4dbd1df3-2cdc-4996-a874-d94ee89c9983-adopt-distributed-tracing-for-public-api-observability-internal-private-use.md (docs) - Update docs/adr/c81e5425-fcc1-4177-9246-bd167e4322af-adopt-distributed-tracing-for-public-api-observability-implement-sampling-strategies.md (docs) - Update docs/adr/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76-adopt-distributed-tracing-for-public-api-observability-tracing-instrumentation-not.md (docs) - Update docs/adr/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7-adopt-distributed-tracing-for-public-api-observability-trace-spans-include.md (docs) - Update docs/adr/372c20f8-a9ae-4489-b3a3-4825d3343af4-adopt-distributed-tracing-for-public-api-observability-implementations-create-child.md (docs) - Update docs/adr/19fbda95-be0a-43ce-88bd-b2cbb1c4075a-adopt-distributed-tracing-for-public-api-observability-each-trace-span.md (docs) - Update docs/adr/36c17245-9a45-4855-affd-3e23712797d8-adopt-distributed-tracing-for-public-api-observability-tracing-implementations-propagate.md (docs) - Update docs/adr/1267d2fe-4cfa-4c45-9004-ada782af479d-adopt-distributed-tracing-for-public-api-observability-public-external-endpoints.md (docs) - Update docs/adr/57fb44e9-c959-408b-9e6a-87b058b1ecc7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-support.md (docs) - Update docs/adr/0ba59d63-b106-4111-a75e-cb18a23f7c56-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-properly.md (docs) - Update docs/adr/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-connections-implement.md (docs) - Update docs/adr/f5680bc5-1992-44e9-92a3-65214350b7a7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-event-streams.md (docs) - Update docs/adr/e5731ea0-1ffc-4e87-b5f7-07551952f609-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-that.md (docs) - Update docs/adr/5b2054c7-ba4e-44f2-b4ee-faacb67305d6-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-utilize.md (docs) - Update docs/adr/f1ba5186-840f-4423-9876-d7cb8ec75614-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-runtime-state-updates.md (docs) --- .../0ba59d63-b106-4111-a75e-cb18a23f7c56.md | 32 +++++ .../0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498.md | 32 +++++ .../1267d2fe-4cfa-4c45-9004-ada782af479d.md | 53 +++++++ .../19fbda95-be0a-43ce-88bd-b2cbb1c4075a.md | 55 ++++++++ .../36c17245-9a45-4855-affd-3e23712797d8.md | 58 ++++++++ .../372c20f8-a9ae-4489-b3a3-4825d3343af4.md | 46 ++++++ .../4dbd1df3-2cdc-4996-a874-d94ee89c9983.md | 42 ++++++ .../57fb44e9-c959-408b-9e6a-87b058b1ecc7.md | 40 ++++++ .../5b2054c7-ba4e-44f2-b4ee-faacb67305d6.md | 38 +++++ .../67d628ca-c9de-4922-93cb-1fd54450f492.md | 60 ++++++++ .../860946d2-f81f-45ed-b3f2-25c8d5f4956f.md | 60 ++++++++ .../8743e5ed-c257-4d49-8295-a4481f8523b5.md | 61 ++++++++ .../9607dcb3-91d3-44b9-85f9-ccb1fa25bd76.md | 55 ++++++++ .../a4786ef4-01e0-4b65-9c9f-0c3175726c86.md | 62 +++++++++ .../b1a53926-9804-45b7-9543-634d35c15734.md | 62 +++++++++ .../c81e5425-fcc1-4177-9246-bd167e4322af.md | 55 ++++++++ .../da815559-806f-4816-ac62-f2bcba187047.md | 61 ++++++++ .../e5731ea0-1ffc-4e87-b5f7-07551952f609.md | 44 ++++++ .../f1ba5186-840f-4423-9876-d7cb8ec75614.md | 32 +++++ .../f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7.md | 55 ++++++++ .../f5680bc5-1992-44e9-92a3-65214350b7a7.md | 31 +++++ .../fcd6a21d-756b-4d3f-98d7-f42d21122afd.md | 63 +++++++++ AGENTS.md | 38 +++++ CLAUDE.md | 40 +++++- ...e-state-updates-route-handlers-properly.md | 115 +++++++++++++++ ...state-updates-sse-connections-implement.md | 115 +++++++++++++++ ...observability-public-external-endpoints.md | 123 ++++++++++++++++ ...ublic-api-observability-each-trace-span.md | 123 ++++++++++++++++ ...ility-tracing-implementations-propagate.md | 123 ++++++++++++++++ ...ervability-implementations-create-child.md | 123 ++++++++++++++++ ...-api-observability-internal-private-use.md | 123 ++++++++++++++++ ...ate-updates-sse-implementations-support.md | 115 +++++++++++++++ ...ate-updates-sse-implementations-utilize.md | 115 +++++++++++++++ ...for-observability-service-boundary-logs.md | 131 ++++++++++++++++++ ...ng-for-observability-log-levels-service.md | 131 ++++++++++++++++++ ...-observability-service-boundary-logging.md | 131 ++++++++++++++++++ ...servability-tracing-instrumentation-not.md | 123 ++++++++++++++++ ...for-observability-service-boundary-exit.md | 131 ++++++++++++++++++ ...or-observability-service-boundary-entry.md | 131 ++++++++++++++++++ ...rvability-implement-sampling-strategies.md | 123 ++++++++++++++++ ...for-observability-service-boundary-logs.md | 131 ++++++++++++++++++ ...ntime-state-updates-route-handlers-that.md | 115 +++++++++++++++ ...ime-state-updates-runtime-state-updates.md | 115 +++++++++++++++ ...c-api-observability-trace-spans-include.md | 123 ++++++++++++++++ ...runtime-state-updates-sse-event-streams.md | 115 +++++++++++++++ ...ability-service-implementations-include.md | 131 ++++++++++++++++++ 46 files changed, 3880 insertions(+), 1 deletion(-) create mode 100644 .claude/rules/0ba59d63-b106-4111-a75e-cb18a23f7c56.md create mode 100644 .claude/rules/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498.md create mode 100644 .claude/rules/1267d2fe-4cfa-4c45-9004-ada782af479d.md create mode 100644 .claude/rules/19fbda95-be0a-43ce-88bd-b2cbb1c4075a.md create mode 100644 .claude/rules/36c17245-9a45-4855-affd-3e23712797d8.md create mode 100644 .claude/rules/372c20f8-a9ae-4489-b3a3-4825d3343af4.md create mode 100644 .claude/rules/4dbd1df3-2cdc-4996-a874-d94ee89c9983.md create mode 100644 .claude/rules/57fb44e9-c959-408b-9e6a-87b058b1ecc7.md create mode 100644 .claude/rules/5b2054c7-ba4e-44f2-b4ee-faacb67305d6.md create mode 100644 .claude/rules/67d628ca-c9de-4922-93cb-1fd54450f492.md create mode 100644 .claude/rules/860946d2-f81f-45ed-b3f2-25c8d5f4956f.md create mode 100644 .claude/rules/8743e5ed-c257-4d49-8295-a4481f8523b5.md create mode 100644 .claude/rules/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76.md create mode 100644 .claude/rules/a4786ef4-01e0-4b65-9c9f-0c3175726c86.md create mode 100644 .claude/rules/b1a53926-9804-45b7-9543-634d35c15734.md create mode 100644 .claude/rules/c81e5425-fcc1-4177-9246-bd167e4322af.md create mode 100644 .claude/rules/da815559-806f-4816-ac62-f2bcba187047.md create mode 100644 .claude/rules/e5731ea0-1ffc-4e87-b5f7-07551952f609.md create mode 100644 .claude/rules/f1ba5186-840f-4423-9876-d7cb8ec75614.md create mode 100644 .claude/rules/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7.md create mode 100644 .claude/rules/f5680bc5-1992-44e9-92a3-65214350b7a7.md create mode 100644 .claude/rules/fcd6a21d-756b-4d3f-98d7-f42d21122afd.md mode change 120000 => 100644 CLAUDE.md create mode 100644 docs/adr/0ba59d63-b106-4111-a75e-cb18a23f7c56-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-properly.md create mode 100644 docs/adr/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-connections-implement.md create mode 100644 docs/adr/1267d2fe-4cfa-4c45-9004-ada782af479d-adopt-distributed-tracing-for-public-api-observability-public-external-endpoints.md create mode 100644 docs/adr/19fbda95-be0a-43ce-88bd-b2cbb1c4075a-adopt-distributed-tracing-for-public-api-observability-each-trace-span.md create mode 100644 docs/adr/36c17245-9a45-4855-affd-3e23712797d8-adopt-distributed-tracing-for-public-api-observability-tracing-implementations-propagate.md create mode 100644 docs/adr/372c20f8-a9ae-4489-b3a3-4825d3343af4-adopt-distributed-tracing-for-public-api-observability-implementations-create-child.md create mode 100644 docs/adr/4dbd1df3-2cdc-4996-a874-d94ee89c9983-adopt-distributed-tracing-for-public-api-observability-internal-private-use.md create mode 100644 docs/adr/57fb44e9-c959-408b-9e6a-87b058b1ecc7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-support.md create mode 100644 docs/adr/5b2054c7-ba4e-44f2-b4ee-faacb67305d6-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-utilize.md create mode 100644 docs/adr/67d628ca-c9de-4922-93cb-1fd54450f492-standardize-service-boundary-logging-for-observability-service-boundary-logs.md create mode 100644 docs/adr/860946d2-f81f-45ed-b3f2-25c8d5f4956f-standardize-service-boundary-logging-for-observability-log-levels-service.md create mode 100644 docs/adr/8743e5ed-c257-4d49-8295-a4481f8523b5-standardize-service-boundary-logging-for-observability-service-boundary-logging.md create mode 100644 docs/adr/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76-adopt-distributed-tracing-for-public-api-observability-tracing-instrumentation-not.md create mode 100644 docs/adr/a4786ef4-01e0-4b65-9c9f-0c3175726c86-standardize-service-boundary-logging-for-observability-service-boundary-exit.md create mode 100644 docs/adr/b1a53926-9804-45b7-9543-634d35c15734-standardize-service-boundary-logging-for-observability-service-boundary-entry.md create mode 100644 docs/adr/c81e5425-fcc1-4177-9246-bd167e4322af-adopt-distributed-tracing-for-public-api-observability-implement-sampling-strategies.md create mode 100644 docs/adr/da815559-806f-4816-ac62-f2bcba187047-standardize-service-boundary-logging-for-observability-service-boundary-logs.md create mode 100644 docs/adr/e5731ea0-1ffc-4e87-b5f7-07551952f609-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-that.md create mode 100644 docs/adr/f1ba5186-840f-4423-9876-d7cb8ec75614-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-runtime-state-updates.md create mode 100644 docs/adr/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7-adopt-distributed-tracing-for-public-api-observability-trace-spans-include.md create mode 100644 docs/adr/f5680bc5-1992-44e9-92a3-65214350b7a7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-event-streams.md create mode 100644 docs/adr/fcd6a21d-756b-4d3f-98d7-f42d21122afd-standardize-service-boundary-logging-for-observability-service-implementations-include.md diff --git a/.claude/rules/0ba59d63-b106-4111-a75e-cb18a23f7c56.md b/.claude/rules/0ba59d63-b106-4111-a75e-cb18a23f7c56.md new file mode 100644 index 000000000..79eac4c44 --- /dev/null +++ b/.claude/rules/0ba59d63-b106-4111-a75e-cb18a23f7c56.md @@ -0,0 +1,32 @@ + +These rules are ALWAYS ACTIVE for all files matching route handlers and SSE implementations that expose runtime execution progress, configuration management endpoints, and real-time monitoring interfaces. + + +### Rules + +- **R-SSE-001** MUST: Route handlers MUST properly close SSE connections and clean up resources when clients disconnect or execution completes. + +### Verify + +```bash +# Verify SSE usage across route handlers +grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l + +# Verify centralized SSE utility imports +grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules + +# Verify SSE utility module exists +find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" +``` + +**Accept when:** +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion +- Route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Structured event payloads use consistent format (e.g., {type: string, data: object, timestamp: number}) +- Connection lifecycle logging and metrics are implemented to monitor SSE health + + +Claude Code MUST NOT skip or defer verification of SSE connection lifecycle management and resource cleanup. All route handlers with SSE connections missing proper cleanup must be flagged before merge. + \ No newline at end of file diff --git a/.claude/rules/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498.md b/.claude/rules/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498.md new file mode 100644 index 000000000..aab871cae --- /dev/null +++ b/.claude/rules/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498.md @@ -0,0 +1,32 @@ + +These rules are ALWAYS ACTIVE for all files implementing real-time configuration and runtime state updates using Server-Sent Events. + + +### Rules + +- **R-SSE-001** SHOULD: SSE connections SHOULD implement heartbeat or keep-alive mechanisms to detect and handle connection failures + +### Verify + +```bash +# Verify SSE usage across route handlers +grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l + +# Verify SSE utility imports in route handlers +grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules + +# Verify centralized SSE utility module exists +find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" +``` + +**Accept when:** +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion +- SSE connections implement heartbeat or keep-alive mechanisms to detect and handle connection failures +- Appropriate headers are set: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Structured event payloads use consistent format (e.g., {type: string, data: object, timestamp: number}) + + +Claude Code MUST NOT skip or defer verification of SSE implementation patterns. All new real-time endpoints must be verified against these rules before acceptance. + \ No newline at end of file diff --git a/.claude/rules/1267d2fe-4cfa-4c45-9004-ada782af479d.md b/.claude/rules/1267d2fe-4cfa-4c45-9004-ada782af479d.md new file mode 100644 index 000000000..12fd8bd84 --- /dev/null +++ b/.claude/rules/1267d2fe-4cfa-4c45-9004-ada782af479d.md @@ -0,0 +1,53 @@ + +These rules are ALWAYS ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + + +### Rules + +- **R-TRACING-001** MUST: All public/external API endpoints MUST instrument entry and exit points with distributed tracing spans. + +**In scope:** +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +**Out of scope:** +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +**Exceptions:** +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +### Verify + +```bash +# Count tracing instrumentation patterns +grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l + +# Search for obs.tracing facet and OpenTelemetry patterns +grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' + +# Verify tracing in test suite +cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' +``` + +**Accept when:** +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +### Implementation Guidance + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + + +Claude Code MUST NOT skip or defer verification. All public API endpoints MUST be instrumented with distributed tracing spans. Violations are caught by automated CI pipeline checks, code review checklists, and integration test validation. Quarterly audits identify non-compliant APIs with mandatory remediation within 30 days. + \ No newline at end of file diff --git a/.claude/rules/19fbda95-be0a-43ce-88bd-b2cbb1c4075a.md b/.claude/rules/19fbda95-be0a-43ce-88bd-b2cbb1c4075a.md new file mode 100644 index 000000000..849113d23 --- /dev/null +++ b/.claude/rules/19fbda95-be0a-43ce-88bd-b2cbb1c4075a.md @@ -0,0 +1,55 @@ + +These rules are ALWAYS ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + + +### Rules + +- **R-TRACING-001** MUST: Each trace span MUST include standardized attributes: service.name, operation.name, http.method, http.status_code, and error.type (if applicable). + +### Scope + +**In scope:** +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +**Out of scope:** +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +**Exceptions:** +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +### Verify + +```bash +# Count tracing instrumentation patterns +grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l + +# Search for obs.tracing facet and OpenTelemetry patterns +grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' + +# Run tracing-specific tests +cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' +``` + +**Accept when:** +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +### Implementation Guidance + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + + +Claude Code MUST NOT skip or defer verification. Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis. Code review checklist includes verification of distributed tracing implementation. Integration test suite validates trace context propagation and span attribute completeness. Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation. Code review blocks merge if tracing requirements are not met without documented exception. Quarterly audits identify non-compliant APIs with remediation plans required within 30 days. Exception process requires submission to architecture review board with performance benchmarks or deprecation timeline, documentation in API specification and architectural decision log, and quarterly re-evaluation. + \ No newline at end of file diff --git a/.claude/rules/36c17245-9a45-4855-affd-3e23712797d8.md b/.claude/rules/36c17245-9a45-4855-affd-3e23712797d8.md new file mode 100644 index 000000000..b88f62383 --- /dev/null +++ b/.claude/rules/36c17245-9a45-4855-affd-3e23712797d8.md @@ -0,0 +1,58 @@ + +These rules are ALWAYS ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + + +### Rules + +- **R-TRACING-001** MUST: Tracing implementations MUST propagate trace context across service boundaries using W3C Trace Context or OpenTelemetry standards. + +### Scope + +**In scope:** +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +**Out of scope:** +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +**Exceptions:** +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +### Verify + +```bash +# Check for tracing instrumentation patterns +grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l + +# Verify obs.tracing facet and OpenTelemetry usage +grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' + +# Run tracing-related tests +cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' +``` + +**Accept when:** +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy +- Trace IDs are included in API error responses (e.g., X-Trace-Id header) for external consumer reference + +### Implementation Guidance + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues +- Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests +- Establish retention policies (e.g., 7 days for all traces, 30 days for error traces) + + +Claude Code MUST NOT skip or defer verification. All public API endpoints MUST implement distributed tracing instrumentation. Violations are caught by automated CI pipeline checks, code review checklists, and integration test validation. Exceptions require architecture review board approval with documented performance benchmarks or deprecation timelines. + \ No newline at end of file diff --git a/.claude/rules/372c20f8-a9ae-4489-b3a3-4825d3343af4.md b/.claude/rules/372c20f8-a9ae-4489-b3a3-4825d3343af4.md new file mode 100644 index 000000000..dcfb22eb4 --- /dev/null +++ b/.claude/rules/372c20f8-a9ae-4489-b3a3-4825d3343af4.md @@ -0,0 +1,46 @@ + +These rules are ALWAYS ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + + +### Rules + +- **R-TRACING-001** SHOULD: API implementations SHOULD create child spans for significant internal operations (database queries, external service calls, computation-heavy operations). + +**In scope:** +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +**Out of scope:** +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +**Exceptions:** +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +### Verify + +```bash +# Check for tracing instrumentation patterns +grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l + +# Verify obs.tracing facet and OpenTelemetry usage +grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' + +# Run tracing-related tests +cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' +``` + +**Accept when:** +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + + +Claude Code MUST NOT skip or defer verification. Verification is mandatory via automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis, code review checklist verification, and integration test suite validation of trace context propagation. + \ No newline at end of file diff --git a/.claude/rules/4dbd1df3-2cdc-4996-a874-d94ee89c9983.md b/.claude/rules/4dbd1df3-2cdc-4996-a874-d94ee89c9983.md new file mode 100644 index 000000000..1206cfd7f --- /dev/null +++ b/.claude/rules/4dbd1df3-2cdc-4996-a874-d94ee89c9983.md @@ -0,0 +1,42 @@ + +These rules are ALWAYS ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + + +### Rules + +- **R-TRACING-001** MUST: Implement distributed tracing instrumentation on all HTTP/REST API endpoints exposed to external consumers. +- **R-TRACING-002** MUST: Implement distributed tracing instrumentation on all GraphQL APIs and gRPC services with external clients. +- **R-TRACING-003** MUST: Implement distributed tracing instrumentation on all webhook handlers and callback endpoints. +- **R-TRACING-004** MUST: Implement distributed tracing instrumentation on all public SDK methods that initiate cross-service operations. +- **R-TRACING-005** MUST: Implement distributed tracing instrumentation on all integration adapters and third-party service connectors. +- **R-TRACING-006** MUST: Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues. +- **R-TRACING-007** SHOULD: Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support. +- **R-TRACING-008** SHOULD: Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes. +- **R-TRACING-009** SHOULD: Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues). +- **R-TRACING-010** MAY: Internal/private APIs MAY use simplified tracing instrumentation with reduced attribute sets. +- **R-TRACING-EXC-001** EXCEPTION: Performance-critical hot paths where tracing overhead exceeds 5% of request latency may be exempted with documented benchmarks. +- **R-TRACING-EXC-002** EXCEPTION: Legacy APIs scheduled for deprecation within 6 months may be exempted with documented deprecation timeline. + +### Verify + +```bash +# Count tracing instrumentation patterns in codebase +grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l + +# Search for obs.tracing facet and OpenTelemetry patterns +grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' + +# Verify tracing in test suite +cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' +``` + +**Accept when:** +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy +- Trace context is included in error responses for external consumer reference +- Shared tracing libraries or middleware are in place for common instrumentation patterns + + +Claude Code MUST verify distributed tracing implementation on all public API endpoints. Verification is mandatory via automated CI pipeline checks, code review checklist, and integration test suite. Violations result in CI pipeline failure, code review blocks, and quarterly audit remediation requirements. Exceptions require architecture review board approval with documented performance benchmarks or deprecation timelines. + \ No newline at end of file diff --git a/.claude/rules/57fb44e9-c959-408b-9e6a-87b058b1ecc7.md b/.claude/rules/57fb44e9-c959-408b-9e6a-87b058b1ecc7.md new file mode 100644 index 000000000..b6ba34ae6 --- /dev/null +++ b/.claude/rules/57fb44e9-c959-408b-9e6a-87b058b1ecc7.md @@ -0,0 +1,40 @@ + +These rules are ALWAYS ACTIVE for all files matching web route handlers, configuration management endpoints, and real-time monitoring interfaces that require push-based state synchronization. + + +### Rules + +- **R-SSE-001** MAY: SSE implementations MAY support event filtering or subscription patterns to reduce bandwidth for clients interested in specific event types. +- **R-SSE-002** MUST: All route handlers exposing runtime state use SSE with text/event-stream Content-Type header. +- **R-SSE-003** MUST: A centralized SSE utility module exists and is imported by route handlers requiring real-time updates. +- **R-SSE-004** MUST: SSE connections implement proper cleanup on client disconnect or execution completion. +- **R-SSE-005** MUST: Route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive. +- **R-SSE-006** SHOULD: Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}). +- **R-SSE-007** SHOULD: Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks. +- **R-SSE-008** SHOULD: Implement periodic heartbeat messages to prevent connection termination by intermediate proxies. +- **R-SSE-009** SHOULD: Multiplex multiple event types over a single SSE connection where possible to respect browser connection limits. + +### Verify + +```bash +# Count SSE implementations using text/event-stream header +grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l + +# Verify centralized SSE utility imports +grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules + +# Locate centralized SSE utility module +find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" +``` + +**Accept when:** +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion +- Route handlers set appropriate headers (Content-Type, Cache-Control, Connection) +- Structured event payloads follow consistent format with type, data, and timestamp fields +- Connection lifecycle logging and metrics are implemented for monitoring + + +Claude Code MUST NOT skip or defer verification. All SSE implementations MUST be reviewed against these rules before acceptance. Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable. Violations detected in code review must be flagged and require revision before merge. + \ No newline at end of file diff --git a/.claude/rules/5b2054c7-ba4e-44f2-b4ee-faacb67305d6.md b/.claude/rules/5b2054c7-ba4e-44f2-b4ee-faacb67305d6.md new file mode 100644 index 000000000..524dc8202 --- /dev/null +++ b/.claude/rules/5b2054c7-ba4e-44f2-b4ee-faacb67305d6.md @@ -0,0 +1,38 @@ + +These rules are ALWAYS ACTIVE for all files matching `**/*`. SSE implementations MUST utilize a centralized library module to ensure consistent event formatting and connection management across web route handlers, configuration endpoints, and real-time monitoring interfaces. + + +### Rules + +- **R-SSE-001** MUST: SSE implementations MUST utilize a centralized library module (e.g., sse.ts) to ensure consistent event formatting and connection management +- **R-SSE-002** MUST: Route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- **R-SSE-003** MUST: Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- **R-SSE-004** MUST: Implement robust connection lifecycle management with cleanup handlers on client disconnect or execution completion +- **R-SSE-005** SHOULD: Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- **R-SSE-006** SHOULD: Implement periodic heartbeat messages to prevent connection termination by intermediate proxies or load balancers +- **R-SSE-007** SHOULD: Multiplex multiple event types over a single SSE connection where possible to work within browser connection limits +- **R-SSE-008** SHOULD: Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +### Verify + +```bash +# Count SSE implementations using text/event-stream header +grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l + +# Verify centralized SSE module imports in route handlers +grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules + +# Locate centralized SSE utility module +find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" +``` + +**Accept when:** +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion +- Structured event payloads follow consistent format with type, data, and timestamp fields +- Connection lifecycle logging is present for monitoring SSE health + + +Claude Code MUST NOT skip or defer verification. All SSE implementations MUST be reviewed against R-SSE-001 through R-SSE-008. Violations must be flagged and require revision before acceptance. Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable. Route handlers with SSE connections missing proper cleanup must be fixed before merge. + \ No newline at end of file diff --git a/.claude/rules/67d628ca-c9de-4922-93cb-1fd54450f492.md b/.claude/rules/67d628ca-c9de-4922-93cb-1fd54450f492.md new file mode 100644 index 000000000..b6351380b --- /dev/null +++ b/.claude/rules/67d628ca-c9de-4922-93cb-1fd54450f492.md @@ -0,0 +1,60 @@ + +These rules are ALWAYS ACTIVE for all service implementations and apply to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + + +### Rules + +- **R-SBL-001** MUST: Service boundary logs MUST include correlation identifiers (request IDs, trace IDs) to enable distributed tracing across service calls. + +### Scope + +**In scope:** +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +**Out of scope:** +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +**Exceptions:** +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +### Verify + +```bash +# Verify boundary logging patterns in service files +grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l + +# Verify logging statements in handlers and queries +rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' + +# Verify correlation identifiers in service boundaries +find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' +``` + +**Accept when:** +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +### Implementation Guidance + +- Create shared logging utilities or macros that encapsulate boundary logging patterns +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters +- Establish standard log fields: timestamp, service_name, operation, request_id, duration_ms, status, error_message +- Configure log aggregation systems to parse and index boundary logs +- Document logging patterns in service templates and starter kits + + +Claude Code MUST verify service boundary logging compliance using the provided bash commands. Verification is mandatory before approving changes to service handlers, API endpoints, or query interfaces. Violations require code review feedback and addition of boundary logging before merge approval. + \ No newline at end of file diff --git a/.claude/rules/860946d2-f81f-45ed-b3f2-25c8d5f4956f.md b/.claude/rules/860946d2-f81f-45ed-b3f2-25c8d5f4956f.md new file mode 100644 index 000000000..23eeb23bb --- /dev/null +++ b/.claude/rules/860946d2-f81f-45ed-b3f2-25c8d5f4956f.md @@ -0,0 +1,60 @@ + +These rules are ALWAYS ACTIVE for all service implementations and apply to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + + +### Rules + +- **R-SBL-001** SHOULD: Log levels at service boundaries SHOULD follow standard conventions: INFO for successful operations, WARN for degraded operations, ERROR for failures. + +### Scope + +**In scope:** +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +**Out of scope:** +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +**Exceptions:** +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +### Verify + +```bash +# Verify boundary logging patterns in service files +grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l + +# Verify logging statements in handlers and queries +rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' + +# Verify correlation identifiers in boundary logs +find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' +``` + +**Accept when:** +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +### Implementation Guidance + +- Create shared logging utilities or macros that encapsulate boundary logging patterns +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters +- Establish standard log fields: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems to parse and index boundary logs +- Document logging patterns in service templates and starter kits + + +Claude Code MUST verify boundary logging compliance using the provided bash commands. Verification is mandatory before accepting service boundary implementations. Violations require code review feedback and addition of boundary logging before merge approval. + \ No newline at end of file diff --git a/.claude/rules/8743e5ed-c257-4d49-8295-a4481f8523b5.md b/.claude/rules/8743e5ed-c257-4d49-8295-a4481f8523b5.md new file mode 100644 index 000000000..4b15948ef --- /dev/null +++ b/.claude/rules/8743e5ed-c257-4d49-8295-a4481f8523b5.md @@ -0,0 +1,61 @@ + +These rules are ALWAYS ACTIVE for all service implementations and apply to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + + +### Rules + +- **R-SBL-001** SHOULD: Service boundary logging SHOULD use structured logging formats (JSON, key-value pairs) rather than unstructured text to enable automated log analysis. + +### In Scope + +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +### Out of Scope + +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +### Exceptions + +- **EXC-001**: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- **EXC-002**: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +### Verify + +```bash +# Verify boundary logging patterns in service files +grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l + +# Verify structured logging in handlers +rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' + +# Verify correlation identifiers in boundary logs +find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' +``` + +**Accept when:** +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +### Implementation Guidance + +- Create shared logging utilities or macros that encapsulate boundary logging patterns +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters +- Establish standard log fields: timestamp, service_name, operation, request_id, duration_ms, status, error_message +- Use async logging frameworks to minimize performance impact +- Implement log sampling for high-frequency endpoints + + +Claude Code MUST verify boundary logging compliance using the provided bash commands. Verification is mandatory before accepting service boundary implementations. Violations require code review feedback and remediation before merge approval. + \ No newline at end of file diff --git a/.claude/rules/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76.md b/.claude/rules/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76.md new file mode 100644 index 000000000..71ce64872 --- /dev/null +++ b/.claude/rules/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76.md @@ -0,0 +1,55 @@ + +These rules are ALWAYS ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + + +### Rules + +- **R-OBS-TRACING-001** MUST_NOT: Tracing instrumentation MUST NOT log sensitive data (passwords, tokens, PII) in span attributes or events. + +### Scope + +**In scope:** +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +**Out of scope:** +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +**Exceptions:** +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +### Verify + +```bash +# Check for tracing instrumentation patterns +grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l + +# Verify obs.tracing facet and OpenTelemetry usage +grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' + +# Run tracing-related tests +cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' +``` + +**Accept when:** +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +### Implementation Guidance + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + + +Claude Code MUST verify distributed tracing compliance for all public API endpoints. Verification is mandatory and MUST NOT be skipped or deferred. Code review MUST check for tracing instrumentation and MUST block merge if requirements are not met without documented exception. CI pipeline MUST fail if new public API endpoints lack tracing instrumentation. + \ No newline at end of file diff --git a/.claude/rules/a4786ef4-01e0-4b65-9c9f-0c3175726c86.md b/.claude/rules/a4786ef4-01e0-4b65-9c9f-0c3175726c86.md new file mode 100644 index 000000000..1b9aa0a14 --- /dev/null +++ b/.claude/rules/a4786ef4-01e0-4b65-9c9f-0c3175726c86.md @@ -0,0 +1,62 @@ + +These rules are ALWAYS ACTIVE for all service implementations and apply to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + + +### Rules + +- **R-SBL-001** MUST: All service boundary exit points MUST log operation completion status (success/failure), duration, and relevant result metadata. + +### In Scope + +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +### Out of Scope + +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +### Exceptions + +- **EXC-001**: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- **EXC-002**: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +### Verify + +```bash +# Verify boundary logging patterns in service files +grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l + +# Verify logging statements in handlers and queries +rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' + +# Verify correlation identifiers in boundary logs +find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' +``` + +**Accept when:** +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs +- Standard log fields are present: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) + +### Implementation Guidance + +- Create shared logging utilities or macros that encapsulate boundary logging patterns +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters +- Establish standard log fields: timestamp, service_name, operation, request_id, duration_ms, status, error_message +- Configure log aggregation systems to parse and index boundary logs +- Document logging patterns in service templates and starter kits + + +Claude Code MUST verify all service boundary exit points contain required logging before approving changes. Verification is mandatory and cannot be deferred. Use automated checks and manual review to confirm compliance with R-SBL-001. + \ No newline at end of file diff --git a/.claude/rules/b1a53926-9804-45b7-9543-634d35c15734.md b/.claude/rules/b1a53926-9804-45b7-9543-634d35c15734.md new file mode 100644 index 000000000..5288cd706 --- /dev/null +++ b/.claude/rules/b1a53926-9804-45b7-9543-634d35c15734.md @@ -0,0 +1,62 @@ + +These rules are ALWAYS ACTIVE for all service implementations and apply to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + + +### Rules + +- **R-SBL-001** MUST: All service boundary entry points (API handlers, RPC endpoints, query interfaces) MUST log incoming requests with sufficient context to identify the operation, caller, and key parameters. + +### In Scope + +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +### Out of Scope + +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +### Exceptions + +- **EXC-001**: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- **EXC-002**: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +### Verify + +```bash +# Verify boundary logging patterns in Rust and TypeScript service files +grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l + +# Verify logging statements in handler and query files +rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' + +# Verify presence of correlation identifiers in service boundary files +find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' +``` + +**Accept when:** +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs +- Standard log fields are present: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) + +### Implementation Guidance + +- Create shared logging utilities or macros that encapsulate boundary logging patterns +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters +- Establish a standard set of log fields for service boundaries: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems to parse and index boundary logs for efficient querying and alerting +- Document logging patterns in service templates and starter kits + + +Claude Code MUST verify boundary logging compliance using the provided bash commands before approving service boundary implementations. Verification is mandatory and cannot be deferred. Violations require code review feedback and addition of boundary logging before merge approval. + \ No newline at end of file diff --git a/.claude/rules/c81e5425-fcc1-4177-9246-bd167e4322af.md b/.claude/rules/c81e5425-fcc1-4177-9246-bd167e4322af.md new file mode 100644 index 000000000..fcfe1c33c --- /dev/null +++ b/.claude/rules/c81e5425-fcc1-4177-9246-bd167e4322af.md @@ -0,0 +1,55 @@ + +These rules are ALWAYS ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + + +### Rules + +- **R-TRACING-001** SHOULD: APIs SHOULD implement sampling strategies to balance observability needs with performance overhead (e.g., 100% for errors, 10% for successful requests). + +### Scope + +**In scope:** +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +**Out of scope:** +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +**Exceptions:** +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +### Verify + +```bash +# Count tracing instrumentation patterns +grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l + +# Search for obs.tracing facet and OpenTelemetry patterns +grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' + +# Run tracing-related tests +cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' +``` + +**Accept when:** +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +### Implementation Guidance + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + + +Clause Code MUST verify distributed tracing implementation for all public API endpoints. Verification is mandatory and MUST NOT be skipped or deferred. CI pipeline MUST fail if new public API endpoints lack tracing instrumentation. Code review MUST block merge if tracing requirements are not met without documented exception. + \ No newline at end of file diff --git a/.claude/rules/da815559-806f-4816-ac62-f2bcba187047.md b/.claude/rules/da815559-806f-4816-ac62-f2bcba187047.md new file mode 100644 index 000000000..b99726b3d --- /dev/null +++ b/.claude/rules/da815559-806f-4816-ac62-f2bcba187047.md @@ -0,0 +1,61 @@ + +These rules are ALWAYS ACTIVE for all service implementations and apply to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + + +### Rules + +- **R-SBL-001** MUST_NOT: Service boundary logs MUST NOT include sensitive data (passwords, tokens, PII) unless explicitly redacted or masked. + +### Scope + +**In scope:** +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +**Out of scope:** +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +**Exceptions:** +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +### Verify + +```bash +# Verify boundary logging patterns in service files +grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l + +# Verify logging statements in handlers and queries +rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' + +# Verify correlation identifiers in boundary logs +find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' +``` + +**Accept when:** +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs +- Service boundary logs do not contain passwords, tokens, or unredacted PII + +### Implementation Guidance + +- Create shared logging utilities or macros that encapsulate boundary logging patterns +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters +- Establish standard log fields: timestamp, service_name, operation, request_id, duration_ms, status, error_message +- Use async logging frameworks to minimize performance overhead +- Implement log sampling for high-frequency endpoints + + +Claude Code MUST verify service boundary logging compliance using the provided bash commands. Violations of R-SBL-001 (sensitive data in logs) are blocking. Missing boundary logging in new service handlers requires code review feedback before merge approval. Performance-critical exceptions require documented benchmarks and quarterly review. + \ No newline at end of file diff --git a/.claude/rules/e5731ea0-1ffc-4e87-b5f7-07551952f609.md b/.claude/rules/e5731ea0-1ffc-4e87-b5f7-07551952f609.md new file mode 100644 index 000000000..4d51ad243 --- /dev/null +++ b/.claude/rules/e5731ea0-1ffc-4e87-b5f7-07551952f609.md @@ -0,0 +1,44 @@ + +These rules are ALWAYS ACTIVE for all files matching route handlers that expose runtime execution state, configuration management endpoints, and real-time monitoring interfaces. + + +### Rules + +- **R-SSE-001** MUST: Route handlers that expose runtime execution state MUST establish SSE connections with appropriate Content-Type headers (text/event-stream) +- **R-SSE-002** MUST: SSE route handlers MUST set Cache-Control: no-cache and Connection: keep-alive headers +- **R-SSE-003** MUST: SSE connections MUST implement proper cleanup on client disconnect or execution completion +- **R-SSE-004** SHOULD: Use a centralized SSE utility module (e.g., lib/sse.ts) for connection setup, event formatting, and cleanup helpers +- **R-SSE-005** SHOULD: Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- **R-SSE-006** SHOULD: Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- **R-SSE-007** MAY: Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +### Verify + +```bash +# Count SSE implementations in route handlers +grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l + +# Verify SSE utility imports in route handlers +grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules + +# Locate centralized SSE utility module +find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" + +# Verify Content-Type headers are set +grep -r "text/event-stream" apps/fabro-web/app/routes/ | grep -v node_modules + +# Check for cleanup handlers in SSE connections +grep -r "cleanup\|disconnect\|close" apps/fabro-web/app/routes/ | grep -i sse +``` + +**Accept when:** +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion +- Cache-Control and Connection headers are set appropriately on all SSE responses +- Structured event payloads follow consistent format across all SSE endpoints +- Connection lifecycle logging is implemented for monitoring and debugging + + +Claude Code MUST NOT skip or defer verification. All SSE route handlers MUST be reviewed against R-SSE-001 through R-SSE-007. Violations must be flagged and require revision before merge. Exceptions require technical justification, architectural decision comments in code, and approval from technical lead. + \ No newline at end of file diff --git a/.claude/rules/f1ba5186-840f-4423-9876-d7cb8ec75614.md b/.claude/rules/f1ba5186-840f-4423-9876-d7cb8ec75614.md new file mode 100644 index 000000000..ab90d5e09 --- /dev/null +++ b/.claude/rules/f1ba5186-840f-4423-9876-d7cb8ec75614.md @@ -0,0 +1,32 @@ + +These rules are ALWAYS ACTIVE for all files matching real-time configuration and runtime state update endpoints, particularly web route handlers and SSE utility modules. + + +### Rules + +- **R-SSE-001** MUST: Runtime state updates and configuration changes MUST be delivered to clients using Server-Sent Events (SSE) protocol. + +### Verify + +```bash +# Verify SSE Content-Type header usage in route handlers +grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l + +# Verify SSE utility imports in route handlers +grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules + +# Verify centralized SSE utility module exists +find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" +``` + +**Accept when:** +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion +- Structured event payloads follow consistent format (e.g., {type: string, data: object, timestamp: number}) +- Appropriate headers are set: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Connection lifecycle logging and metrics are implemented to monitor SSE health + + +Claude Code MUST NOT skip or defer verification of SSE implementation requirements. All new real-time endpoints must be verified against these rules before acceptance. Violations require justification or architectural exception approval. + \ No newline at end of file diff --git a/.claude/rules/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7.md b/.claude/rules/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7.md new file mode 100644 index 000000000..95ac33d5e --- /dev/null +++ b/.claude/rules/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7.md @@ -0,0 +1,55 @@ + +These rules are ALWAYS ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + + +### Rules + +- **R-TRACING-001** SHOULD: Trace spans SHOULD include business-relevant metadata (user_id, tenant_id, request_id) as span attributes for correlation. + +### Scope + +**In scope:** +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +**Out of scope:** +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +**Exceptions:** +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +### Verify + +```bash +# Count tracing instrumentation patterns +grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l + +# Search for obs.tracing facet and OpenTelemetry patterns +grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' + +# Run tracing-related tests +cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' +``` + +**Accept when:** +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +### Implementation Guidance + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + + +Claude Code MUST NOT skip or defer verification. Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis. Code review checklist includes verification of distributed tracing implementation. Integration test suite validates trace context propagation and span attribute completeness. Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation. Code review blocks merge if tracing requirements are not met without documented exception. Quarterly audits identify non-compliant APIs with remediation plans required within 30 days. + \ No newline at end of file diff --git a/.claude/rules/f5680bc5-1992-44e9-92a3-65214350b7a7.md b/.claude/rules/f5680bc5-1992-44e9-92a3-65214350b7a7.md new file mode 100644 index 000000000..e149066dc --- /dev/null +++ b/.claude/rules/f5680bc5-1992-44e9-92a3-65214350b7a7.md @@ -0,0 +1,31 @@ + +These rules are ALWAYS ACTIVE for all files matching web route handlers, configuration management endpoints, and real-time monitoring interfaces that expose runtime execution progress and state synchronization. + + +### Rules + +- **R-SSE-001** SHOULD: SSE event streams SHOULD include structured data payloads (JSON) with type identifiers to enable client-side event routing. + +### Verify + +```bash +# Verify SSE usage across route handlers +grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l + +# Verify SSE imports in route handlers +grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules + +# Verify centralized SSE utility module exists +find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" +``` + +**Accept when:** +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion +- Structured event payloads follow consistent format with type, data, and timestamp fields +- Appropriate headers are set: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive + + +Claude Code MUST NOT skip or defer verification. All route handlers introducing real-time updates must be verified against SSE patterns. Violations must be flagged in code review and require revision before merge. Exceptions require technical justification, architectural decision comments in code, and approval from technical lead. + \ No newline at end of file diff --git a/.claude/rules/fcd6a21d-756b-4d3f-98d7-f42d21122afd.md b/.claude/rules/fcd6a21d-756b-4d3f-98d7-f42d21122afd.md new file mode 100644 index 000000000..bce7a01c5 --- /dev/null +++ b/.claude/rules/fcd6a21d-756b-4d3f-98d7-f42d21122afd.md @@ -0,0 +1,63 @@ + +These rules are ALWAYS ACTIVE for all service implementations and apply to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + + +### Rules + +- **R-SBL-001** MAY: Service implementations MAY include additional context-specific logging within service boundaries for detailed debugging, but this does not replace boundary logging requirements. + +### In Scope + +Logging is required at these service boundary points: +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +### Out of Scope + +Logging is NOT required for: +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +### Exceptions + +- **EXC-001**: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- **EXC-002**: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +### Verify + +```bash +# Verify logging patterns in service boundary files +grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l + +# Verify structured logging in handlers +rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' + +# Verify correlation identifiers present +find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' +``` + +**Accept when:** +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +### Implementation Guidance + +- Create shared logging utilities or macros that encapsulate boundary logging patterns +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters +- Establish standard log fields: timestamp, service_name, operation, request_id, duration_ms, status, error_message +- Configure log aggregation systems to parse and index boundary logs +- Document logging patterns in service templates and starter kits + + +Claude Code MUST verify boundary logging presence using the provided bash commands. Verification is mandatory before accepting service implementations. Missing boundary logging at service entry/exit points MUST be flagged for remediation. + \ No newline at end of file diff --git a/AGENTS.md b/AGENTS.md index 61b329f67..697442849 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,3 +1,41 @@ +--- + +## Architecture Decision Records + + +ADRs govern validated architectural standards for this project. +Full ADR documents: @docs/adr/ + + + +These directives are ALWAYS ACTIVE. All AI coding agents MUST apply all rules in this +document to every code generation, modification, and review action within this +project. No exceptions unless explicitly noted per-rule. + + +--- + +### Verification Protocol + + +All rules in this document follow the **Verify → Fix → Repeat** loop. + + +After generating or modifying code for any rule, the agent MUST: + +1. **RUN** the targeted verification command(s) in the rule's **Verify** block. +2. **CAPTURE** the full command output (stdout + stderr). +3. **EVALUATE** whether the **Accept when** criteria are satisfied. +4. **IF FAILING:** diagnose the root cause, apply a fix, and re-run from step 1. +5. **IF PASSING:** include the passing output as inline evidence before proposing further changes. +6. **MAX ITERATIONS:** 5 attempts per rule. If still failing after 5 attempts, STOP and report the failure with all captured outputs. + + +Compliance is not optional. Agents must not skip verification steps, assume +correctness, or defer verification to a later task. Evidence of a passing +verification run must accompany every code change that touches a governed area. + + # CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. diff --git a/CLAUDE.md b/CLAUDE.md deleted file mode 120000 index 47dc3e3d8..000000000 --- a/CLAUDE.md +++ /dev/null @@ -1 +0,0 @@ -AGENTS.md \ No newline at end of file diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 000000000..567061855 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,39 @@ +--- + +## Architecture Decision Records + + +ADRs govern validated architectural standards for this project. +Full ADR documents: @docs/adr/ + + + +These directives are ALWAYS ACTIVE. Claude Code MUST apply all rules in this +document to every code generation, modification, and review action within this +project. No exceptions unless explicitly noted per-rule. + + +--- + +### Verification Protocol + + +All rules in this document follow the **Verify → Fix → Repeat** loop. + + +After generating or modifying code for any rule, Claude Code MUST: + +1. **RUN** the targeted verification command(s) in the rule's **Verify** block. +2. **CAPTURE** the full command output (stdout + stderr). +3. **EVALUATE** whether the **Accept when** criteria are satisfied. +4. **IF FAILING:** diagnose the root cause, apply a fix, and re-run from step 1. +5. **IF PASSING:** include the passing output as inline evidence before proposing further changes. +6. **MAX ITERATIONS:** 5 attempts per rule. If still failing after 5 attempts, STOP and report the failure with all captured outputs. + + +Compliance is not optional. Claude Code must not skip verification steps, assume +correctness, or defer verification to a later task. Evidence of a passing +verification run must accompany every code change that touches a governed area. + + +AGENTS.md diff --git a/docs/adr/0ba59d63-b106-4111-a75e-cb18a23f7c56-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-properly.md b/docs/adr/0ba59d63-b106-4111-a75e-cb18a23f7c56-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-properly.md new file mode 100644 index 000000000..07a19a7b5 --- /dev/null +++ b/docs/adr/0ba59d63-b106-4111-a75e-cb18a23f7c56-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-properly.md @@ -0,0 +1,115 @@ +# Standardize Server-Sent Events (SSE) for Real-Time Configuration and Runtime State Updates: Route Handlers Properly + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Context + +- The application requires real-time updates of runtime execution state and configuration changes to be pushed to client interfaces without polling +- Multiple route handlers (run-stages.tsx, run-files.tsx) need to stream execution progress and file processing status to web clients +- A centralized SSE implementation pattern (sse.ts) has emerged to handle persistent connections and event streaming across different runtime contexts +- The pattern provides a consistent mechanism for broadcasting configuration changes and environment state updates during execution workflows + +## Problem Statement + +Applications need a reliable, standardized mechanism to push real-time runtime state, configuration updates, and execution progress to clients without the overhead and latency of polling-based approaches, while maintaining consistent event handling across multiple execution contexts. + +## Decision + +1. MUST: Route handlers MUST properly close SSE connections and clean up resources when clients disconnect or execution completes + +## Policy Block + +- MUST Route handlers MUST properly close SSE connections and clean up resources when clients disconnect or execution completes + +In scope: +- Web route handlers that expose runtime execution progress (run-stages, run-files, etc.) +- Configuration management endpoints that broadcast environment or setting changes +- Real-time monitoring interfaces for execution workflows +- Client-facing APIs that require push-based state synchronization + +Out of scope: +- Bidirectional communication patterns requiring client-to-server messaging (use WebSockets instead) +- Binary data streaming or large file transfers +- Internal service-to-service communication (use message queues or gRPC) +- Static configuration loading at application startup + +## Rationale + +- Pattern detected across 3 files with 91.07% confidence indicates a deliberate architectural choice for real-time state communication +- SSE provides a lightweight, HTTP-based protocol that works through firewalls and proxies without requiring WebSocket infrastructure +- Centralized SSE library (sse.ts) ensures consistent event formatting, error handling, and connection lifecycle management across multiple route handlers +- The pattern aligns with unidirectional data flow from server to client, which matches the use case of broadcasting runtime state and configuration updates + +## Consequences + +Positive: +- Clients receive real-time updates without polling overhead, reducing server load and improving responsiveness +- Consistent SSE implementation across routes reduces code duplication and simplifies maintenance +- HTTP-based protocol ensures compatibility with existing infrastructure (load balancers, proxies, CDNs) +- Automatic reconnection support in browsers provides resilience against transient network failures + +Negative: +- SSE connections are unidirectional, requiring separate HTTP requests for client-to-server communication +- Long-lived connections may require infrastructure tuning (timeouts, connection limits) to support many concurrent clients +- Browser connection limits (typically 6 per domain) may constrain applications with multiple simultaneous SSE streams +- Debugging and monitoring SSE connections can be more complex than traditional request-response patterns + +## Alternatives + +- Use polling-based approach with periodic HTTP requests to fetch runtime state updates (rejected) + Rejected because: Polling introduces latency, increases server load with redundant requests, and provides poor user experience for real-time execution monitoring + When valid: Acceptable for low-frequency updates (>30 seconds) or when SSE infrastructure is unavailable +- Implement WebSocket-based bidirectional communication for all real-time features (rejected) + Rejected because: WebSockets add complexity and infrastructure requirements when only unidirectional server-to-client communication is needed for configuration and state updates + When valid: Appropriate when bidirectional real-time communication is required (e.g., collaborative editing, chat) +- Use GraphQL subscriptions for real-time data updates (rejected) + Rejected because: Adds GraphQL infrastructure overhead and complexity when simple event streaming suffices for runtime state broadcasting + When valid: Valid when already using GraphQL extensively and need subscription-based data synchronization + +## Risks + +- SSE connections may be terminated by intermediate proxies or load balancers with aggressive timeout policies + Mitigation: Implement periodic heartbeat messages and configure infrastructure timeouts appropriately (e.g., 60+ seconds). Document required infrastructure settings. + Owner: Engineering team and DevOps +- Memory leaks or resource exhaustion if SSE connections are not properly cleaned up on client disconnect + Mitigation: Implement robust connection lifecycle management with cleanup handlers. Add monitoring for open connection counts and memory usage. + Owner: Engineering team +- Browser connection limits may prevent multiple SSE streams from functioning simultaneously + Mitigation: Multiplex multiple event types over a single SSE connection where possible. Document connection usage and provide guidance on connection management. + Owner: Engineering team + +## Implementation Notes + +- Create a centralized SSE utility module (e.g., lib/sse.ts) that provides connection setup, event formatting, and cleanup helpers +- Ensure route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- Document SSE usage patterns and provide examples for common scenarios (execution progress, configuration updates) +- Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +## Continuation Context + + +Verify commands: +- grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l +- grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules +- find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" + +Accept when: +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion + +## Enforcement + +- Verified by: Code review checklist verifying SSE usage for new real-time endpoints +- Verified by: Automated linting rules to detect missing Content-Type headers on streaming routes +- Verified by: Integration tests validating SSE connection lifecycle and event delivery +- Violation handling: Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable +- Violation handling: Route handlers with SSE connections missing proper cleanup must be fixed before merge +- Violation handling: Violations detected in code review are flagged and require revision +- Exception process: Document technical justification for alternative approach (e.g., WebSocket requirement, infrastructure constraints) +- Exception process: Obtain approval from technical lead or architect +- Exception process: Add architectural decision comment in code explaining exception rationale \ No newline at end of file diff --git a/docs/adr/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-connections-implement.md b/docs/adr/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-connections-implement.md new file mode 100644 index 000000000..f8d882cc3 --- /dev/null +++ b/docs/adr/0c99d9c0-10fb-40b3-bc7d-6f69d4cd5498-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-connections-implement.md @@ -0,0 +1,115 @@ +# Standardize Server-Sent Events (SSE) for Real-Time Configuration and Runtime State Updates: Sse Connections Implement + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Context + +- The application requires real-time updates of runtime execution state and configuration changes to be pushed to client interfaces without polling +- Multiple route handlers (run-stages.tsx, run-files.tsx) need to stream execution progress and file processing status to web clients +- A centralized SSE implementation pattern (sse.ts) has emerged to handle persistent connections and event streaming across different runtime contexts +- The pattern provides a consistent mechanism for broadcasting configuration changes and environment state updates during execution workflows + +## Problem Statement + +Applications need a reliable, standardized mechanism to push real-time runtime state, configuration updates, and execution progress to clients without the overhead and latency of polling-based approaches, while maintaining consistent event handling across multiple execution contexts. + +## Decision + +1. SHOULD: SSE connections SHOULD implement heartbeat or keep-alive mechanisms to detect and handle connection failures + +## Policy Block + +- SHOULD SSE connections SHOULD implement heartbeat or keep-alive mechanisms to detect and handle connection failures + +In scope: +- Web route handlers that expose runtime execution progress (run-stages, run-files, etc.) +- Configuration management endpoints that broadcast environment or setting changes +- Real-time monitoring interfaces for execution workflows +- Client-facing APIs that require push-based state synchronization + +Out of scope: +- Bidirectional communication patterns requiring client-to-server messaging (use WebSockets instead) +- Binary data streaming or large file transfers +- Internal service-to-service communication (use message queues or gRPC) +- Static configuration loading at application startup + +## Rationale + +- Pattern detected across 3 files with 91.07% confidence indicates a deliberate architectural choice for real-time state communication +- SSE provides a lightweight, HTTP-based protocol that works through firewalls and proxies without requiring WebSocket infrastructure +- Centralized SSE library (sse.ts) ensures consistent event formatting, error handling, and connection lifecycle management across multiple route handlers +- The pattern aligns with unidirectional data flow from server to client, which matches the use case of broadcasting runtime state and configuration updates + +## Consequences + +Positive: +- Clients receive real-time updates without polling overhead, reducing server load and improving responsiveness +- Consistent SSE implementation across routes reduces code duplication and simplifies maintenance +- HTTP-based protocol ensures compatibility with existing infrastructure (load balancers, proxies, CDNs) +- Automatic reconnection support in browsers provides resilience against transient network failures + +Negative: +- SSE connections are unidirectional, requiring separate HTTP requests for client-to-server communication +- Long-lived connections may require infrastructure tuning (timeouts, connection limits) to support many concurrent clients +- Browser connection limits (typically 6 per domain) may constrain applications with multiple simultaneous SSE streams +- Debugging and monitoring SSE connections can be more complex than traditional request-response patterns + +## Alternatives + +- Use polling-based approach with periodic HTTP requests to fetch runtime state updates (rejected) + Rejected because: Polling introduces latency, increases server load with redundant requests, and provides poor user experience for real-time execution monitoring + When valid: Acceptable for low-frequency updates (>30 seconds) or when SSE infrastructure is unavailable +- Implement WebSocket-based bidirectional communication for all real-time features (rejected) + Rejected because: WebSockets add complexity and infrastructure requirements when only unidirectional server-to-client communication is needed for configuration and state updates + When valid: Appropriate when bidirectional real-time communication is required (e.g., collaborative editing, chat) +- Use GraphQL subscriptions for real-time data updates (rejected) + Rejected because: Adds GraphQL infrastructure overhead and complexity when simple event streaming suffices for runtime state broadcasting + When valid: Valid when already using GraphQL extensively and need subscription-based data synchronization + +## Risks + +- SSE connections may be terminated by intermediate proxies or load balancers with aggressive timeout policies + Mitigation: Implement periodic heartbeat messages and configure infrastructure timeouts appropriately (e.g., 60+ seconds). Document required infrastructure settings. + Owner: Engineering team and DevOps +- Memory leaks or resource exhaustion if SSE connections are not properly cleaned up on client disconnect + Mitigation: Implement robust connection lifecycle management with cleanup handlers. Add monitoring for open connection counts and memory usage. + Owner: Engineering team +- Browser connection limits may prevent multiple SSE streams from functioning simultaneously + Mitigation: Multiplex multiple event types over a single SSE connection where possible. Document connection usage and provide guidance on connection management. + Owner: Engineering team + +## Implementation Notes + +- Create a centralized SSE utility module (e.g., lib/sse.ts) that provides connection setup, event formatting, and cleanup helpers +- Ensure route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- Document SSE usage patterns and provide examples for common scenarios (execution progress, configuration updates) +- Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +## Continuation Context + + +Verify commands: +- grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l +- grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules +- find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" + +Accept when: +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion + +## Enforcement + +- Verified by: Code review checklist verifying SSE usage for new real-time endpoints +- Verified by: Automated linting rules to detect missing Content-Type headers on streaming routes +- Verified by: Integration tests validating SSE connection lifecycle and event delivery +- Violation handling: Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable +- Violation handling: Route handlers with SSE connections missing proper cleanup must be fixed before merge +- Violation handling: Violations detected in code review are flagged and require revision +- Exception process: Document technical justification for alternative approach (e.g., WebSocket requirement, infrastructure constraints) +- Exception process: Obtain approval from technical lead or architect +- Exception process: Add architectural decision comment in code explaining exception rationale \ No newline at end of file diff --git a/docs/adr/1267d2fe-4cfa-4c45-9004-ada782af479d-adopt-distributed-tracing-for-public-api-observability-public-external-endpoints.md b/docs/adr/1267d2fe-4cfa-4c45-9004-ada782af479d-adopt-distributed-tracing-for-public-api-observability-public-external-endpoints.md new file mode 100644 index 000000000..e0c174f76 --- /dev/null +++ b/docs/adr/1267d2fe-4cfa-4c45-9004-ada782af479d-adopt-distributed-tracing-for-public-api-observability-public-external-endpoints.md @@ -0,0 +1,123 @@ +# Adopt Distributed Tracing for Public API Observability: Public External Endpoints + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + +## Context + +- Public and external APIs require comprehensive observability to diagnose issues across distributed system boundaries where direct debugging is not feasible +- The pattern was detected with 90.27% confidence across 3 critical infrastructure files (terminal.rs, lib.rs, stall.rs) in the fabro crate ecosystem, indicating a systematic approach to tracing +- The facet 'obs.tracing' suggests this pattern specifically addresses observability through distributed tracing instrumentation +- External API consumers and integration partners require visibility into request flows, latency bottlenecks, and error propagation paths +- Without standardized tracing, debugging cross-service issues becomes prohibitively expensive and time-consuming + +## Problem Statement + +Public and external APIs lack consistent observability instrumentation, making it difficult to diagnose performance issues, trace request flows across service boundaries, and provide actionable debugging information to external consumers. This creates operational blind spots and increases mean time to resolution (MTTR) for production incidents. + +## Decision + +1. MUST: All public/external API endpoints MUST instrument entry and exit points with distributed tracing spans + +## Policy Block + +- MUST All public/external API endpoints MUST instrument entry and exit points with distributed tracing spans + +In scope: +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +Out of scope: +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +Exceptions: +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +## Rationale + +- The pattern detection across 3 files with 90.27% confidence indicates this is an established architectural practice in the codebase, not an isolated implementation +- Distributed tracing provides end-to-end visibility that is essential for debugging issues in microservices architectures where requests span multiple services +- Standardizing on obs.tracing facet ensures consistent instrumentation patterns across the organization, reducing cognitive load and improving debugging efficiency +- External API consumers benefit from trace IDs in error responses, enabling them to provide actionable information when reporting issues + +## Consequences + +Positive: +- Reduced mean time to resolution (MTTR) for production incidents through comprehensive request flow visibility +- Improved developer experience with standardized tracing patterns across all public APIs +- Enhanced external consumer support with traceable request identifiers for issue reporting +- Better capacity planning and performance optimization through latency distribution analysis + +Negative: +- Increased operational complexity with additional infrastructure for trace collection and storage +- Minor performance overhead (typically 1-3% latency increase) from tracing instrumentation +- Development time investment required to instrument existing APIs and train teams on tracing best practices +- Potential for trace data explosion requiring careful sampling strategy and retention policies + +## Alternatives + +- Structured logging only without distributed tracing (rejected) + Rejected because: Logs lack the causal relationships and timing information needed to reconstruct request flows across service boundaries. Correlation requires manual log aggregation which is error-prone and time-consuming. + When valid: May be sufficient for monolithic applications with single-service request handling +- Metrics-based observability with counters and histograms (rejected) + Rejected because: Metrics provide aggregate statistics but cannot trace individual request paths or identify specific failure scenarios. Complementary to tracing but insufficient as sole observability strategy. + When valid: Appropriate for high-level health monitoring and alerting, should be used alongside tracing +- Vendor-specific tracing solutions (e.g., AWS X-Ray, Datadog APM) (deferred) + Rejected because: Not rejected but deferred to implementation phase. Vendor choice should align with existing infrastructure while maintaining standard trace context propagation. + When valid: Valid if vendor solution supports OpenTelemetry or W3C Trace Context standards for interoperability + +## Risks + +- Trace data volume may exceed storage capacity or budget constraints, leading to incomplete observability + Mitigation: Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests. Establish retention policies (e.g., 7 days for all traces, 30 days for error traces). + Owner: Platform Engineering Team +- Inconsistent tracing implementation across teams may result in fragmented observability + Mitigation: Provide shared tracing libraries and middleware with sensible defaults. Conduct code reviews specifically checking for tracing compliance. Create runbooks and training materials. + Owner: Engineering Team +- Performance overhead from tracing may impact latency-sensitive APIs + Mitigation: Benchmark tracing overhead during implementation. Use asynchronous trace export to minimize request latency impact. Allow exceptions for proven performance-critical paths. + Owner: API Development Team + +## Implementation Notes + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + +## Continuation Context + + +Verify commands: +- grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l +- grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' +- cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' + +Accept when: +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +## Enforcement + +- Verified by: Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis +- Verified by: Code review checklist includes verification of distributed tracing implementation +- Verified by: Integration test suite validates trace context propagation and span attribute completeness +- Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation +- Violation handling: Code review blocks merge if tracing requirements are not met without documented exception +- Violation handling: Quarterly audits identify non-compliant APIs with remediation plans required within 30 days +- Exception process: Submit exception request to architecture review board with performance benchmarks or deprecation timeline +- Exception process: Document exception in API specification and architectural decision log +- Exception process: Re-evaluate exceptions quarterly to determine if circumstances have changed \ No newline at end of file diff --git a/docs/adr/19fbda95-be0a-43ce-88bd-b2cbb1c4075a-adopt-distributed-tracing-for-public-api-observability-each-trace-span.md b/docs/adr/19fbda95-be0a-43ce-88bd-b2cbb1c4075a-adopt-distributed-tracing-for-public-api-observability-each-trace-span.md new file mode 100644 index 000000000..d698189b5 --- /dev/null +++ b/docs/adr/19fbda95-be0a-43ce-88bd-b2cbb1c4075a-adopt-distributed-tracing-for-public-api-observability-each-trace-span.md @@ -0,0 +1,123 @@ +# Adopt Distributed Tracing for Public API Observability: Each Trace Span + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + +## Context + +- Public and external APIs require comprehensive observability to diagnose issues across distributed system boundaries where direct debugging is not feasible +- The pattern was detected with 90.27% confidence across 3 critical infrastructure files (terminal.rs, lib.rs, stall.rs) in the fabro crate ecosystem, indicating a systematic approach to tracing +- The facet 'obs.tracing' suggests this pattern specifically addresses observability through distributed tracing instrumentation +- External API consumers and integration partners require visibility into request flows, latency bottlenecks, and error propagation paths +- Without standardized tracing, debugging cross-service issues becomes prohibitively expensive and time-consuming + +## Problem Statement + +Public and external APIs lack consistent observability instrumentation, making it difficult to diagnose performance issues, trace request flows across service boundaries, and provide actionable debugging information to external consumers. This creates operational blind spots and increases mean time to resolution (MTTR) for production incidents. + +## Decision + +1. MUST: Each trace span MUST include standardized attributes: service.name, operation.name, http.method, http.status_code, and error.type (if applicable) + +## Policy Block + +- MUST Each trace span MUST include standardized attributes: service.name, operation.name, http.method, http.status_code, and error.type (if applicable) + +In scope: +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +Out of scope: +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +Exceptions: +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +## Rationale + +- The pattern detection across 3 files with 90.27% confidence indicates this is an established architectural practice in the codebase, not an isolated implementation +- Distributed tracing provides end-to-end visibility that is essential for debugging issues in microservices architectures where requests span multiple services +- Standardizing on obs.tracing facet ensures consistent instrumentation patterns across the organization, reducing cognitive load and improving debugging efficiency +- External API consumers benefit from trace IDs in error responses, enabling them to provide actionable information when reporting issues + +## Consequences + +Positive: +- Reduced mean time to resolution (MTTR) for production incidents through comprehensive request flow visibility +- Improved developer experience with standardized tracing patterns across all public APIs +- Enhanced external consumer support with traceable request identifiers for issue reporting +- Better capacity planning and performance optimization through latency distribution analysis + +Negative: +- Increased operational complexity with additional infrastructure for trace collection and storage +- Minor performance overhead (typically 1-3% latency increase) from tracing instrumentation +- Development time investment required to instrument existing APIs and train teams on tracing best practices +- Potential for trace data explosion requiring careful sampling strategy and retention policies + +## Alternatives + +- Structured logging only without distributed tracing (rejected) + Rejected because: Logs lack the causal relationships and timing information needed to reconstruct request flows across service boundaries. Correlation requires manual log aggregation which is error-prone and time-consuming. + When valid: May be sufficient for monolithic applications with single-service request handling +- Metrics-based observability with counters and histograms (rejected) + Rejected because: Metrics provide aggregate statistics but cannot trace individual request paths or identify specific failure scenarios. Complementary to tracing but insufficient as sole observability strategy. + When valid: Appropriate for high-level health monitoring and alerting, should be used alongside tracing +- Vendor-specific tracing solutions (e.g., AWS X-Ray, Datadog APM) (deferred) + Rejected because: Not rejected but deferred to implementation phase. Vendor choice should align with existing infrastructure while maintaining standard trace context propagation. + When valid: Valid if vendor solution supports OpenTelemetry or W3C Trace Context standards for interoperability + +## Risks + +- Trace data volume may exceed storage capacity or budget constraints, leading to incomplete observability + Mitigation: Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests. Establish retention policies (e.g., 7 days for all traces, 30 days for error traces). + Owner: Platform Engineering Team +- Inconsistent tracing implementation across teams may result in fragmented observability + Mitigation: Provide shared tracing libraries and middleware with sensible defaults. Conduct code reviews specifically checking for tracing compliance. Create runbooks and training materials. + Owner: Engineering Team +- Performance overhead from tracing may impact latency-sensitive APIs + Mitigation: Benchmark tracing overhead during implementation. Use asynchronous trace export to minimize request latency impact. Allow exceptions for proven performance-critical paths. + Owner: API Development Team + +## Implementation Notes + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + +## Continuation Context + + +Verify commands: +- grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l +- grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' +- cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' + +Accept when: +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +## Enforcement + +- Verified by: Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis +- Verified by: Code review checklist includes verification of distributed tracing implementation +- Verified by: Integration test suite validates trace context propagation and span attribute completeness +- Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation +- Violation handling: Code review blocks merge if tracing requirements are not met without documented exception +- Violation handling: Quarterly audits identify non-compliant APIs with remediation plans required within 30 days +- Exception process: Submit exception request to architecture review board with performance benchmarks or deprecation timeline +- Exception process: Document exception in API specification and architectural decision log +- Exception process: Re-evaluate exceptions quarterly to determine if circumstances have changed \ No newline at end of file diff --git a/docs/adr/36c17245-9a45-4855-affd-3e23712797d8-adopt-distributed-tracing-for-public-api-observability-tracing-implementations-propagate.md b/docs/adr/36c17245-9a45-4855-affd-3e23712797d8-adopt-distributed-tracing-for-public-api-observability-tracing-implementations-propagate.md new file mode 100644 index 000000000..0cd2bba0c --- /dev/null +++ b/docs/adr/36c17245-9a45-4855-affd-3e23712797d8-adopt-distributed-tracing-for-public-api-observability-tracing-implementations-propagate.md @@ -0,0 +1,123 @@ +# Adopt Distributed Tracing for Public API Observability: Tracing Implementations Propagate + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + +## Context + +- Public and external APIs require comprehensive observability to diagnose issues across distributed system boundaries where direct debugging is not feasible +- The pattern was detected with 90.27% confidence across 3 critical infrastructure files (terminal.rs, lib.rs, stall.rs) in the fabro crate ecosystem, indicating a systematic approach to tracing +- The facet 'obs.tracing' suggests this pattern specifically addresses observability through distributed tracing instrumentation +- External API consumers and integration partners require visibility into request flows, latency bottlenecks, and error propagation paths +- Without standardized tracing, debugging cross-service issues becomes prohibitively expensive and time-consuming + +## Problem Statement + +Public and external APIs lack consistent observability instrumentation, making it difficult to diagnose performance issues, trace request flows across service boundaries, and provide actionable debugging information to external consumers. This creates operational blind spots and increases mean time to resolution (MTTR) for production incidents. + +## Decision + +1. MUST: Tracing implementations MUST propagate trace context across service boundaries using W3C Trace Context or OpenTelemetry standards + +## Policy Block + +- MUST Tracing implementations MUST propagate trace context across service boundaries using W3C Trace Context or OpenTelemetry standards + +In scope: +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +Out of scope: +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +Exceptions: +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +## Rationale + +- The pattern detection across 3 files with 90.27% confidence indicates this is an established architectural practice in the codebase, not an isolated implementation +- Distributed tracing provides end-to-end visibility that is essential for debugging issues in microservices architectures where requests span multiple services +- Standardizing on obs.tracing facet ensures consistent instrumentation patterns across the organization, reducing cognitive load and improving debugging efficiency +- External API consumers benefit from trace IDs in error responses, enabling them to provide actionable information when reporting issues + +## Consequences + +Positive: +- Reduced mean time to resolution (MTTR) for production incidents through comprehensive request flow visibility +- Improved developer experience with standardized tracing patterns across all public APIs +- Enhanced external consumer support with traceable request identifiers for issue reporting +- Better capacity planning and performance optimization through latency distribution analysis + +Negative: +- Increased operational complexity with additional infrastructure for trace collection and storage +- Minor performance overhead (typically 1-3% latency increase) from tracing instrumentation +- Development time investment required to instrument existing APIs and train teams on tracing best practices +- Potential for trace data explosion requiring careful sampling strategy and retention policies + +## Alternatives + +- Structured logging only without distributed tracing (rejected) + Rejected because: Logs lack the causal relationships and timing information needed to reconstruct request flows across service boundaries. Correlation requires manual log aggregation which is error-prone and time-consuming. + When valid: May be sufficient for monolithic applications with single-service request handling +- Metrics-based observability with counters and histograms (rejected) + Rejected because: Metrics provide aggregate statistics but cannot trace individual request paths or identify specific failure scenarios. Complementary to tracing but insufficient as sole observability strategy. + When valid: Appropriate for high-level health monitoring and alerting, should be used alongside tracing +- Vendor-specific tracing solutions (e.g., AWS X-Ray, Datadog APM) (deferred) + Rejected because: Not rejected but deferred to implementation phase. Vendor choice should align with existing infrastructure while maintaining standard trace context propagation. + When valid: Valid if vendor solution supports OpenTelemetry or W3C Trace Context standards for interoperability + +## Risks + +- Trace data volume may exceed storage capacity or budget constraints, leading to incomplete observability + Mitigation: Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests. Establish retention policies (e.g., 7 days for all traces, 30 days for error traces). + Owner: Platform Engineering Team +- Inconsistent tracing implementation across teams may result in fragmented observability + Mitigation: Provide shared tracing libraries and middleware with sensible defaults. Conduct code reviews specifically checking for tracing compliance. Create runbooks and training materials. + Owner: Engineering Team +- Performance overhead from tracing may impact latency-sensitive APIs + Mitigation: Benchmark tracing overhead during implementation. Use asynchronous trace export to minimize request latency impact. Allow exceptions for proven performance-critical paths. + Owner: API Development Team + +## Implementation Notes + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + +## Continuation Context + + +Verify commands: +- grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l +- grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' +- cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' + +Accept when: +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +## Enforcement + +- Verified by: Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis +- Verified by: Code review checklist includes verification of distributed tracing implementation +- Verified by: Integration test suite validates trace context propagation and span attribute completeness +- Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation +- Violation handling: Code review blocks merge if tracing requirements are not met without documented exception +- Violation handling: Quarterly audits identify non-compliant APIs with remediation plans required within 30 days +- Exception process: Submit exception request to architecture review board with performance benchmarks or deprecation timeline +- Exception process: Document exception in API specification and architectural decision log +- Exception process: Re-evaluate exceptions quarterly to determine if circumstances have changed \ No newline at end of file diff --git a/docs/adr/372c20f8-a9ae-4489-b3a3-4825d3343af4-adopt-distributed-tracing-for-public-api-observability-implementations-create-child.md b/docs/adr/372c20f8-a9ae-4489-b3a3-4825d3343af4-adopt-distributed-tracing-for-public-api-observability-implementations-create-child.md new file mode 100644 index 000000000..5330dae27 --- /dev/null +++ b/docs/adr/372c20f8-a9ae-4489-b3a3-4825d3343af4-adopt-distributed-tracing-for-public-api-observability-implementations-create-child.md @@ -0,0 +1,123 @@ +# Adopt Distributed Tracing for Public API Observability: Implementations Create Child + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + +## Context + +- Public and external APIs require comprehensive observability to diagnose issues across distributed system boundaries where direct debugging is not feasible +- The pattern was detected with 90.27% confidence across 3 critical infrastructure files (terminal.rs, lib.rs, stall.rs) in the fabro crate ecosystem, indicating a systematic approach to tracing +- The facet 'obs.tracing' suggests this pattern specifically addresses observability through distributed tracing instrumentation +- External API consumers and integration partners require visibility into request flows, latency bottlenecks, and error propagation paths +- Without standardized tracing, debugging cross-service issues becomes prohibitively expensive and time-consuming + +## Problem Statement + +Public and external APIs lack consistent observability instrumentation, making it difficult to diagnose performance issues, trace request flows across service boundaries, and provide actionable debugging information to external consumers. This creates operational blind spots and increases mean time to resolution (MTTR) for production incidents. + +## Decision + +1. SHOULD: API implementations SHOULD create child spans for significant internal operations (database queries, external service calls, computation-heavy operations) + +## Policy Block + +- SHOULD API implementations SHOULD create child spans for significant internal operations (database queries, external service calls, computation-heavy operations) + +In scope: +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +Out of scope: +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +Exceptions: +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +## Rationale + +- The pattern detection across 3 files with 90.27% confidence indicates this is an established architectural practice in the codebase, not an isolated implementation +- Distributed tracing provides end-to-end visibility that is essential for debugging issues in microservices architectures where requests span multiple services +- Standardizing on obs.tracing facet ensures consistent instrumentation patterns across the organization, reducing cognitive load and improving debugging efficiency +- External API consumers benefit from trace IDs in error responses, enabling them to provide actionable information when reporting issues + +## Consequences + +Positive: +- Reduced mean time to resolution (MTTR) for production incidents through comprehensive request flow visibility +- Improved developer experience with standardized tracing patterns across all public APIs +- Enhanced external consumer support with traceable request identifiers for issue reporting +- Better capacity planning and performance optimization through latency distribution analysis + +Negative: +- Increased operational complexity with additional infrastructure for trace collection and storage +- Minor performance overhead (typically 1-3% latency increase) from tracing instrumentation +- Development time investment required to instrument existing APIs and train teams on tracing best practices +- Potential for trace data explosion requiring careful sampling strategy and retention policies + +## Alternatives + +- Structured logging only without distributed tracing (rejected) + Rejected because: Logs lack the causal relationships and timing information needed to reconstruct request flows across service boundaries. Correlation requires manual log aggregation which is error-prone and time-consuming. + When valid: May be sufficient for monolithic applications with single-service request handling +- Metrics-based observability with counters and histograms (rejected) + Rejected because: Metrics provide aggregate statistics but cannot trace individual request paths or identify specific failure scenarios. Complementary to tracing but insufficient as sole observability strategy. + When valid: Appropriate for high-level health monitoring and alerting, should be used alongside tracing +- Vendor-specific tracing solutions (e.g., AWS X-Ray, Datadog APM) (deferred) + Rejected because: Not rejected but deferred to implementation phase. Vendor choice should align with existing infrastructure while maintaining standard trace context propagation. + When valid: Valid if vendor solution supports OpenTelemetry or W3C Trace Context standards for interoperability + +## Risks + +- Trace data volume may exceed storage capacity or budget constraints, leading to incomplete observability + Mitigation: Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests. Establish retention policies (e.g., 7 days for all traces, 30 days for error traces). + Owner: Platform Engineering Team +- Inconsistent tracing implementation across teams may result in fragmented observability + Mitigation: Provide shared tracing libraries and middleware with sensible defaults. Conduct code reviews specifically checking for tracing compliance. Create runbooks and training materials. + Owner: Engineering Team +- Performance overhead from tracing may impact latency-sensitive APIs + Mitigation: Benchmark tracing overhead during implementation. Use asynchronous trace export to minimize request latency impact. Allow exceptions for proven performance-critical paths. + Owner: API Development Team + +## Implementation Notes + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + +## Continuation Context + + +Verify commands: +- grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l +- grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' +- cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' + +Accept when: +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +## Enforcement + +- Verified by: Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis +- Verified by: Code review checklist includes verification of distributed tracing implementation +- Verified by: Integration test suite validates trace context propagation and span attribute completeness +- Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation +- Violation handling: Code review blocks merge if tracing requirements are not met without documented exception +- Violation handling: Quarterly audits identify non-compliant APIs with remediation plans required within 30 days +- Exception process: Submit exception request to architecture review board with performance benchmarks or deprecation timeline +- Exception process: Document exception in API specification and architectural decision log +- Exception process: Re-evaluate exceptions quarterly to determine if circumstances have changed \ No newline at end of file diff --git a/docs/adr/4dbd1df3-2cdc-4996-a874-d94ee89c9983-adopt-distributed-tracing-for-public-api-observability-internal-private-use.md b/docs/adr/4dbd1df3-2cdc-4996-a874-d94ee89c9983-adopt-distributed-tracing-for-public-api-observability-internal-private-use.md new file mode 100644 index 000000000..19ad5ff08 --- /dev/null +++ b/docs/adr/4dbd1df3-2cdc-4996-a874-d94ee89c9983-adopt-distributed-tracing-for-public-api-observability-internal-private-use.md @@ -0,0 +1,123 @@ +# Adopt Distributed Tracing for Public API Observability: Internal Private Use + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + +## Context + +- Public and external APIs require comprehensive observability to diagnose issues across distributed system boundaries where direct debugging is not feasible +- The pattern was detected with 90.27% confidence across 3 critical infrastructure files (terminal.rs, lib.rs, stall.rs) in the fabro crate ecosystem, indicating a systematic approach to tracing +- The facet 'obs.tracing' suggests this pattern specifically addresses observability through distributed tracing instrumentation +- External API consumers and integration partners require visibility into request flows, latency bottlenecks, and error propagation paths +- Without standardized tracing, debugging cross-service issues becomes prohibitively expensive and time-consuming + +## Problem Statement + +Public and external APIs lack consistent observability instrumentation, making it difficult to diagnose performance issues, trace request flows across service boundaries, and provide actionable debugging information to external consumers. This creates operational blind spots and increases mean time to resolution (MTTR) for production incidents. + +## Decision + +1. MAY: Internal/private APIs MAY use simplified tracing instrumentation with reduced attribute sets + +## Policy Block + +- MAY Internal/private APIs MAY use simplified tracing instrumentation with reduced attribute sets + +In scope: +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +Out of scope: +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +Exceptions: +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +## Rationale + +- The pattern detection across 3 files with 90.27% confidence indicates this is an established architectural practice in the codebase, not an isolated implementation +- Distributed tracing provides end-to-end visibility that is essential for debugging issues in microservices architectures where requests span multiple services +- Standardizing on obs.tracing facet ensures consistent instrumentation patterns across the organization, reducing cognitive load and improving debugging efficiency +- External API consumers benefit from trace IDs in error responses, enabling them to provide actionable information when reporting issues + +## Consequences + +Positive: +- Reduced mean time to resolution (MTTR) for production incidents through comprehensive request flow visibility +- Improved developer experience with standardized tracing patterns across all public APIs +- Enhanced external consumer support with traceable request identifiers for issue reporting +- Better capacity planning and performance optimization through latency distribution analysis + +Negative: +- Increased operational complexity with additional infrastructure for trace collection and storage +- Minor performance overhead (typically 1-3% latency increase) from tracing instrumentation +- Development time investment required to instrument existing APIs and train teams on tracing best practices +- Potential for trace data explosion requiring careful sampling strategy and retention policies + +## Alternatives + +- Structured logging only without distributed tracing (rejected) + Rejected because: Logs lack the causal relationships and timing information needed to reconstruct request flows across service boundaries. Correlation requires manual log aggregation which is error-prone and time-consuming. + When valid: May be sufficient for monolithic applications with single-service request handling +- Metrics-based observability with counters and histograms (rejected) + Rejected because: Metrics provide aggregate statistics but cannot trace individual request paths or identify specific failure scenarios. Complementary to tracing but insufficient as sole observability strategy. + When valid: Appropriate for high-level health monitoring and alerting, should be used alongside tracing +- Vendor-specific tracing solutions (e.g., AWS X-Ray, Datadog APM) (deferred) + Rejected because: Not rejected but deferred to implementation phase. Vendor choice should align with existing infrastructure while maintaining standard trace context propagation. + When valid: Valid if vendor solution supports OpenTelemetry or W3C Trace Context standards for interoperability + +## Risks + +- Trace data volume may exceed storage capacity or budget constraints, leading to incomplete observability + Mitigation: Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests. Establish retention policies (e.g., 7 days for all traces, 30 days for error traces). + Owner: Platform Engineering Team +- Inconsistent tracing implementation across teams may result in fragmented observability + Mitigation: Provide shared tracing libraries and middleware with sensible defaults. Conduct code reviews specifically checking for tracing compliance. Create runbooks and training materials. + Owner: Engineering Team +- Performance overhead from tracing may impact latency-sensitive APIs + Mitigation: Benchmark tracing overhead during implementation. Use asynchronous trace export to minimize request latency impact. Allow exceptions for proven performance-critical paths. + Owner: API Development Team + +## Implementation Notes + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + +## Continuation Context + + +Verify commands: +- grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l +- grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' +- cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' + +Accept when: +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +## Enforcement + +- Verified by: Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis +- Verified by: Code review checklist includes verification of distributed tracing implementation +- Verified by: Integration test suite validates trace context propagation and span attribute completeness +- Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation +- Violation handling: Code review blocks merge if tracing requirements are not met without documented exception +- Violation handling: Quarterly audits identify non-compliant APIs with remediation plans required within 30 days +- Exception process: Submit exception request to architecture review board with performance benchmarks or deprecation timeline +- Exception process: Document exception in API specification and architectural decision log +- Exception process: Re-evaluate exceptions quarterly to determine if circumstances have changed \ No newline at end of file diff --git a/docs/adr/57fb44e9-c959-408b-9e6a-87b058b1ecc7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-support.md b/docs/adr/57fb44e9-c959-408b-9e6a-87b058b1ecc7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-support.md new file mode 100644 index 000000000..a7c05af8f --- /dev/null +++ b/docs/adr/57fb44e9-c959-408b-9e6a-87b058b1ecc7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-support.md @@ -0,0 +1,115 @@ +# Standardize Server-Sent Events (SSE) for Real-Time Configuration and Runtime State Updates: Sse Implementations Support + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Context + +- The application requires real-time updates of runtime execution state and configuration changes to be pushed to client interfaces without polling +- Multiple route handlers (run-stages.tsx, run-files.tsx) need to stream execution progress and file processing status to web clients +- A centralized SSE implementation pattern (sse.ts) has emerged to handle persistent connections and event streaming across different runtime contexts +- The pattern provides a consistent mechanism for broadcasting configuration changes and environment state updates during execution workflows + +## Problem Statement + +Applications need a reliable, standardized mechanism to push real-time runtime state, configuration updates, and execution progress to clients without the overhead and latency of polling-based approaches, while maintaining consistent event handling across multiple execution contexts. + +## Decision + +1. MAY: SSE implementations MAY support event filtering or subscription patterns to reduce bandwidth for clients interested in specific event types + +## Policy Block + +- MAY SSE implementations MAY support event filtering or subscription patterns to reduce bandwidth for clients interested in specific event types + +In scope: +- Web route handlers that expose runtime execution progress (run-stages, run-files, etc.) +- Configuration management endpoints that broadcast environment or setting changes +- Real-time monitoring interfaces for execution workflows +- Client-facing APIs that require push-based state synchronization + +Out of scope: +- Bidirectional communication patterns requiring client-to-server messaging (use WebSockets instead) +- Binary data streaming or large file transfers +- Internal service-to-service communication (use message queues or gRPC) +- Static configuration loading at application startup + +## Rationale + +- Pattern detected across 3 files with 91.07% confidence indicates a deliberate architectural choice for real-time state communication +- SSE provides a lightweight, HTTP-based protocol that works through firewalls and proxies without requiring WebSocket infrastructure +- Centralized SSE library (sse.ts) ensures consistent event formatting, error handling, and connection lifecycle management across multiple route handlers +- The pattern aligns with unidirectional data flow from server to client, which matches the use case of broadcasting runtime state and configuration updates + +## Consequences + +Positive: +- Clients receive real-time updates without polling overhead, reducing server load and improving responsiveness +- Consistent SSE implementation across routes reduces code duplication and simplifies maintenance +- HTTP-based protocol ensures compatibility with existing infrastructure (load balancers, proxies, CDNs) +- Automatic reconnection support in browsers provides resilience against transient network failures + +Negative: +- SSE connections are unidirectional, requiring separate HTTP requests for client-to-server communication +- Long-lived connections may require infrastructure tuning (timeouts, connection limits) to support many concurrent clients +- Browser connection limits (typically 6 per domain) may constrain applications with multiple simultaneous SSE streams +- Debugging and monitoring SSE connections can be more complex than traditional request-response patterns + +## Alternatives + +- Use polling-based approach with periodic HTTP requests to fetch runtime state updates (rejected) + Rejected because: Polling introduces latency, increases server load with redundant requests, and provides poor user experience for real-time execution monitoring + When valid: Acceptable for low-frequency updates (>30 seconds) or when SSE infrastructure is unavailable +- Implement WebSocket-based bidirectional communication for all real-time features (rejected) + Rejected because: WebSockets add complexity and infrastructure requirements when only unidirectional server-to-client communication is needed for configuration and state updates + When valid: Appropriate when bidirectional real-time communication is required (e.g., collaborative editing, chat) +- Use GraphQL subscriptions for real-time data updates (rejected) + Rejected because: Adds GraphQL infrastructure overhead and complexity when simple event streaming suffices for runtime state broadcasting + When valid: Valid when already using GraphQL extensively and need subscription-based data synchronization + +## Risks + +- SSE connections may be terminated by intermediate proxies or load balancers with aggressive timeout policies + Mitigation: Implement periodic heartbeat messages and configure infrastructure timeouts appropriately (e.g., 60+ seconds). Document required infrastructure settings. + Owner: Engineering team and DevOps +- Memory leaks or resource exhaustion if SSE connections are not properly cleaned up on client disconnect + Mitigation: Implement robust connection lifecycle management with cleanup handlers. Add monitoring for open connection counts and memory usage. + Owner: Engineering team +- Browser connection limits may prevent multiple SSE streams from functioning simultaneously + Mitigation: Multiplex multiple event types over a single SSE connection where possible. Document connection usage and provide guidance on connection management. + Owner: Engineering team + +## Implementation Notes + +- Create a centralized SSE utility module (e.g., lib/sse.ts) that provides connection setup, event formatting, and cleanup helpers +- Ensure route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- Document SSE usage patterns and provide examples for common scenarios (execution progress, configuration updates) +- Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +## Continuation Context + + +Verify commands: +- grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l +- grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules +- find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" + +Accept when: +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion + +## Enforcement + +- Verified by: Code review checklist verifying SSE usage for new real-time endpoints +- Verified by: Automated linting rules to detect missing Content-Type headers on streaming routes +- Verified by: Integration tests validating SSE connection lifecycle and event delivery +- Violation handling: Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable +- Violation handling: Route handlers with SSE connections missing proper cleanup must be fixed before merge +- Violation handling: Violations detected in code review are flagged and require revision +- Exception process: Document technical justification for alternative approach (e.g., WebSocket requirement, infrastructure constraints) +- Exception process: Obtain approval from technical lead or architect +- Exception process: Add architectural decision comment in code explaining exception rationale \ No newline at end of file diff --git a/docs/adr/5b2054c7-ba4e-44f2-b4ee-faacb67305d6-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-utilize.md b/docs/adr/5b2054c7-ba4e-44f2-b4ee-faacb67305d6-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-utilize.md new file mode 100644 index 000000000..cd430c156 --- /dev/null +++ b/docs/adr/5b2054c7-ba4e-44f2-b4ee-faacb67305d6-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-implementations-utilize.md @@ -0,0 +1,115 @@ +# Standardize Server-Sent Events (SSE) for Real-Time Configuration and Runtime State Updates: Sse Implementations Utilize + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Context + +- The application requires real-time updates of runtime execution state and configuration changes to be pushed to client interfaces without polling +- Multiple route handlers (run-stages.tsx, run-files.tsx) need to stream execution progress and file processing status to web clients +- A centralized SSE implementation pattern (sse.ts) has emerged to handle persistent connections and event streaming across different runtime contexts +- The pattern provides a consistent mechanism for broadcasting configuration changes and environment state updates during execution workflows + +## Problem Statement + +Applications need a reliable, standardized mechanism to push real-time runtime state, configuration updates, and execution progress to clients without the overhead and latency of polling-based approaches, while maintaining consistent event handling across multiple execution contexts. + +## Decision + +1. MUST: SSE implementations MUST utilize a centralized library module (e.g., sse.ts) to ensure consistent event formatting and connection management + +## Policy Block + +- MUST SSE implementations MUST utilize a centralized library module (e.g., sse.ts) to ensure consistent event formatting and connection management + +In scope: +- Web route handlers that expose runtime execution progress (run-stages, run-files, etc.) +- Configuration management endpoints that broadcast environment or setting changes +- Real-time monitoring interfaces for execution workflows +- Client-facing APIs that require push-based state synchronization + +Out of scope: +- Bidirectional communication patterns requiring client-to-server messaging (use WebSockets instead) +- Binary data streaming or large file transfers +- Internal service-to-service communication (use message queues or gRPC) +- Static configuration loading at application startup + +## Rationale + +- Pattern detected across 3 files with 91.07% confidence indicates a deliberate architectural choice for real-time state communication +- SSE provides a lightweight, HTTP-based protocol that works through firewalls and proxies without requiring WebSocket infrastructure +- Centralized SSE library (sse.ts) ensures consistent event formatting, error handling, and connection lifecycle management across multiple route handlers +- The pattern aligns with unidirectional data flow from server to client, which matches the use case of broadcasting runtime state and configuration updates + +## Consequences + +Positive: +- Clients receive real-time updates without polling overhead, reducing server load and improving responsiveness +- Consistent SSE implementation across routes reduces code duplication and simplifies maintenance +- HTTP-based protocol ensures compatibility with existing infrastructure (load balancers, proxies, CDNs) +- Automatic reconnection support in browsers provides resilience against transient network failures + +Negative: +- SSE connections are unidirectional, requiring separate HTTP requests for client-to-server communication +- Long-lived connections may require infrastructure tuning (timeouts, connection limits) to support many concurrent clients +- Browser connection limits (typically 6 per domain) may constrain applications with multiple simultaneous SSE streams +- Debugging and monitoring SSE connections can be more complex than traditional request-response patterns + +## Alternatives + +- Use polling-based approach with periodic HTTP requests to fetch runtime state updates (rejected) + Rejected because: Polling introduces latency, increases server load with redundant requests, and provides poor user experience for real-time execution monitoring + When valid: Acceptable for low-frequency updates (>30 seconds) or when SSE infrastructure is unavailable +- Implement WebSocket-based bidirectional communication for all real-time features (rejected) + Rejected because: WebSockets add complexity and infrastructure requirements when only unidirectional server-to-client communication is needed for configuration and state updates + When valid: Appropriate when bidirectional real-time communication is required (e.g., collaborative editing, chat) +- Use GraphQL subscriptions for real-time data updates (rejected) + Rejected because: Adds GraphQL infrastructure overhead and complexity when simple event streaming suffices for runtime state broadcasting + When valid: Valid when already using GraphQL extensively and need subscription-based data synchronization + +## Risks + +- SSE connections may be terminated by intermediate proxies or load balancers with aggressive timeout policies + Mitigation: Implement periodic heartbeat messages and configure infrastructure timeouts appropriately (e.g., 60+ seconds). Document required infrastructure settings. + Owner: Engineering team and DevOps +- Memory leaks or resource exhaustion if SSE connections are not properly cleaned up on client disconnect + Mitigation: Implement robust connection lifecycle management with cleanup handlers. Add monitoring for open connection counts and memory usage. + Owner: Engineering team +- Browser connection limits may prevent multiple SSE streams from functioning simultaneously + Mitigation: Multiplex multiple event types over a single SSE connection where possible. Document connection usage and provide guidance on connection management. + Owner: Engineering team + +## Implementation Notes + +- Create a centralized SSE utility module (e.g., lib/sse.ts) that provides connection setup, event formatting, and cleanup helpers +- Ensure route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- Document SSE usage patterns and provide examples for common scenarios (execution progress, configuration updates) +- Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +## Continuation Context + + +Verify commands: +- grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l +- grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules +- find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" + +Accept when: +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion + +## Enforcement + +- Verified by: Code review checklist verifying SSE usage for new real-time endpoints +- Verified by: Automated linting rules to detect missing Content-Type headers on streaming routes +- Verified by: Integration tests validating SSE connection lifecycle and event delivery +- Violation handling: Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable +- Violation handling: Route handlers with SSE connections missing proper cleanup must be fixed before merge +- Violation handling: Violations detected in code review are flagged and require revision +- Exception process: Document technical justification for alternative approach (e.g., WebSocket requirement, infrastructure constraints) +- Exception process: Obtain approval from technical lead or architect +- Exception process: Add architectural decision comment in code explaining exception rationale \ No newline at end of file diff --git a/docs/adr/67d628ca-c9de-4922-93cb-1fd54450f492-standardize-service-boundary-logging-for-observability-service-boundary-logs.md b/docs/adr/67d628ca-c9de-4922-93cb-1fd54450f492-standardize-service-boundary-logging-for-observability-service-boundary-logs.md new file mode 100644 index 000000000..bd002f8aa --- /dev/null +++ b/docs/adr/67d628ca-c9de-4922-93cb-1fd54450f492-standardize-service-boundary-logging-for-observability-service-boundary-logs.md @@ -0,0 +1,131 @@ +# Standardize Service Boundary Logging for Observability: Service Boundary Logs + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all service implementations and applies to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + +## Context + +- The codebase exhibits a consistent pattern of logging at service boundaries across multiple components (fabro-server and fabro-web), indicating an architectural decision to instrument entry and exit points of service operations +- Service boundary logging is critical for distributed system observability, enabling request tracing, performance monitoring, and debugging across service interactions +- The pattern appears in both backend Rust server implementations and frontend TypeScript query layers, suggesting a cross-stack architectural concern +- With 3 files showing this pattern at 90.17% confidence, this represents a deliberate architectural choice rather than ad-hoc logging practices +- The facet 'boundaries.service_definitions' indicates this pattern specifically targets the interfaces where services are defined and exposed + +## Problem Statement + +Without standardized logging at service boundaries, distributed systems suffer from poor observability, making it difficult to trace requests across services, diagnose performance bottlenecks, identify failure points, and understand system behavior in production. Ad-hoc logging approaches lead to inconsistent log formats, missing context, and gaps in observability coverage. + +## Decision + +1. MUST: Service boundary logs MUST include correlation identifiers (request IDs, trace IDs) to enable distributed tracing across service calls + +## Policy Block + +- MUST Service boundary logs MUST include correlation identifiers (request IDs, trace IDs) to enable distributed tracing across service calls + +In scope: +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +Out of scope: +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +Exceptions: +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +## Rationale + +- The detection of this pattern across 3 files with 90.17% confidence indicates this is an established architectural practice in the codebase, representing proven value in production systems +- Service boundary logging provides the minimal necessary observability coverage for distributed systems without requiring invasive instrumentation throughout the codebase +- Standardizing logging at boundaries creates consistent observability across heterogeneous technology stacks (Rust backend, TypeScript frontend), enabling unified monitoring and debugging workflows +- This approach balances observability needs with performance concerns by focusing logging on key interaction points rather than instrumenting every function + +## Consequences + +Positive: +- Improved ability to trace requests through distributed system components, reducing mean time to resolution (MTTR) for production issues +- Consistent log structure across services enables automated log aggregation, analysis, and alerting +- Clear observability boundaries make it easier for developers to understand where logging is required versus optional +- Correlation identifiers at service boundaries enable powerful distributed tracing capabilities without complex instrumentation + +Negative: +- Additional logging overhead at service boundaries may impact latency for high-frequency operations, requiring careful performance monitoring +- Developers must remember to implement boundary logging for all new service endpoints, creating potential for human error +- Log volume may increase significantly in high-traffic systems, requiring investment in log management infrastructure +- Maintaining consistent logging patterns across multiple technology stacks (Rust, TypeScript) requires ongoing coordination and code review discipline + +## Alternatives + +- Comprehensive instrumentation logging throughout the entire codebase, not just at service boundaries (rejected) + Rejected because: Creates excessive log volume, performance overhead, and maintenance burden. Most internal function calls do not provide actionable observability value compared to boundary logging. + When valid: May be appropriate for specific critical subsystems during active debugging or performance optimization efforts, but not as a general architectural pattern +- Rely solely on distributed tracing frameworks (OpenTelemetry, Jaeger) without explicit boundary logging (rejected) + Rejected because: Tracing frameworks add complexity and dependencies, may not be available in all environments, and logs provide complementary information (detailed context, error messages) that traces don't capture well + When valid: Could be adopted in addition to boundary logging as systems mature, but should not replace foundational logging practices +- Implement logging through middleware/interceptors only, without explicit logging in service code (deferred) + Rejected because: While middleware can handle many boundary logging concerns, some service-specific context and business logic details are best logged within the service implementation itself + When valid: Middleware-based logging should be used where possible for consistency, with explicit service logging for context that middleware cannot capture + +## Risks + +- Inconsistent implementation across teams and services leads to gaps in observability coverage + Mitigation: Implement automated linting rules to detect missing boundary logging, provide code templates and examples, include boundary logging checks in code review checklists + Owner: Engineering team leads and platform team +- Sensitive data leakage through logs if developers are not careful about what context they include + Mitigation: Implement automated PII detection in logs, provide sanitization utilities, conduct security training on logging best practices, enable log redaction in production + Owner: Security team and engineering team +- Performance degradation in high-throughput services due to synchronous logging overhead + Mitigation: Use async logging frameworks, implement log sampling for high-frequency endpoints, monitor logging performance impact, establish performance budgets for logging overhead + Owner: Platform team and service owners + +## Implementation Notes + +- Create shared logging utilities or macros that encapsulate boundary logging patterns, making it easy for developers to add consistent logging with minimal boilerplate +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' that provide zero-cost abstractions and async logging capabilities +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters for consistency with backend logs +- Establish a standard set of log fields for service boundaries: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems (ELK, Splunk, CloudWatch) to parse and index boundary logs for efficient querying and alerting +- Document logging patterns in service templates and starter kits to ensure new services follow established practices from the beginning + +## Continuation Context + + +Verify commands: +- grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l +- rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' +- find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' + +Accept when: +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +## Enforcement + +- Verified by: Automated code review checks using static analysis tools (clippy for Rust, ESLint for TypeScript) with custom rules for boundary logging +- Verified by: Manual code review checklist items requiring reviewers to verify boundary logging presence +- Verified by: CI pipeline integration tests that verify log output from service endpoints contains required fields +- Verified by: Periodic observability audits reviewing log coverage across services +- Violation handling: CI pipeline warnings for missing boundary logging patterns (non-blocking initially, blocking after grace period) +- Violation handling: Code review feedback requiring addition of boundary logging before merge approval +- Violation handling: Quarterly observability reports identifying services with insufficient logging coverage +- Violation handling: Escalation to architecture review for repeated violations or services with poor observability +- Exception process: Submit exception request to team lead or architect with documented rationale (performance impact, alternative observability approach) +- Exception process: For performance-critical paths, provide benchmark data demonstrating logging overhead exceeds acceptable latency budget +- Exception process: Document approved exceptions in service README and architecture decision log +- Exception process: Review exceptions quarterly to determine if mitigations (async logging, sampling) can eliminate the need for exception \ No newline at end of file diff --git a/docs/adr/860946d2-f81f-45ed-b3f2-25c8d5f4956f-standardize-service-boundary-logging-for-observability-log-levels-service.md b/docs/adr/860946d2-f81f-45ed-b3f2-25c8d5f4956f-standardize-service-boundary-logging-for-observability-log-levels-service.md new file mode 100644 index 000000000..81c2e3da5 --- /dev/null +++ b/docs/adr/860946d2-f81f-45ed-b3f2-25c8d5f4956f-standardize-service-boundary-logging-for-observability-log-levels-service.md @@ -0,0 +1,131 @@ +# Standardize Service Boundary Logging for Observability: Log Levels Service + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all service implementations and applies to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + +## Context + +- The codebase exhibits a consistent pattern of logging at service boundaries across multiple components (fabro-server and fabro-web), indicating an architectural decision to instrument entry and exit points of service operations +- Service boundary logging is critical for distributed system observability, enabling request tracing, performance monitoring, and debugging across service interactions +- The pattern appears in both backend Rust server implementations and frontend TypeScript query layers, suggesting a cross-stack architectural concern +- With 3 files showing this pattern at 90.17% confidence, this represents a deliberate architectural choice rather than ad-hoc logging practices +- The facet 'boundaries.service_definitions' indicates this pattern specifically targets the interfaces where services are defined and exposed + +## Problem Statement + +Without standardized logging at service boundaries, distributed systems suffer from poor observability, making it difficult to trace requests across services, diagnose performance bottlenecks, identify failure points, and understand system behavior in production. Ad-hoc logging approaches lead to inconsistent log formats, missing context, and gaps in observability coverage. + +## Decision + +1. SHOULD: Log levels at service boundaries SHOULD follow standard conventions: INFO for successful operations, WARN for degraded operations, ERROR for failures + +## Policy Block + +- SHOULD Log levels at service boundaries SHOULD follow standard conventions: INFO for successful operations, WARN for degraded operations, ERROR for failures + +In scope: +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +Out of scope: +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +Exceptions: +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +## Rationale + +- The detection of this pattern across 3 files with 90.17% confidence indicates this is an established architectural practice in the codebase, representing proven value in production systems +- Service boundary logging provides the minimal necessary observability coverage for distributed systems without requiring invasive instrumentation throughout the codebase +- Standardizing logging at boundaries creates consistent observability across heterogeneous technology stacks (Rust backend, TypeScript frontend), enabling unified monitoring and debugging workflows +- This approach balances observability needs with performance concerns by focusing logging on key interaction points rather than instrumenting every function + +## Consequences + +Positive: +- Improved ability to trace requests through distributed system components, reducing mean time to resolution (MTTR) for production issues +- Consistent log structure across services enables automated log aggregation, analysis, and alerting +- Clear observability boundaries make it easier for developers to understand where logging is required versus optional +- Correlation identifiers at service boundaries enable powerful distributed tracing capabilities without complex instrumentation + +Negative: +- Additional logging overhead at service boundaries may impact latency for high-frequency operations, requiring careful performance monitoring +- Developers must remember to implement boundary logging for all new service endpoints, creating potential for human error +- Log volume may increase significantly in high-traffic systems, requiring investment in log management infrastructure +- Maintaining consistent logging patterns across multiple technology stacks (Rust, TypeScript) requires ongoing coordination and code review discipline + +## Alternatives + +- Comprehensive instrumentation logging throughout the entire codebase, not just at service boundaries (rejected) + Rejected because: Creates excessive log volume, performance overhead, and maintenance burden. Most internal function calls do not provide actionable observability value compared to boundary logging. + When valid: May be appropriate for specific critical subsystems during active debugging or performance optimization efforts, but not as a general architectural pattern +- Rely solely on distributed tracing frameworks (OpenTelemetry, Jaeger) without explicit boundary logging (rejected) + Rejected because: Tracing frameworks add complexity and dependencies, may not be available in all environments, and logs provide complementary information (detailed context, error messages) that traces don't capture well + When valid: Could be adopted in addition to boundary logging as systems mature, but should not replace foundational logging practices +- Implement logging through middleware/interceptors only, without explicit logging in service code (deferred) + Rejected because: While middleware can handle many boundary logging concerns, some service-specific context and business logic details are best logged within the service implementation itself + When valid: Middleware-based logging should be used where possible for consistency, with explicit service logging for context that middleware cannot capture + +## Risks + +- Inconsistent implementation across teams and services leads to gaps in observability coverage + Mitigation: Implement automated linting rules to detect missing boundary logging, provide code templates and examples, include boundary logging checks in code review checklists + Owner: Engineering team leads and platform team +- Sensitive data leakage through logs if developers are not careful about what context they include + Mitigation: Implement automated PII detection in logs, provide sanitization utilities, conduct security training on logging best practices, enable log redaction in production + Owner: Security team and engineering team +- Performance degradation in high-throughput services due to synchronous logging overhead + Mitigation: Use async logging frameworks, implement log sampling for high-frequency endpoints, monitor logging performance impact, establish performance budgets for logging overhead + Owner: Platform team and service owners + +## Implementation Notes + +- Create shared logging utilities or macros that encapsulate boundary logging patterns, making it easy for developers to add consistent logging with minimal boilerplate +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' that provide zero-cost abstractions and async logging capabilities +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters for consistency with backend logs +- Establish a standard set of log fields for service boundaries: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems (ELK, Splunk, CloudWatch) to parse and index boundary logs for efficient querying and alerting +- Document logging patterns in service templates and starter kits to ensure new services follow established practices from the beginning + +## Continuation Context + + +Verify commands: +- grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l +- rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' +- find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' + +Accept when: +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +## Enforcement + +- Verified by: Automated code review checks using static analysis tools (clippy for Rust, ESLint for TypeScript) with custom rules for boundary logging +- Verified by: Manual code review checklist items requiring reviewers to verify boundary logging presence +- Verified by: CI pipeline integration tests that verify log output from service endpoints contains required fields +- Verified by: Periodic observability audits reviewing log coverage across services +- Violation handling: CI pipeline warnings for missing boundary logging patterns (non-blocking initially, blocking after grace period) +- Violation handling: Code review feedback requiring addition of boundary logging before merge approval +- Violation handling: Quarterly observability reports identifying services with insufficient logging coverage +- Violation handling: Escalation to architecture review for repeated violations or services with poor observability +- Exception process: Submit exception request to team lead or architect with documented rationale (performance impact, alternative observability approach) +- Exception process: For performance-critical paths, provide benchmark data demonstrating logging overhead exceeds acceptable latency budget +- Exception process: Document approved exceptions in service README and architecture decision log +- Exception process: Review exceptions quarterly to determine if mitigations (async logging, sampling) can eliminate the need for exception \ No newline at end of file diff --git a/docs/adr/8743e5ed-c257-4d49-8295-a4481f8523b5-standardize-service-boundary-logging-for-observability-service-boundary-logging.md b/docs/adr/8743e5ed-c257-4d49-8295-a4481f8523b5-standardize-service-boundary-logging-for-observability-service-boundary-logging.md new file mode 100644 index 000000000..8e3306669 --- /dev/null +++ b/docs/adr/8743e5ed-c257-4d49-8295-a4481f8523b5-standardize-service-boundary-logging-for-observability-service-boundary-logging.md @@ -0,0 +1,131 @@ +# Standardize Service Boundary Logging for Observability: Service Boundary Logging + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all service implementations and applies to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + +## Context + +- The codebase exhibits a consistent pattern of logging at service boundaries across multiple components (fabro-server and fabro-web), indicating an architectural decision to instrument entry and exit points of service operations +- Service boundary logging is critical for distributed system observability, enabling request tracing, performance monitoring, and debugging across service interactions +- The pattern appears in both backend Rust server implementations and frontend TypeScript query layers, suggesting a cross-stack architectural concern +- With 3 files showing this pattern at 90.17% confidence, this represents a deliberate architectural choice rather than ad-hoc logging practices +- The facet 'boundaries.service_definitions' indicates this pattern specifically targets the interfaces where services are defined and exposed + +## Problem Statement + +Without standardized logging at service boundaries, distributed systems suffer from poor observability, making it difficult to trace requests across services, diagnose performance bottlenecks, identify failure points, and understand system behavior in production. Ad-hoc logging approaches lead to inconsistent log formats, missing context, and gaps in observability coverage. + +## Decision + +1. SHOULD: Service boundary logging SHOULD use structured logging formats (JSON, key-value pairs) rather than unstructured text to enable automated log analysis + +## Policy Block + +- SHOULD Service boundary logging SHOULD use structured logging formats (JSON, key-value pairs) rather than unstructured text to enable automated log analysis + +In scope: +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +Out of scope: +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +Exceptions: +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +## Rationale + +- The detection of this pattern across 3 files with 90.17% confidence indicates this is an established architectural practice in the codebase, representing proven value in production systems +- Service boundary logging provides the minimal necessary observability coverage for distributed systems without requiring invasive instrumentation throughout the codebase +- Standardizing logging at boundaries creates consistent observability across heterogeneous technology stacks (Rust backend, TypeScript frontend), enabling unified monitoring and debugging workflows +- This approach balances observability needs with performance concerns by focusing logging on key interaction points rather than instrumenting every function + +## Consequences + +Positive: +- Improved ability to trace requests through distributed system components, reducing mean time to resolution (MTTR) for production issues +- Consistent log structure across services enables automated log aggregation, analysis, and alerting +- Clear observability boundaries make it easier for developers to understand where logging is required versus optional +- Correlation identifiers at service boundaries enable powerful distributed tracing capabilities without complex instrumentation + +Negative: +- Additional logging overhead at service boundaries may impact latency for high-frequency operations, requiring careful performance monitoring +- Developers must remember to implement boundary logging for all new service endpoints, creating potential for human error +- Log volume may increase significantly in high-traffic systems, requiring investment in log management infrastructure +- Maintaining consistent logging patterns across multiple technology stacks (Rust, TypeScript) requires ongoing coordination and code review discipline + +## Alternatives + +- Comprehensive instrumentation logging throughout the entire codebase, not just at service boundaries (rejected) + Rejected because: Creates excessive log volume, performance overhead, and maintenance burden. Most internal function calls do not provide actionable observability value compared to boundary logging. + When valid: May be appropriate for specific critical subsystems during active debugging or performance optimization efforts, but not as a general architectural pattern +- Rely solely on distributed tracing frameworks (OpenTelemetry, Jaeger) without explicit boundary logging (rejected) + Rejected because: Tracing frameworks add complexity and dependencies, may not be available in all environments, and logs provide complementary information (detailed context, error messages) that traces don't capture well + When valid: Could be adopted in addition to boundary logging as systems mature, but should not replace foundational logging practices +- Implement logging through middleware/interceptors only, without explicit logging in service code (deferred) + Rejected because: While middleware can handle many boundary logging concerns, some service-specific context and business logic details are best logged within the service implementation itself + When valid: Middleware-based logging should be used where possible for consistency, with explicit service logging for context that middleware cannot capture + +## Risks + +- Inconsistent implementation across teams and services leads to gaps in observability coverage + Mitigation: Implement automated linting rules to detect missing boundary logging, provide code templates and examples, include boundary logging checks in code review checklists + Owner: Engineering team leads and platform team +- Sensitive data leakage through logs if developers are not careful about what context they include + Mitigation: Implement automated PII detection in logs, provide sanitization utilities, conduct security training on logging best practices, enable log redaction in production + Owner: Security team and engineering team +- Performance degradation in high-throughput services due to synchronous logging overhead + Mitigation: Use async logging frameworks, implement log sampling for high-frequency endpoints, monitor logging performance impact, establish performance budgets for logging overhead + Owner: Platform team and service owners + +## Implementation Notes + +- Create shared logging utilities or macros that encapsulate boundary logging patterns, making it easy for developers to add consistent logging with minimal boilerplate +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' that provide zero-cost abstractions and async logging capabilities +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters for consistency with backend logs +- Establish a standard set of log fields for service boundaries: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems (ELK, Splunk, CloudWatch) to parse and index boundary logs for efficient querying and alerting +- Document logging patterns in service templates and starter kits to ensure new services follow established practices from the beginning + +## Continuation Context + + +Verify commands: +- grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l +- rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' +- find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' + +Accept when: +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +## Enforcement + +- Verified by: Automated code review checks using static analysis tools (clippy for Rust, ESLint for TypeScript) with custom rules for boundary logging +- Verified by: Manual code review checklist items requiring reviewers to verify boundary logging presence +- Verified by: CI pipeline integration tests that verify log output from service endpoints contains required fields +- Verified by: Periodic observability audits reviewing log coverage across services +- Violation handling: CI pipeline warnings for missing boundary logging patterns (non-blocking initially, blocking after grace period) +- Violation handling: Code review feedback requiring addition of boundary logging before merge approval +- Violation handling: Quarterly observability reports identifying services with insufficient logging coverage +- Violation handling: Escalation to architecture review for repeated violations or services with poor observability +- Exception process: Submit exception request to team lead or architect with documented rationale (performance impact, alternative observability approach) +- Exception process: For performance-critical paths, provide benchmark data demonstrating logging overhead exceeds acceptable latency budget +- Exception process: Document approved exceptions in service README and architecture decision log +- Exception process: Review exceptions quarterly to determine if mitigations (async logging, sampling) can eliminate the need for exception \ No newline at end of file diff --git a/docs/adr/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76-adopt-distributed-tracing-for-public-api-observability-tracing-instrumentation-not.md b/docs/adr/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76-adopt-distributed-tracing-for-public-api-observability-tracing-instrumentation-not.md new file mode 100644 index 000000000..ed9aa68d8 --- /dev/null +++ b/docs/adr/9607dcb3-91d3-44b9-85f9-ccb1fa25bd76-adopt-distributed-tracing-for-public-api-observability-tracing-instrumentation-not.md @@ -0,0 +1,123 @@ +# Adopt Distributed Tracing for Public API Observability: Tracing Instrumentation Not + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + +## Context + +- Public and external APIs require comprehensive observability to diagnose issues across distributed system boundaries where direct debugging is not feasible +- The pattern was detected with 90.27% confidence across 3 critical infrastructure files (terminal.rs, lib.rs, stall.rs) in the fabro crate ecosystem, indicating a systematic approach to tracing +- The facet 'obs.tracing' suggests this pattern specifically addresses observability through distributed tracing instrumentation +- External API consumers and integration partners require visibility into request flows, latency bottlenecks, and error propagation paths +- Without standardized tracing, debugging cross-service issues becomes prohibitively expensive and time-consuming + +## Problem Statement + +Public and external APIs lack consistent observability instrumentation, making it difficult to diagnose performance issues, trace request flows across service boundaries, and provide actionable debugging information to external consumers. This creates operational blind spots and increases mean time to resolution (MTTR) for production incidents. + +## Decision + +1. MUST_NOT: Tracing instrumentation MUST NOT log sensitive data (passwords, tokens, PII) in span attributes or events + +## Policy Block + +- MUST_NOT Tracing instrumentation MUST NOT log sensitive data (passwords, tokens, PII) in span attributes or events + +In scope: +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +Out of scope: +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +Exceptions: +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +## Rationale + +- The pattern detection across 3 files with 90.27% confidence indicates this is an established architectural practice in the codebase, not an isolated implementation +- Distributed tracing provides end-to-end visibility that is essential for debugging issues in microservices architectures where requests span multiple services +- Standardizing on obs.tracing facet ensures consistent instrumentation patterns across the organization, reducing cognitive load and improving debugging efficiency +- External API consumers benefit from trace IDs in error responses, enabling them to provide actionable information when reporting issues + +## Consequences + +Positive: +- Reduced mean time to resolution (MTTR) for production incidents through comprehensive request flow visibility +- Improved developer experience with standardized tracing patterns across all public APIs +- Enhanced external consumer support with traceable request identifiers for issue reporting +- Better capacity planning and performance optimization through latency distribution analysis + +Negative: +- Increased operational complexity with additional infrastructure for trace collection and storage +- Minor performance overhead (typically 1-3% latency increase) from tracing instrumentation +- Development time investment required to instrument existing APIs and train teams on tracing best practices +- Potential for trace data explosion requiring careful sampling strategy and retention policies + +## Alternatives + +- Structured logging only without distributed tracing (rejected) + Rejected because: Logs lack the causal relationships and timing information needed to reconstruct request flows across service boundaries. Correlation requires manual log aggregation which is error-prone and time-consuming. + When valid: May be sufficient for monolithic applications with single-service request handling +- Metrics-based observability with counters and histograms (rejected) + Rejected because: Metrics provide aggregate statistics but cannot trace individual request paths or identify specific failure scenarios. Complementary to tracing but insufficient as sole observability strategy. + When valid: Appropriate for high-level health monitoring and alerting, should be used alongside tracing +- Vendor-specific tracing solutions (e.g., AWS X-Ray, Datadog APM) (deferred) + Rejected because: Not rejected but deferred to implementation phase. Vendor choice should align with existing infrastructure while maintaining standard trace context propagation. + When valid: Valid if vendor solution supports OpenTelemetry or W3C Trace Context standards for interoperability + +## Risks + +- Trace data volume may exceed storage capacity or budget constraints, leading to incomplete observability + Mitigation: Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests. Establish retention policies (e.g., 7 days for all traces, 30 days for error traces). + Owner: Platform Engineering Team +- Inconsistent tracing implementation across teams may result in fragmented observability + Mitigation: Provide shared tracing libraries and middleware with sensible defaults. Conduct code reviews specifically checking for tracing compliance. Create runbooks and training materials. + Owner: Engineering Team +- Performance overhead from tracing may impact latency-sensitive APIs + Mitigation: Benchmark tracing overhead during implementation. Use asynchronous trace export to minimize request latency impact. Allow exceptions for proven performance-critical paths. + Owner: API Development Team + +## Implementation Notes + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + +## Continuation Context + + +Verify commands: +- grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l +- grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' +- cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' + +Accept when: +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +## Enforcement + +- Verified by: Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis +- Verified by: Code review checklist includes verification of distributed tracing implementation +- Verified by: Integration test suite validates trace context propagation and span attribute completeness +- Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation +- Violation handling: Code review blocks merge if tracing requirements are not met without documented exception +- Violation handling: Quarterly audits identify non-compliant APIs with remediation plans required within 30 days +- Exception process: Submit exception request to architecture review board with performance benchmarks or deprecation timeline +- Exception process: Document exception in API specification and architectural decision log +- Exception process: Re-evaluate exceptions quarterly to determine if circumstances have changed \ No newline at end of file diff --git a/docs/adr/a4786ef4-01e0-4b65-9c9f-0c3175726c86-standardize-service-boundary-logging-for-observability-service-boundary-exit.md b/docs/adr/a4786ef4-01e0-4b65-9c9f-0c3175726c86-standardize-service-boundary-logging-for-observability-service-boundary-exit.md new file mode 100644 index 000000000..50cdc607f --- /dev/null +++ b/docs/adr/a4786ef4-01e0-4b65-9c9f-0c3175726c86-standardize-service-boundary-logging-for-observability-service-boundary-exit.md @@ -0,0 +1,131 @@ +# Standardize Service Boundary Logging for Observability: Service Boundary Exit + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all service implementations and applies to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + +## Context + +- The codebase exhibits a consistent pattern of logging at service boundaries across multiple components (fabro-server and fabro-web), indicating an architectural decision to instrument entry and exit points of service operations +- Service boundary logging is critical for distributed system observability, enabling request tracing, performance monitoring, and debugging across service interactions +- The pattern appears in both backend Rust server implementations and frontend TypeScript query layers, suggesting a cross-stack architectural concern +- With 3 files showing this pattern at 90.17% confidence, this represents a deliberate architectural choice rather than ad-hoc logging practices +- The facet 'boundaries.service_definitions' indicates this pattern specifically targets the interfaces where services are defined and exposed + +## Problem Statement + +Without standardized logging at service boundaries, distributed systems suffer from poor observability, making it difficult to trace requests across services, diagnose performance bottlenecks, identify failure points, and understand system behavior in production. Ad-hoc logging approaches lead to inconsistent log formats, missing context, and gaps in observability coverage. + +## Decision + +1. MUST: All service boundary exit points MUST log operation completion status (success/failure), duration, and relevant result metadata + +## Policy Block + +- MUST All service boundary exit points MUST log operation completion status (success/failure), duration, and relevant result metadata + +In scope: +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +Out of scope: +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +Exceptions: +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +## Rationale + +- The detection of this pattern across 3 files with 90.17% confidence indicates this is an established architectural practice in the codebase, representing proven value in production systems +- Service boundary logging provides the minimal necessary observability coverage for distributed systems without requiring invasive instrumentation throughout the codebase +- Standardizing logging at boundaries creates consistent observability across heterogeneous technology stacks (Rust backend, TypeScript frontend), enabling unified monitoring and debugging workflows +- This approach balances observability needs with performance concerns by focusing logging on key interaction points rather than instrumenting every function + +## Consequences + +Positive: +- Improved ability to trace requests through distributed system components, reducing mean time to resolution (MTTR) for production issues +- Consistent log structure across services enables automated log aggregation, analysis, and alerting +- Clear observability boundaries make it easier for developers to understand where logging is required versus optional +- Correlation identifiers at service boundaries enable powerful distributed tracing capabilities without complex instrumentation + +Negative: +- Additional logging overhead at service boundaries may impact latency for high-frequency operations, requiring careful performance monitoring +- Developers must remember to implement boundary logging for all new service endpoints, creating potential for human error +- Log volume may increase significantly in high-traffic systems, requiring investment in log management infrastructure +- Maintaining consistent logging patterns across multiple technology stacks (Rust, TypeScript) requires ongoing coordination and code review discipline + +## Alternatives + +- Comprehensive instrumentation logging throughout the entire codebase, not just at service boundaries (rejected) + Rejected because: Creates excessive log volume, performance overhead, and maintenance burden. Most internal function calls do not provide actionable observability value compared to boundary logging. + When valid: May be appropriate for specific critical subsystems during active debugging or performance optimization efforts, but not as a general architectural pattern +- Rely solely on distributed tracing frameworks (OpenTelemetry, Jaeger) without explicit boundary logging (rejected) + Rejected because: Tracing frameworks add complexity and dependencies, may not be available in all environments, and logs provide complementary information (detailed context, error messages) that traces don't capture well + When valid: Could be adopted in addition to boundary logging as systems mature, but should not replace foundational logging practices +- Implement logging through middleware/interceptors only, without explicit logging in service code (deferred) + Rejected because: While middleware can handle many boundary logging concerns, some service-specific context and business logic details are best logged within the service implementation itself + When valid: Middleware-based logging should be used where possible for consistency, with explicit service logging for context that middleware cannot capture + +## Risks + +- Inconsistent implementation across teams and services leads to gaps in observability coverage + Mitigation: Implement automated linting rules to detect missing boundary logging, provide code templates and examples, include boundary logging checks in code review checklists + Owner: Engineering team leads and platform team +- Sensitive data leakage through logs if developers are not careful about what context they include + Mitigation: Implement automated PII detection in logs, provide sanitization utilities, conduct security training on logging best practices, enable log redaction in production + Owner: Security team and engineering team +- Performance degradation in high-throughput services due to synchronous logging overhead + Mitigation: Use async logging frameworks, implement log sampling for high-frequency endpoints, monitor logging performance impact, establish performance budgets for logging overhead + Owner: Platform team and service owners + +## Implementation Notes + +- Create shared logging utilities or macros that encapsulate boundary logging patterns, making it easy for developers to add consistent logging with minimal boilerplate +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' that provide zero-cost abstractions and async logging capabilities +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters for consistency with backend logs +- Establish a standard set of log fields for service boundaries: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems (ELK, Splunk, CloudWatch) to parse and index boundary logs for efficient querying and alerting +- Document logging patterns in service templates and starter kits to ensure new services follow established practices from the beginning + +## Continuation Context + + +Verify commands: +- grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l +- rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' +- find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' + +Accept when: +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +## Enforcement + +- Verified by: Automated code review checks using static analysis tools (clippy for Rust, ESLint for TypeScript) with custom rules for boundary logging +- Verified by: Manual code review checklist items requiring reviewers to verify boundary logging presence +- Verified by: CI pipeline integration tests that verify log output from service endpoints contains required fields +- Verified by: Periodic observability audits reviewing log coverage across services +- Violation handling: CI pipeline warnings for missing boundary logging patterns (non-blocking initially, blocking after grace period) +- Violation handling: Code review feedback requiring addition of boundary logging before merge approval +- Violation handling: Quarterly observability reports identifying services with insufficient logging coverage +- Violation handling: Escalation to architecture review for repeated violations or services with poor observability +- Exception process: Submit exception request to team lead or architect with documented rationale (performance impact, alternative observability approach) +- Exception process: For performance-critical paths, provide benchmark data demonstrating logging overhead exceeds acceptable latency budget +- Exception process: Document approved exceptions in service README and architecture decision log +- Exception process: Review exceptions quarterly to determine if mitigations (async logging, sampling) can eliminate the need for exception \ No newline at end of file diff --git a/docs/adr/b1a53926-9804-45b7-9543-634d35c15734-standardize-service-boundary-logging-for-observability-service-boundary-entry.md b/docs/adr/b1a53926-9804-45b7-9543-634d35c15734-standardize-service-boundary-logging-for-observability-service-boundary-entry.md new file mode 100644 index 000000000..b4feb82e9 --- /dev/null +++ b/docs/adr/b1a53926-9804-45b7-9543-634d35c15734-standardize-service-boundary-logging-for-observability-service-boundary-entry.md @@ -0,0 +1,131 @@ +# Standardize Service Boundary Logging for Observability: Service Boundary Entry + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all service implementations and applies to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + +## Context + +- The codebase exhibits a consistent pattern of logging at service boundaries across multiple components (fabro-server and fabro-web), indicating an architectural decision to instrument entry and exit points of service operations +- Service boundary logging is critical for distributed system observability, enabling request tracing, performance monitoring, and debugging across service interactions +- The pattern appears in both backend Rust server implementations and frontend TypeScript query layers, suggesting a cross-stack architectural concern +- With 3 files showing this pattern at 90.17% confidence, this represents a deliberate architectural choice rather than ad-hoc logging practices +- The facet 'boundaries.service_definitions' indicates this pattern specifically targets the interfaces where services are defined and exposed + +## Problem Statement + +Without standardized logging at service boundaries, distributed systems suffer from poor observability, making it difficult to trace requests across services, diagnose performance bottlenecks, identify failure points, and understand system behavior in production. Ad-hoc logging approaches lead to inconsistent log formats, missing context, and gaps in observability coverage. + +## Decision + +1. MUST: All service boundary entry points (API handlers, RPC endpoints, query interfaces) MUST log incoming requests with sufficient context to identify the operation, caller, and key parameters + +## Policy Block + +- MUST All service boundary entry points (API handlers, RPC endpoints, query interfaces) MUST log incoming requests with sufficient context to identify the operation, caller, and key parameters + +In scope: +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +Out of scope: +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +Exceptions: +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +## Rationale + +- The detection of this pattern across 3 files with 90.17% confidence indicates this is an established architectural practice in the codebase, representing proven value in production systems +- Service boundary logging provides the minimal necessary observability coverage for distributed systems without requiring invasive instrumentation throughout the codebase +- Standardizing logging at boundaries creates consistent observability across heterogeneous technology stacks (Rust backend, TypeScript frontend), enabling unified monitoring and debugging workflows +- This approach balances observability needs with performance concerns by focusing logging on key interaction points rather than instrumenting every function + +## Consequences + +Positive: +- Improved ability to trace requests through distributed system components, reducing mean time to resolution (MTTR) for production issues +- Consistent log structure across services enables automated log aggregation, analysis, and alerting +- Clear observability boundaries make it easier for developers to understand where logging is required versus optional +- Correlation identifiers at service boundaries enable powerful distributed tracing capabilities without complex instrumentation + +Negative: +- Additional logging overhead at service boundaries may impact latency for high-frequency operations, requiring careful performance monitoring +- Developers must remember to implement boundary logging for all new service endpoints, creating potential for human error +- Log volume may increase significantly in high-traffic systems, requiring investment in log management infrastructure +- Maintaining consistent logging patterns across multiple technology stacks (Rust, TypeScript) requires ongoing coordination and code review discipline + +## Alternatives + +- Comprehensive instrumentation logging throughout the entire codebase, not just at service boundaries (rejected) + Rejected because: Creates excessive log volume, performance overhead, and maintenance burden. Most internal function calls do not provide actionable observability value compared to boundary logging. + When valid: May be appropriate for specific critical subsystems during active debugging or performance optimization efforts, but not as a general architectural pattern +- Rely solely on distributed tracing frameworks (OpenTelemetry, Jaeger) without explicit boundary logging (rejected) + Rejected because: Tracing frameworks add complexity and dependencies, may not be available in all environments, and logs provide complementary information (detailed context, error messages) that traces don't capture well + When valid: Could be adopted in addition to boundary logging as systems mature, but should not replace foundational logging practices +- Implement logging through middleware/interceptors only, without explicit logging in service code (deferred) + Rejected because: While middleware can handle many boundary logging concerns, some service-specific context and business logic details are best logged within the service implementation itself + When valid: Middleware-based logging should be used where possible for consistency, with explicit service logging for context that middleware cannot capture + +## Risks + +- Inconsistent implementation across teams and services leads to gaps in observability coverage + Mitigation: Implement automated linting rules to detect missing boundary logging, provide code templates and examples, include boundary logging checks in code review checklists + Owner: Engineering team leads and platform team +- Sensitive data leakage through logs if developers are not careful about what context they include + Mitigation: Implement automated PII detection in logs, provide sanitization utilities, conduct security training on logging best practices, enable log redaction in production + Owner: Security team and engineering team +- Performance degradation in high-throughput services due to synchronous logging overhead + Mitigation: Use async logging frameworks, implement log sampling for high-frequency endpoints, monitor logging performance impact, establish performance budgets for logging overhead + Owner: Platform team and service owners + +## Implementation Notes + +- Create shared logging utilities or macros that encapsulate boundary logging patterns, making it easy for developers to add consistent logging with minimal boilerplate +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' that provide zero-cost abstractions and async logging capabilities +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters for consistency with backend logs +- Establish a standard set of log fields for service boundaries: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems (ELK, Splunk, CloudWatch) to parse and index boundary logs for efficient querying and alerting +- Document logging patterns in service templates and starter kits to ensure new services follow established practices from the beginning + +## Continuation Context + + +Verify commands: +- grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l +- rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' +- find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' + +Accept when: +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +## Enforcement + +- Verified by: Automated code review checks using static analysis tools (clippy for Rust, ESLint for TypeScript) with custom rules for boundary logging +- Verified by: Manual code review checklist items requiring reviewers to verify boundary logging presence +- Verified by: CI pipeline integration tests that verify log output from service endpoints contains required fields +- Verified by: Periodic observability audits reviewing log coverage across services +- Violation handling: CI pipeline warnings for missing boundary logging patterns (non-blocking initially, blocking after grace period) +- Violation handling: Code review feedback requiring addition of boundary logging before merge approval +- Violation handling: Quarterly observability reports identifying services with insufficient logging coverage +- Violation handling: Escalation to architecture review for repeated violations or services with poor observability +- Exception process: Submit exception request to team lead or architect with documented rationale (performance impact, alternative observability approach) +- Exception process: For performance-critical paths, provide benchmark data demonstrating logging overhead exceeds acceptable latency budget +- Exception process: Document approved exceptions in service README and architecture decision log +- Exception process: Review exceptions quarterly to determine if mitigations (async logging, sampling) can eliminate the need for exception \ No newline at end of file diff --git a/docs/adr/c81e5425-fcc1-4177-9246-bd167e4322af-adopt-distributed-tracing-for-public-api-observability-implement-sampling-strategies.md b/docs/adr/c81e5425-fcc1-4177-9246-bd167e4322af-adopt-distributed-tracing-for-public-api-observability-implement-sampling-strategies.md new file mode 100644 index 000000000..4e24d1baa --- /dev/null +++ b/docs/adr/c81e5425-fcc1-4177-9246-bd167e4322af-adopt-distributed-tracing-for-public-api-observability-implement-sampling-strategies.md @@ -0,0 +1,123 @@ +# Adopt Distributed Tracing for Public API Observability: Implement Sampling Strategies + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + +## Context + +- Public and external APIs require comprehensive observability to diagnose issues across distributed system boundaries where direct debugging is not feasible +- The pattern was detected with 90.27% confidence across 3 critical infrastructure files (terminal.rs, lib.rs, stall.rs) in the fabro crate ecosystem, indicating a systematic approach to tracing +- The facet 'obs.tracing' suggests this pattern specifically addresses observability through distributed tracing instrumentation +- External API consumers and integration partners require visibility into request flows, latency bottlenecks, and error propagation paths +- Without standardized tracing, debugging cross-service issues becomes prohibitively expensive and time-consuming + +## Problem Statement + +Public and external APIs lack consistent observability instrumentation, making it difficult to diagnose performance issues, trace request flows across service boundaries, and provide actionable debugging information to external consumers. This creates operational blind spots and increases mean time to resolution (MTTR) for production incidents. + +## Decision + +1. SHOULD: APIs SHOULD implement sampling strategies to balance observability needs with performance overhead (e.g., 100% for errors, 10% for successful requests) + +## Policy Block + +- SHOULD APIs SHOULD implement sampling strategies to balance observability needs with performance overhead (e.g., 100% for errors, 10% for successful requests) + +In scope: +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +Out of scope: +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +Exceptions: +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +## Rationale + +- The pattern detection across 3 files with 90.27% confidence indicates this is an established architectural practice in the codebase, not an isolated implementation +- Distributed tracing provides end-to-end visibility that is essential for debugging issues in microservices architectures where requests span multiple services +- Standardizing on obs.tracing facet ensures consistent instrumentation patterns across the organization, reducing cognitive load and improving debugging efficiency +- External API consumers benefit from trace IDs in error responses, enabling them to provide actionable information when reporting issues + +## Consequences + +Positive: +- Reduced mean time to resolution (MTTR) for production incidents through comprehensive request flow visibility +- Improved developer experience with standardized tracing patterns across all public APIs +- Enhanced external consumer support with traceable request identifiers for issue reporting +- Better capacity planning and performance optimization through latency distribution analysis + +Negative: +- Increased operational complexity with additional infrastructure for trace collection and storage +- Minor performance overhead (typically 1-3% latency increase) from tracing instrumentation +- Development time investment required to instrument existing APIs and train teams on tracing best practices +- Potential for trace data explosion requiring careful sampling strategy and retention policies + +## Alternatives + +- Structured logging only without distributed tracing (rejected) + Rejected because: Logs lack the causal relationships and timing information needed to reconstruct request flows across service boundaries. Correlation requires manual log aggregation which is error-prone and time-consuming. + When valid: May be sufficient for monolithic applications with single-service request handling +- Metrics-based observability with counters and histograms (rejected) + Rejected because: Metrics provide aggregate statistics but cannot trace individual request paths or identify specific failure scenarios. Complementary to tracing but insufficient as sole observability strategy. + When valid: Appropriate for high-level health monitoring and alerting, should be used alongside tracing +- Vendor-specific tracing solutions (e.g., AWS X-Ray, Datadog APM) (deferred) + Rejected because: Not rejected but deferred to implementation phase. Vendor choice should align with existing infrastructure while maintaining standard trace context propagation. + When valid: Valid if vendor solution supports OpenTelemetry or W3C Trace Context standards for interoperability + +## Risks + +- Trace data volume may exceed storage capacity or budget constraints, leading to incomplete observability + Mitigation: Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests. Establish retention policies (e.g., 7 days for all traces, 30 days for error traces). + Owner: Platform Engineering Team +- Inconsistent tracing implementation across teams may result in fragmented observability + Mitigation: Provide shared tracing libraries and middleware with sensible defaults. Conduct code reviews specifically checking for tracing compliance. Create runbooks and training materials. + Owner: Engineering Team +- Performance overhead from tracing may impact latency-sensitive APIs + Mitigation: Benchmark tracing overhead during implementation. Use asynchronous trace export to minimize request latency impact. Allow exceptions for proven performance-critical paths. + Owner: API Development Team + +## Implementation Notes + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + +## Continuation Context + + +Verify commands: +- grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l +- grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' +- cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' + +Accept when: +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +## Enforcement + +- Verified by: Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis +- Verified by: Code review checklist includes verification of distributed tracing implementation +- Verified by: Integration test suite validates trace context propagation and span attribute completeness +- Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation +- Violation handling: Code review blocks merge if tracing requirements are not met without documented exception +- Violation handling: Quarterly audits identify non-compliant APIs with remediation plans required within 30 days +- Exception process: Submit exception request to architecture review board with performance benchmarks or deprecation timeline +- Exception process: Document exception in API specification and architectural decision log +- Exception process: Re-evaluate exceptions quarterly to determine if circumstances have changed \ No newline at end of file diff --git a/docs/adr/da815559-806f-4816-ac62-f2bcba187047-standardize-service-boundary-logging-for-observability-service-boundary-logs.md b/docs/adr/da815559-806f-4816-ac62-f2bcba187047-standardize-service-boundary-logging-for-observability-service-boundary-logs.md new file mode 100644 index 000000000..800e08fd3 --- /dev/null +++ b/docs/adr/da815559-806f-4816-ac62-f2bcba187047-standardize-service-boundary-logging-for-observability-service-boundary-logs.md @@ -0,0 +1,131 @@ +# Standardize Service Boundary Logging for Observability: Service Boundary Logs + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all service implementations and applies to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + +## Context + +- The codebase exhibits a consistent pattern of logging at service boundaries across multiple components (fabro-server and fabro-web), indicating an architectural decision to instrument entry and exit points of service operations +- Service boundary logging is critical for distributed system observability, enabling request tracing, performance monitoring, and debugging across service interactions +- The pattern appears in both backend Rust server implementations and frontend TypeScript query layers, suggesting a cross-stack architectural concern +- With 3 files showing this pattern at 90.17% confidence, this represents a deliberate architectural choice rather than ad-hoc logging practices +- The facet 'boundaries.service_definitions' indicates this pattern specifically targets the interfaces where services are defined and exposed + +## Problem Statement + +Without standardized logging at service boundaries, distributed systems suffer from poor observability, making it difficult to trace requests across services, diagnose performance bottlenecks, identify failure points, and understand system behavior in production. Ad-hoc logging approaches lead to inconsistent log formats, missing context, and gaps in observability coverage. + +## Decision + +1. MUST_NOT: Service boundary logs MUST NOT include sensitive data (passwords, tokens, PII) unless explicitly redacted or masked + +## Policy Block + +- MUST_NOT Service boundary logs MUST NOT include sensitive data (passwords, tokens, PII) unless explicitly redacted or masked + +In scope: +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +Out of scope: +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +Exceptions: +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +## Rationale + +- The detection of this pattern across 3 files with 90.17% confidence indicates this is an established architectural practice in the codebase, representing proven value in production systems +- Service boundary logging provides the minimal necessary observability coverage for distributed systems without requiring invasive instrumentation throughout the codebase +- Standardizing logging at boundaries creates consistent observability across heterogeneous technology stacks (Rust backend, TypeScript frontend), enabling unified monitoring and debugging workflows +- This approach balances observability needs with performance concerns by focusing logging on key interaction points rather than instrumenting every function + +## Consequences + +Positive: +- Improved ability to trace requests through distributed system components, reducing mean time to resolution (MTTR) for production issues +- Consistent log structure across services enables automated log aggregation, analysis, and alerting +- Clear observability boundaries make it easier for developers to understand where logging is required versus optional +- Correlation identifiers at service boundaries enable powerful distributed tracing capabilities without complex instrumentation + +Negative: +- Additional logging overhead at service boundaries may impact latency for high-frequency operations, requiring careful performance monitoring +- Developers must remember to implement boundary logging for all new service endpoints, creating potential for human error +- Log volume may increase significantly in high-traffic systems, requiring investment in log management infrastructure +- Maintaining consistent logging patterns across multiple technology stacks (Rust, TypeScript) requires ongoing coordination and code review discipline + +## Alternatives + +- Comprehensive instrumentation logging throughout the entire codebase, not just at service boundaries (rejected) + Rejected because: Creates excessive log volume, performance overhead, and maintenance burden. Most internal function calls do not provide actionable observability value compared to boundary logging. + When valid: May be appropriate for specific critical subsystems during active debugging or performance optimization efforts, but not as a general architectural pattern +- Rely solely on distributed tracing frameworks (OpenTelemetry, Jaeger) without explicit boundary logging (rejected) + Rejected because: Tracing frameworks add complexity and dependencies, may not be available in all environments, and logs provide complementary information (detailed context, error messages) that traces don't capture well + When valid: Could be adopted in addition to boundary logging as systems mature, but should not replace foundational logging practices +- Implement logging through middleware/interceptors only, without explicit logging in service code (deferred) + Rejected because: While middleware can handle many boundary logging concerns, some service-specific context and business logic details are best logged within the service implementation itself + When valid: Middleware-based logging should be used where possible for consistency, with explicit service logging for context that middleware cannot capture + +## Risks + +- Inconsistent implementation across teams and services leads to gaps in observability coverage + Mitigation: Implement automated linting rules to detect missing boundary logging, provide code templates and examples, include boundary logging checks in code review checklists + Owner: Engineering team leads and platform team +- Sensitive data leakage through logs if developers are not careful about what context they include + Mitigation: Implement automated PII detection in logs, provide sanitization utilities, conduct security training on logging best practices, enable log redaction in production + Owner: Security team and engineering team +- Performance degradation in high-throughput services due to synchronous logging overhead + Mitigation: Use async logging frameworks, implement log sampling for high-frequency endpoints, monitor logging performance impact, establish performance budgets for logging overhead + Owner: Platform team and service owners + +## Implementation Notes + +- Create shared logging utilities or macros that encapsulate boundary logging patterns, making it easy for developers to add consistent logging with minimal boilerplate +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' that provide zero-cost abstractions and async logging capabilities +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters for consistency with backend logs +- Establish a standard set of log fields for service boundaries: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems (ELK, Splunk, CloudWatch) to parse and index boundary logs for efficient querying and alerting +- Document logging patterns in service templates and starter kits to ensure new services follow established practices from the beginning + +## Continuation Context + + +Verify commands: +- grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l +- rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' +- find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' + +Accept when: +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +## Enforcement + +- Verified by: Automated code review checks using static analysis tools (clippy for Rust, ESLint for TypeScript) with custom rules for boundary logging +- Verified by: Manual code review checklist items requiring reviewers to verify boundary logging presence +- Verified by: CI pipeline integration tests that verify log output from service endpoints contains required fields +- Verified by: Periodic observability audits reviewing log coverage across services +- Violation handling: CI pipeline warnings for missing boundary logging patterns (non-blocking initially, blocking after grace period) +- Violation handling: Code review feedback requiring addition of boundary logging before merge approval +- Violation handling: Quarterly observability reports identifying services with insufficient logging coverage +- Violation handling: Escalation to architecture review for repeated violations or services with poor observability +- Exception process: Submit exception request to team lead or architect with documented rationale (performance impact, alternative observability approach) +- Exception process: For performance-critical paths, provide benchmark data demonstrating logging overhead exceeds acceptable latency budget +- Exception process: Document approved exceptions in service README and architecture decision log +- Exception process: Review exceptions quarterly to determine if mitigations (async logging, sampling) can eliminate the need for exception \ No newline at end of file diff --git a/docs/adr/e5731ea0-1ffc-4e87-b5f7-07551952f609-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-that.md b/docs/adr/e5731ea0-1ffc-4e87-b5f7-07551952f609-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-that.md new file mode 100644 index 000000000..2c91fd3f6 --- /dev/null +++ b/docs/adr/e5731ea0-1ffc-4e87-b5f7-07551952f609-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-route-handlers-that.md @@ -0,0 +1,115 @@ +# Standardize Server-Sent Events (SSE) for Real-Time Configuration and Runtime State Updates: Route Handlers That + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Context + +- The application requires real-time updates of runtime execution state and configuration changes to be pushed to client interfaces without polling +- Multiple route handlers (run-stages.tsx, run-files.tsx) need to stream execution progress and file processing status to web clients +- A centralized SSE implementation pattern (sse.ts) has emerged to handle persistent connections and event streaming across different runtime contexts +- The pattern provides a consistent mechanism for broadcasting configuration changes and environment state updates during execution workflows + +## Problem Statement + +Applications need a reliable, standardized mechanism to push real-time runtime state, configuration updates, and execution progress to clients without the overhead and latency of polling-based approaches, while maintaining consistent event handling across multiple execution contexts. + +## Decision + +1. MUST: Route handlers that expose runtime execution state MUST establish SSE connections with appropriate Content-Type headers (text/event-stream) + +## Policy Block + +- MUST Route handlers that expose runtime execution state MUST establish SSE connections with appropriate Content-Type headers (text/event-stream) + +In scope: +- Web route handlers that expose runtime execution progress (run-stages, run-files, etc.) +- Configuration management endpoints that broadcast environment or setting changes +- Real-time monitoring interfaces for execution workflows +- Client-facing APIs that require push-based state synchronization + +Out of scope: +- Bidirectional communication patterns requiring client-to-server messaging (use WebSockets instead) +- Binary data streaming or large file transfers +- Internal service-to-service communication (use message queues or gRPC) +- Static configuration loading at application startup + +## Rationale + +- Pattern detected across 3 files with 91.07% confidence indicates a deliberate architectural choice for real-time state communication +- SSE provides a lightweight, HTTP-based protocol that works through firewalls and proxies without requiring WebSocket infrastructure +- Centralized SSE library (sse.ts) ensures consistent event formatting, error handling, and connection lifecycle management across multiple route handlers +- The pattern aligns with unidirectional data flow from server to client, which matches the use case of broadcasting runtime state and configuration updates + +## Consequences + +Positive: +- Clients receive real-time updates without polling overhead, reducing server load and improving responsiveness +- Consistent SSE implementation across routes reduces code duplication and simplifies maintenance +- HTTP-based protocol ensures compatibility with existing infrastructure (load balancers, proxies, CDNs) +- Automatic reconnection support in browsers provides resilience against transient network failures + +Negative: +- SSE connections are unidirectional, requiring separate HTTP requests for client-to-server communication +- Long-lived connections may require infrastructure tuning (timeouts, connection limits) to support many concurrent clients +- Browser connection limits (typically 6 per domain) may constrain applications with multiple simultaneous SSE streams +- Debugging and monitoring SSE connections can be more complex than traditional request-response patterns + +## Alternatives + +- Use polling-based approach with periodic HTTP requests to fetch runtime state updates (rejected) + Rejected because: Polling introduces latency, increases server load with redundant requests, and provides poor user experience for real-time execution monitoring + When valid: Acceptable for low-frequency updates (>30 seconds) or when SSE infrastructure is unavailable +- Implement WebSocket-based bidirectional communication for all real-time features (rejected) + Rejected because: WebSockets add complexity and infrastructure requirements when only unidirectional server-to-client communication is needed for configuration and state updates + When valid: Appropriate when bidirectional real-time communication is required (e.g., collaborative editing, chat) +- Use GraphQL subscriptions for real-time data updates (rejected) + Rejected because: Adds GraphQL infrastructure overhead and complexity when simple event streaming suffices for runtime state broadcasting + When valid: Valid when already using GraphQL extensively and need subscription-based data synchronization + +## Risks + +- SSE connections may be terminated by intermediate proxies or load balancers with aggressive timeout policies + Mitigation: Implement periodic heartbeat messages and configure infrastructure timeouts appropriately (e.g., 60+ seconds). Document required infrastructure settings. + Owner: Engineering team and DevOps +- Memory leaks or resource exhaustion if SSE connections are not properly cleaned up on client disconnect + Mitigation: Implement robust connection lifecycle management with cleanup handlers. Add monitoring for open connection counts and memory usage. + Owner: Engineering team +- Browser connection limits may prevent multiple SSE streams from functioning simultaneously + Mitigation: Multiplex multiple event types over a single SSE connection where possible. Document connection usage and provide guidance on connection management. + Owner: Engineering team + +## Implementation Notes + +- Create a centralized SSE utility module (e.g., lib/sse.ts) that provides connection setup, event formatting, and cleanup helpers +- Ensure route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- Document SSE usage patterns and provide examples for common scenarios (execution progress, configuration updates) +- Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +## Continuation Context + + +Verify commands: +- grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l +- grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules +- find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" + +Accept when: +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion + +## Enforcement + +- Verified by: Code review checklist verifying SSE usage for new real-time endpoints +- Verified by: Automated linting rules to detect missing Content-Type headers on streaming routes +- Verified by: Integration tests validating SSE connection lifecycle and event delivery +- Violation handling: Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable +- Violation handling: Route handlers with SSE connections missing proper cleanup must be fixed before merge +- Violation handling: Violations detected in code review are flagged and require revision +- Exception process: Document technical justification for alternative approach (e.g., WebSocket requirement, infrastructure constraints) +- Exception process: Obtain approval from technical lead or architect +- Exception process: Add architectural decision comment in code explaining exception rationale \ No newline at end of file diff --git a/docs/adr/f1ba5186-840f-4423-9876-d7cb8ec75614-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-runtime-state-updates.md b/docs/adr/f1ba5186-840f-4423-9876-d7cb8ec75614-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-runtime-state-updates.md new file mode 100644 index 000000000..11b8077fc --- /dev/null +++ b/docs/adr/f1ba5186-840f-4423-9876-d7cb8ec75614-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-runtime-state-updates.md @@ -0,0 +1,115 @@ +# Standardize Server-Sent Events (SSE) for Real-Time Configuration and Runtime State Updates: Runtime State Updates + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Context + +- The application requires real-time updates of runtime execution state and configuration changes to be pushed to client interfaces without polling +- Multiple route handlers (run-stages.tsx, run-files.tsx) need to stream execution progress and file processing status to web clients +- A centralized SSE implementation pattern (sse.ts) has emerged to handle persistent connections and event streaming across different runtime contexts +- The pattern provides a consistent mechanism for broadcasting configuration changes and environment state updates during execution workflows + +## Problem Statement + +Applications need a reliable, standardized mechanism to push real-time runtime state, configuration updates, and execution progress to clients without the overhead and latency of polling-based approaches, while maintaining consistent event handling across multiple execution contexts. + +## Decision + +1. MUST: Runtime state updates and configuration changes MUST be delivered to clients using Server-Sent Events (SSE) protocol + +## Policy Block + +- MUST Runtime state updates and configuration changes MUST be delivered to clients using Server-Sent Events (SSE) protocol + +In scope: +- Web route handlers that expose runtime execution progress (run-stages, run-files, etc.) +- Configuration management endpoints that broadcast environment or setting changes +- Real-time monitoring interfaces for execution workflows +- Client-facing APIs that require push-based state synchronization + +Out of scope: +- Bidirectional communication patterns requiring client-to-server messaging (use WebSockets instead) +- Binary data streaming or large file transfers +- Internal service-to-service communication (use message queues or gRPC) +- Static configuration loading at application startup + +## Rationale + +- Pattern detected across 3 files with 91.07% confidence indicates a deliberate architectural choice for real-time state communication +- SSE provides a lightweight, HTTP-based protocol that works through firewalls and proxies without requiring WebSocket infrastructure +- Centralized SSE library (sse.ts) ensures consistent event formatting, error handling, and connection lifecycle management across multiple route handlers +- The pattern aligns with unidirectional data flow from server to client, which matches the use case of broadcasting runtime state and configuration updates + +## Consequences + +Positive: +- Clients receive real-time updates without polling overhead, reducing server load and improving responsiveness +- Consistent SSE implementation across routes reduces code duplication and simplifies maintenance +- HTTP-based protocol ensures compatibility with existing infrastructure (load balancers, proxies, CDNs) +- Automatic reconnection support in browsers provides resilience against transient network failures + +Negative: +- SSE connections are unidirectional, requiring separate HTTP requests for client-to-server communication +- Long-lived connections may require infrastructure tuning (timeouts, connection limits) to support many concurrent clients +- Browser connection limits (typically 6 per domain) may constrain applications with multiple simultaneous SSE streams +- Debugging and monitoring SSE connections can be more complex than traditional request-response patterns + +## Alternatives + +- Use polling-based approach with periodic HTTP requests to fetch runtime state updates (rejected) + Rejected because: Polling introduces latency, increases server load with redundant requests, and provides poor user experience for real-time execution monitoring + When valid: Acceptable for low-frequency updates (>30 seconds) or when SSE infrastructure is unavailable +- Implement WebSocket-based bidirectional communication for all real-time features (rejected) + Rejected because: WebSockets add complexity and infrastructure requirements when only unidirectional server-to-client communication is needed for configuration and state updates + When valid: Appropriate when bidirectional real-time communication is required (e.g., collaborative editing, chat) +- Use GraphQL subscriptions for real-time data updates (rejected) + Rejected because: Adds GraphQL infrastructure overhead and complexity when simple event streaming suffices for runtime state broadcasting + When valid: Valid when already using GraphQL extensively and need subscription-based data synchronization + +## Risks + +- SSE connections may be terminated by intermediate proxies or load balancers with aggressive timeout policies + Mitigation: Implement periodic heartbeat messages and configure infrastructure timeouts appropriately (e.g., 60+ seconds). Document required infrastructure settings. + Owner: Engineering team and DevOps +- Memory leaks or resource exhaustion if SSE connections are not properly cleaned up on client disconnect + Mitigation: Implement robust connection lifecycle management with cleanup handlers. Add monitoring for open connection counts and memory usage. + Owner: Engineering team +- Browser connection limits may prevent multiple SSE streams from functioning simultaneously + Mitigation: Multiplex multiple event types over a single SSE connection where possible. Document connection usage and provide guidance on connection management. + Owner: Engineering team + +## Implementation Notes + +- Create a centralized SSE utility module (e.g., lib/sse.ts) that provides connection setup, event formatting, and cleanup helpers +- Ensure route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- Document SSE usage patterns and provide examples for common scenarios (execution progress, configuration updates) +- Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +## Continuation Context + + +Verify commands: +- grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l +- grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules +- find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" + +Accept when: +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion + +## Enforcement + +- Verified by: Code review checklist verifying SSE usage for new real-time endpoints +- Verified by: Automated linting rules to detect missing Content-Type headers on streaming routes +- Verified by: Integration tests validating SSE connection lifecycle and event delivery +- Violation handling: Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable +- Violation handling: Route handlers with SSE connections missing proper cleanup must be fixed before merge +- Violation handling: Violations detected in code review are flagged and require revision +- Exception process: Document technical justification for alternative approach (e.g., WebSocket requirement, infrastructure constraints) +- Exception process: Obtain approval from technical lead or architect +- Exception process: Add architectural decision comment in code explaining exception rationale \ No newline at end of file diff --git a/docs/adr/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7-adopt-distributed-tracing-for-public-api-observability-trace-spans-include.md b/docs/adr/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7-adopt-distributed-tracing-for-public-api-observability-trace-spans-include.md new file mode 100644 index 000000000..2039046e2 --- /dev/null +++ b/docs/adr/f39ea92e-c1c0-46b6-b4ff-9c07cc3b42b7-adopt-distributed-tracing-for-public-api-observability-trace-spans-include.md @@ -0,0 +1,123 @@ +# Adopt Distributed Tracing for Public API Observability: Trace Spans Include + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all public/external API implementations and integration patterns. All new public API endpoints and external integrations MUST implement distributed tracing as specified herein. + +## Context + +- Public and external APIs require comprehensive observability to diagnose issues across distributed system boundaries where direct debugging is not feasible +- The pattern was detected with 90.27% confidence across 3 critical infrastructure files (terminal.rs, lib.rs, stall.rs) in the fabro crate ecosystem, indicating a systematic approach to tracing +- The facet 'obs.tracing' suggests this pattern specifically addresses observability through distributed tracing instrumentation +- External API consumers and integration partners require visibility into request flows, latency bottlenecks, and error propagation paths +- Without standardized tracing, debugging cross-service issues becomes prohibitively expensive and time-consuming + +## Problem Statement + +Public and external APIs lack consistent observability instrumentation, making it difficult to diagnose performance issues, trace request flows across service boundaries, and provide actionable debugging information to external consumers. This creates operational blind spots and increases mean time to resolution (MTTR) for production incidents. + +## Decision + +1. SHOULD: Trace spans SHOULD include business-relevant metadata (user_id, tenant_id, request_id) as span attributes for correlation + +## Policy Block + +- SHOULD Trace spans SHOULD include business-relevant metadata (user_id, tenant_id, request_id) as span attributes for correlation + +In scope: +- All HTTP/REST API endpoints exposed to external consumers +- GraphQL APIs and gRPC services with external clients +- Webhook handlers and callback endpoints +- Public SDK methods that initiate cross-service operations +- Integration adapters and third-party service connectors + +Out of scope: +- Internal library functions that do not cross service boundaries +- Pure computation functions without I/O operations +- Development and test environments (may use mock tracers) +- CLI tools and batch jobs (unless they interact with public APIs) + +Exceptions: +- EXC-001: Performance-critical hot paths where tracing overhead exceeds 5% of request latency +- EXC-002: Legacy APIs scheduled for deprecation within 6 months + +## Rationale + +- The pattern detection across 3 files with 90.27% confidence indicates this is an established architectural practice in the codebase, not an isolated implementation +- Distributed tracing provides end-to-end visibility that is essential for debugging issues in microservices architectures where requests span multiple services +- Standardizing on obs.tracing facet ensures consistent instrumentation patterns across the organization, reducing cognitive load and improving debugging efficiency +- External API consumers benefit from trace IDs in error responses, enabling them to provide actionable information when reporting issues + +## Consequences + +Positive: +- Reduced mean time to resolution (MTTR) for production incidents through comprehensive request flow visibility +- Improved developer experience with standardized tracing patterns across all public APIs +- Enhanced external consumer support with traceable request identifiers for issue reporting +- Better capacity planning and performance optimization through latency distribution analysis + +Negative: +- Increased operational complexity with additional infrastructure for trace collection and storage +- Minor performance overhead (typically 1-3% latency increase) from tracing instrumentation +- Development time investment required to instrument existing APIs and train teams on tracing best practices +- Potential for trace data explosion requiring careful sampling strategy and retention policies + +## Alternatives + +- Structured logging only without distributed tracing (rejected) + Rejected because: Logs lack the causal relationships and timing information needed to reconstruct request flows across service boundaries. Correlation requires manual log aggregation which is error-prone and time-consuming. + When valid: May be sufficient for monolithic applications with single-service request handling +- Metrics-based observability with counters and histograms (rejected) + Rejected because: Metrics provide aggregate statistics but cannot trace individual request paths or identify specific failure scenarios. Complementary to tracing but insufficient as sole observability strategy. + When valid: Appropriate for high-level health monitoring and alerting, should be used alongside tracing +- Vendor-specific tracing solutions (e.g., AWS X-Ray, Datadog APM) (deferred) + Rejected because: Not rejected but deferred to implementation phase. Vendor choice should align with existing infrastructure while maintaining standard trace context propagation. + When valid: Valid if vendor solution supports OpenTelemetry or W3C Trace Context standards for interoperability + +## Risks + +- Trace data volume may exceed storage capacity or budget constraints, leading to incomplete observability + Mitigation: Implement adaptive sampling strategies with 100% sampling for errors and configurable rates for successful requests. Establish retention policies (e.g., 7 days for all traces, 30 days for error traces). + Owner: Platform Engineering Team +- Inconsistent tracing implementation across teams may result in fragmented observability + Mitigation: Provide shared tracing libraries and middleware with sensible defaults. Conduct code reviews specifically checking for tracing compliance. Create runbooks and training materials. + Owner: Engineering Team +- Performance overhead from tracing may impact latency-sensitive APIs + Mitigation: Benchmark tracing overhead during implementation. Use asynchronous trace export to minimize request latency impact. Allow exceptions for proven performance-critical paths. + Owner: API Development Team + +## Implementation Notes + +- Use OpenTelemetry SDK for language-agnostic tracing instrumentation with broad ecosystem support +- Implement tracing middleware at the API gateway/framework level to automatically instrument all endpoints with minimal code changes +- Create shared libraries or decorators that encapsulate tracing logic for common patterns (database access, HTTP clients, message queues) +- Include trace_id in API error responses (e.g., in X-Trace-Id header) to enable external consumers to reference specific requests when reporting issues + +## Continuation Context + + +Verify commands: +- grep -r 'tracing::instrument\|#\[instrument\]\|tracer\.start_span' lib/crates/*/src/ | wc -l +- grep -r 'obs\.tracing\|opentelemetry\|trace_context' lib/crates/fabro-*/src/ --include='*.rs' +- cargo test --package fabro-test -- tracing --nocapture 2>&1 | grep -i 'span\|trace' + +Accept when: +- All public API endpoints in fabro-sandbox, fabro-test, and fabro-core crates demonstrate tracing instrumentation with entry/exit spans +- Verification commands show tracing patterns present in at least 80% of public API implementation files +- Integration tests successfully propagate trace context across service boundaries and validate span hierarchy + +## Enforcement + +- Verified by: Automated CI pipeline checks for tracing instrumentation in new API endpoints using static analysis +- Verified by: Code review checklist includes verification of distributed tracing implementation +- Verified by: Integration test suite validates trace context propagation and span attribute completeness +- Violation handling: CI pipeline fails if new public API endpoints lack tracing instrumentation +- Violation handling: Code review blocks merge if tracing requirements are not met without documented exception +- Violation handling: Quarterly audits identify non-compliant APIs with remediation plans required within 30 days +- Exception process: Submit exception request to architecture review board with performance benchmarks or deprecation timeline +- Exception process: Document exception in API specification and architectural decision log +- Exception process: Re-evaluate exceptions quarterly to determine if circumstances have changed \ No newline at end of file diff --git a/docs/adr/f5680bc5-1992-44e9-92a3-65214350b7a7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-event-streams.md b/docs/adr/f5680bc5-1992-44e9-92a3-65214350b7a7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-event-streams.md new file mode 100644 index 000000000..40f3ff3e7 --- /dev/null +++ b/docs/adr/f5680bc5-1992-44e9-92a3-65214350b7a7-standardize-server-sent-events-sse-for-real-time-configuration-and-runtime-state-updates-sse-event-streams.md @@ -0,0 +1,115 @@ +# Standardize Server-Sent Events (SSE) for Real-Time Configuration and Runtime State Updates: Sse Event Streams + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Context + +- The application requires real-time updates of runtime execution state and configuration changes to be pushed to client interfaces without polling +- Multiple route handlers (run-stages.tsx, run-files.tsx) need to stream execution progress and file processing status to web clients +- A centralized SSE implementation pattern (sse.ts) has emerged to handle persistent connections and event streaming across different runtime contexts +- The pattern provides a consistent mechanism for broadcasting configuration changes and environment state updates during execution workflows + +## Problem Statement + +Applications need a reliable, standardized mechanism to push real-time runtime state, configuration updates, and execution progress to clients without the overhead and latency of polling-based approaches, while maintaining consistent event handling across multiple execution contexts. + +## Decision + +1. SHOULD: SSE event streams SHOULD include structured data payloads (JSON) with type identifiers to enable client-side event routing + +## Policy Block + +- SHOULD SSE event streams SHOULD include structured data payloads (JSON) with type identifiers to enable client-side event routing + +In scope: +- Web route handlers that expose runtime execution progress (run-stages, run-files, etc.) +- Configuration management endpoints that broadcast environment or setting changes +- Real-time monitoring interfaces for execution workflows +- Client-facing APIs that require push-based state synchronization + +Out of scope: +- Bidirectional communication patterns requiring client-to-server messaging (use WebSockets instead) +- Binary data streaming or large file transfers +- Internal service-to-service communication (use message queues or gRPC) +- Static configuration loading at application startup + +## Rationale + +- Pattern detected across 3 files with 91.07% confidence indicates a deliberate architectural choice for real-time state communication +- SSE provides a lightweight, HTTP-based protocol that works through firewalls and proxies without requiring WebSocket infrastructure +- Centralized SSE library (sse.ts) ensures consistent event formatting, error handling, and connection lifecycle management across multiple route handlers +- The pattern aligns with unidirectional data flow from server to client, which matches the use case of broadcasting runtime state and configuration updates + +## Consequences + +Positive: +- Clients receive real-time updates without polling overhead, reducing server load and improving responsiveness +- Consistent SSE implementation across routes reduces code duplication and simplifies maintenance +- HTTP-based protocol ensures compatibility with existing infrastructure (load balancers, proxies, CDNs) +- Automatic reconnection support in browsers provides resilience against transient network failures + +Negative: +- SSE connections are unidirectional, requiring separate HTTP requests for client-to-server communication +- Long-lived connections may require infrastructure tuning (timeouts, connection limits) to support many concurrent clients +- Browser connection limits (typically 6 per domain) may constrain applications with multiple simultaneous SSE streams +- Debugging and monitoring SSE connections can be more complex than traditional request-response patterns + +## Alternatives + +- Use polling-based approach with periodic HTTP requests to fetch runtime state updates (rejected) + Rejected because: Polling introduces latency, increases server load with redundant requests, and provides poor user experience for real-time execution monitoring + When valid: Acceptable for low-frequency updates (>30 seconds) or when SSE infrastructure is unavailable +- Implement WebSocket-based bidirectional communication for all real-time features (rejected) + Rejected because: WebSockets add complexity and infrastructure requirements when only unidirectional server-to-client communication is needed for configuration and state updates + When valid: Appropriate when bidirectional real-time communication is required (e.g., collaborative editing, chat) +- Use GraphQL subscriptions for real-time data updates (rejected) + Rejected because: Adds GraphQL infrastructure overhead and complexity when simple event streaming suffices for runtime state broadcasting + When valid: Valid when already using GraphQL extensively and need subscription-based data synchronization + +## Risks + +- SSE connections may be terminated by intermediate proxies or load balancers with aggressive timeout policies + Mitigation: Implement periodic heartbeat messages and configure infrastructure timeouts appropriately (e.g., 60+ seconds). Document required infrastructure settings. + Owner: Engineering team and DevOps +- Memory leaks or resource exhaustion if SSE connections are not properly cleaned up on client disconnect + Mitigation: Implement robust connection lifecycle management with cleanup handlers. Add monitoring for open connection counts and memory usage. + Owner: Engineering team +- Browser connection limits may prevent multiple SSE streams from functioning simultaneously + Mitigation: Multiplex multiple event types over a single SSE connection where possible. Document connection usage and provide guidance on connection management. + Owner: Engineering team + +## Implementation Notes + +- Create a centralized SSE utility module (e.g., lib/sse.ts) that provides connection setup, event formatting, and cleanup helpers +- Ensure route handlers set appropriate headers: Content-Type: text/event-stream, Cache-Control: no-cache, Connection: keep-alive +- Implement structured event payloads with consistent format (e.g., {type: string, data: object, timestamp: number}) +- Add connection lifecycle logging and metrics to monitor SSE health and detect connection leaks +- Document SSE usage patterns and provide examples for common scenarios (execution progress, configuration updates) +- Consider implementing event replay or catch-up mechanisms for clients that reconnect after disconnection + +## Continuation Context + + +Verify commands: +- grep -r "text/event-stream" apps/fabro-web/app/routes/ | wc -l +- grep -r "import.*sse" apps/fabro-web/app/routes/ | grep -v node_modules +- find apps/fabro-web/app/lib -name "sse.ts" -o -name "sse.js" + +Accept when: +- All route handlers exposing runtime state use SSE with text/event-stream Content-Type header +- A centralized SSE utility module exists and is imported by route handlers requiring real-time updates +- SSE connections implement proper cleanup on client disconnect or execution completion + +## Enforcement + +- Verified by: Code review checklist verifying SSE usage for new real-time endpoints +- Verified by: Automated linting rules to detect missing Content-Type headers on streaming routes +- Verified by: Integration tests validating SSE connection lifecycle and event delivery +- Violation handling: Pull requests introducing polling-based approaches for real-time updates must justify why SSE is not suitable +- Violation handling: Route handlers with SSE connections missing proper cleanup must be fixed before merge +- Violation handling: Violations detected in code review are flagged and require revision +- Exception process: Document technical justification for alternative approach (e.g., WebSocket requirement, infrastructure constraints) +- Exception process: Obtain approval from technical lead or architect +- Exception process: Add architectural decision comment in code explaining exception rationale \ No newline at end of file diff --git a/docs/adr/fcd6a21d-756b-4d3f-98d7-f42d21122afd-standardize-service-boundary-logging-for-observability-service-implementations-include.md b/docs/adr/fcd6a21d-756b-4d3f-98d7-f42d21122afd-standardize-service-boundary-logging-for-observability-service-implementations-include.md new file mode 100644 index 000000000..f1731386b --- /dev/null +++ b/docs/adr/fcd6a21d-756b-4d3f-98d7-f42d21122afd-standardize-service-boundary-logging-for-observability-service-implementations-include.md @@ -0,0 +1,131 @@ +# Standardize Service Boundary Logging for Observability: Service Implementations Include + +Status: proposed +Date: 2024-01-15 +Deciders: Detection Pipeline (automated) + +## Activation + +This ADR is ACTIVE for all service implementations and applies to all components that define service boundaries, including API handlers, server implementations, and client query interfaces. + +## Context + +- The codebase exhibits a consistent pattern of logging at service boundaries across multiple components (fabro-server and fabro-web), indicating an architectural decision to instrument entry and exit points of service operations +- Service boundary logging is critical for distributed system observability, enabling request tracing, performance monitoring, and debugging across service interactions +- The pattern appears in both backend Rust server implementations and frontend TypeScript query layers, suggesting a cross-stack architectural concern +- With 3 files showing this pattern at 90.17% confidence, this represents a deliberate architectural choice rather than ad-hoc logging practices +- The facet 'boundaries.service_definitions' indicates this pattern specifically targets the interfaces where services are defined and exposed + +## Problem Statement + +Without standardized logging at service boundaries, distributed systems suffer from poor observability, making it difficult to trace requests across services, diagnose performance bottlenecks, identify failure points, and understand system behavior in production. Ad-hoc logging approaches lead to inconsistent log formats, missing context, and gaps in observability coverage. + +## Decision + +1. MAY: Service implementations MAY include additional context-specific logging within service boundaries for detailed debugging, but this does not replace boundary logging requirements + +## Policy Block + +- MAY Service implementations MAY include additional context-specific logging within service boundaries for detailed debugging, but this does not replace boundary logging requirements + +In scope: +- HTTP API handlers and endpoints +- RPC service method implementations +- GraphQL resolvers and query handlers +- Message queue consumers and producers +- Database query interfaces that represent service boundaries +- External service client wrappers +- WebSocket connection handlers + +Out of scope: +- Internal utility functions that do not represent service boundaries +- Private helper methods within a service implementation +- Data transformation and validation logic +- Pure business logic functions without external interactions +- Unit test code and test fixtures + +Exceptions: +- EXC-001: High-frequency, low-value operations (health checks, metrics endpoints) where logging every request would create excessive log volume +- EXC-002: Performance-critical hot paths where logging overhead is measured and documented to exceed acceptable latency budgets + +## Rationale + +- The detection of this pattern across 3 files with 90.17% confidence indicates this is an established architectural practice in the codebase, representing proven value in production systems +- Service boundary logging provides the minimal necessary observability coverage for distributed systems without requiring invasive instrumentation throughout the codebase +- Standardizing logging at boundaries creates consistent observability across heterogeneous technology stacks (Rust backend, TypeScript frontend), enabling unified monitoring and debugging workflows +- This approach balances observability needs with performance concerns by focusing logging on key interaction points rather than instrumenting every function + +## Consequences + +Positive: +- Improved ability to trace requests through distributed system components, reducing mean time to resolution (MTTR) for production issues +- Consistent log structure across services enables automated log aggregation, analysis, and alerting +- Clear observability boundaries make it easier for developers to understand where logging is required versus optional +- Correlation identifiers at service boundaries enable powerful distributed tracing capabilities without complex instrumentation + +Negative: +- Additional logging overhead at service boundaries may impact latency for high-frequency operations, requiring careful performance monitoring +- Developers must remember to implement boundary logging for all new service endpoints, creating potential for human error +- Log volume may increase significantly in high-traffic systems, requiring investment in log management infrastructure +- Maintaining consistent logging patterns across multiple technology stacks (Rust, TypeScript) requires ongoing coordination and code review discipline + +## Alternatives + +- Comprehensive instrumentation logging throughout the entire codebase, not just at service boundaries (rejected) + Rejected because: Creates excessive log volume, performance overhead, and maintenance burden. Most internal function calls do not provide actionable observability value compared to boundary logging. + When valid: May be appropriate for specific critical subsystems during active debugging or performance optimization efforts, but not as a general architectural pattern +- Rely solely on distributed tracing frameworks (OpenTelemetry, Jaeger) without explicit boundary logging (rejected) + Rejected because: Tracing frameworks add complexity and dependencies, may not be available in all environments, and logs provide complementary information (detailed context, error messages) that traces don't capture well + When valid: Could be adopted in addition to boundary logging as systems mature, but should not replace foundational logging practices +- Implement logging through middleware/interceptors only, without explicit logging in service code (deferred) + Rejected because: While middleware can handle many boundary logging concerns, some service-specific context and business logic details are best logged within the service implementation itself + When valid: Middleware-based logging should be used where possible for consistency, with explicit service logging for context that middleware cannot capture + +## Risks + +- Inconsistent implementation across teams and services leads to gaps in observability coverage + Mitigation: Implement automated linting rules to detect missing boundary logging, provide code templates and examples, include boundary logging checks in code review checklists + Owner: Engineering team leads and platform team +- Sensitive data leakage through logs if developers are not careful about what context they include + Mitigation: Implement automated PII detection in logs, provide sanitization utilities, conduct security training on logging best practices, enable log redaction in production + Owner: Security team and engineering team +- Performance degradation in high-throughput services due to synchronous logging overhead + Mitigation: Use async logging frameworks, implement log sampling for high-frequency endpoints, monitor logging performance impact, establish performance budgets for logging overhead + Owner: Platform team and service owners + +## Implementation Notes + +- Create shared logging utilities or macros that encapsulate boundary logging patterns, making it easy for developers to add consistent logging with minimal boilerplate +- For Rust services, leverage structured logging crates like 'tracing' or 'slog' that provide zero-cost abstractions and async logging capabilities +- For TypeScript/JavaScript services, use structured logging libraries like 'pino' or 'winston' with JSON formatters for consistency with backend logs +- Establish a standard set of log fields for service boundaries: timestamp, service_name, operation, request_id, duration_ms, status, error_message (if applicable) +- Configure log aggregation systems (ELK, Splunk, CloudWatch) to parse and index boundary logs for efficient querying and alerting +- Document logging patterns in service templates and starter kits to ensure new services follow established practices from the beginning + +## Continuation Context + + +Verify commands: +- grep -r 'log.*request' --include='*.rs' --include='*.ts' lib/crates/fabro-server/src/server/ apps/fabro-web/app/lib/ | wc -l +- rg '(info!|warn!|error!).*\(|console\.(log|info|warn|error)' --type rust --type ts -g '*/handler/*' -g '*/queries.*' +- find . -name 'server.rs' -o -name 'handler*.rs' -o -name 'queries.ts' | xargs grep -l 'trace_id\|request_id\|correlation' + +Accept when: +- All service handler files (server.rs, handler/*.rs, queries.ts) contain logging statements at operation entry and exit points +- Grep commands identify logging patterns in at least 80% of service boundary files +- Code review confirms presence of correlation identifiers (request_id, trace_id) in boundary logs + +## Enforcement + +- Verified by: Automated code review checks using static analysis tools (clippy for Rust, ESLint for TypeScript) with custom rules for boundary logging +- Verified by: Manual code review checklist items requiring reviewers to verify boundary logging presence +- Verified by: CI pipeline integration tests that verify log output from service endpoints contains required fields +- Verified by: Periodic observability audits reviewing log coverage across services +- Violation handling: CI pipeline warnings for missing boundary logging patterns (non-blocking initially, blocking after grace period) +- Violation handling: Code review feedback requiring addition of boundary logging before merge approval +- Violation handling: Quarterly observability reports identifying services with insufficient logging coverage +- Violation handling: Escalation to architecture review for repeated violations or services with poor observability +- Exception process: Submit exception request to team lead or architect with documented rationale (performance impact, alternative observability approach) +- Exception process: For performance-critical paths, provide benchmark data demonstrating logging overhead exceeds acceptable latency budget +- Exception process: Document approved exceptions in service README and architecture decision log +- Exception process: Review exceptions quarterly to determine if mitigations (async logging, sampling) can eliminate the need for exception \ No newline at end of file