diff --git a/docs/security-scanning.md b/docs/security-scanning.md index e5e09d5d..b659f81a 100644 --- a/docs/security-scanning.md +++ b/docs/security-scanning.md @@ -1,8 +1,14 @@ # Skill Scanner Backend Runtime Guide +> **Document status:** This is a historical backend implementation note, updated with the +> current Scanner 2.1.0 runtime and rollout constraints. It does not expand the supported +> analyzer or deployment contract. + ## Overview SkillHub now supports a backend-only security scanning chain around `skill-scanner`. +The Scanner image pins `cisco-ai-skill-scanner==2.1.0` and uses a glibc-based Linux runtime. +Published images support `linux/amd64` and `linux/arm64`. The publish flow changes are: 1. publish request enters `SkillPublishService` @@ -20,14 +26,17 @@ Frontend is intentionally out of scope here. The frontend should fetch audit det Two runtime modes are supported: - `local` - Use `POST /scan` and pass a filesystem path. This only works when SkillHub and `skill-scanner` can see the same files. + Use `POST /scan` and pass a filesystem path. SkillHub and `skill-scanner` must mount the same + directory at the same path. The Scanner must also allow that root; for the standard path, set + `SKILL_SCANNER_ALLOWED_ROOTS=/tmp/skillhub-scans`. - `upload` - Use `POST /scan-upload` and upload the package archive. This is the safer default for split deployments. + Use `POST /scan-upload` and upload the package archive. This is the mode used by the official + Compose and Kubernetes deployments. Recommended usage: - local development with shared filesystem: `local` -- Kubernetes or any split-service deployment: `upload` +- official Compose, Kubernetes, or any split-service deployment: `upload` ## Backend Configuration @@ -70,6 +79,7 @@ Scanner-side optional environment variables: - `SKILL_SCANNER_LLM_MODEL` - `SKILLHUB_SCANNER_MAX_CONCURRENT_SCANS` (default `1`) - `SKILLHUB_SCANNER_HARD_TIMEOUT_SECONDS` (default `930`) +- `SKILLHUB_SCANNER_MAX_UPLOAD_SIZE_BYTES` (default `110100480`, or 105 MiB) If the LLM variables are absent, the scanner should still run with non-LLM analyzers. The default timeout ordering is server read timeout (900 seconds), scanner hard timeout @@ -95,6 +105,21 @@ Relevant manifests: The scanner service is internal-only by default and is consumed by the backend through cluster DNS. +## Rolling Upgrade to Scanner 2.1.0 + +The Server and Scanner HTTP contracts must be upgraded in this order: + +1. deploy the compatibility Server release while the old Scanner is still running +2. drain and remove every old Server instance, including in-flight scan requests +3. upgrade the Scanner to 2.1.0 +4. verify `/health` and an upload-mode scan before restoring normal traffic + +Do not run an old Server against Scanner 2.1.0. During a mixed-version rollout in `upload` mode, +keep AI Defense disabled (`SKILLHUB_SCANNER_USE_AI_DEFENSE=false`, the default). If AI Defense must +remain enabled before the old Scanner is retired, configure its credential directly in the old +Scanner environment using the variable supported by that Scanner version. Never place an AI Defense +key in URL query parameters. + ## Verification Verify the scanner service itself: @@ -133,6 +158,10 @@ Response fields include: - `scannedAt` - `createdAt` +`isSafe: true` means the scan found no high-risk issue. It does not mean that the scan produced no +findings: lower-severity findings may still be present and `findingsCount` may be non-zero. The UI +therefore renders this state as **No high-risk findings**, not as an unconditional safety guarantee. + ## Failure Semantics - scan task retries are handled by `AbstractStreamConsumer` diff --git a/docs/skillhub/en/faq.md b/docs/skillhub/en/faq.md index acfe6d89..ecae1d0d 100644 --- a/docs/skillhub/en/faq.md +++ b/docs/skillhub/en/faq.md @@ -200,7 +200,23 @@ A: SkillHub has built-in security scanning. The scanner integration, task orches ## Q: Which version of cisco-ai-skill-scanner does SkillHub use? -A: `scanner/Dockerfile` runs `pip install cisco-ai-skill-scanner` directly without pinning a version, so the latest version on PyPI is pulled when the image is built. To pin a version, do so yourself when customizing the build. +A: `scanner/Dockerfile` pins `cisco-ai-skill-scanner==2.1.0`. The Scanner image uses glibc Linux and supports `linux/amd64` and `linux/arm64`. + +## Q: Should the Scanner use upload mode or local mode? + +A: The official Compose and Kubernetes deployments use `upload` mode and send skill packages through `POST /scan-upload`. `SKILLHUB_SCANNER_MAX_UPLOAD_SIZE_BYTES` controls the upload limit; its default is `110100480` bytes (105 MiB). + +Use `local` mode only when the Server and Scanner can see the same directory at the **same path**. Both services must share the mount, and the Scanner must allow that root; for the standard path, set `SKILL_SCANNER_ALLOWED_ROOTS=/tmp/skillhub-scans`. + +## Q: Does “No high-risk findings” mean that a security scan returned no findings? + +A: No. A Scanner response with `is_safe=true`, rendered in the UI as “No high-risk findings,” only means that no high-risk issue was found. Lower-severity findings may still exist and `findingsCount` may be greater than zero. Review the finding details instead of treating this state as an unconditional safety guarantee. + +## Q: How should I roll out the Scanner 2.1.0 upgrade? + +A: First deploy the Server release that is compatible with the 2.1.0 protocol. While the old Scanner is still running, drain and remove every old Server instance and its in-flight scans. Then upgrade the Scanner and verify `/health` plus one upload-mode scan. Do not connect an old Server to Scanner 2.1.0. + +During the mixed-version window, keep AI Defense disabled in upload mode (the default is `SKILLHUB_SCANNER_USE_AI_DEFENSE=false`). If AI Defense must remain enabled before the upgrade, configure its credential directly in the old Scanner environment using the variable supported by that Scanner version; never put an AI Defense key in URL query parameters. ## Q: How do I troubleshoot a `registry returned 400` error from `skillhub publish` (CLI)? diff --git a/docs/skillhub/faq.md b/docs/skillhub/faq.md index 8906150c..386cec1d 100644 --- a/docs/skillhub/faq.md +++ b/docs/skillhub/faq.md @@ -200,7 +200,23 @@ A: SkillHub 内置安全扫描能力。其中扫描接入、任务编排、审 ## Q: SkillHub 使用的 cisco-ai-skill-scanner 是哪个版本? -A: `scanner/Dockerfile` 中直接执行 `pip install cisco-ai-skill-scanner`,未锁定版本,因此构建镜像时会拉取 PyPI 上的最新版本。如需固定版本,可在二次开发时自行锁定。 +A: `scanner/Dockerfile` 已固定使用 `cisco-ai-skill-scanner==2.1.0`。Scanner 镜像基于 glibc Linux,支持 `linux/amd64` 和 `linux/arm64`。 + +## Q: Scanner 应该使用 upload mode 还是 local mode? + +A: 官方 Compose 和 Kubernetes 部署使用 `upload` mode,通过 `POST /scan-upload` 上传技能包。上传大小上限由 `SKILLHUB_SCANNER_MAX_UPLOAD_SIZE_BYTES` 控制,默认是 `110100480` 字节(105 MiB)。 + +`local` mode 仅适用于 Server 与 Scanner 能在**相同路径**看到同一目录的部署。双方需要共享挂载路径,并在 Scanner 中设置允许的根目录;使用标准路径时配置 `SKILL_SCANNER_ALLOWED_ROOTS=/tmp/skillhub-scans`。 + +## Q: 安全审计显示“未发现高风险问题”,是否代表没有任何 findings? + +A: 不是。Scanner 返回 `is_safe=true`、UI 显示“未发现高风险问题”,仅表示没有发现高风险问题;低等级 findings 仍可能存在,`findingsCount` 也可能大于 0。请继续查看 findings 明细,而不要把该状态理解为无条件安全保证。 + +## Q: 如何滚动升级到 Scanner 2.1.0? + +A: 必须先部署兼容 2.1.0 协议的 Server,在旧 Scanner 仍运行时排空并下线所有旧 Server 实例及其进行中的扫描,然后再升级 Scanner,最后验证 `/health` 和一次 upload mode 扫描。不要让旧 Server 连接 Scanner 2.1.0。 + +混合版本期间,upload mode 应保持 AI Defense 关闭(默认 `SKILLHUB_SCANNER_USE_AI_DEFENSE=false`)。如果升级前必须继续使用 AI Defense,应按旧 Scanner 版本支持的环境变量把凭据直接配置到旧 Scanner 环境中;不要把 AI Defense key 放入 URL query 参数。 ## Q: 使用 CLI `skillhub publish` 报错 `registry returned 400` 怎么排查?