mirror of
https://github.com/HKUDS/OpenSpace.git
synced 2026-10-08 03:07:51 +00:00
Add split embedding routing for local Codex sidecar
This commit is contained in:
parent
659d2da39e
commit
d62af9ecb5
12 changed files with 385 additions and 90 deletions
19
AGENTS.md
19
AGENTS.md
|
|
@ -1,5 +1,24 @@
|
|||
# AGENTS
|
||||
|
||||
## Project Skill Bucket
|
||||
|
||||
For this repository, the project-scoped OpenSpace skill bucket is:
|
||||
|
||||
- `~/.codex/projects/openspace/skills`
|
||||
- index: `~/.codex/projects/openspace/SKILL_INDEX.md`
|
||||
|
||||
Routing preference for work inside this repo:
|
||||
|
||||
1. project bucket `openspace`
|
||||
2. shared local bucket `default`
|
||||
3. common global skills
|
||||
|
||||
Mirror OpenSpace's own pattern:
|
||||
- first run `/Users/admin/.codex/tools/route_codex_skills_via_openspace.py`
|
||||
- prefilter by skill header metadata first
|
||||
- only open the most likely 1-2 `SKILL.md` files
|
||||
- avoid scanning every project skill file unless the user explicitly asks
|
||||
|
||||
## Codex Desktop Sidecar Evolution
|
||||
|
||||
Use this workflow when the user is coding in Codex Desktop with their normal subscription login and wants OpenSpace to do post-task skill capture through the isolated `openspace_evolution` sidecar.
|
||||
|
|
|
|||
|
|
@ -28,6 +28,31 @@ Instead, the final implementation uses **process-level split routing**:
|
|||
|
||||
API compatibility work was still necessary, but it is only one part of the solution.
|
||||
|
||||
## Embedding Split Routing
|
||||
|
||||
The sidecar now also supports a separate skill-embedding route from the main LLM.
|
||||
|
||||
Recommended setup:
|
||||
|
||||
```bash
|
||||
OPENSPACE_MODEL=gpt-5.4
|
||||
OPENSPACE_LLM_API_KEY=sk-xxx
|
||||
OPENSPACE_LLM_API_BASE=http://127.0.0.1:8080/v1
|
||||
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND=local
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL=BAAI/bge-small-en-v1.5
|
||||
```
|
||||
|
||||
If you want a dedicated remote endpoint for skill embeddings instead of local
|
||||
fastembed, set:
|
||||
|
||||
```bash
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND=remote
|
||||
OPENSPACE_SKILL_EMBEDDING_API_KEY=sk-embed-xxx
|
||||
OPENSPACE_SKILL_EMBEDDING_API_BASE=https://example.com/v1
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL=openai/text-embedding-3-small
|
||||
```
|
||||
|
||||
## What Was Implemented
|
||||
|
||||
### 1. OpenAI-compatible provider bridge for OpenSpace
|
||||
|
|
|
|||
81
docs/current-routing-flow.md
Normal file
81
docs/current-routing-flow.md
Normal file
|
|
@ -0,0 +1,81 @@
|
|||
# Current Routing Flow
|
||||
|
||||
This document records the current OpenSpace routing setup for this local environment.
|
||||
|
||||
## Effective Split Routing
|
||||
|
||||
- Main LLM:
|
||||
- model: `gpt-5.4`
|
||||
- API base: `http://127.0.0.1:8080/v1`
|
||||
- source: `OPENSPACE_LLM_*`
|
||||
- Skill embeddings:
|
||||
- backend: `local`
|
||||
- model: `BAAI/bge-small-en-v1.5`
|
||||
- source: `OPENSPACE_SKILL_EMBEDDING_*`
|
||||
|
||||
This means:
|
||||
|
||||
- normal OpenSpace generation and tool-calling still use the OpenAI-compatible provider path
|
||||
- skill-router semantic re-rank does not depend on remote `/v1/embeddings`
|
||||
- Codex Desktop main session remains isolated from the sidecar/provider env
|
||||
|
||||
## Flow 1: OpenSpace CLI
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A["User runs ./scripts/openspace.sh"] --> B["Load openspace/.env"]
|
||||
B --> C["Set OPENSPACE_LLM_*"]
|
||||
B --> D["Set OPENSPACE_SKILL_EMBEDDING_*"]
|
||||
C --> E["LLM client"]
|
||||
D --> F["SkillRanker"]
|
||||
E --> G["sub2api / local OpenAI-compatible gateway<br/>http://127.0.0.1:8080/v1"]
|
||||
F --> H["fastembed local model<br/>BAAI/bge-small-en-v1.5"]
|
||||
G --> I["GroundingAgent execution"]
|
||||
H --> J["BM25 + vector prefilter"]
|
||||
J --> I
|
||||
```
|
||||
|
||||
## Flow 2: Codex Desktop With OpenSpace Sidecar
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A["User runs ./scripts/codex-desktop-evolution app"] --> B["Create isolated CODEX_HOME overlay"]
|
||||
B --> C["Main Codex Desktop session"]
|
||||
B --> D["openspace_evolution MCP sidecar"]
|
||||
C --> E["Normal Codex subscription/API workflow"]
|
||||
D --> F["OpenSpace evolution server"]
|
||||
F --> G["OPENSPACE_LLM_* -> gpt-5.4 via http://127.0.0.1:8080/v1"]
|
||||
F --> H["OPENSPACE_SKILL_EMBEDDING_* -> local fastembed"]
|
||||
G --> I["Evolution / skill capture"]
|
||||
H --> I
|
||||
```
|
||||
|
||||
## Flow 3: Skill Routing Internals
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A["Task text"] --> B["Early abstain check"]
|
||||
B --> C["BM25 rough rank"]
|
||||
C --> D["Local embedding re-rank"]
|
||||
D --> E["Top candidate skills"]
|
||||
E --> F["Optional LLM selection"]
|
||||
F --> G["Injected / selected skills"]
|
||||
```
|
||||
|
||||
## Key Config Inputs
|
||||
|
||||
- `OPENSPACE_LLM_API_KEY`
|
||||
- `OPENSPACE_LLM_API_BASE`
|
||||
- `OPENSPACE_LLM_OPENAI_STREAM_COMPAT`
|
||||
- `OPENSPACE_SKILL_EMBEDDING_BACKEND`
|
||||
- `OPENSPACE_SKILL_EMBEDDING_MODEL`
|
||||
|
||||
## Operational Notes
|
||||
|
||||
- If the provider does not expose `/v1/embeddings`, the main LLM path still works.
|
||||
- With the current setup, skill embeddings stay local, so router prefilter remains available.
|
||||
- If needed later, skill embeddings can be moved to a separate remote endpoint by setting:
|
||||
- `OPENSPACE_SKILL_EMBEDDING_BACKEND=remote`
|
||||
- `OPENSPACE_SKILL_EMBEDDING_API_KEY`
|
||||
- `OPENSPACE_SKILL_EMBEDDING_API_BASE`
|
||||
- `OPENSPACE_SKILL_EMBEDDING_MODEL`
|
||||
|
|
@ -38,6 +38,18 @@ OPENROUTER_API_KEY=
|
|||
# OPENSPACE_LLM_API_KEY=sk-xxx
|
||||
# OPENSPACE_LLM_API_BASE=https://openrouter.ai/api/v1
|
||||
|
||||
# --- Recommended split routing for OpenSpace itself ---
|
||||
# Keep the main LLM on your OpenAI-compatible provider,
|
||||
# but route skill embeddings separately.
|
||||
#
|
||||
# Example: LLM via sub2api / local gateway, skill embeddings via local fastembed
|
||||
#
|
||||
# OPENSPACE_MODEL=gpt-5.4
|
||||
# OPENSPACE_LLM_API_KEY=sk-xxx
|
||||
# OPENSPACE_LLM_API_BASE=http://127.0.0.1:8080/v1
|
||||
# OPENSPACE_SKILL_EMBEDDING_BACKEND=local
|
||||
# OPENSPACE_SKILL_EMBEDDING_MODEL=BAAI/bge-small-en-v1.5
|
||||
|
||||
# ── OpenSpace Cloud (optional) ──────────────────────────────
|
||||
# Register at https://open-space.cloud to get your key.
|
||||
# Enables cloud skill search & upload; local features work without it.
|
||||
|
|
@ -51,9 +63,30 @@ OPENSPACE_API_KEY=sk_xxxxxxxxxxxxxxxx
|
|||
# Optional backup key for rate limit fallback:
|
||||
# ANTHROPIC_API_KEY_BACKUP=
|
||||
|
||||
# ── Embedding (optional) ────────────────────────────────────
|
||||
# For remote embedding API instead of local model.
|
||||
# If not set, OpenSpace uses a local embedding model (BAAI/bge-small-en-v1.5).
|
||||
# ── Skill Embedding (optional, router-only) ─────────────────
|
||||
# Controls the skill-router semantic re-rank path independently
|
||||
# from the main LLM provider.
|
||||
#
|
||||
# OPENSPACE_SKILL_EMBEDDING_BACKEND=auto
|
||||
# - auto → prefer explicit remote embedding config, then legacy OpenAI-compatible env, then local fastembed
|
||||
# - local → force local fastembed model
|
||||
# - remote → force remote OpenAI-compatible /embeddings endpoint
|
||||
#
|
||||
# OPENSPACE_SKILL_EMBEDDING_BACKEND=local
|
||||
# OPENSPACE_SKILL_EMBEDDING_MODEL=BAAI/bge-small-en-v1.5
|
||||
#
|
||||
# Or use a dedicated remote embedding endpoint:
|
||||
# OPENSPACE_SKILL_EMBEDDING_BACKEND=remote
|
||||
# OPENSPACE_SKILL_EMBEDDING_API_KEY=sk-xxx
|
||||
# OPENSPACE_SKILL_EMBEDDING_API_BASE=https://example.com/v1
|
||||
# OPENSPACE_SKILL_EMBEDDING_MODEL=openai/text-embedding-3-small
|
||||
|
||||
# ── Tool / Generic Embedding (optional) ─────────────────────
|
||||
# Used by tool search's semantic retrieval. Can also act as a fallback
|
||||
# remote embedding endpoint for the skill router when the dedicated
|
||||
# OPENSPACE_SKILL_EMBEDDING_* vars are not set.
|
||||
#
|
||||
# If not set, tool search uses a local embedding model (BAAI/bge-small-en-v1.5).
|
||||
# EMBEDDING_BASE_URL=
|
||||
# EMBEDDING_API_KEY=
|
||||
# EMBEDDING_MODEL=openai/text-embedding-3-small
|
||||
|
|
@ -68,4 +101,4 @@ OPENSPACE_API_KEY=sk_xxxxxxxxxxxxxxxx
|
|||
# LOCAL_SERVER_URL=http://127.0.0.1:5000
|
||||
|
||||
# ---- Debug (Optional) ----
|
||||
# OPENSPACE_DEBUG=true
|
||||
# OPENSPACE_DEBUG=true
|
||||
|
|
|
|||
|
|
@ -1,4 +1,13 @@
|
|||
"""Embedding generation via OpenAI-compatible API."""
|
||||
"""Embedding generation for skill routing.
|
||||
|
||||
Supports a dedicated skill-embedding path that can be routed
|
||||
independently from the main LLM:
|
||||
|
||||
- ``OPENSPACE_SKILL_EMBEDDING_BACKEND=local`` → local fastembed model
|
||||
- ``OPENSPACE_SKILL_EMBEDDING_BACKEND=remote`` → dedicated OpenAI-compatible endpoint
|
||||
- ``OPENSPACE_SKILL_EMBEDDING_BACKEND=auto`` → prefer dedicated/generic remote config,
|
||||
then fall back to legacy OpenAI-compatible env vars, then local fastembed
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
|
|
@ -11,26 +20,54 @@ from typing import List, Optional, Tuple
|
|||
|
||||
logger = logging.getLogger("openspace.cloud")
|
||||
|
||||
# Constants (duplicated here to avoid top-level import of skill_ranker)
|
||||
SKILL_EMBEDDING_MODEL = "openai/text-embedding-3-small"
|
||||
# Defaults
|
||||
SKILL_REMOTE_EMBEDDING_MODEL = "openai/text-embedding-3-small"
|
||||
SKILL_LOCAL_EMBEDDING_MODEL = "BAAI/bge-small-en-v1.5"
|
||||
SKILL_EMBEDDING_MAX_CHARS = 12_000
|
||||
SKILL_EMBEDDING_DIMENSIONS = 1536
|
||||
|
||||
_OPENROUTER_BASE = "https://openrouter.ai/api/v1"
|
||||
_OPENAI_BASE = "https://api.openai.com/v1"
|
||||
_VALID_BACKENDS = {"auto", "local", "remote"}
|
||||
_LOCAL_EMBEDDER = None
|
||||
_LOCAL_EMBEDDER_MODEL = None
|
||||
|
||||
|
||||
def resolve_embedding_api() -> Tuple[Optional[str], str]:
|
||||
"""Resolve API key and base URL for embedding requests.
|
||||
def resolve_skill_embedding_backend() -> str:
|
||||
"""Resolve skill-embedding backend mode."""
|
||||
value = os.environ.get("OPENSPACE_SKILL_EMBEDDING_BACKEND", "auto").strip().lower()
|
||||
if value in _VALID_BACKENDS:
|
||||
return value
|
||||
return "auto"
|
||||
|
||||
Priority:
|
||||
1. ``OPENROUTER_API_KEY`` → OpenRouter base URL
|
||||
2. ``OPENAI_API_KEY`` + ``OPENAI_BASE_URL`` (default ``api.openai.com``)
|
||||
3. host-agent config (nanobot / openclaw)
|
||||
|
||||
Returns:
|
||||
``(api_key, base_url)`` — *api_key* may be ``None`` when no key is found.
|
||||
"""
|
||||
def resolve_skill_embedding_model(backend: Optional[str] = None) -> str:
|
||||
"""Resolve the model name for skill embeddings."""
|
||||
backend = backend or resolve_skill_embedding_backend()
|
||||
explicit = os.environ.get("OPENSPACE_SKILL_EMBEDDING_MODEL", "").strip()
|
||||
if explicit:
|
||||
return explicit
|
||||
if backend == "local":
|
||||
return SKILL_LOCAL_EMBEDDING_MODEL
|
||||
if backend == "auto":
|
||||
remote_key, _ = _resolve_remote_embedding_api()
|
||||
if not remote_key:
|
||||
return SKILL_LOCAL_EMBEDDING_MODEL
|
||||
return SKILL_REMOTE_EMBEDDING_MODEL
|
||||
|
||||
|
||||
def _resolve_remote_embedding_api() -> Tuple[Optional[str], str]:
|
||||
"""Resolve remote embedding credentials/base URL for skill routing."""
|
||||
dedicated_key = os.environ.get("OPENSPACE_SKILL_EMBEDDING_API_KEY")
|
||||
dedicated_base = os.environ.get("OPENSPACE_SKILL_EMBEDDING_API_BASE")
|
||||
if dedicated_key and dedicated_base:
|
||||
return dedicated_key, dedicated_base.rstrip("/")
|
||||
|
||||
generic_key = os.environ.get("EMBEDDING_API_KEY")
|
||||
generic_base = os.environ.get("EMBEDDING_BASE_URL")
|
||||
if generic_key and generic_base:
|
||||
return generic_key, generic_base.rstrip("/")
|
||||
|
||||
or_key = os.environ.get("OPENROUTER_API_KEY")
|
||||
if or_key:
|
||||
return or_key, _OPENROUTER_BASE
|
||||
|
|
@ -42,6 +79,7 @@ def resolve_embedding_api() -> Tuple[Optional[str], str]:
|
|||
|
||||
try:
|
||||
from openspace.host_detection import get_openai_api_key
|
||||
|
||||
host_key = get_openai_api_key()
|
||||
if host_key:
|
||||
base = os.environ.get("OPENAI_BASE_URL", _OPENAI_BASE).rstrip("/")
|
||||
|
|
@ -52,6 +90,22 @@ def resolve_embedding_api() -> Tuple[Optional[str], str]:
|
|||
return None, _OPENAI_BASE
|
||||
|
||||
|
||||
def resolve_embedding_api() -> Tuple[Optional[str], str]:
|
||||
"""Resolve API key and base URL for remote embedding requests.
|
||||
|
||||
Priority:
|
||||
1. ``OPENSPACE_SKILL_EMBEDDING_API_*`` dedicated skill-router endpoint
|
||||
2. ``EMBEDDING_*`` generic embedding endpoint
|
||||
3. ``OPENROUTER_API_KEY`` → OpenRouter base URL
|
||||
4. ``OPENAI_API_KEY`` + ``OPENAI_BASE_URL`` (default ``api.openai.com``)
|
||||
5. host-agent config (nanobot / openclaw)
|
||||
|
||||
Returns:
|
||||
``(api_key, base_url)`` — *api_key* may be ``None`` when no key is found.
|
||||
"""
|
||||
return _resolve_remote_embedding_api()
|
||||
|
||||
|
||||
def cosine_similarity(a: List[float], b: List[float]) -> float:
|
||||
"""Compute cosine similarity between two vectors."""
|
||||
if len(a) != len(b) or not a:
|
||||
|
|
@ -81,33 +135,83 @@ def build_skill_embedding_text(
|
|||
return raw[:max_chars]
|
||||
|
||||
|
||||
def _load_local_embedder(model_name: str):
|
||||
"""Load and cache the local embedding model."""
|
||||
global _LOCAL_EMBEDDER, _LOCAL_EMBEDDER_MODEL
|
||||
|
||||
if _LOCAL_EMBEDDER is not None and _LOCAL_EMBEDDER_MODEL == model_name:
|
||||
return _LOCAL_EMBEDDER
|
||||
|
||||
try:
|
||||
from fastembed import TextEmbedding
|
||||
except ImportError:
|
||||
logger.warning(
|
||||
"Local skill embeddings requested but fastembed is not installed. "
|
||||
"Install it with `pip install fastembed`."
|
||||
)
|
||||
return None
|
||||
|
||||
try:
|
||||
logger.info("Loading local skill embedding model: %s", model_name)
|
||||
_LOCAL_EMBEDDER = TextEmbedding(model_name=model_name)
|
||||
_LOCAL_EMBEDDER_MODEL = model_name
|
||||
return _LOCAL_EMBEDDER
|
||||
except Exception as exc:
|
||||
logger.warning("Failed to load local skill embedding model %s: %s", model_name, exc)
|
||||
return None
|
||||
|
||||
|
||||
def _generate_local_embedding(text: str, model_name: str) -> Optional[List[float]]:
|
||||
embedder = _load_local_embedder(model_name)
|
||||
if embedder is None:
|
||||
return None
|
||||
|
||||
try:
|
||||
vector = next(iter(embedder.embed([text])))
|
||||
if hasattr(vector, "tolist"):
|
||||
return vector.tolist()
|
||||
return list(vector)
|
||||
except Exception as exc:
|
||||
logger.warning("Local skill embedding generation failed: %s", exc)
|
||||
return None
|
||||
|
||||
|
||||
def generate_embedding(text: str, api_key: Optional[str] = None) -> Optional[List[float]]:
|
||||
"""Generate embedding using OpenAI-compatible API.
|
||||
"""Generate skill embedding using the configured local/remote backend.
|
||||
|
||||
When *api_key* is ``None``, credentials are resolved automatically via
|
||||
:func:`resolve_embedding_api` (``OPENROUTER_API_KEY`` → ``OPENAI_API_KEY``
|
||||
→ host-agent config).
|
||||
:func:`resolve_embedding_api`.
|
||||
|
||||
This is a **synchronous** call (uses urllib). In async contexts,
|
||||
wrap with ``asyncio.to_thread()``.
|
||||
Local mode uses ``fastembed``.
|
||||
Remote mode uses an OpenAI-compatible ``/embeddings`` endpoint.
|
||||
|
||||
Args:
|
||||
text: The text to embed.
|
||||
api_key: Explicit API key. When provided, base URL is still resolved
|
||||
from environment (``OPENROUTER_API_KEY`` presence determines
|
||||
the endpoint).
|
||||
api_key: Explicit API key for remote mode.
|
||||
|
||||
Returns:
|
||||
Embedding vector, or None on failure.
|
||||
"""
|
||||
backend = resolve_skill_embedding_backend()
|
||||
model_name = resolve_skill_embedding_model(backend)
|
||||
|
||||
if backend == "local":
|
||||
return _generate_local_embedding(text, model_name)
|
||||
|
||||
resolved_key, base_url = resolve_embedding_api()
|
||||
if api_key is None:
|
||||
api_key = resolved_key
|
||||
|
||||
if not api_key:
|
||||
return None
|
||||
if backend == "remote":
|
||||
logger.warning(
|
||||
"Remote skill embeddings requested but no embedding API key/base was resolved."
|
||||
)
|
||||
return None
|
||||
return _generate_local_embedding(text, SKILL_LOCAL_EMBEDDING_MODEL)
|
||||
|
||||
body = json.dumps({
|
||||
"model": SKILL_EMBEDDING_MODEL,
|
||||
"model": model_name,
|
||||
"input": text,
|
||||
}).encode("utf-8")
|
||||
|
||||
|
|
@ -125,5 +229,7 @@ def generate_embedding(text: str, api_key: Optional[str] = None) -> Optional[Lis
|
|||
data = json.loads(resp.read().decode("utf-8"))
|
||||
return data.get("data", [{}])[0].get("embedding")
|
||||
except Exception as e:
|
||||
logger.warning("Embedding generation failed: %s", e)
|
||||
logger.warning("Remote skill embedding generation failed: %s", e)
|
||||
if backend == "auto":
|
||||
return _generate_local_embedding(text, SKILL_LOCAL_EMBEDDING_MODEL)
|
||||
return None
|
||||
|
|
|
|||
|
|
@ -35,6 +35,13 @@ Set via `.env`, MCP config `env` block, or system environment.
|
|||
| `OPENSPACE_LLM_API_BASE` | LLM API base URL | — |
|
||||
| `OPENSPACE_LLM_EXTRA_HEADERS` | Extra LLM headers (JSON) | — |
|
||||
| `OPENSPACE_LLM_CONFIG` | Arbitrary litellm kwargs (JSON) | — |
|
||||
| `OPENSPACE_SKILL_EMBEDDING_BACKEND` | Skill-router embedding backend: `auto`, `local`, or `remote` | `auto` |
|
||||
| `OPENSPACE_SKILL_EMBEDDING_MODEL` | Skill-router embedding model | `BAAI/bge-small-en-v1.5` in local mode, `openai/text-embedding-3-small` in remote mode |
|
||||
| `OPENSPACE_SKILL_EMBEDDING_API_KEY` | Dedicated remote embedding API key for skill routing | — |
|
||||
| `OPENSPACE_SKILL_EMBEDDING_API_BASE` | Dedicated remote embedding API base for skill routing | — |
|
||||
| `EMBEDDING_API_KEY` | Generic embedding API key (tool search, optional skill-router fallback) | — |
|
||||
| `EMBEDDING_BASE_URL` | Generic embedding API base URL | — |
|
||||
| `EMBEDDING_MODEL` | Generic embedding model for tool search | `BAAI/bge-small-en-v1.5` |
|
||||
| `OPENSPACE_API_KEY` | Cloud API key ([open-space.cloud](https://open-space.cloud)) | — |
|
||||
| `OPENSPACE_MAX_ITERATIONS` | Max agent iterations per task | `20` |
|
||||
| `OPENSPACE_BACKEND_SCOPE` | Enabled backends (comma-separated) | `shell,gui,mcp,web,system` |
|
||||
|
|
@ -47,6 +54,29 @@ Set via `.env`, MCP config `env` block, or system environment.
|
|||
| `OPENSPACE_ENABLE_RECORDING` | Record execution traces | `true` |
|
||||
| `OPENSPACE_LOG_LEVEL` | Log level | `INFO` |
|
||||
|
||||
### Split-routing example
|
||||
|
||||
Keep the main LLM on an OpenAI-compatible provider, but force the
|
||||
skill-router embedding path to stay local:
|
||||
|
||||
```bash
|
||||
OPENSPACE_MODEL=gpt-5.4
|
||||
OPENSPACE_LLM_API_KEY=sk-xxx
|
||||
OPENSPACE_LLM_API_BASE=http://127.0.0.1:8080/v1
|
||||
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND=local
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL=BAAI/bge-small-en-v1.5
|
||||
```
|
||||
|
||||
Or send skill embeddings to a separate endpoint:
|
||||
|
||||
```bash
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND=remote
|
||||
OPENSPACE_SKILL_EMBEDDING_API_KEY=sk-embed-xxx
|
||||
OPENSPACE_SKILL_EMBEDDING_API_BASE=https://example.com/v1
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL=openai/text-embedding-3-small
|
||||
```
|
||||
|
||||
## 3. MCP Servers (`config_mcp.json`)
|
||||
|
||||
Register external MCP servers that OpenSpace connects to as a **client** (e.g. GitHub, Slack, databases):
|
||||
|
|
|
|||
|
|
@ -7,7 +7,9 @@ Provides a two-stage retrieval pipeline for skill selection:
|
|||
Embedding strategy:
|
||||
- Text = ``name + description + SKILL.md body`` (consistent with MCP
|
||||
``search_skills`` and the clawhub cloud platform)
|
||||
- Model: ``qwen/qwen3-embedding-8b`` via OpenRouter API
|
||||
- Backend is configurable via ``OPENSPACE_SKILL_EMBEDDING_BACKEND``
|
||||
and can use either a local fastembed model or a remote
|
||||
OpenAI-compatible embedding endpoint
|
||||
- Embeddings are cached in-memory keyed by ``skill_id`` and optionally
|
||||
persisted to a pickle file for cross-session reuse
|
||||
|
||||
|
|
@ -32,8 +34,6 @@ from openspace.utils.logging import Logger
|
|||
|
||||
logger = Logger.get_logger(__name__)
|
||||
|
||||
# Embedding model — must match clawhub platform for vector-space compatibility
|
||||
SKILL_EMBEDDING_MODEL = "openai/text-embedding-3-small"
|
||||
SKILL_EMBEDDING_MAX_CHARS = 12_000
|
||||
|
||||
# Pre-filter threshold: when local skills exceed this count, BM25 pre-filter
|
||||
|
|
@ -44,7 +44,7 @@ PREFILTER_THRESHOLD = 10
|
|||
BM25_CANDIDATES_MULTIPLIER = 3 # top_k * 3
|
||||
|
||||
# Cache version — increment when format changes
|
||||
_CACHE_VERSION = 1
|
||||
_CACHE_VERSION = 2
|
||||
|
||||
|
||||
@dataclass
|
||||
|
|
@ -238,13 +238,6 @@ class SkillRanker:
|
|||
|
||||
return ranked[:top_k]
|
||||
|
||||
@staticmethod
|
||||
def _get_openai_api_key() -> Optional[str]:
|
||||
"""Resolve OpenAI-compatible API key for embedding requests."""
|
||||
from openspace.cloud.embedding import resolve_embedding_api
|
||||
api_key, _ = resolve_embedding_api()
|
||||
return api_key
|
||||
|
||||
@staticmethod
|
||||
def _build_embedding_text(candidate: SkillCandidate) -> str:
|
||||
"""Build text for embedding, consistent with MCP search_skills."""
|
||||
|
|
@ -264,12 +257,8 @@ class SkillRanker:
|
|||
top_k: int,
|
||||
) -> List[SkillCandidate]:
|
||||
"""Rank candidates using embedding cosine similarity."""
|
||||
api_key = self._get_openai_api_key()
|
||||
if not api_key:
|
||||
return []
|
||||
|
||||
# Generate query embedding
|
||||
query_emb = self._generate_embedding(query, api_key=api_key)
|
||||
query_emb = self._generate_embedding(query)
|
||||
if not query_emb:
|
||||
return []
|
||||
|
||||
|
|
@ -281,7 +270,7 @@ class SkillRanker:
|
|||
c.embedding = cached
|
||||
else:
|
||||
text = self._build_embedding_text(c)
|
||||
emb = self._generate_embedding(text, api_key=api_key)
|
||||
emb = self._generate_embedding(text)
|
||||
if emb:
|
||||
c.embedding = emb
|
||||
self._embedding_cache[c.skill_id] = emb
|
||||
|
|
@ -305,53 +294,20 @@ class SkillRanker:
|
|||
text: str,
|
||||
api_key: Optional[str] = None,
|
||||
) -> Optional[List[float]]:
|
||||
"""Generate embedding via OpenAI-compatible API (text-embedding-3-small).
|
||||
"""Generate embedding via the configured skill-embedding backend.
|
||||
|
||||
Delegates credential / base-URL resolution to
|
||||
:func:`openspace.cloud.embedding.resolve_embedding_api`.
|
||||
Delegates backend/model resolution to :mod:`openspace.cloud.embedding`.
|
||||
"""
|
||||
from openspace.cloud.embedding import resolve_embedding_api
|
||||
from openspace.cloud.embedding import generate_embedding
|
||||
|
||||
resolved_key, base_url = resolve_embedding_api()
|
||||
if not api_key:
|
||||
api_key = resolved_key
|
||||
if not api_key:
|
||||
return None
|
||||
|
||||
import urllib.request
|
||||
|
||||
body = json.dumps({
|
||||
"model": SKILL_EMBEDDING_MODEL,
|
||||
"input": text,
|
||||
}).encode("utf-8")
|
||||
|
||||
req = urllib.request.Request(
|
||||
f"{base_url}/embeddings",
|
||||
data=body,
|
||||
headers={
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {api_key}",
|
||||
},
|
||||
method="POST",
|
||||
)
|
||||
import time
|
||||
last_err = None
|
||||
for attempt in range(3):
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=15) as resp:
|
||||
data = json.loads(resp.read().decode("utf-8"))
|
||||
return data.get("data", [{}])[0].get("embedding")
|
||||
except Exception as e:
|
||||
last_err = e
|
||||
if attempt < 2:
|
||||
delay = 2 * (attempt + 1)
|
||||
logger.debug("Embedding request failed (attempt %d/3), retrying in %ds: %s", attempt + 1, delay, e)
|
||||
time.sleep(delay)
|
||||
logger.warning("Skill embedding generation failed after 3 attempts: %s", last_err)
|
||||
return None
|
||||
return generate_embedding(text, api_key=api_key)
|
||||
|
||||
def _cache_file(self) -> Path:
|
||||
return self._cache_dir / f"skill_embeddings_v{_CACHE_VERSION}.pkl"
|
||||
from openspace.cloud.embedding import resolve_skill_embedding_model
|
||||
|
||||
model_name = resolve_skill_embedding_model()
|
||||
safe_model_name = re.sub(r"[^a-zA-Z0-9_.-]+", "_", model_name)
|
||||
return self._cache_dir / f"skill_embeddings_{safe_model_name}_v{_CACHE_VERSION}.pkl"
|
||||
|
||||
def _load_cache(self) -> None:
|
||||
"""Load embedding cache from disk."""
|
||||
|
|
@ -376,7 +332,7 @@ class SkillRanker:
|
|||
self._cache_dir.mkdir(parents=True, exist_ok=True)
|
||||
data = {
|
||||
"version": _CACHE_VERSION,
|
||||
"model": SKILL_EMBEDDING_MODEL,
|
||||
"model": self._cache_file().stem,
|
||||
"last_updated": datetime.now().isoformat(),
|
||||
"embeddings": self._embedding_cache,
|
||||
}
|
||||
|
|
@ -412,4 +368,3 @@ def build_skill_embedding_text(
|
|||
if len(raw) <= max_chars:
|
||||
return raw
|
||||
return raw[:max_chars]
|
||||
|
||||
|
|
|
|||
|
|
@ -17,6 +17,7 @@ dependencies = [
|
|||
"litellm>=1.70.0,<1.82.7", # pinned to avoid PYSEC-2026-2 supply-chain compromise (1.82.7/1.82.8 were malicious)
|
||||
"python-dotenv>=1.0.0",
|
||||
"openai>=1.0.0",
|
||||
"fastembed>=0.8.0",
|
||||
"jsonschema>=4.25.0",
|
||||
"mcp>=1.0.0",
|
||||
"websockets>=15.0.0",
|
||||
|
|
|
|||
|
|
@ -2,6 +2,7 @@
|
|||
litellm>=1.70.0,<1.82.7 # pinned to avoid PYSEC-2026-2 supply-chain compromise (1.82.7/1.82.8 were malicious)
|
||||
python-dotenv>=1.0.0
|
||||
openai>=1.0.0
|
||||
fastembed>=0.8.0
|
||||
jsonschema>=4.25.0
|
||||
mcp>=1.0.0
|
||||
websockets>=15.0.0
|
||||
|
|
|
|||
|
|
@ -53,6 +53,15 @@ OPENSPACE_LLM_API_BASE="${OPENSPACE_LLM_API_BASE:-$(read_env_value OPENSPACE_LLM
|
|||
OPENSPACE_LLM_API_BASE="${OPENSPACE_LLM_API_BASE:-https://codexapi.space/v1}"
|
||||
OPENSPACE_LLM_OPENAI_STREAM_COMPAT="${OPENSPACE_LLM_OPENAI_STREAM_COMPAT:-$(read_env_value OPENSPACE_LLM_OPENAI_STREAM_COMPAT)}"
|
||||
OPENSPACE_LLM_OPENAI_STREAM_COMPAT="${OPENSPACE_LLM_OPENAI_STREAM_COMPAT:-true}"
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND="${OPENSPACE_SKILL_EMBEDDING_BACKEND:-$(read_env_value OPENSPACE_SKILL_EMBEDDING_BACKEND)}"
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND="${OPENSPACE_SKILL_EMBEDDING_BACKEND:-local}"
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL="${OPENSPACE_SKILL_EMBEDDING_MODEL:-$(read_env_value OPENSPACE_SKILL_EMBEDDING_MODEL)}"
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL="${OPENSPACE_SKILL_EMBEDDING_MODEL:-BAAI/bge-small-en-v1.5}"
|
||||
OPENSPACE_SKILL_EMBEDDING_API_KEY="${OPENSPACE_SKILL_EMBEDDING_API_KEY:-$(read_env_value OPENSPACE_SKILL_EMBEDDING_API_KEY)}"
|
||||
OPENSPACE_SKILL_EMBEDDING_API_BASE="${OPENSPACE_SKILL_EMBEDDING_API_BASE:-$(read_env_value OPENSPACE_SKILL_EMBEDDING_API_BASE)}"
|
||||
EMBEDDING_API_KEY="${EMBEDDING_API_KEY:-$(read_env_value EMBEDDING_API_KEY)}"
|
||||
EMBEDDING_BASE_URL="${EMBEDDING_BASE_URL:-$(read_env_value EMBEDDING_BASE_URL)}"
|
||||
EMBEDDING_MODEL="${EMBEDDING_MODEL:-$(read_env_value EMBEDDING_MODEL)}"
|
||||
|
||||
sync_profile_dir() {
|
||||
local name="$1"
|
||||
|
|
@ -78,7 +87,7 @@ bootstrap_profile_home() {
|
|||
|
||||
cp "$PRIMARY_CODEX_HOME/auth.json" "$PROFILE_HOME/auth.json"
|
||||
|
||||
python3 - <<'PY' "$PRIMARY_CODEX_HOME/config.toml" "$PROFILE_HOME/config.toml" "$REPO_ROOT" "$REPO_PYTHON" "$PROJECT_SKILL_DIR" "$PROFILE_HOME/skills" "$OPENSPACE_MODEL" "$OPENSPACE_LLM_API_KEY" "$OPENSPACE_LLM_API_BASE" "$OPENSPACE_LLM_OPENAI_STREAM_COMPAT"
|
||||
python3 - <<'PY' "$PRIMARY_CODEX_HOME/config.toml" "$PROFILE_HOME/config.toml" "$REPO_ROOT" "$REPO_PYTHON" "$PROJECT_SKILL_DIR" "$PROFILE_HOME/skills" "$OPENSPACE_MODEL" "$OPENSPACE_LLM_API_KEY" "$OPENSPACE_LLM_API_BASE" "$OPENSPACE_LLM_OPENAI_STREAM_COMPAT" "$OPENSPACE_SKILL_EMBEDDING_BACKEND" "$OPENSPACE_SKILL_EMBEDDING_MODEL" "$OPENSPACE_SKILL_EMBEDDING_API_KEY" "$OPENSPACE_SKILL_EMBEDDING_API_BASE" "$EMBEDDING_API_KEY" "$EMBEDDING_BASE_URL" "$EMBEDDING_MODEL"
|
||||
from pathlib import Path
|
||||
import os
|
||||
import sys
|
||||
|
|
@ -93,6 +102,13 @@ model = sys.argv[7]
|
|||
api_key = sys.argv[8]
|
||||
api_base = sys.argv[9]
|
||||
stream_compat = sys.argv[10]
|
||||
skill_embedding_backend = sys.argv[11]
|
||||
skill_embedding_model = sys.argv[12]
|
||||
skill_embedding_api_key = sys.argv[13]
|
||||
skill_embedding_api_base = sys.argv[14]
|
||||
embedding_api_key = sys.argv[15]
|
||||
embedding_base_url = sys.argv[16]
|
||||
embedding_model = sys.argv[17]
|
||||
|
||||
def strip_tables(text: str, table_names: set[str]) -> str:
|
||||
kept = []
|
||||
|
|
@ -134,6 +150,13 @@ OPENSPACE_MODEL = "{model}"
|
|||
OPENSPACE_LLM_API_KEY = "{api_key}"
|
||||
OPENSPACE_LLM_API_BASE = "{api_base}"
|
||||
OPENSPACE_LLM_OPENAI_STREAM_COMPAT = "{stream_compat}"
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND = "{skill_embedding_backend}"
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL = "{skill_embedding_model}"
|
||||
OPENSPACE_SKILL_EMBEDDING_API_KEY = "{skill_embedding_api_key}"
|
||||
OPENSPACE_SKILL_EMBEDDING_API_BASE = "{skill_embedding_api_base}"
|
||||
EMBEDDING_API_KEY = "{embedding_api_key}"
|
||||
EMBEDDING_BASE_URL = "{embedding_base_url}"
|
||||
EMBEDDING_MODEL = "{embedding_model}"
|
||||
OPENSPACE_ENABLE_RECORDING = "false"
|
||||
OPENSPACE_BACKEND_SCOPE = "shell,system"
|
||||
'''
|
||||
|
|
@ -149,6 +172,9 @@ clear_openspace_env() {
|
|||
for var in ${!OPENSPACE_@}; do
|
||||
unset "$var"
|
||||
done
|
||||
unset EMBEDDING_API_KEY
|
||||
unset EMBEDDING_BASE_URL
|
||||
unset EMBEDDING_MODEL
|
||||
}
|
||||
|
||||
if [[ -z "$OPENSPACE_LLM_API_KEY" ]]; then
|
||||
|
|
|
|||
|
|
@ -24,6 +24,8 @@ fi
|
|||
OPENSPACE_MODEL="${OPENSPACE_MODEL:-gpt-5.4}"
|
||||
OPENSPACE_LLM_API_BASE="${OPENSPACE_LLM_API_BASE:-https://codexapi.space/v1}"
|
||||
OPENSPACE_LLM_OPENAI_STREAM_COMPAT="${OPENSPACE_LLM_OPENAI_STREAM_COMPAT:-true}"
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND="${OPENSPACE_SKILL_EMBEDDING_BACKEND:-local}"
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL="${OPENSPACE_SKILL_EMBEDDING_MODEL:-BAAI/bge-small-en-v1.5}"
|
||||
|
||||
sync_profile_dir() {
|
||||
local name="$1"
|
||||
|
|
@ -102,6 +104,13 @@ OPENSPACE_MODEL = "$OPENSPACE_MODEL"
|
|||
OPENSPACE_LLM_API_KEY = "$OPENSPACE_LLM_API_KEY"
|
||||
OPENSPACE_LLM_API_BASE = "$OPENSPACE_LLM_API_BASE"
|
||||
OPENSPACE_LLM_OPENAI_STREAM_COMPAT = "$OPENSPACE_LLM_OPENAI_STREAM_COMPAT"
|
||||
OPENSPACE_SKILL_EMBEDDING_BACKEND = "$OPENSPACE_SKILL_EMBEDDING_BACKEND"
|
||||
OPENSPACE_SKILL_EMBEDDING_MODEL = "$OPENSPACE_SKILL_EMBEDDING_MODEL"
|
||||
OPENSPACE_SKILL_EMBEDDING_API_KEY = "${OPENSPACE_SKILL_EMBEDDING_API_KEY:-}"
|
||||
OPENSPACE_SKILL_EMBEDDING_API_BASE = "${OPENSPACE_SKILL_EMBEDDING_API_BASE:-}"
|
||||
EMBEDDING_API_KEY = "${EMBEDDING_API_KEY:-}"
|
||||
EMBEDDING_BASE_URL = "${EMBEDDING_BASE_URL:-}"
|
||||
EMBEDDING_MODEL = "${EMBEDDING_MODEL:-}"
|
||||
|
||||
[model_providers.codexapi]
|
||||
name = "codexapi"
|
||||
|
|
|
|||
|
|
@ -5,6 +5,13 @@ SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
|||
REPO_ROOT="$(cd -- "$SCRIPT_DIR/.." && pwd)"
|
||||
ALT_HOME="${CODEX_HOME:-$HOME/.codex-openspace}"
|
||||
AUTH_FILE="${OPENSPACE_AUTH_FILE:-$ALT_HOME/auth.json}"
|
||||
ENV_FILE="${OPENSPACE_ENV_FILE:-$REPO_ROOT/openspace/.env}"
|
||||
|
||||
if [[ -f "$ENV_FILE" ]]; then
|
||||
set -a
|
||||
source "$ENV_FILE"
|
||||
set +a
|
||||
fi
|
||||
|
||||
api_key="${OPENSPACE_LLM_API_KEY:-}"
|
||||
if [[ -z "$api_key" ]]; then
|
||||
|
|
@ -28,8 +35,10 @@ fi
|
|||
|
||||
export OPENSPACE_MODEL="${OPENSPACE_MODEL:-gpt-5.4}"
|
||||
export OPENSPACE_LLM_API_KEY="$api_key"
|
||||
export OPENSPACE_LLM_API_BASE="${OPENSPACE_LLM_API_BASE:-https://codexapi.space/v1}"
|
||||
export OPENSPACE_LLM_API_BASE="${OPENSPACE_LLM_API_BASE:-http://127.0.0.1:8080/v1}"
|
||||
export OPENSPACE_LLM_OPENAI_STREAM_COMPAT="${OPENSPACE_LLM_OPENAI_STREAM_COMPAT:-true}"
|
||||
export OPENSPACE_SKILL_EMBEDDING_BACKEND="${OPENSPACE_SKILL_EMBEDDING_BACKEND:-local}"
|
||||
export OPENSPACE_SKILL_EMBEDDING_MODEL="${OPENSPACE_SKILL_EMBEDDING_MODEL:-BAAI/bge-small-en-v1.5}"
|
||||
export OPENSPACE_HOST_SKILL_DIRS="${OPENSPACE_HOST_SKILL_DIRS:-$ALT_HOME/skills}"
|
||||
export OPENSPACE_WORKSPACE="${OPENSPACE_WORKSPACE:-$REPO_ROOT}"
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue