* feat: Add safeWriteJson utility for atomic file operations
Implements a robust JSON file writing utility that:
- Prevents concurrent writes to the same file using in-memory locks
- Ensures atomic operations with temporary file and backup strategies
- Handles error cases with proper rollback mechanisms
- Cleans up temporary files even when operations fail
- Provides comprehensive test coverage for success and failure scenarios
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* fix: use safeWriteJson for all JSON file writes
This change refactors all direct JSON file writes to use the safeWriteJson
utility, which implements atomic file writes to prevent data corruption
during write operations.
- Modified safeWriteJson to accept optional replacer and space arguments
- Updated tests to verify correct behavior with the new implementation
Fixes: #722
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: Implement inter-process file locking for safeWriteJson
Replaces the previous in-memory lock in `safeWriteJson` with
`proper-lockfile` to provide robust, cross-process advisory file
locking. This enhances safety when multiple processes might attempt
concurrent writes to the same JSON file.
Key changes:
- Added `proper-lockfile` and `@types/proper-lockfile` dependencies.
- `safeWriteJson` now uses `proper-lockfile.lock()` with configured
retries, staleness checks (31s), and lock update intervals (10s).
- An `onCompromised` handler is included to manage scenarios where
the lock state is unexpectedly altered.
- Logging and comments within `safeWriteJson` have been refined for
clarity, ensuring error logs include backtraces.
- The test suite `safeWriteJson.test.ts` has been significantly
updated to:
- Use real timers (`jest.useRealTimers()`).
- Employ a more comprehensive mock for `fs/promises`.
- Correctly manage file pre-existence for various scenarios.
- Simulate lock contention by mocking `proper-lockfile.lock()`
using `jest.doMock` and a dynamic require for the SUT.
- Verify lock release by checking for the absence of the `.lock`
file.
All tests are passing with these changes.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: implement streaming JSON write in safeWriteJson
Refactor safeWriteJson to use stream-json for memory-efficient JSON serialization:
- Replace in-memory string creation with streaming pipeline
- Add Disassembler and Stringer from stream-json library
- Extract streaming logic to a dedicated helper function
- Add proper-lockfile and stream-json dependencies
This implementation reduces memory usage when writing large JSON objects.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* fix: improve safeWriteJson locking mechanism
- Use file path itself for locking instead of separate lock file
- Improve error handling and clarity of code
- Enhance cleanup of temporary files
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* test: fix safeWriteJson test failures
- Ensure test file exists before locking
- Add proper mocking for fs.createWriteStream
- Fix test assertions to match expected behavior
- Improve test comments to follow project guidelines
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* test: update tests to work with safeWriteJson
Updated tests to work with safeWriteJson instead of direct fs.writeFile calls:
- Updated importExport.test.ts to expect safeWriteJson calls instead of fs.writeFile
- Fixed McpHub.test.ts by properly mocking fs/promises module:
- Moved jest.mock() to the top of the file before any imports
- Added mock implementations for all fs functions used by safeWriteJson
- Updated the test setup to work with the mocked fs module
All tests now pass successfully.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: replace JSON.stringify with safeWriteJson for file operations
Replace all non-test instances of JSON.stringify used for writing to JSON files with safeWriteJson to ensure safer file operations with proper locking, error handling, and atomic writes.
- Updated src/services/mcp/McpHub.ts
- Updated src/services/code-index/cache-manager.ts
- Updated src/api/providers/fetchers/modelEndpointCache.ts
- Updated src/api/providers/fetchers/modelCache.ts
- Updated tests to match the new implementation
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* docs: add rules for using safeWriteJson
Add concise rules for using safeWriteJson instead of JSON.stringify with file operations to ensure atomic writes and prevent data corruption.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
---------
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
Co-authored-by: Eric Wheeler <roo-code@z.ewheeler.org>
Co-authored-by: Daniel <57051444+daniel-lxs@users.noreply.github.com>
* feat: Add OpenAI Compatible embedder for codebase indexing
- Implement OpenAiCompatibleEmbedder with batching and retry logic
- Add configuration support for base URL and API key
- Update UI with provider selection and input fields
- Add comprehensive test coverage
- Support for all OpenAI-compatible endpoints (LiteLLM, LMStudio, Ollama, etc.)
- Add internationalization for 17 languages
* fix: Update CodeIndexSettings tests for OpenAI Compatible provider
- Fix field count expectations (4 fields including Qdrant)
- Use specific test IDs for button selection
- Fix input handling with clear() before type()
- Use toHaveBeenLastCalledWith for better assertions
- Fix status text matching with regex pattern
* fix: resolve UI test failures and ESLint errors
- Remove unused waitFor import to fix ESLint error
- Fix test expectations to match actual component behavior for input fields
- Simplify provider selection test by removing complex mock interactions
- All CodeIndexSettings tests now pass (20/20)
* feat: add custom model infrastructure for OpenAI-compatible embedder
- Add manual model ID and embedding dimension configuration
- Enable custom model input via text field in settings UI
- Add modelDimension parameter to OpenAiCompatibleEmbedder
- Update configuration management to persist dimension setting
- Prioritize manual dimension over hardcoded model profiles
- Add comprehensive test coverage for new functionality
This allows users to specify any custom embedding model and its
dimension for OpenAI-compatible providers, removing dependency
on hardcoded model profiles.
* Add missing translations for OpenAI-compatible model dimension settings in all locales
* refactor: remove unused modelDimension parameter from OpenAiCompatibleEmbedder
- Remove modelDimension property and constructor parameter from OpenAiCompatibleEmbedder class
- Update ServiceFactory to not pass dimension to embedder constructor
- Update tests to match new constructor signature
- The dimension is still used for QdrantVectorStore configuration
* chore: bot suggestion
Co-authored-by: ellipsis-dev[bot] <65095814+ellipsis-dev[bot]@users.noreply.github.com>
* chore: bot suggestion
Co-authored-by: ellipsis-dev[bot] <65095814+ellipsis-dev[bot]@users.noreply.github.com>
* refactor: rename OpenAiCompatibleEmbedder to OpenAICompatibleEmbedder for consistency
* feat: add model dimension validation for OpenAI-compatible settings
* refactor: improve default model ID retrieval logic for embedding providers
* feat: add default model ID retrieval for openai-compatible provider
* refactor: update default model ID retrieval to use shared utility function
* fix: Remove unnecessary type assertion in OpenAICompatibleEmbedder
* feat: add model dimension input for openai-compatible provider
---------
Co-authored-by: Daniel Riccio <ricciodaniel98@gmail.com>
Co-authored-by: Daniel <57051444+daniel-lxs@users.noreply.github.com>
Co-authored-by: ellipsis-dev[bot] <65095814+ellipsis-dev[bot]@users.noreply.github.com>
- Replace manual PATH and HOME env var handling with getDefaultEnvironment()
- Improves consistency and reliability of MCP client environment setup
- Leverages SDK's built-in environment configuration
* add support for mcp server instructions
* Update McpView.tsx
* feat(mcp): add instructions field to MCP localization files and update UI to display instructions
---------
Co-authored-by: huixin <yuanhx@cffex.com.cn>
Co-authored-by: Daniel Riccio <ricciodaniel98@gmail.com>
* Implement support for streamable-http transport type mcp servers
* add streamable-http mock in same fashion as sse - which does not seem to currently be actually leveraged
* rename mock to resolve kebabcase vs camelCase
* fix (seemingly unrelatd) test failure in writeToFileTool.test.ts
* fix tests
* feat: Make checkpoint on new task
Ensures that invoking the `newTaskTool` always creates a checkpoint, even if no files have changed. This provides a consistent state snapshot before a sub-task is initiated.
* refactor: remove delay
---------
Co-authored-by: Daniel <57051444+daniel-lxs@users.noreply.github.com>
* feat: Enhance configuration change handling and model dimension checks
* fix: Update embedder creation to use modelId from config
* test: Add unit tests for ServiceFactory embedder and vector store creation
* test: Add comprehensive restart detection tests for configuration changes
* feat: Implement model ID selection logic for provider changes in CodeIndexSettings
* fix: Initialize configuration on constructor to prevent false restart triggers
* fix: Enhance API key handling and restart logic in CodeIndexConfigManager
* fix: Improve handling of external settings changes and automatic indexing in webviewMessageHandler
* fix: Ensure handleExternalSettingsChange only restarts service when manager is initialized
* refactor: remove console logs
* fix: Load configuration during initialization to ensure correct state and restart requirements
* feat: apply changes from local main
* fix: add missing types
* feat: deduplicate code blocks coming out of parser
* feat: implement a cache manager to improve cache handling
* refactor: move code index service initialization to extension and remove await from indexing process
* fix: return undefined instead of throwing if no workspace is detected
* feat: allow auto approve if it is active for read tools
* refactor: improve UI of the results and allow opening the ranges directly in the editor
* refactor: use dependency injection to improve performance
* feat: implement result filtering by directory path
* refactor: centralize path normalization logic
* refactor: remove unnecessary barrel file
* refactor: prevent restarting the service if no settings change
* fix: the indexing process should never be awaited
* refactor: cleanup unused method
* refactor: remove batch limits for ollama
* refactor(parser): simplify method signatures and improve chunking logic
- Remove redundant min/max chars parameters
- Add better handling for oversized lines
- Improve chunking logic with segment handling
- Clean up method signatures and parameter ordering
* fix(settings): make select inputs full width in CodeIndexSettings
* refactor: increase max list file limit
* feat(ui): improve codebase search result display formatting
* test: add tests for cache and config managers
* test: create unit tests for parser and scanner
* feat(parser): improve segment hash uniqueness
- Added startCharIndex to segment hash calculation in _chunkTextByLines
- Track character position when splitting oversized lines
- Ensures unique identification of segments from same line
* feat(file-watcher): add error logging and optional ignoreController injection
* fix: allow getting the state if the service is disabled
* fix: set the embedding models when cline provider is initialized
* feat: use zod to validate form
* feat(file-watcher): enhance file watcher for batched deletions and improved vector store interactions
Improve file watcher to handle file deletions in batches and optimize vector store operations.
* feat(CodeIndexSettings): move OpenAI key input to a conditional rendering block
* feat(CodeIndexSettings): update button visibility based on indexing status
* feat(file-watcher): refactor vscode mock and enhance file watcher tests
* fix(CodeIndexManager): do not await startIndexing on configuration changes
* feat(types): add codeIndexOpenAiKey and codeIndexQdrantApiKey to ProviderSettings and IpcMessage
* feat(FileWatcher): enhance file processing with batch operations and new status handling
* fix(webviewMessageHandler): handle errors during CodeIndexManager initialization
* refactor(CodeIndexManager): streamline service creation by consolidating into a single method
* feat(CodeIndex): implement minimum search score configuration and update search methods
* refactor(CodeIndexSettings): replace ApiConfiguration with ProviderSettings and update related methods
* refactor: move contants to centralized file
* refactor(constants): rename CODEBASE_INDEX_SEARCH_MIN_SCORE to SEARCH_MIN_SCORE
* feat(QdrantVectorStore): enhance search functionality with new query structure and indexing
* feat(FileWatcher): implement batch processing and retry logic for upserting points
* fix(CodeIndexSettings): rename setProviderSettingsField to setApiConfigurationField and move model label
* fix(ChatRow): remove limit from search query messages
* refactor(CodebaseSearchResult): remove unused props from component
* feat: implement batch processing for file events in FileWatcher
- Introduced a new mechanism to accumulate file events (create, change, delete) and process them in batches.
- Added debounce functionality to optimize processing frequency.
- Emitted events for batch processing start, progress updates, and completion with detailed summaries.
- Refactored existing processing logic to handle batch deletions and upserts efficiently.
- Enhanced error handling and logging for better traceability during batch operations.
* feat(CodeIndex): implement batch processing and update progress reporting
* fix: define a default url for qdrant
* feat(CodeIndexManager): add initialization check and update startIndexing logic
* feat(CodeIndexSettings): validate Qdrant URL and update settings commitment logic
* feat: refactor progress calculation and update progress bar rendering
* refactor: remove webview provider and related methods
* fix: simplify indexing status update by directly using update values
* feat: integrate .gitignore support into file processing and scanning logic
* fix: update clearCacheFile method to write an empty object instead of deleting the cache file
* Revert this
* Run prettier
* fix: add new dependencies for qdrant client and directory scanner
* feat: add codebase search functionality to localization files
* feat: add localization strings for codebase indexing settings
* feat: integrate CodeIndexSettings into ExperimentalSettings and update settings localization
* refactor: remove console logs from various components for cleaner output
* feat: enhance capabilities section and codebase search tool description
* feat: add code indexing localization for multiple languages
* fix: correct indentation for CodeIndexSettings component in ExperimentalSettings
* refactor: update unit tests to properly test current functionality
* feat: add mock implementation for p-limit and update Jest config
* feat: track file creation, change, and deletion events in accumulatedEvents
* refactor: simplify file watcher tests by removing waitForFileProcessingToFinish and using direct event accumulation
* refactor: mock ContextProxy's getValue method to return current config name in ClineProvider tests
* refactor: mock missing properties required by codebase indexing manager
---------
Co-authored-by: cte <cestreich@gmail.com>
Co-authored-by: ellipsis-dev[bot] <65095814+ellipsis-dev[bot]@users.noreply.github.com>
Co-authored-by: Matt Rubens <mrubens@users.noreply.github.com>
Co-authored-by: Your Name <you@example.com>
* refactor: standardize C# Tree-Sitter parser tests
This commit standardizes the C# Tree-Sitter parser tests to ensure:
- Variable/function names in sample code match the tests
- Each test only verifies specific identifiers relevant to that test
- Test names clearly indicate what they're testing
- No skipped tests remain
- Only one test exists per structure or combination of structures
- All debug output uses debugLog from helpers.ts
The standardization improves maintainability and clarity while preserving
all existing functionality.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: standardize C Tree-Sitter parser tests
This commit standardizes the C Tree-Sitter parser tests to ensure:
- Variable/function names in sample code match the tests (using snake_case)
- Each test only verifies specific identifiers relevant to that test
- Test names clearly indicate what they're testing
- No skipped tests remain - tests for unsupported features are removed
- Only one test exists per structure or combination of structures
- All debug output uses debugLog from helpers.ts
The standardization improves maintainability and clarity while preserving
all existing functionality. Documentation has been added for C language
constructs that are not currently supported by the parser.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: standardize Kotlin Tree-Sitter parser
Standardized Kotlin Tree-Sitter parser to ensure consistent naming and structure:
- Renamed identifiers to clearly indicate what they're testing
- Ensured all code sections are at least 4 lines long
- Organized query patterns logically with clear comments
- Removed skipped tests and unused query patterns
- Ensured only one test per structure or combination of structures
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: standardize PHP Tree-Sitter parser
Standardized PHP Tree-Sitter parser to ensure consistent naming and structure:
- Renamed identifiers to clearly indicate what they're testing
- Ensured all code sections are at least 4 lines long
- Organized query patterns logically with clear comments
- Consolidated related tests to ensure only one test per structure
- Ensured all debug output uses debugLog from helpers.ts
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: standardize Ruby Tree-Sitter parser
Standardized Ruby Tree-Sitter parser to ensure consistent naming and structure:
- Expanded query to support many more Ruby language constructs
- Added comprehensive test file with clear section headers
- Ensured all code sections are at least 4 lines long
- Added tests for each language construct
- Used consistent naming conventions for test identifiers
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: standardize Swift Tree-Sitter parser
Standardized Swift Tree-Sitter parser to ensure consistent naming and structure:
- Renamed identifiers to clearly indicate what they're testing
- Ensured all code sections are at least 4 lines long
- Organized query patterns logically with clear comments
- Removed skipped tests and unused query patterns
- Ensured only one test per structure or combination of structures
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: dynamically include all tree-sitter WASM files in build process
- Modified esbuild.js to dynamically read and copy all WASM files from tree-sitter-wasms package
- Added error handling and logging for better debugging
- Ensures all language parsers are automatically included without requiring manual updates
* refactor: move test mocks and fixtures to shared locations
Move test mocks to helpers.ts to be shared between test files
Create fixtures directory for shared test data
Move sample TSX content to fixtures/sample-tsx.ts
Update test files to use shared mocks and fixtures
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: simplify tree-sitter debug output
- Remove detailed node inspection to reduce noise
- Keep only essential tree structure output
- Remove duplicate content logging
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: standardize C# tree-sitter test structure
Move C# sample code to dedicated fixture file and create inspectCSharp test
following the same pattern as other language tests. This improves test
organization and maintainability by:
- Extracting sample C# code to fixtures/sample-c-sharp.ts
- Adding inspectCSharp.test.ts for tree structure inspection
- Updating parseSourceCodeDefinitions.c-sharp.test.ts to use fixture
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: add language-specific capture processing
- Add language parameter to processCaptures for language-specific handling
- Implement selective HTML filtering for jsx/tsx files
- Update all call sites to pass correct language parameter
- Fix type safety by properly passing language strings
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: move tree-sitter test samples to fixtures
Move sample code from parseSourceCodeDefinitions test files into dedicated
fixtures directory for better organization and reuse:
- Created fixtures directory for language samples
- Moved sample code for C#, C, JSON, Kotlin, PHP, Ruby, Rust, and Swift
- Created inspect test files for each language
- Updated original test files to import from fixtures
- Standardized test options across languages
- No behavior changes, all tests passing
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: standardize tree-sitter test code across languages
Standardize test code structure and naming across all language parsers:
- Rename identifiers to clearly indicate test purpose
- Ensure each code section has descriptive comments
- Group related code sections together
- Maintain language-specific naming conventions
- Keep test structure consistent across languages
- Remove tests for unsupported features
Changes made to:
- C, JSON, Kotlin, PHP, Ruby, Rust, and Swift test files
- Sample code fixtures
- Parser definition tests
- Inspect structure tests
No functional changes - all tests passing with improved maintainability.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: standardize tree-sitter parser tests across all languages
- Ensure all test structures span at least 4 lines for better parsing
- Create exactly one consolidated test per data structure type
- Use descriptive names that clearly indicate test purpose
- Improve query pattern organization and documentation
- Simplify inspect test files to focus on structure validation
- Implement result caching in test files for better performance
- Remove duplicate and skipped tests
- Follow consistent naming conventions across all languages
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* first pass all but elm-bash
* test: inspectTreeStructure helper now returns a string
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: establish single source of truth for MIN_COMPONENT_LINES
- Move minComponentLines parameter to a file-level constant
- Remove parameter passing to processCaptures function
- Update all references to use the global constant
- Remove redundant local constant declaration in parseFile
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* refactor: improve MIN_COMPONENT_LINES implementation
- Changed MIN_COMPONENT_LINES constant to use getter/setter functions
- Updated references to use getMinComponentLines() function
- Modified test helper to set value to 0 during tests
- This establishes a single source of truth while making testing easier
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: standardize tree-sitter test implementations
Standardized tree-sitter test implementations across all supported languages:
- Removed debugLog calls from test files
- Implemented consistent line number pattern matching
- Ensured proper test structure and execution order
- Added comprehensive test coverage for supported structures
- Maintained 1:1 mapping between queries, tests and samples
- Documented unsupported features in TODO sections
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* docs: document tree-sitter parseable but unsupported structures
Added TODO sections in test files documenting structures that can be parsed
by tree-sitter but currently lack query pattern support:
- Python: f-strings, complex type annotations, pattern matching
- C++: virtual methods, field initializers, base class clauses
- Java: import declarations, field declarations with modifiers
- TSX: React hooks, context providers, event handlers
Added examples and clarifying comments to existing TODO sections.
Enhanced Java query patterns for better structure capture including
lambda expressions, field declarations, and type parameters.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* fix: use correct distDir variable for WASM file copying
The copyWasmFiles plugin was using an undefined targetDir variable when copying
tree-sitter WASM files. Changed to use the correctly defined distDir variable,
fixing the build process.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
* lint: remove unused imports and variables from tree-sitter tests
Removed unused imports and variables from tree-sitter test files to fix linting errors:
- Removed unused goQuery import from helpers.ts
- Removed unused imports and mockedFs variables from language-specific test files
- Cleaned up test files to only import what they use
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
---------
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
Co-authored-by: Eric Wheeler <roo-code@z.ewheeler.org>
* feat: support environment variables reference in mcp `env` config
* tests(src/utils/config): add test for `injectEnv`
* fix(injectEnv): use `env == null` and `??` check instead of `!env`, `||`
* refactor: remove unnecessary type declare
* chore!: simplify regexp, remove replacement for env vars with dots
This commit significantly enhances the Python tree-sitter parser to support a comprehensive range of Python language constructs, enabling more accurate and detailed code analysis.
Key improvements:
- Added support for method definitions (instance, class, and static methods)
- Added support for decorators on functions and classes
- Added support for module-level variables and constants
- Added support for async functions and methods
- Added support for property getters/setters
- Added support for type annotations in various contexts
- Added support for dataclasses
- Added support for nested functions and classes
- Added support for generator functions
- Added support for list/dict/set comprehensions
- Added support for lambda functions
- Added support for abstract base classes and methods
The parser now handles Python's rich feature set more comprehensively, including special Python patterns like decorators, type annotations, and various comprehension types. This enables better code navigation, understanding, and analysis for Python codebases.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>
This enhancement significantly expands the Java parser's capabilities to recognize and parse a wide range of Java language constructs:
- Added support for enum declarations and enum constants
- Added support for annotation type declarations and elements
- Added support for field declarations
- Added support for constructor declarations
- Added support for lambda expressions
- Added support for inner and anonymous classes
- Added support for type parameters (generics)
- Added support for package and import declarations
These improvements enable more comprehensive code analysis for Java projects, providing better definition extraction and navigation capabilities.
Signed-off-by: Eric Wheeler <roo-code@z.ewheeler.org>