* fix: add embedder validation to prevent misleading status indicators (#4398)
* fix: address PR feedback and fix critical issues
- Fixed settings-save flow to save before validation
- Fixed Error constructor usage in scanner.ts
- Fixed segment identification in file-watcher.ts
- Added missing translation keys for embedder validation errors
* fix: add missing Ollama translation keys
- Added missing ollama.title, description, and settings keys
- Fixed translation check failure in CI/CD pipeline
- Synchronized all 17 non-English locale files
* feat: add proactive embedder validation on provider switch
- Validate embedder connection when switching providers
- Prevent misleading 'Indexed' status when embedder is unavailable
- Show immediate error feedback for invalid configurations
- Add comprehensive test coverage for validation flow
This ensures users get immediate feedback when configuring embedders,
preventing confusion when providers like Ollama are not accessible.
* fix: improve error handling and validation in code indexing process
* refactor: extract common embedder validation and error handling logic
- Created shared/validation-helpers.ts with centralized error handling utilities
- Refactored OpenAI, OpenAI-Compatible, and Ollama embedders to use shared helpers
- Eliminated duplicate error handling code across embedders
- Improved maintainability and consistency of error handling
- Fixed test compatibility in manager.spec.ts
- All 2721 tests passing
* refactor: simplify validation helpers by removing unnecessary wrapper functions
- Removed getErrorMessageForConnectionError and inlined logic into handleValidationError
- Removed isRateLimitError, logRateLimitRetry, and logEmbeddingError wrapper functions
- Updated openai.ts and openai-compatible.ts to inline rate limit checking and logging
- Reduced code complexity while maintaining all functionality
- All 311 tests continue to pass
* fix: add missing invalidResponse i18n key and fix French translation
- Added missing 'invalidResponse' key to all locale files
- Fixed French translation: changed 'and accessible' to 'et accessible'
- Ensures proper error messages are displayed when embedder returns invalid responses
* fix: restore removed score settings in webviewMessageHandler
- Restored codebaseIndexSearchMaxResults and codebaseIndexSearchMinScore settings that were unintentionally removed
- Keep embedder validation related changes
* fix: revert unintended changes to file-watcher and scanner
- Reverted point ID generation back to using line numbers instead of segmentHash
- Restored { cause: deleteError } parameter in scanner error handling
- These changes were unrelated to the embedder validation feature
---------
Co-authored-by: Daniel Riccio <ricciodaniel98@gmail.com>
* feat: add markdown support to codebase indexing (#4660)
* fix: implement chunking for large markdown sections and fix Qdrant deduplication issue (#4660)
- Modified parseMarkdownContent to chunk large sections (>1150 chars)
- Added support for chunking header-less markdown files
- Fixed _chunkTextByLines to handle oversized lines properly
- Added defensive check for parseMarkdown returning undefined
- Fixed Qdrant ID generation to use segmentHash instead of file:line
- This was the root cause: chunks were being deduplicated
- Each chunk now gets a unique ID even from the same line
- Added comprehensive tests for all edge cases
- Ensures all markdown content is properly indexed in Qdrant
* fix: remove redundant supported-extensions test file
As identified in PR review, the supported-extensions.spec.ts file only tests
the contents of an array, which is already implicitly covered by the functional
tests in parser.spec.ts. Removing this reduces maintenance overhead without
sacrificing test quality.
* test: address PR review feedback
- Remove redundant test 'should handle large markdown documentation folders efficiently'
that only verified the scanner could iterate over mocked files
- Add test to verify unique point IDs are generated for each block from the same file,
ensuring the segmentHash-based ID generation prevents collisions
* fix: add segmentHash to vector point payload and refactor tests
- Add segmentHash to payload in scanner.ts to fix vector point ID generation
- Split parser.spec.ts tests into focused unit tests (mocked dependencies)
- Move integration tests to new markdownIntegration.spec.ts file
- Each test suite now has clear, distinct responsibilities
- Fixes issue #4660: vector point ID collisions for large Markdown files
* refactor: move redundant tests
* feat: enhance markdown processing with consistent chunking logic and segment hashing
---------
Co-authored-by: Daniel Riccio <ricciodaniel98@gmail.com>
* feat: apply changes from local main
* fix: add missing types
* feat: deduplicate code blocks coming out of parser
* feat: implement a cache manager to improve cache handling
* refactor: move code index service initialization to extension and remove await from indexing process
* fix: return undefined instead of throwing if no workspace is detected
* feat: allow auto approve if it is active for read tools
* refactor: improve UI of the results and allow opening the ranges directly in the editor
* refactor: use dependency injection to improve performance
* feat: implement result filtering by directory path
* refactor: centralize path normalization logic
* refactor: remove unnecessary barrel file
* refactor: prevent restarting the service if no settings change
* fix: the indexing process should never be awaited
* refactor: cleanup unused method
* refactor: remove batch limits for ollama
* refactor(parser): simplify method signatures and improve chunking logic
- Remove redundant min/max chars parameters
- Add better handling for oversized lines
- Improve chunking logic with segment handling
- Clean up method signatures and parameter ordering
* fix(settings): make select inputs full width in CodeIndexSettings
* refactor: increase max list file limit
* feat(ui): improve codebase search result display formatting
* test: add tests for cache and config managers
* test: create unit tests for parser and scanner
* feat(parser): improve segment hash uniqueness
- Added startCharIndex to segment hash calculation in _chunkTextByLines
- Track character position when splitting oversized lines
- Ensures unique identification of segments from same line
* feat(file-watcher): add error logging and optional ignoreController injection
* fix: allow getting the state if the service is disabled
* fix: set the embedding models when cline provider is initialized
* feat: use zod to validate form
* feat(file-watcher): enhance file watcher for batched deletions and improved vector store interactions
Improve file watcher to handle file deletions in batches and optimize vector store operations.
* feat(CodeIndexSettings): move OpenAI key input to a conditional rendering block
* feat(CodeIndexSettings): update button visibility based on indexing status
* feat(file-watcher): refactor vscode mock and enhance file watcher tests
* fix(CodeIndexManager): do not await startIndexing on configuration changes
* feat(types): add codeIndexOpenAiKey and codeIndexQdrantApiKey to ProviderSettings and IpcMessage
* feat(FileWatcher): enhance file processing with batch operations and new status handling
* fix(webviewMessageHandler): handle errors during CodeIndexManager initialization
* refactor(CodeIndexManager): streamline service creation by consolidating into a single method
* feat(CodeIndex): implement minimum search score configuration and update search methods
* refactor(CodeIndexSettings): replace ApiConfiguration with ProviderSettings and update related methods
* refactor: move contants to centralized file
* refactor(constants): rename CODEBASE_INDEX_SEARCH_MIN_SCORE to SEARCH_MIN_SCORE
* feat(QdrantVectorStore): enhance search functionality with new query structure and indexing
* feat(FileWatcher): implement batch processing and retry logic for upserting points
* fix(CodeIndexSettings): rename setProviderSettingsField to setApiConfigurationField and move model label
* fix(ChatRow): remove limit from search query messages
* refactor(CodebaseSearchResult): remove unused props from component
* feat: implement batch processing for file events in FileWatcher
- Introduced a new mechanism to accumulate file events (create, change, delete) and process them in batches.
- Added debounce functionality to optimize processing frequency.
- Emitted events for batch processing start, progress updates, and completion with detailed summaries.
- Refactored existing processing logic to handle batch deletions and upserts efficiently.
- Enhanced error handling and logging for better traceability during batch operations.
* feat(CodeIndex): implement batch processing and update progress reporting
* fix: define a default url for qdrant
* feat(CodeIndexManager): add initialization check and update startIndexing logic
* feat(CodeIndexSettings): validate Qdrant URL and update settings commitment logic
* feat: refactor progress calculation and update progress bar rendering
* refactor: remove webview provider and related methods
* fix: simplify indexing status update by directly using update values
* feat: integrate .gitignore support into file processing and scanning logic
* fix: update clearCacheFile method to write an empty object instead of deleting the cache file
* Revert this
* Run prettier
* fix: add new dependencies for qdrant client and directory scanner
* feat: add codebase search functionality to localization files
* feat: add localization strings for codebase indexing settings
* feat: integrate CodeIndexSettings into ExperimentalSettings and update settings localization
* refactor: remove console logs from various components for cleaner output
* feat: enhance capabilities section and codebase search tool description
* feat: add code indexing localization for multiple languages
* fix: correct indentation for CodeIndexSettings component in ExperimentalSettings
* refactor: update unit tests to properly test current functionality
* feat: add mock implementation for p-limit and update Jest config
* feat: track file creation, change, and deletion events in accumulatedEvents
* refactor: simplify file watcher tests by removing waitForFileProcessingToFinish and using direct event accumulation
* refactor: mock ContextProxy's getValue method to return current config name in ClineProvider tests
* refactor: mock missing properties required by codebase indexing manager
---------
Co-authored-by: cte <cestreich@gmail.com>