* fix: race conditions in subtask delegation system
Comprehensive fix for race conditions and error handling gaps in the subtask
delegation system. Addresses multiple failure modes that could leave parent
tasks permanently stuck in 'delegated' status, causing nested subtasks to hang.
Key fixes:
- Remove initialStatus from taskMetadata rebuild (eliminates status overwrites)
- Persist delegation metadata to per-task files (resolves globalState eviction)
- Add delegationInProgress mutex guard (prevents concurrent delegation ops)
- TOCTOU race fixes with fresh re-reads before writes
- Abort-aware pWaitFor predicate (prevents false 60s timeout on user input)
- Remove silent .catch(() => {}) — all errors now logged unless task is aborting
- Single-attempt delegation with parent repair on failure (no retry band-aids)
- Cancel debouncedEmitTokenUsage in dispose() (prevents zombie callbacks)
- new_task isolation truncation for parallel tool calls
* fix: write all 6 delegation fields in every saveDelegationMeta call site
* fix: align delegation tests with single-attempt implementation (no retry)
* feat: migrate LiteLLM provider to AI SDK (@ai-sdk/openai-compatible)
- Replace raw OpenAI SDK (RouterProvider) with Vercel AI SDK's
createOpenAICompatible via OpenAICompatibleHandler base class
- Retain dynamic model fetching from LiteLLM server via /v1/model/info
- Use centralized getModelMaxOutputTokens() to cap output tokens at 20%
of context window, preventing overflow errors
- Remove LiteLLM-specific workarounds (Gemini thought signature injection,
prompt cache control headers) now handled by the proxy or AI SDK
- Rewrite tests to mock AI SDK (streamText, generateText) instead of
raw OpenAI SDK
* fix: call fetchModel() in completePrompt() before execution
Addresses review feedback - completePrompt() now fetches models
before executing to ensure correct model info for token limits,
matching the behavior of createMessage().
* fix: make removeClineFromStack() delegation-aware to prevent orphaned parent tasks
When a delegated child task is removed via removeClineFromStack() (e.g., Clear
Task, navigate to history, start new task), the parent task was left orphaned
in "delegated" status with a stale awaitingChildId. This made the parent
unresumable without manual history repair.
This fix captures parentTaskId and childTaskId before abort/dispose, then
repairs the parent metadata (status -> active, clear awaitingChildId) when
the popped task is a delegated child and awaitingChildId matches.
Parent lookup + updateTaskHistory are wrapped in try/catch so failures are
non-fatal (logged but do not block the pop).
Closes#11301
* fix: add skipDelegationRepair opt-out to removeClineFromStack() for nested delegation
---------
Co-authored-by: Roo Code <roomote@roocode.com>
* refactor: migrate LM Studio provider to Vercel AI SDK
Migrate LmStudioHandler from raw OpenAI SDK to Vercel AI SDK via
OpenAICompatibleHandler base class.
Changes:
- Extend OpenAICompatibleHandler instead of BaseProvider
- Use createOpenAICompatible from @ai-sdk/openai-compatible
- Use streamText/generateText from ai package
- Add extractReasoningMiddleware for <think> tag extraction parity
- Pass draft_model via providerOptions for speculative decoding
- Remove unused getLmStudioModels function (active version in fetchers/)
- Update all tests to mock AI SDK instead of OpenAI SDK
* fix: wrap completePrompt with handleAiSdkError for consistent error handling
---------
Co-authored-by: Roo Code <roomote@roocode.com>
Replace @anthropic-ai/vertex-sdk with @ai-sdk/google-vertex/anthropic,
using streamText/generateText from the Vercel AI SDK for consistent
provider behavior.
Changes:
- Use createVertexAnthropic from @ai-sdk/google-vertex/anthropic
- Use streamText/generateText instead of direct Anthropic API calls
- Add AI SDK transform utilities for message/tool conversion
- Handle cache control via AI SDK providerOptions
- Handle thinking/reasoning via providerOptions.anthropic.thinking
- Add thought signature and redacted thinking block tracking
- Set isAiSdkProvider() to return true
- Remove unused deps: @anthropic-ai/vertex-sdk, google-auth-library
- Rewrite tests to mock AI SDK instead of @anthropic-ai/vertex-sdk
* Latest main branch snapshot from API
* feat: add dedicated Azure OpenAI provider using @ai-sdk/azure package
* feat: add Azure provider UI component and translations
* feat: add Azure provider translations for all locales
* chore: add missing Azure placeholder translations
* Delete .changeset/azure-ai-sdk-migration.md
* fix: add Azure provider validation for onboarding workflow
- Add azureApiKey to SECRET_STATE_KEYS for proper configuration detection
- Add Azure validation case in validateModelsAndKeysProvided
- Add validation translations for azureResourceName and azureDeploymentName across all 18 locales
This fixes the issue where the Finish button does nothing when setting up Azure provider in the onboarding workflow.
* feat(azure): add model metadata, model picker, rename to Azure AI Foundry
- Add static model metadata for 29 Azure models (from models.dev)
with Roo-specific flags (reasoning, tools, verbosity) matching
openAiNativeModels
- Add model picker dropdown to Azure provider settings for model
capability detection (context window, max tokens, pricing)
- Rename provider label from 'Azure OpenAI' to 'Azure AI Foundry'
across all 18 locales
- Make API key optional (supports Azure managed identity / Entra ID)
- Update default API version from 2024-08-01-preview to 2025-04-01-preview
- Fix maxOutputTokens validation (filter invalid values <= 0)
- Handler separates deployment name (API calls) from model ID
(capability lookup) with azureDefaultModelInfo (gpt-4o) fallback
- Remove unhelpful 'Get Azure AI Foundry Access' button
- Prevent stale model IDs from other providers carrying over
- Suppress validation errors on fresh provider selection
* fix(azure): add missing isAiSdkProvider() override for reasoning block preservation
* Azure Fixes for Hannes
* Quick Fix for Respones API Only (for Hannes)
* fix: use explicit azureOpenAiDefaultApiVersion fallback when apiVersion is empty
Addresses review feedback: the UI placeholder shows '2025-04-01-preview' via
azureOpenAiDefaultApiVersion, so the handler should use the same constant as
fallback instead of silently deferring to the SDK's internal default.
* fix: remove stale Cerebras references (retired provider)
* fix: add missing retiredProviderMessage translations for all locales
* fix: do not map promptCacheMissTokens to cacheWriteTokens for Azure
Azure uses OpenAI-compatible caching which does not report cache write
tokens separately. promptCacheMissTokens represents tokens NOT found in
cache (processed from scratch), not tokens written to cache. This aligns
the Azure handler with the OpenAI native handler behavior.
---------
Co-authored-by: Hannes Rudolph <hrudolph@gmail.com>
Co-authored-by: Roo Code <roomote@roocode.com>
Co-authored-by: daniel-lxs <ricciodaniel98@gmail.com>
Co-authored-by: Matt Rubens <mrubens@users.noreply.github.com>
* feat: migrate OpenAI Native provider to @ai-sdk/openai
Replace the raw OpenAI SDK (openai) usage in OpenAiNativeHandler with
@ai-sdk/openai and AI SDK's streamText/generateText, following the same
pattern used by other migrated providers (Groq, xAI, Fireworks, etc.).
Key changes:
- Use createOpenAI from @ai-sdk/openai with provider.responses() for
the Responses API
- Use streamText/generateText from ai for streaming and completions
- Pass OpenAI-specific features via providerOptions.openai (store,
reasoningEffort, reasoningSummary, textVerbosity, serviceTier,
promptCacheRetention, parallelToolCalls, include)
- Capture responseId, serviceTier, and encrypted reasoning content
from providerMetadata after streaming
- Preserve getEncryptedContent() and getResponseId() for Task.ts
- Preserve service tier pricing adjustment in cost calculation
- Mark as isAiSdkProvider: true
- Eliminate ~1100 lines of manual SSE parsing, raw fetch fallback,
and event handling code
- Rewrite all 3 test files to use AI SDK mocking pattern
* fix: remove non-existent cacheWriteTokens from providerMetadata
The OpenAI Responses API does not report cache write tokens separately.
Remove the reference to providerMetadata?.openai?.cacheWriteTokens which
does not exist in the @ai-sdk/openai provider metadata schema.
* fix: filter standalone encrypted reasoning items from messages
Task.ts buildCleanConversationHistory injects standalone reasoning items
with { type: 'reasoning', encrypted_content: '...' } into the messages
array. These have no 'role' property and would be silently dropped by
convertToAiSdkMessages. Filter them explicitly to prevent confusion.
Note: Encrypted reasoning content round-tripping for stateless continuity
is a known limitation of the AI SDK migration. The @ai-sdk/openai
provider does not support injecting raw Responses API reasoning items.
Plain-text reasoning round-tripping works correctly via isAiSdkProvider().
* fix: restore reasoning round-trip for OpenAI Responses API via AI SDK
- Strip plain-text reasoning blocks from assistant messages before
convertToAiSdkMessages() to eliminate 'Non-OpenAI reasoning parts'
warnings from @ai-sdk/openai Responses provider
- Re-inject encrypted reasoning items as AI SDK reasoning parts with
providerOptions.openai.itemId and reasoningEncryptedContent, restoring
reasoning continuity that was silently broken after the migration
- Restructure createMessage() into a 5-step pipeline:
collect → filter → strip → convert → inject
- Add 21 new tests for both plain-text stripping and encrypted
reasoning injection
---------
Co-authored-by: Hannes Rudolph <hrudolph@gmail.com>
* fix: resolve race condition in new_task delegation that loses parent task history
When delegateParentAndOpenChild creates a child task via createTask(), the
Task constructor fires startTask() as a fire-and-forget async call. The child
immediately begins its task loop and eventually calls saveClineMessages() →
updateTaskHistory(), which reads globalState, modifies it, and writes back.
Meanwhile, delegateParentAndOpenChild persists the parent's delegation
metadata (status: 'delegated', delegatedToId, awaitingChildId, childIds) via
a separate updateTaskHistory() call AFTER createTask() returns.
These two concurrent read-modify-write operations on globalState race: the
last writer wins, overwriting the other's changes. When the child's write
lands last, the parent's delegation fields are lost, making the parent task
unresumable when the child finishes.
Fix: create the child task with startTask: false, persist the parent's
delegation metadata first, then manually call child.start(). This ensures
the parent metadata is safely in globalState before the child begins writing.
* docs: clarify Task.start() only handles new tasks, not history resume