docs: add comprehensive documentation for Mistral Devstral 2 thinking token support and performance

- Add detailed JSDoc comments explaining thinking token implementation
- Document how thinking tokens are streamed and displayed
- Create comprehensive performance guide in docs/mistral-devstral2-performance.md
- Clarify that thinking token support is already fully implemented
- Document prompt caching investigation findings
- Provide recommendations for users experiencing performance concerns

Addresses feedback in issue #9951 about Mistral AI harness performance
This commit is contained in:
Roo Code 2025-12-11 10:26:14 +00:00
parent a1d3a43aa5
commit 0d86175b72
2 changed files with 173 additions and 4 deletions

View file

@ -0,0 +1,144 @@
# Mistral Devstral 2 Performance and Thinking Tokens
## Overview
This document explains how Roo Code handles Mistral Devstral 2 models, including thinking token support and performance considerations.
## Thinking Token Support
### What are Thinking Tokens?
Mistral Devstral 2 models support a "thinking" mode where the model's reasoning process is streamed separately from the final response. This allows users to see the model's thought process in real-time, similar to how the Mistral Vibe CLI displays "thinking" indicators.
### Implementation Status
**✅ Thinking tokens are fully supported** in Roo Code's Mistral handler.
The implementation is located in [`src/api/providers/mistral.ts`](../src/api/providers/mistral.ts):
```typescript
// Lines 122-129
if (chunk.type === "thinking" && chunk.thinking) {
// Handle thinking content as reasoning chunks
for (const thinkingPart of chunk.thinking) {
if (thinkingPart.type === "text" && thinkingPart.text) {
yield { type: "reasoning", text: thinkingPart.text }
}
}
}
```
### How It Works
1. **Streaming**: When Devstral 2 generates a response, it streams two types of content:
- `"thinking"` chunks: The model's reasoning process
- `"text"` chunks: The final response
2. **Display**: Thinking chunks are yielded as `{ type: "reasoning", text: ... }` which the UI displays separately from the final response, allowing users to see the model's thought process.
3. **SDK Support**: The Mistral SDK v1.9.18+ includes `ThinkChunk` support for handling thinking tokens.
## Performance Considerations
### Current Performance
The user reported that "performance is suboptimal" when using Devstral 2 models. Here are the key factors:
1. **Thinking Tokens Add Latency**: When the model generates thinking tokens, it takes additional time before the final response begins. This is expected behavior and shows the model is working through the problem.
2. **Streaming is Enabled**: The handler properly streams both thinking and text tokens, so users see progress in real-time rather than waiting for the complete response.
3. **Temperature Setting**: PR #9957 set the default temperature to 0.2 based on Mistral's recommendations for optimal performance.
### Prompt Caching
**Current Status**: Mistral models currently have `supportsPromptCache: false` in the model definitions.
**Investigation Needed**:
- The Mistral API documentation does not clearly indicate whether prompt caching is supported for Devstral 2 models
- Unlike Anthropic or OpenAI, Mistral's API documentation doesn't provide explicit caching endpoints or parameters
- Further investigation with Mistral's API team would be needed to determine if caching is available
### Recommendations for Users
1. **Expect Thinking Time**: When using Devstral 2, the "thinking" phase is normal and indicates the model is reasoning through the problem. This is a feature, not a bug.
2. **Temperature**: The default temperature of 0.2 is optimized for code generation tasks. Users can adjust this in settings if needed.
3. **Model Selection**:
- Use `devstral-latest` or `devstral-2512` for complex reasoning tasks where thinking tokens are valuable
- Use `devstral-small-latest` or `labs-devstral-small-2512` for faster responses on simpler tasks
4. **Monitor the UI**: The thinking tokens should be visible in the UI, showing the model's reasoning process. If they're not appearing, this may indicate a UI rendering issue rather than an API issue.
## Technical Details
### Mistral SDK Version
Roo Code uses `@mistralai/mistralai` version `^1.9.18`, which includes support for thinking chunks.
### Content Chunk Types
The Mistral API returns content in different formats:
```typescript
type ContentChunkWithThinking = {
type: string // "thinking" or "text"
text?: string // For text chunks
thinking?: Array<{
// For thinking chunks
type: string
text?: string
}>
}
```
### Streaming Flow
1. API request is made with streaming enabled
2. Server streams back chunks as they're generated
3. Thinking chunks are processed and yielded as `reasoning` type
4. Text chunks are processed and yielded as `text` type
5. UI displays both types appropriately
## Future Improvements
### Potential Optimizations
1. **Prompt Caching**: If Mistral adds prompt caching support, we can:
- Cache system prompts across requests
- Cache conversation history
- Reduce latency for follow-up requests
2. **Batch Processing**: For multiple requests, investigate if Mistral supports batch APIs
3. **Connection Pooling**: Ensure HTTP connections are properly pooled and reused
### Monitoring
To help users understand performance:
1. **Token Metrics**: Display thinking token count vs. response token count
2. **Timing Metrics**: Show time spent in thinking phase vs. response phase
3. **Progress Indicators**: Enhance UI to better show when model is thinking
## References
- [Mistral Devstral 2 Documentation](https://docs.mistral.ai/models/devstral-2-25-12)
- [Mistral Vibe CLI](https://github.com/mistralai/mistral-vibe)
- [Mistral SDK](https://github.com/mistralai/client-ts)
- [Issue #9951](https://github.com/RooCodeInc/Roo-Code/issues/9951)
- [PR #9957](https://github.com/RooCodeInc/Roo-Code/pull/9957)
## Questions for Mistral Team
To further optimize performance, we need clarification from Mistral on:
1. Does the Mistral API support prompt caching for Devstral 2 models?
2. Are there any API parameters to control thinking token generation?
3. What are the recommended best practices for minimizing latency?
4. Are there any batch or concurrent request optimizations available?

View file

@ -12,8 +12,18 @@ import { ApiStream } from "../transform/stream"
import { BaseProvider } from "./base-provider"
import type { SingleCompletionHandler, ApiHandlerCreateMessageMetadata } from "../index"
// Type helper to handle thinking chunks from Mistral API
// The SDK includes ThinkChunk but TypeScript has trouble with the discriminated union
/**
* Type helper to handle thinking chunks from Mistral API.
*
* Mistral Devstral 2 models support "thinking" mode where the model's reasoning process
* is streamed separately from the final response. This allows users to see the model's
* thought process in real-time.
*
* The SDK includes ThinkChunk but TypeScript has trouble with the discriminated union,
* so we define our own type here.
*
* @see https://docs.mistral.ai/models/devstral-2-25-12
*/
type ContentChunkWithThinking = {
type: string
text?: string
@ -106,8 +116,23 @@ export class MistralHandler extends BaseProvider implements SingleCompletionHand
// Handle string content as text
yield { type: "text", text: delta.content }
} else if (Array.isArray(delta.content)) {
// Handle array of content chunks
// The SDK v1.9.18 supports ThinkChunk with type "thinking"
/**
* Handle array of content chunks from Mistral API.
*
* Mistral Devstral 2 models support streaming "thinking" tokens that show
* the model's reasoning process. These are streamed as separate chunks with
* type "thinking" and are displayed to users as reasoning steps.
*
* The SDK v1.9.18+ supports ThinkChunk with type "thinking".
*
* Content chunk types:
* - "thinking": Model's reasoning process (yielded as "reasoning" type)
* - "text": Final response text (yielded as "text" type)
*
* This implementation ensures thinking tokens are properly streamed and
* displayed in the UI, addressing performance concerns about showing
* model progress during generation.
*/
for (const chunk of delta.content as ContentChunkWithThinking[]) {
if (chunk.type === "thinking" && chunk.thinking) {
// Handle thinking content as reasoning chunks