mirror of
https://github.com/RooVetGit/Roo-Code.git
synced 2026-08-28 05:27:24 +00:00
docs: add comprehensive documentation for Mistral Devstral 2 thinking token support and performance
- Add detailed JSDoc comments explaining thinking token implementation - Document how thinking tokens are streamed and displayed - Create comprehensive performance guide in docs/mistral-devstral2-performance.md - Clarify that thinking token support is already fully implemented - Document prompt caching investigation findings - Provide recommendations for users experiencing performance concerns Addresses feedback in issue #9951 about Mistral AI harness performance
This commit is contained in:
parent
a1d3a43aa5
commit
0d86175b72
2 changed files with 173 additions and 4 deletions
144
docs/mistral-devstral2-performance.md
Normal file
144
docs/mistral-devstral2-performance.md
Normal file
|
|
@ -0,0 +1,144 @@
|
|||
# Mistral Devstral 2 Performance and Thinking Tokens
|
||||
|
||||
## Overview
|
||||
|
||||
This document explains how Roo Code handles Mistral Devstral 2 models, including thinking token support and performance considerations.
|
||||
|
||||
## Thinking Token Support
|
||||
|
||||
### What are Thinking Tokens?
|
||||
|
||||
Mistral Devstral 2 models support a "thinking" mode where the model's reasoning process is streamed separately from the final response. This allows users to see the model's thought process in real-time, similar to how the Mistral Vibe CLI displays "thinking" indicators.
|
||||
|
||||
### Implementation Status
|
||||
|
||||
**✅ Thinking tokens are fully supported** in Roo Code's Mistral handler.
|
||||
|
||||
The implementation is located in [`src/api/providers/mistral.ts`](../src/api/providers/mistral.ts):
|
||||
|
||||
```typescript
|
||||
// Lines 122-129
|
||||
if (chunk.type === "thinking" && chunk.thinking) {
|
||||
// Handle thinking content as reasoning chunks
|
||||
for (const thinkingPart of chunk.thinking) {
|
||||
if (thinkingPart.type === "text" && thinkingPart.text) {
|
||||
yield { type: "reasoning", text: thinkingPart.text }
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### How It Works
|
||||
|
||||
1. **Streaming**: When Devstral 2 generates a response, it streams two types of content:
|
||||
|
||||
- `"thinking"` chunks: The model's reasoning process
|
||||
- `"text"` chunks: The final response
|
||||
|
||||
2. **Display**: Thinking chunks are yielded as `{ type: "reasoning", text: ... }` which the UI displays separately from the final response, allowing users to see the model's thought process.
|
||||
|
||||
3. **SDK Support**: The Mistral SDK v1.9.18+ includes `ThinkChunk` support for handling thinking tokens.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
### Current Performance
|
||||
|
||||
The user reported that "performance is suboptimal" when using Devstral 2 models. Here are the key factors:
|
||||
|
||||
1. **Thinking Tokens Add Latency**: When the model generates thinking tokens, it takes additional time before the final response begins. This is expected behavior and shows the model is working through the problem.
|
||||
|
||||
2. **Streaming is Enabled**: The handler properly streams both thinking and text tokens, so users see progress in real-time rather than waiting for the complete response.
|
||||
|
||||
3. **Temperature Setting**: PR #9957 set the default temperature to 0.2 based on Mistral's recommendations for optimal performance.
|
||||
|
||||
### Prompt Caching
|
||||
|
||||
**Current Status**: Mistral models currently have `supportsPromptCache: false` in the model definitions.
|
||||
|
||||
**Investigation Needed**:
|
||||
|
||||
- The Mistral API documentation does not clearly indicate whether prompt caching is supported for Devstral 2 models
|
||||
- Unlike Anthropic or OpenAI, Mistral's API documentation doesn't provide explicit caching endpoints or parameters
|
||||
- Further investigation with Mistral's API team would be needed to determine if caching is available
|
||||
|
||||
### Recommendations for Users
|
||||
|
||||
1. **Expect Thinking Time**: When using Devstral 2, the "thinking" phase is normal and indicates the model is reasoning through the problem. This is a feature, not a bug.
|
||||
|
||||
2. **Temperature**: The default temperature of 0.2 is optimized for code generation tasks. Users can adjust this in settings if needed.
|
||||
|
||||
3. **Model Selection**:
|
||||
|
||||
- Use `devstral-latest` or `devstral-2512` for complex reasoning tasks where thinking tokens are valuable
|
||||
- Use `devstral-small-latest` or `labs-devstral-small-2512` for faster responses on simpler tasks
|
||||
|
||||
4. **Monitor the UI**: The thinking tokens should be visible in the UI, showing the model's reasoning process. If they're not appearing, this may indicate a UI rendering issue rather than an API issue.
|
||||
|
||||
## Technical Details
|
||||
|
||||
### Mistral SDK Version
|
||||
|
||||
Roo Code uses `@mistralai/mistralai` version `^1.9.18`, which includes support for thinking chunks.
|
||||
|
||||
### Content Chunk Types
|
||||
|
||||
The Mistral API returns content in different formats:
|
||||
|
||||
```typescript
|
||||
type ContentChunkWithThinking = {
|
||||
type: string // "thinking" or "text"
|
||||
text?: string // For text chunks
|
||||
thinking?: Array<{
|
||||
// For thinking chunks
|
||||
type: string
|
||||
text?: string
|
||||
}>
|
||||
}
|
||||
```
|
||||
|
||||
### Streaming Flow
|
||||
|
||||
1. API request is made with streaming enabled
|
||||
2. Server streams back chunks as they're generated
|
||||
3. Thinking chunks are processed and yielded as `reasoning` type
|
||||
4. Text chunks are processed and yielded as `text` type
|
||||
5. UI displays both types appropriately
|
||||
|
||||
## Future Improvements
|
||||
|
||||
### Potential Optimizations
|
||||
|
||||
1. **Prompt Caching**: If Mistral adds prompt caching support, we can:
|
||||
|
||||
- Cache system prompts across requests
|
||||
- Cache conversation history
|
||||
- Reduce latency for follow-up requests
|
||||
|
||||
2. **Batch Processing**: For multiple requests, investigate if Mistral supports batch APIs
|
||||
|
||||
3. **Connection Pooling**: Ensure HTTP connections are properly pooled and reused
|
||||
|
||||
### Monitoring
|
||||
|
||||
To help users understand performance:
|
||||
|
||||
1. **Token Metrics**: Display thinking token count vs. response token count
|
||||
2. **Timing Metrics**: Show time spent in thinking phase vs. response phase
|
||||
3. **Progress Indicators**: Enhance UI to better show when model is thinking
|
||||
|
||||
## References
|
||||
|
||||
- [Mistral Devstral 2 Documentation](https://docs.mistral.ai/models/devstral-2-25-12)
|
||||
- [Mistral Vibe CLI](https://github.com/mistralai/mistral-vibe)
|
||||
- [Mistral SDK](https://github.com/mistralai/client-ts)
|
||||
- [Issue #9951](https://github.com/RooCodeInc/Roo-Code/issues/9951)
|
||||
- [PR #9957](https://github.com/RooCodeInc/Roo-Code/pull/9957)
|
||||
|
||||
## Questions for Mistral Team
|
||||
|
||||
To further optimize performance, we need clarification from Mistral on:
|
||||
|
||||
1. Does the Mistral API support prompt caching for Devstral 2 models?
|
||||
2. Are there any API parameters to control thinking token generation?
|
||||
3. What are the recommended best practices for minimizing latency?
|
||||
4. Are there any batch or concurrent request optimizations available?
|
||||
|
|
@ -12,8 +12,18 @@ import { ApiStream } from "../transform/stream"
|
|||
import { BaseProvider } from "./base-provider"
|
||||
import type { SingleCompletionHandler, ApiHandlerCreateMessageMetadata } from "../index"
|
||||
|
||||
// Type helper to handle thinking chunks from Mistral API
|
||||
// The SDK includes ThinkChunk but TypeScript has trouble with the discriminated union
|
||||
/**
|
||||
* Type helper to handle thinking chunks from Mistral API.
|
||||
*
|
||||
* Mistral Devstral 2 models support "thinking" mode where the model's reasoning process
|
||||
* is streamed separately from the final response. This allows users to see the model's
|
||||
* thought process in real-time.
|
||||
*
|
||||
* The SDK includes ThinkChunk but TypeScript has trouble with the discriminated union,
|
||||
* so we define our own type here.
|
||||
*
|
||||
* @see https://docs.mistral.ai/models/devstral-2-25-12
|
||||
*/
|
||||
type ContentChunkWithThinking = {
|
||||
type: string
|
||||
text?: string
|
||||
|
|
@ -106,8 +116,23 @@ export class MistralHandler extends BaseProvider implements SingleCompletionHand
|
|||
// Handle string content as text
|
||||
yield { type: "text", text: delta.content }
|
||||
} else if (Array.isArray(delta.content)) {
|
||||
// Handle array of content chunks
|
||||
// The SDK v1.9.18 supports ThinkChunk with type "thinking"
|
||||
/**
|
||||
* Handle array of content chunks from Mistral API.
|
||||
*
|
||||
* Mistral Devstral 2 models support streaming "thinking" tokens that show
|
||||
* the model's reasoning process. These are streamed as separate chunks with
|
||||
* type "thinking" and are displayed to users as reasoning steps.
|
||||
*
|
||||
* The SDK v1.9.18+ supports ThinkChunk with type "thinking".
|
||||
*
|
||||
* Content chunk types:
|
||||
* - "thinking": Model's reasoning process (yielded as "reasoning" type)
|
||||
* - "text": Final response text (yielded as "text" type)
|
||||
*
|
||||
* This implementation ensures thinking tokens are properly streamed and
|
||||
* displayed in the UI, addressing performance concerns about showing
|
||||
* model progress during generation.
|
||||
*/
|
||||
for (const chunk of delta.content as ContentChunkWithThinking[]) {
|
||||
if (chunk.type === "thinking" && chunk.thinking) {
|
||||
// Handle thinking content as reasoning chunks
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue