- Add detailed JSDoc comments explaining thinking token implementation - Document how thinking tokens are streamed and displayed - Create comprehensive performance guide in docs/mistral-devstral2-performance.md - Clarify that thinking token support is already fully implemented - Document prompt caching investigation findings - Provide recommendations for users experiencing performance concerns Addresses feedback in issue #9951 about Mistral AI harness performance
5.6 KiB
Mistral Devstral 2 Performance and Thinking Tokens
Overview
This document explains how Roo Code handles Mistral Devstral 2 models, including thinking token support and performance considerations.
Thinking Token Support
What are Thinking Tokens?
Mistral Devstral 2 models support a "thinking" mode where the model's reasoning process is streamed separately from the final response. This allows users to see the model's thought process in real-time, similar to how the Mistral Vibe CLI displays "thinking" indicators.
Implementation Status
✅ Thinking tokens are fully supported in Roo Code's Mistral handler.
The implementation is located in src/api/providers/mistral.ts:
// Lines 122-129
if (chunk.type === "thinking" && chunk.thinking) {
// Handle thinking content as reasoning chunks
for (const thinkingPart of chunk.thinking) {
if (thinkingPart.type === "text" && thinkingPart.text) {
yield { type: "reasoning", text: thinkingPart.text }
}
}
}
How It Works
-
Streaming: When Devstral 2 generates a response, it streams two types of content:
"thinking"chunks: The model's reasoning process"text"chunks: The final response
-
Display: Thinking chunks are yielded as
{ type: "reasoning", text: ... }which the UI displays separately from the final response, allowing users to see the model's thought process. -
SDK Support: The Mistral SDK v1.9.18+ includes
ThinkChunksupport for handling thinking tokens.
Performance Considerations
Current Performance
The user reported that "performance is suboptimal" when using Devstral 2 models. Here are the key factors:
-
Thinking Tokens Add Latency: When the model generates thinking tokens, it takes additional time before the final response begins. This is expected behavior and shows the model is working through the problem.
-
Streaming is Enabled: The handler properly streams both thinking and text tokens, so users see progress in real-time rather than waiting for the complete response.
-
Temperature Setting: PR #9957 set the default temperature to 0.2 based on Mistral's recommendations for optimal performance.
Prompt Caching
Current Status: Mistral models currently have supportsPromptCache: false in the model definitions.
Investigation Needed:
- The Mistral API documentation does not clearly indicate whether prompt caching is supported for Devstral 2 models
- Unlike Anthropic or OpenAI, Mistral's API documentation doesn't provide explicit caching endpoints or parameters
- Further investigation with Mistral's API team would be needed to determine if caching is available
Recommendations for Users
-
Expect Thinking Time: When using Devstral 2, the "thinking" phase is normal and indicates the model is reasoning through the problem. This is a feature, not a bug.
-
Temperature: The default temperature of 0.2 is optimized for code generation tasks. Users can adjust this in settings if needed.
-
Model Selection:
- Use
devstral-latestordevstral-2512for complex reasoning tasks where thinking tokens are valuable - Use
devstral-small-latestorlabs-devstral-small-2512for faster responses on simpler tasks
- Use
-
Monitor the UI: The thinking tokens should be visible in the UI, showing the model's reasoning process. If they're not appearing, this may indicate a UI rendering issue rather than an API issue.
Technical Details
Mistral SDK Version
Roo Code uses @mistralai/mistralai version ^1.9.18, which includes support for thinking chunks.
Content Chunk Types
The Mistral API returns content in different formats:
type ContentChunkWithThinking = {
type: string // "thinking" or "text"
text?: string // For text chunks
thinking?: Array<{
// For thinking chunks
type: string
text?: string
}>
}
Streaming Flow
- API request is made with streaming enabled
- Server streams back chunks as they're generated
- Thinking chunks are processed and yielded as
reasoningtype - Text chunks are processed and yielded as
texttype - UI displays both types appropriately
Future Improvements
Potential Optimizations
-
Prompt Caching: If Mistral adds prompt caching support, we can:
- Cache system prompts across requests
- Cache conversation history
- Reduce latency for follow-up requests
-
Batch Processing: For multiple requests, investigate if Mistral supports batch APIs
-
Connection Pooling: Ensure HTTP connections are properly pooled and reused
Monitoring
To help users understand performance:
- Token Metrics: Display thinking token count vs. response token count
- Timing Metrics: Show time spent in thinking phase vs. response phase
- Progress Indicators: Enhance UI to better show when model is thinking
References
Questions for Mistral Team
To further optimize performance, we need clarification from Mistral on:
- Does the Mistral API support prompt caching for Devstral 2 models?
- Are there any API parameters to control thinking token generation?
- What are the recommended best practices for minimizing latency?
- Are there any batch or concurrent request optimizations available?