Mistral AI Slashes Prices 60% With New Mistral Small 3.1 24B Model
The French AI startup that popularized open-weight models is now competing on price. Mistral AI has released Mistral Small 3.1 24B, a new model priced 40 to 60 percent lower than comparable offerings from OpenAI and Google.
The model costs $0.10 per million input tokens and $0.30 per million output tokens via the Mistral API. That undercuts GPT-4o mini ($0.15/$0.60) and Gemini 2.0 Flash ($0.10/$0.40) on output pricing.
Mistral claims the new 24-billion-parameter model outperforms GPT-4o mini, Gemini 2.0 Flash, and Claude 3.5 Haiku on key benchmarks. It also matches the performance of the much larger Llama 3.1 70B and Qwen 2.5 72B.
A Strategic Pivot From Open Weights to Volume
Mistral built its reputation by releasing open-weight models under the Apache 2.0 license. The company made open-weights mainstream with models like Mistral 7B and Mixtral 8x7B.
Now the startup is shifting strategy. Instead of only pushing frontier performance, Mistral is betting on aggressive pricing to drive API adoption and revenue.
“We are not just competing on performance anymore. We are competing on economics,” a Mistral spokesperson said.
The new pricing model reflects a broader industry trend. AI companies are racing to lower costs as enterprises demand cheaper inference for large-scale deployments.
Mistral Small 3.1 24B: Key Specifications and Capabilities
The model features a 128,000-token context window — double the size of many competitors. That allows processing long documents, codebases, or multi-turn conversations without truncation.
It supports text-only input and output with an emphasis on low latency. Mistral designed the model for real-time applications like chatbots, code assistants, and content generation.
The 24B parameter count places it in the “small” category by modern standards. That enables deployment on consumer GPUs with efficient memory usage.
Benchmark scores show the model competing well:
- MMLU (knowledge reasoning): 81.1% — beats GPT-4o mini (79.9%)
- HumanEval (code generation): 82.3% — matches Llama 3.1 70B
- GSM8K (math word problems): 91.5% — exceeds Gemini 2.0 Flash
How the Pricing Stacks Up Against Competitors
The cost advantage is most visible on output tokens, which typically dominate inference budgets for generative tasks.
| Model | Input cost per million tokens | Output cost per million tokens |
|---|---|---|
| Mistral Small 3.1 24B | $0.10 | $0.30 |
| GPT-4o mini | $0.15 | $0.60 |
| Gemini 2.0 Flash | $0.10 | $0.40 |
| Claude 3.5 Haiku | $0.25 | $1.25 |
Mistral also offers a free tier with rate limits for developers testing the model. Enterprise customers can negotiate volume discounts.
Open Weights Remain on the Table
Despite the aggressive API pricing, Mistral has not abandoned its open-weight philosophy. The company plans to release the model weights under an open license in the coming weeks.
That move allows developers to self-host the model, avoiding API costs entirely. It also undercuts proprietary models that lock users into vendor-specific infrastructure.
“Open weights give developers freedom. Competitive pricing gives them affordability. Both are necessary for widespread AI adoption,” Mistral noted in its announcement.
The combination of open availability and low-cost inference could pressure competitors to lower their own prices or release more capable open models.
What This Means for the AI Market
Mistral’s strategy signals a maturing market where differentiation shifts from raw benchmark scores to cost-per-task. Enterprises increasingly evaluate models on total cost of ownership, not just accuracy.
For developers, the choice becomes simpler: use Mistral Small 3.1 24B for cost-sensitive workloads, or reserve premium models for tasks requiring frontier reasoning.
The model is available now via the Mistral API and will land on platforms like Hugging Face, AWS, and Azure in the coming days.
Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.
What are your thoughts on this? I’d love to hear about your own experiences in the comments below.