GLM, Mistral, and Deepseek

Randoms thoughts and impressions from using the Mistral and the GLM models hosted by Mistral. No benchmarks were used. This has just been from my personal experience.

Z.Ai's GLM 5.2 and 5.3 through Mistral

Got to hand it to Mistral for setting up these models from Z.ai on Mistral's own infrastructure because these models have been awesome. This model family has been my daily driver and has been churning through 1M token autonomous Pi harness sessions without any problems

My current Pi harness may be overstuffed with MCP servers which has caused Deekseek Flash v4.1 to time out on occassion. Not a problem for GLM5.3. These models are fast and smart and I have yet to hit a limit on the Mistral $15/month plan. For such a cheap subscription, that is a ginourmous bang for my buck.

GLM 5.3 has been my primary driver for building out this dark software factory as well as this website. It has helped me with significant data science type research work including setting up many econometrics experiments. There has only been 1-2 instances where i have needed to swap out (Gemini / Opus) to help solve a problem and arguably that was more my fault.

The other Mistral models

Mistral Medium and Mistral Large are several generations behind at this point and are noticeably inferior to the competition. Mistral Small 4 is still decent, but ChatGPT Luna 6 (especialy at that new price) is also dramatically better.

When I have used the mistral models (and they have worked), I have found that I prefer their "style" of responding and chat the best. They are all terse and professional and I don't have to waste time with all of the superflous flowery language. Now if we could only get a new generation of their generalist models....

The Mistral product family overall is quite interesting. It is clear they are playing the long game by building out products that cater to enterprises and especially those with high legal and regulatory requirements (ex: European companies, military suppliers, healthcare). I have yet to try the Mistral OCR models but I will probably look there first when I start experimenting with that technology. I can also see myself embedding Mistral's Lean4 model into my own software factory for formal verification unit testing (going to need a Rust-> Lean4 conversion step there). Honestly it feels like Mistral seems like the only adult in the room compared to some of the other AI labs. Yeah, they aren't releasing the flashiest models, but they are building out the tools and models one would actually use and reach for in critical professional settings.

Deepseek

While benchmarking sites put Deepseek Flash v4.1 slightly ahead of GLM 5.3 that has not been my experience. That being said, The Deepseek models are very good and honestly so cheap as to be functionally free. I've been using it extensively (for me) and have spent all of $13 on the Deepseek API. It feels functionally free.

The model is insanely fast, which is good, because oh lord does it think... and think ... and think some more and then be incredibly verbose when it does answer. The answers are usually good. But there are occassions where the model is so fast, the commands it issued to the system haven't completed yet and it thinks there must be an issue (since LLMs don't have a concept of time).

Still the quality is high (anything as smart or smarter than Sonnet 4.6 is my general standard) and so Deepseek has been my backup driver for when I feel concerned about Mistral limits or when I know the task is simple enough that I can trust Deepseek to see it through the entire way.

← All posts