Nate puts Chinese models (DeepSeek, Qwen, GLM, Kimi, MiniMax) to the test in a real multi-agent system

🕒 Published on Zendoric: July 28, 2026 · 00:38
The newsletter takes on the debate over whether it is worth moving agentic AI workloads to Chinese models (DeepSeek, Qwen, GLM, Kimi, MiniMax) given that they cost a fraction of the price of Western frontier models.
By Nate from Nate's Substack · July 27, 2026.
The email tackles the debate over whether it is worth moving agentic AI workloads to Chinese models (DeepSeek, Qwen, GLM, Kimi, MiniMax) given that they cost a fraction of the price of Western frontier models. Nate argues the discussion usually boils down to anecdotes: one good answer from one of these models prompts claims that they have already caught up with the US frontier, and one censorship failure or a faulty citation prompts writing off the entire Chinese stack at a stroke. In his view, neither extreme answers the question that really matters, which is whether a specific model is good for a specific job.
His starting position is that serious AI users should test the Chinese models. He uses them selectively himself — aggressively in some workflows — but insists that the job to be done, the model chosen and the deployment path are three separate decisions that should not be conflated.
Nate says that while building a system of his own called Ringer, he tried Qwen as one of the 'workers' and found it useful, though he explicitly acknowledges that he does not have a full record of the exact Qwen checkpoint used, the task mix, the hit rate or the cost of that particular experiment; he prefers to admit that gap rather than blend separate experiments to make the headline sound neater. He clarifies that the experiment for which he does have complete logs — a run of 34 tasks in Ringer — was carried out with GLM-5.2, GPT-5.5, Grok 4.5 and Composer 2.5 Fast, the latter two through the Grok Build CLI.
The email doubles as promotion for a paid guide previously announced in a video, and previews its contents without fully developing them:
- An analysis of that 34-task run in Ringer that, according to Nate, changed how he assigns models to tasks: one of the workers reported 213 verified verbatim citations, 13 of which turned out to be 'stitched together' from fragments.
- An explanation of what an accepted result actually costs, proposing an alternative metric to price per token, and the idea that a cheaper model can end up doubling the human review time required.
- Recommendations on where to start with each model family (DeepSeek, Qwen, GLM, Kimi and MiniMax), noting that there is a specific failure to watch for in each one, though the email does not spell out what those failures are.
- A 'bakeoff kit' made up of a validator, a manifest, a scorecard and two test fixtures designed to check that the verification system does indeed reject faulty work.
Nate closes the introduction by noting that the full experiment behind these conclusions cost roughly $8. Access to the full content — the detailed analysis and the guide — is reserved for paying subscribers, who also get membership in his Slack community.
🔗 Related on Zendoric
- Claude Sonnet 5 makes agentic AI cheaper: the battle is no longer the benchmark, it's who integrates best · 2026-07-17
- Hugging Face turned to an open Chinese model to analyze its own breach because OpenAI and Anthropic refused to help · 2026-07-22
- The price of intelligence and the manufacturing of experience · 2026-07-25


