Zendoric
← Back to the day · July 30, 2026

Laguna S 2.1: Poolside's 118B model that beats trillion-parameter giants

🕒 Published on Zendoric: July 30, 2026 · 00:20

This week's "AI of the Week" article starts from a simple visual exercise: take every open-weights model that discloses its parameter count, plot total parameters on a logarithmic X axis and the Terminal-Bench 2.1 score on the Y axis.

By TheSequence · July 29, 2026.

This week's "AI of the Week" article starts from a simple visual exercise: take every open-weights model that discloses its parameter count, plot total parameters on a logarithmic X axis and the Terminal-Bench 2.1 score on the Y axis. The expected result, and the one actually observed in most cases, is a cloud of points rising from left to right: the bigger the model, the better the score. It is the relationship the industry has taken as the norm.

There is, however, one point that breaks that trend strikingly: at 118 billion parameters (118B), well above the fit line the rest of the models follow, sits Laguna S 2.1, with a score of 70.2% on Terminal-Bench 2.1. To put that in perspective, the newsletter compares that result with three far larger models: DeepSeek-V4-Pro-Max, with 1.6 trillion parameters, which scores 64.0 points; Inkling, with 975 billion parameters, which reaches 63.8; and Nemotron 3 Ultra, with 550 billion parameters, which stops at 56.4.

In other words, a model a fraction the size of its competitors not only matches but beats models between 4.6 and 13.5 times bigger than it on this benchmark.

The gap widens even further on DeepSWE, described in the newsletter as a harder and less saturated benchmark than Terminal-Bench. There, Laguna S 2.1 posts 40.4 points against DeepSeek-V4-Pro-Max's 9.0, a gap the text itself flags as no longer subtle at all: almost 4.5 times the performance with 13 times fewer parameters.

The newsletter stresses that this kind of result — a 13x parameter deficit combined with a 4x score advantage — is usually the sort of data point that raises suspicions the benchmark is badly designed or has been gamed ('someone broke the eval'). According to the text, the company behind Laguna, Poolside, seemed to anticipate that skeptical reaction, and for that reason published all the trajectories of every trial in the final evaluation set, so that anyone can review exactly what the model did step by step. The article notes that this decision in favor of total transparency reveals a great deal in itself about how this launch was designed.

The body of the newsletter cuts off just as it begins to introduce the company responsible, with the line 'The company you have probably never heard of...', indicating that the rest of the analysis — presumably about Poolside, its approach and more details on the Laguna model — continues in the full version of the article on Substack, which was not included in the emailed text.

More broadly, as industry context, Terminal-Bench and DeepSWE are benchmarks used to measure language models' capabilities on terminal and agentic software-engineering tasks, and the efficiency comparison (performance per parameter) is a recurring theme in evaluating open models against large-scale models such as DeepSeek or Nemotron.

🔗 Related on Zendoric

Sources & references