Zendoric
← Back to the day · July 30, 2026

Agentic AI shifts the bottleneck from the GPU to memory — and that's where Penguin Solutions' risk lies

🕒 Published on Zendoric: July 30, 2026 · 00:20

Penguin Solutions is up 123% in 2026 and Micron 188%, driven by one thesis: agentic AI, which works nonstop, turns memory —not the GPU— into the bottleneck of AI factories. The numbers are real; so is the concentration of risk.

By Zendoric · July 30, 2026.

Penguin Solutions CEO Kash Shaikh summed it up on CNBC on July 29 with a line that suits his business very well: "memory is the new compute, especially with agentic AI." Behind the line is a company that has gained 123% since January 2026, reaching a market value of about $2.47 billion, with a buy consensus among analysts and a price target of $74.29 against the current $43.70, according to data compiled by 24/7 Wall St.

The quarterly numbers back up the narrative. In its fiscal third quarter of 2026, Penguin posted revenue of $478.71 million (+48% year over year) and non-GAAP earnings per share of $0.84, beating consensus by 13.61% on revenue and 49.33% on earnings, according to the company itself. Its memory and AI infrastructure division grew more than 104% year over year and already accounts for more than three quarters of total sales. The company raised its sales growth forecast for the fiscal year to 22% (±2 points) and its non-GAAP earnings-per-share guidance to $2.60 (±0.05). It was also recently named an NVIDIA specialized partner for "AI factories" —the term NVIDIA itself uses for data centers dedicated exclusively to producing intelligence at industrial scale— and Dell's partner of the year for the Americas region.

Shaikh's technical argument makes sense and is worth understanding properly. A conventional AI assistant answers a question and stops there: memory use is a one-off. An autonomous agent, by contrast, operates "24 hours a day, seven days a week," running tasks and workflows without pause. That requires continuously keeping alive long context windows and what technical jargon calls the KV (key-value) cache: the memory where a language model stores the calculations it has already performed on each fragment of processed text, so it does not have to repeat them at every subsequent step. The longer and more persistent the agent's task, the more cache has to be sustained in memory, not in the GPU's compute engine. Penguin's MemoryAI server, based on the CXL interconnect standard (which allows memory to be added to a server as if it were its own, instead of relying only on what each processor brings), is already deployed at a top-tier financial institution, according to the company.

Micron offers larger-scale validation of the same phenomenon. Its fiscal third-quarter 2026 revenue reached $41.46 billion, 345.7% more than a year earlier, with a GAAP gross margin of 84.6%. The company is already shipping its HBM4 memory in volume (stacked high-bandwidth memory designed to feed data to GPUs much faster than conventional memory) and expects revenue of $50 billion (±$1 billion) for the fourth quarter. Its shares are up close to 188% for the year. Against that backdrop, NVIDIA posted revenue of $81.62 billion in its fiscal first quarter of 2027, of which $75.25 billion came from data centers; Jensen Huang has described it as "the largest infrastructure expansion in the history of humanity."

Shaikh's thesis should be read with the usual caution that applies when the person stating it is also the one selling the solution: an executive at a memory infrastructure company has every incentive to declare that memory is the bottleneck of the moment. That does not make it false —the technical argument about persistent agents and KV cache is solid, and it is backed by Micron figures unrelated to Penguin's commercial narrative— but it is worth distinguishing between a real industry trend and the pitch of whoever is capitalizing on it in the stock market.

Where the story gets more fragile is risk concentration. AI already accounts for 74% of Penguin's revenue, but 89% of its operating profit depends on a single segment, memory, with a beta of 2.83 (a stock with that beta moves, in theory, almost three times as much as the market, both up and down) and a 22.68% drop over the past month despite the annual rally. In the days this article was published, 24/7 Wall St.'s own website showed Micron in its markets bar among the session's biggest losers, down close to 10%, despite being up 188% so far this year. It is a good reminder that, broadly speaking, the memory business has been one of the most cyclical sectors in all of semiconductors: the same components that are scarce today and send margins soaring may be in surplus tomorrow if the pace of data center construction slows.

Even with that caveat, the underlying story matters more than one day's ticker. The bottleneck shifting from the GPU to memory confirms something we at Zendoric have been pointing out from another angle: the fight over where and how agents "live" is no longer just a war of software platforms —who hosts the agent, who orchestrates the workflow— but also a war over the physical substrate that sustains their persistent memory. For an agent to work continuously for days without losing the thread, someone has to manufacture, package and integrate the memory that makes it possible. Broadly speaking, the competitors there are, on the manufacturing side, giants such as Micron, SK Hynix and Samsung, and on the integration side, companies like Penguin, which sell that memory already packaged into systems ready for banks, governments and cloud providers.

In the long run, this is exactly the kind of physical, tedious, volatile-margin investment needed for the promise of agentic AI —agents that work tirelessly, freeing people from routine tasks— to stop being a demo and become real, large-scale infrastructure. The abundance we champion as a horizon does not arrive by decree: it arrives by literally building more memory, more bandwidth and more data centers than exist today. But in the short term, what there is is this: a handful of companies heavily exposed to a single business cycle, trading at multiples that already price in much of the optimism —Penguin's forward P/E, at around 12 times, suggests the market itself takes for granted that this growth will moderate— and a stock market that rewards anyone betting big on a single piece of the AI puzzle with extreme volatility.

🔗 Related on Zendoric

Sources & references