Zendoric
← Back to the day · July 27, 2026

Meta's Muse Spark 1.1 beats GPT-5.6 on health at 1/7 the cost — and ships free to 3.56 billion people

🕒 Published on Zendoric: July 27, 2026 · 00:21

Meta says its new Muse Spark 1.1 outscores GPT-5.6 on HealthBench, OpenAI's own health evaluation, at roughly one-seventh the cost. The distribution number is the real story: free access inside products with 3.56 billion daily users.

The claims, per the note: Meta has released Muse Spark 1.1; it surpasses GPT-5.6 on HealthBench — the health-conversation benchmark built by OpenAI — at a cost roughly seven times lower; and it is offered free inside Meta's products, which reach 3.56 billion daily active users.

Two of those three numbers deserve scrutiny, and the third deserves alarm. On the benchmark: HealthBench grades model responses to medical conversations against physician-written rubrics. Beating a rival on one evaluation, self-reported by the challenger, is not the same as being better at health. Our own composite reading still places GPT-5.6 near the top of the frontier on hard, non-saturated tasks. One benchmark win is a marketing artifact until independent replication says otherwise.

On cost: a 7× price gap is the genuinely consequential figure, and it's consistent with what we've tracked all year — inference economics collapsing while total spend rises. Cheap medical-grade conversation is how you get triage advice to people who currently get none. That is the abundance thesis working as intended.

On distribution: 3.56 billion daily users is roughly 40% of humanity, and it is why this launch matters more than its scorecard. Meta is not selling a health model; it is defaulting one into WhatsApp, Instagram and Facebook. The competitive lesson we keep returning to holds — the frontier is won at the plumbing layer, not the leaderboard.

Our reading: free, cheap health guidance at planetary scale is one of the most concretely good things AI can do in this decade, and also the highest-stakes place to ship a model that hasn't been independently audited. Medical advice fails asymmetrically: a marginally better rubric score does not cover a confidently wrong answer given to someone who won't see a doctor afterward. We'd like to see external evaluation, published refusal behavior on emergencies, and clarity on what happens to those conversations. Until then: a real step forward, delivered at a scale that leaves no room for benchmark theater.

🔗 Related on Zendoric

Sources & references