Zendoric
← Back to the day · September 1, 2026

Your cluster is a graph: drift detection in Kubernetes with Drasi and Cypher

🕒 Published on Zendoric: September 1, 2026 · 00:48

✨ AI-generated · how it's made

The article is not about generative AI but about platform engineering: it documents a hands-on experiment with Drasi, an engine that lets you treat the live state of a Kubernetes cluster (specifically AKS) as a graph queryable with Cypher, the query language popularized by graph databases.

The article isn't about generative AI but about platform engineering: it documents a hands-on experiment with Drasi, an engine that makes it possible to treat the live state of a Kubernetes cluster (specifically AKS) as a graph you can query with Cypher, the query language popularized by graph databases. The starting point is honest about the tool's limits: in any cluster running GitOps there are already drift-detection mechanisms — Gatekeeper or Kyverno blocking invalid configurations at admission time, Flux or Argo comparing cluster state against git, and Prometheus firing alerts on metrics sustained over a given period. The author does not present Drasi as a replacement for any of them, but as a piece that fills a specific gap: relational questions that span several resource types at once and that none of those tools can answer on its own, of the kind 'is this Deployment still under-replicated five minutes later?' or 'is this Pod running an unapproved image?'.

Drasi's mechanics consist of "standing" (continuous) Cypher queries over the resource graph, with temporal functions such as trueFor (a condition must remain true throughout a time window) and trueLater (self-scheduling of future checks). The advantage over a periodic polling system is that reactions fire on real state transitions, not on every check tick, and the result set keeps itself up to date. Drasi can be deployed inside the cluster itself or point at an external one through a kubeconfig Secret, and the author recommends limiting that credential's RBAC permissions to list/watch on only the resource types the rules need to read, as if it were the service account of a read-only dashboard.

The core of the article is methodological: the author first wrote six rules based solely on the official documentation, and then tested them against a real cluster with deliberately triggered violations. The result was revealing — only two rules survived unchanged, two needed rewriting and two turned out to be outright unworkable, not because of syntax errors but because of a structural limitation of the platform. That gap between what looks correct on paper and what actually works when it runs is, in the author's own words, the most useful outcome of the exercise.

The structural limitation is key to understanding the rest of the article: Drasi's Kubernetes source watches a fixed list of twelve resource types — Pod, Deployment, ReplicaSet, StatefulSet, DaemonSet, Job, Service, ServiceAccount, Node, Ingress, PersistentVolume and PersistentVolumeClaim — and no additional RBAC permission extends that list. Namespace, NetworkPolicy, PodDisruptionBudget and Secret are simply not observable today, so any rule that depends on them is ruled out from the start, and the only solution is to pick another resource to build on, not to try to fix it with a cleverer query.

Among the rules that did work, the one for sustained under-replicated deployments compares spec.replicas with status.readyReplicas wrapped in a five-minute trueFor, precisely so that the normal fluctuations of a rollout don't trigger false positives. Adding coalesce(status.readyReplicas, 0) was necessary because that field comes back null while the deployment is in progress, and without that adjustment the original query failed outright. The author validated it by provoking a real violation with an unschedulable nodeSelector, shortening the window to five minutes just to be able to watch the transition without long waits.

The rule for images outside the approved registry revealed another limit of the parser: STARTS WITH, CONTAINS, ENDS WITH and regular expressions with =~ don't exist in the supported subset of Cypher, so it had to be rewritten using left(image, N) <> 'prefix'. Once on the real cluster, the query correctly detected a docker.io/library/nginx image that had been deliberately tagged incorrectly, but it also brought to light two real edge cases a simplified test cluster would not have shown: Dapr sidecars are reported as docker.io/daprio/daprd, so that prefix has to be added to the allowed list or every pod with that sidecar would raise a false alert; and Calico components are reported only as a sha256 digest with no registry prefix at all, which leaves any prefix-based rule completely blind to digest-pinned images, a real limitation the author prefers to flag rather than hide.

The three discarded rules — namespaces without a NetworkPolicy, TLS certificate expiry in Secrets, and PodDisruptionBudget exhaustion — share the same underlying problem: the resource types involved (Namespace, NetworkPolicy, Secret, PodDisruptionBudget) are not on the source's watch list. On top of that, the EXISTS { MATCH ... } syntax, meant to express the absence of a relationship (for example, 'this namespace never had a NetworkPolicy associated with it'), isn't supported by the current parser either. The author stresses that the idea of using trueLater to self-schedule certificate expiry checks without needing a cron job scanning secrets remains conceptually valid and pedagogically interesting, even if it can't be demonstrated against Kubernetes Secrets today; he suggests that cert-manager's Certificate CRD could be an alternative route, though he clarifies that he hasn't tried it.

The article also documents three new rules not contemplated in the original plan. The one for sustained node pressure (memory or disk) uses the same trueFor pattern over Node conditions, and was validated on syntax only: the author decided not to force a real violation because the test node shared Drasi's own control plane, and deliberately exhausting its memory would have put the whole environment at risk. The one for containers in CrashLoopBackOff required an "unwind" middleware to extract each Pod's containerStatuses as independent Container nodes, with the practical warning that this middleware block must be nested under the manifest's sources: key and not under spec:, a mistake that fails silently and with no clear message if placed wrongly. This rule was verified with a real violation, staying empty on the healthy cluster and firing the instant the failure state was forced. The third, deployments with no associated ReplicaSet, replaces the (unsupported) EXISTS {} subquery with an OPTIONAL MATCH combined with count(), and stayed empty because every deployment in the test cluster had its corresponding ReplicaSet.

One of the article's most practical ideas is the philosophy of a 'healthy rulebook': on a compliant cluster, the rules that work should be empty the vast majority of the time, showing a row only at the exact moment a real violation crosses its threshold and disappearing again as soon as the problem is resolved, with no polling delay. If a rule fires constantly against a healthy cluster, the recommendation is to review the query first — usually a coalesce is missing or the debounce window is too short for the normal behavior of a rollout. And if a rule never fires no matter what, the reverse recommendation is just as important: don't assume the cluster is fine, but first check whether the resource type in question is even on the source's watch list before trusting an empty result.

The author also summarizes, as a practical reference that isn't officially documented and was deduced from the query engine's own error messages, the real surface of the supported Cypher parser: comparison operators (=, <>, !=, <, <=, >, >=), IN, IS and basic arithmetic do work, as do left(), coalesce(), OPTIONAL MATCH with aggregation functions and the unwind middleware; by contrast, STARTS WITH, CONTAINS, ENDS WITH, regular expressions with =~ and existential EXISTS { MATCH } subqueries are not supported. He also mentions, as a useful aside for anyone wanting to reproduce the tests, that listing the actual images of Drasi's own pods revealed the internal names of its components (query-container-query-host, query-container-view-svc, query-container-publish-api, source-query-api, source-change-dispatcher, source-change-router), which explains why registry probes using simpler names such as query-host or view-svc always return a 404 error.

As for operational notes from a real deployment, the article points out that the kubeconfig should be stored as a Kubernetes Secret, avoiding exec-based configurations in runtime containers; it recommends watching for ResourceVersionTooOld events in the source's namespace, which indicate an out-of-sync watch cache capable of losing changes without raising any obvious warning; it warns that modifying the Kubernetes source forces you to delete it and create it again, along with every query that depends on it, unlike SQL Server sources, which can be cleanly reapplied by name; and it clarifies that default provider registration now happens automatically after running drasi init, so older guides that told you to apply those manifests manually are out of date for the current version of the CLI.

The article closes by reinforcing where this technique should not be used: policy enforcement at admission time remains the responsibility of Gatekeeper or Kyverno — Drasi observes, it doesn't block — and drift detection between git and the cluster still belongs to Flux or Argo. It also recommends keeping one Drasi instance per cluster rather than pointing a single external instance at several clusters at once, a lesson the author says he only half-absorbed from someone else's experience. The overall conclusion is modest and consistent with the tone of the piece: five rules that genuinely work against a real cluster, with the exact syntax that is actually interpreted, are worth more as a starting point than six rules that only looked correct on paper.

🔗 Related on Zendoric

Sources & references