What NASA Knew About If/Else That You Don't
In 1985, NASA built a rule engine called CLIPS. Forty years later, it still has lessons for every developer whose business logic is buried in procedural code.


In 1985, at NASA’s Johnson Space Center, a team of engineers started building something unusual. While most of the industry was chasing expensive Lisp machines for AI development, they wrote an inference engine in plain C.
They called it CLIPS: C Language Integrated Production System.
A Quiet Piece of Engineering History
The context matters. Expert systems were the hot AI technology of the 1980s. Companies were spending millions on specialized hardware and proprietary Lisp environments to build them. NASA needed the same capabilities but couldn’t justify tying critical systems to expensive, single-vendor platforms.
Robert Savely conceived the project. Frank Lopez designed and built version 1.0 in the spring of 1985, in just over two months. Chris Culbert managed the project. And Gary Riley designed and developed the rule-based engine at the heart of CLIPS: the pattern matching, the inference, the core of what makes it work.
CLIPS compiled on any platform with a C compiler. It was fast. It was embeddable. And in 1996, when NASA stopped funding expert system research, Gary kept going. He has maintained and developed CLIPS independently ever since. The codebase has outlived entire programming paradigms, multiple AI winters, and the rise of machine learning. Most developers have never heard of it.
That’s a shame, because what Gary and the original team built encodes an idea that most developers rediscover the hard way: procedural logic buries intent. You can’t inspect it, you can’t explain it, and you can’t change one rule without risking the rest.
CLIPS offered an alternative. What if each rule stated its conditions and its actions, and nothing else? No awareness of other rules. No position in a sequence. No responsibility for chaining or cleanup. You give the engine facts (data the system knows about) and rules (patterns that match facts and trigger actions). The engine watches your facts, finds which rules match, and decides what to execute.
When a rule fires, it can assert new facts. Those new facts can trigger other rules. The cascade emerges from the data, not from your control flow. And when a fact changes, any conclusions that depended on it can be automatically retracted. You describe what should happen. The engine figures out when, in what order, and what to undo.
Who uses this approach? CLIPS itself was deployed across NASA, the U.S. military, and companies like DuPont. The same paradigm powers systems you interact with daily: JP Morgan Chase, HSBC, and Citi use Pega’s rule engine for financial decisioning. FICO’s Blaze Advisor (built on Charles Forgy’s Rete III algorithm) drives credit decisions at major banks worldwide. CERN uses Drools. The approach scales from expert systems to enterprise infrastructure.
The Imperative Version
If you’ve written a Kubernetes operator, you know the pattern. A reconciliation loop polls for state, checks conditions, takes actions, and hopes it caught every edge case.
Here’s a simplified version of what cluster management logic looks like in a typical controller:
def reconcile_cluster(nodes, pods, deployments):
for node in nodes:
if node["memory_percent"] > 90:
alert(f"Memory pressure on {node['name']}")
if node["memory_percent"] > 95:
cordon(node["name"])
for pod in pods:
is_on_node = pod["node"] == node["name"]
if is_on_node and pod["status"] == "running":
evict(pod["name"], pod["namespace"])
for dep in deployments:
same_ns = dep["namespace"] == pod["namespace"]
if same_ns and dep["available"] < dep["replicas"]:
schedule_pod(dep["name"], dep["namespace"])
for pod in pods:
if pod["restarts"] > 5:
alert(f"CrashLoopBackOff: {pod['name']}")
if pod["restarts"] > 10:
evict(pod["name"], pod["namespace"])
This works for six rules. Now picture forty, which is closer to what a real operator handles.
The problems run deeper than code structure:
- You can’t inspect it. When an on-call engineer asks “why was this pod evicted?”, the answer is somewhere in the nesting. You have to trace execution paths to reconstruct the chain of events.
- You can’t explain it. The business intent (“evict pods from nodes under memory pressure”) is buried inside loops, index variables, and branching logic. A product manager can’t read it. Neither can an auditor.
- You can’t change one rule without risking the rest. Adding a new condition means editing this function. The eviction logic for memory pressure and crashloops is duplicated. The order of the nested loops matters: check pods before nodes and you miss the cordon step.
- You can’t trust what you checked. Once a condition passes, the code moves on. By the time you act on it, the state may have changed. The function has no way to know whether what it checked at the top of the loop still holds true three nesting levels later.
The function tangles three concerns: what the rules are, in what order they execute, and how they chain together.
This is the problem CLIPS was built to solve. Let’s see how.
Enter CLIPS, via Python
clipspyx is a Python DSL for CLIPS. CLIPS has its own syntax that looks like Lisp (though it isn’t one). clipspyx lets you skip that entirely: you write annotated Python classes that compile to CLIPS constructs behind the scenes. Let’s rebuild the same cluster management logic.
Want to follow along? Open this notebook in Google Colab and run the code yourself.
First, define your data as templates:
from clipspyx import Environment
from clipspyx.dsl import Template, Rule
class Node(Template):
"""A Kubernetes cluster node."""
name: str
status: str
cpu_percent: float = 0.0
memory_percent: float = 0.0
class Pod(Template):
"""A running pod in the cluster."""
name: str
namespace: str
node: Node # fact-address: which node this pod runs on
status: str
restarts: int = 0
class Deployment(Template):
"""A deployment managing a set of pods."""
name: str
namespace: str
replicas: int
available: int
class Alert(Template):
"""A derived alert, asserted by rules."""
kind: str
target: str
message: str
severity: str
class Cordon(Template):
"""Cordon a node to prevent new pod scheduling."""
node_name: str
reason: str
class Evict(Template):
"""Evict a pod from its current node."""
pod_name: str
reason: str
class Schedule(Template):
"""Schedule replacement pods for a deployment."""
namespace: str
deployment: str
reason: str
Templates are your schema. Node, Pod, and Deployment are facts you assert from cluster state. Alert, Cordon, Evict, and Schedule are derived facts that rules produce. Each one names exactly what it represents.
Now the rules. Notice something as you read them: no rule knows about any other rule. Yet by the end, they’ll chain four levels deep without a single explicit call between them.
class NodeMemoryPressure(Rule):
"""Alert when a node's memory exceeds 90%."""
logical(Node(name=name,
memory_percent=mem and mem > 90))
def __action__(self):
Alert(
__env__=self.__env__,
kind="memory_pressure",
target=self.name,
message=f"Node {self.name} at {self.mem}% memory",
severity="warning",
)
class NodeCriticalMemory(Rule):
"""Cordon a node when memory exceeds 95%."""
n = Node(name=name,
memory_percent=mem and mem > 95)
def __action__(self):
Cordon(
__env__=self.__env__,
node_name=self.name,
reason=f"Memory critical: {self.mem}%",
)
Notice logical() wrapping the Node pattern in NodeMemoryPressure. This tells CLIPS: the Alert this rule asserts is logically supported by that Node fact. If the Node fact changes, the Alert retracts itself. No cleanup code. Remember the fourth bullet: “you can’t trust what you checked”? logical() is the answer. We’ll prove it shortly.
Here’s where it gets interesting. NodeCriticalMemory asserts a Cordon fact. It has no idea what happens next. But another rule is watching:
class EvictPodsFromCordonedNode(Rule):
"""When a node is cordoned, evict its running pods."""
Cordon(node_name=node_name)
n = Node(name=node_name)
p = Pod(node=n, status="running",
name=pod_name, namespace=ns)
def __action__(self):
Evict(
__env__=self.__env__,
pod_name=self.pod_name,
reason=f"Node {self.node_name} cordoned",
)
The chain isn’t done. EvictPodsFromCordonedNode just asserted Evict facts. It doesn’t know this, but another rule is waiting for exactly that:
class RescheduleEvictedPod(Rule):
"""When a pod is evicted, schedule a replacement."""
Evict(pod_name=pod_name)
p = Pod(name=pod_name, namespace=ns)
d = Deployment(namespace=ns, replicas=desired,
available=actual and actual < desired)
def __action__(self):
Schedule(
__env__=self.__env__,
namespace=self.ns,
deployment=self.d.name,
reason=f"{self.actual}/{self.desired} available",
)
That’s three levels of cascading, and no rule called another. Meanwhile, a completely separate chain handles crashlooping pods. These rules have no idea the memory pressure chain exists:
class CrashLoopDetection(Rule):
"""Alert on pods that keep restarting."""
logical(Pod(name=name, namespace=ns,
restarts=restarts and restarts > 5))
def __action__(self):
sev = "critical" if self.restarts > 10 else "warning"
Alert(
__env__=self.__env__,
kind="crashloop",
target=self.name,
message=f"Pod {self.name}: {self.restarts} restarts",
severity=sev,
)
class EvictCrashLoopPod(Rule):
"""Evict pods stuck in a severe crash loop."""
p = Pod(name=name, namespace=ns,
restarts=restarts and restarts > 10)
def __action__(self):
Evict(
__env__=self.__env__,
pod_name=self.name,
reason=f"CrashLoopBackOff: {self.restarts} restarts",
)
Now let’s light the fuse. We assert some cluster state and call env.run() exactly once:
env = Environment()
NodeAssert = env.define(Node)
PodAssert = env.define(Pod)
DeploymentAssert = env.define(Deployment)
env.define(Alert)
env.define(Cordon)
env.define(Evict)
env.define(Schedule)
env.define(NodeMemoryPressure)
env.define(NodeCriticalMemory)
env.define(EvictPodsFromCordonedNode)
env.define(RescheduleEvictedPod)
env.define(CrashLoopDetection)
env.define(EvictCrashLoopPod)
env.reset()
# Assert cluster state
node1 = NodeAssert(
name="node-1", status="ready",
cpu_percent=45.0, memory_percent=97.0,
)
node2 = NodeAssert(
name="node-2", status="ready",
cpu_percent=30.0, memory_percent=60.0,
)
PodAssert(name="api-server-a", namespace="production",
node=node1, status="running")
PodAssert(name="worker-b", namespace="production",
node=node1, status="running")
PodAssert(name="cache-c", namespace="staging",
node=node2, status="running", restarts=12)
DeploymentAssert(name="api-server", namespace="production",
replicas=3, available=2)
DeploymentAssert(name="worker", namespace="production",
replicas=2, available=1)
env.run()
One call to env.run(). No loops. No polling. Here’s what the engine does:
Two independent chains fire from a single env.run(). The memory pressure chain cascades four levels deep: Alert, Cordon, Evict, Schedule. The crashloop chain runs in parallel, unaware the other chain exists. No rule called another rule. Each one reacted to facts.
Why This Is Different
Go back and read EvictPodsFromCordonedNode. Its docstring says “when a node is cordoned, evict its running pods.” Its patterns say the same thing in code. An on-call engineer, a product manager, or an auditor can look at that rule and understand exactly when pods get evicted. The intent is the code.
Now try answering “when do pods get evicted?” from the reconciliation loop. You’d need to trace two separate branches (memory pressure on line 54, crashloop on line 65), understand the nesting that gates each one, and mentally reconstruct the conditions. The intent is buried.
Each rule is also independently changeable. You can add a DrainBeforeCordon rule without touching NodeCriticalMemory. You can remove CrashLoopDetection without breaking eviction. Rules don’t call each other. They react to facts. EvictPodsFromCordonedNode doesn’t know who asserted the Cordon fact. It watches for Cordon facts, and when one appears, it fires. The chain happens because rules produce facts that other rules consume.
Each rule is testable in isolation: assert the facts it cares about, run the engine, check what it produced. No mocking the rest of the controller.
Truth maintenance in action
Remember logical() on the alert rules? Here’s what it buys you. After the initial run, node-1’s memory drops back to 60%. You update the fact:
node1.modify_slots(memory_percent=60.0)
env.run()
The memory pressure alert is gone. You didn’t delete it. You didn’t write cleanup code. You didn’t run a reconciliation pass to find stale alerts. CLIPS retracted it automatically because the fact that logically supported it changed and no longer matches the rule’s conditions.
In a Kubernetes operator, you’d need an explicit “clear alert when condition no longer holds” branch in your reconciliation loop. Every alert type needs a matching cleanup path. Miss one and you have phantom alerts that never resolve. With logical(), the cleanup is structural: change the cause, the effect disappears.
Where CLIPS Stops
CLIPS and clipspyx are remarkable for what they are. They’re also honest about what they aren’t.
No persistence. Facts live in memory. Restart the process and they’re gone. You need to reassert everything on startup.
No distributed execution. CLIPS runs in a single process on a single machine. You can’t spread rule evaluation across a cluster.
No streaming integration. Facts are asserted manually. There’s no built-in way to connect CLIPS to a Kubernetes watch stream, a message queue, or a database changelog.
No scale for large fact bases. The Rete algorithm works well for expert systems with thousands of facts. Other systems have improved on Rete with more efficient inference algorithms, but CLIPS wasn’t designed for millions of facts.
No multi-tenancy, no observability, no audit trail. These are production infrastructure concerns that a 1985 inference engine reasonably doesn’t address.
These aren’t flaws. CLIPS was built for a specific class of problems, and the ideas it proved out shaped everything that followed: that rules and data should live together, that pattern matching can replace control flow, that truth maintenance can replace manual cleanup. Every modern rule engine builds on these foundations.
Gary Riley has maintained and evolved CLIPS for over thirty years since NASA stopped funding the project. That’s extraordinary. The upcoming 7.0 release adds backward chaining (we’ll cover that in a future post). If you want to learn how rule engines think, CLIPS remains one of the best places to start.
So why isn’t everyone using rule engines already?
Because this architecture has never won commercially. The ideas were right in 1985 and they’re right now, but everything around them was wrong. Expert systems needed hardware that didn’t exist yet at reasonable cost. Distributing inference across machines required networks that were still too slow and too expensive. The data volumes that would make rule engines compelling at scale required databases that hadn’t been built. Companies shipped real products to real customers and still couldn’t make the economics work. By the time the AI winter hit, the industry had moved on to problems that felt more tractable.
What’s changed isn’t the ideas. It’s everything else. Hardware is orders of magnitude faster. Networks are cheap and ubiquitous. Databases handle billions of rows without breaking a sweat. Computer science has produced better inference algorithms. And the rise of AI agents creates exactly the kind of demand rule engines were designed for: systems that need to react to changing data, explain their decisions, and compose behavior from independent, inspectable rules. The timing problem that killed the first wave doesn’t exist anymore.
The foundations CLIPS proved out are sound. What was missing was a more capable engine built for today’s reality, and the infrastructure to run it at production scale. That’s a solvable problem now.
That’s at the core of what we’re building at Inferal. A new engine designed from the ground up for the world that exists now: declarative rules, pattern matching, forward chaining, and truth maintenance, with persistent fact storage, distributed evaluation, real-time data integration, and performance that scales elastically. The ideas NASA validated forty years ago, with an engine and infrastructure they deserve. If you’re ready to see what rule engines can do with today’s hardware, let’s talk.