Skip to content

1. The 3 A.M. Call

At three in the morning, every minute feels expensive. On this call, it was. An important chain of systems inside a large enterprise had failed, and each hour of downtime risked millions in damage.

The domain experts could give us one concrete fact: production was on hold. Nothing in front of us showed the technical path that led there.

I did not know the affected systems, their interfaces, or the decisions behind them. That is normal in a large organization. An incident call brings the necessary knowledge into one place. This time, the knowledge we needed had not arrived, and the system could not supply it.

We were trying to reach the person who knew the affected system well enough to tell us where to look. Until then, the rest of the room inspected fragments and exchanged partial theories. The failure was serious, yet investigation depended on locating one human being in the middle of the night.

Hours passed.

Complex systems fail. Dependencies disappear, interfaces return surprising data, and earlier assumptions meet conditions their authors never saw. Reliability has to account for those events. What troubled me was the blindness that followed this one.

The organization had engineers, operators, escalation paths, and an active incident process. A responder still could not follow the execution far enough to identify the responsible operation. The scarce resource was context.

The system expert carried that context: which component mattered, what its interface expected, and which behavior pointed to the fault. The outage contained a technical failure and an explanatory failure. The second one kept a capable room idle while the first continued to cost money.

Incident reviews often answer this problem with more documentation. A better diagram, another runbook section, or a linked architecture decision can all help. Each also describes the system from a particular moment and viewpoint.

The running application keeps changing. A diagram can remain plausible after the execution path has moved. A runbook covers failures its authors anticipated. A ticket records work for the people who performed it. Under delivery pressure, teams update the artifact that makes the system run; secondary explanations depend on time and memory.

Source code stays close to the implementation, although it may still leave an unfamiliar operator to reconstruct configuration, infrastructure, and ownership across several repositories. Logs can name an error without locating it in the larger process. The architecture exists, scattered among code, deployment settings, dashboards, tickets, and people.

I wanted the executing system to carry more of that explanation. Its real operations and boundaries should remain visible. Runtime evidence should point to the operation that produced it. A responder should have a useful starting point before the original author joins the call.

Imagine the same incident represented as a visible Flow. The canvas shows the operations and their connections. Inputs and outputs have types. External calls appear as identifiable nodes. Run evidence belongs to the building block that emitted it.

One node shows the failure.

The graph cannot repair an unavailable dependency or choose the correct business response. It can change the first hour of investigation. The team opens the failed node, inspects what reached it, and follows the surrounding path. Types narrow the plausible causes. Per-node logs connect the runtime symptom to a specific operation.

Flow-Like’s run inspection is the feature that most directly answers that night. A run leads from recorded evidence back to the relevant node on the Board. The application stops looking like one undifferentiated failure.

Large programs also need a dense way to work. FlowScript lets a developer search, compare, review, and edit the same logic as text. During an incident, the canvas offers a map; during a broad code change, the text editor may be faster. Both act on the same Flow.

I believe that structure would have shortened the original investigation. That claim is a counterfactual, not a benchmark. The defensible requirement is smaller: a failure should leave enough connected evidence for a capable responder to know where to begin.

1.4 Domain knowledge belongs in the program

Section titled “1.4 Domain knowledge belongs in the program”

The call exposed another distance. Domain experts understood the operation, its exceptions, its regulatory setting, and the cost of getting it wrong. Conventional delivery often passes that knowledge through several groups before it becomes code. The people who supplied the rule may then struggle to inspect its implementation.

Simpler visual tools help until the application outgrows them. Teams often respond by placing general-purpose code inside a workflow block. The canvas still shows a reassuring rectangle while the real behavior moves behind its label.

My background in game development suggested a stronger visual model. Unreal Engine Blueprints had shown that a graph could be a genuine programming surface and a practical route into software. Enterprise workflows add another requirement: once the graph becomes large, experienced developers need text.

That combination led to the core design. A developer can use FlowScript for a large change. A domain expert can inspect the resulting process on the canvas. An operator can follow a failed run through the same nodes. Their expertise meets in one program.

The idea grew beyond the editor. Existing systems still have to connect to the logic. Applications need data, identity, permissions, versioning, deployment, and evidence. Flow-Like became the platform around those concerns; FlowScript became the text form of its Flow logic.

The standard remains concrete. When a Flow fails at three in the morning, a capable person who did not build it should be able to find a credible starting point before the system expert answers the phone.