
[Continuous Decision Intelligence (CDI)] combines evaluation, decision tracking, and policy enforcement into systems that improve decisions over time. … Bound together and run as a loop, these three elements turn a pile of fast, cheap, individually plausible agent decisions into a system whose decision quality can be measured, governed, and improved rather than one that is merely fast.
From When Agents Decide, pp. 185-186
I met Gene Kim in NYC at a Vibe Coding book signing last fall. We discussed how unclear the future of enterprise software development was, and specifically how this unique disruption would impact experienced engineers. This first conversation led to many more, with one result being, When Agents Decide: Continuous Decision Intelligence and the Substrate Beneath It, published in the Fall 2026 issue of the Enterprise Technology Leadership Journal. There are six other papers in the larger journal, which I highly recommend as timely, in-depth positioning for where our industry stands right now, and a valuable asset for anyone in a leadership role.
If you vibe with this paper and its thought process, you can hear more about it and my next steps at the Enterprise AI Summit in Charlotte, NC Oct 7–8, 2026.
I want to thank the paper’s co-authors, Joseph Enochs (who proposed the topic), Adrian Cockcroft, Jenn Spykerman, and Brian Wald for inspiring conversation and for lending uncommon industry vertical expertise, and Gene for creating the Journal platform and opportunity.
The paper opens at 5:47 a.m. with a CIO getting a call that overnight tariffs have hit and a key supplier has gone dark. The company’s agentic forecasting platform, built eighteen months earlier, is still running. It just answers questions from last week. The chief data officer calls it “a museum.”
The following articles are our attempt to explain why the museum happens and how to build dynamism instead.
The Model Is Rarely the Bottleneck
The paper’s central claim fits in one sentence: “The execution gap is not a gap in intelligence; it is a gap in structure.”
Most conversations about agent quality start with the model. Smarter model, better prompt, bigger context window. We argue the breakdown happens in the decision system around the model. Can the organization route a decision to the right place, act on it, record it, and learn from it before the next one arrives?
We call the layer that does this Continuous Decision Intelligence (CDI). It combines evaluation (asking whether a decision was any good, continuously, rather than once at launch), decision tracking (recording what was decided, on what basis, and under which policy), and policy enforcement (constraining what the system may do before it acts, not after).
Three Parts of the Substrate
Underneath CDI sits what we call the decision substrate:
- Decision topology is the map of what gets decided where, which decisions an agent may make alone, and where authority returns to a human.
- Provenance is the record of what shaped each decision: which sources, who owned them, and whether they could be trusted.
- Governance is the set of constraints applied before action and the accountability applied after.
Many automation efforts stall because leaders skip the difficult step of redefining a process around emerging capabilities. A team buys evaluation tooling and points agents at live decisions without ever drawing the map of who decides what. Without a topology you can’t govern decisions, and governance is required for production scenarios.
The decision topology has a cousin in Mik Kersten’s Output to Outcome, which our paper cites. His Outcome Tree cascades strategy and budget down through an organization and aggregates outcomes back up, with the same loop repeating at every level. CDI applies that self-similar idea to decisions.

R is reinforcing (like compounding interest) and B is balancing (like adhering to a budget).
Autonomy Is a Ladder
We describe four stages. At Stage 1 a human approves every step. At Stage 2 the agent executes within bounds and the human watches KPIs. At Stage 3 the human handles exceptions only. At Stage 4 the workflow is zero-touch, with a human called only at thresholds defined in advance.
Here is an example workflow I call “exceptional trace mining”. I am applying machine learning techniques to agent trajectories to build a reinforcing loop that improves an SDLC or other process. Doing this manually is the first step towards building automated decision-making policies.
The Unhappy Path
Almost none of the failures our paper documents are model failures. They happen in the context that feeds the model – recently I’ve heard this referred to as a “context layer” – and in the threshold that should stop it. My agentic GPS prototype is a step in that direction – towards better automation.
- Upstream context risk. Context gets poisoned at the source. For example, should agents have access to the internet? Stale facts get “laundered” through summarization (many harnesses are laser focused on helping with this), and agents quietly select a non-representative slice of sources. Brian Wald’s software development case proposes a Context Bill of Materials (CxBOM), attached to the merge request, that records what the agent read, who owns it, and how policy scored it. When something breaks, a team can reconstruct what the agent believed instead of guessing.
- Rubber-stamp oversight. A person nominally in the loop who approves everything gives you control on the org chart and none in practice. The paper says a guardrail that is never exercised is worse than none, because it manufactures false confidence.
- System dynamics. Every individual decision can be well governed and the system can still oscillate. Agents reacting to each other’s delayed signals behave like the bullwhip effect in the MIT Beer Game.
The Platform That Runs This
I wrote the developer platforms case. The CxBOM defines what to capture about a decision. My section is about the cost and machinery of capturing it at scale, and covers three pieces:
- Context factory. An MCP server that unifies enterprise tribal knowledge scattered across ticketing, corporate directories, and spreadsheets. It reconciles naming conventions and data formats so agents get structured, machine-parsable answers instead of hallucinating around gaps.
- AI tokenomics. With models commoditizing, the question stops being which model and becomes value per token, an idea I worked through in Tokenomics for Code. A value-based semantic router classifies each request and sends it to the cheapest model that still meets the service-level objective. The KPIs start to look like hard-goods manufacturing. See Value-based routing demo below.
- Hill-climbing harness. A continuous evaluation loop that uses judge-LLMs to interpret and reconcile results from multiple concurrent AI sessions, aggregated asynchronously. This is CDI running as part of the system’s runtime rather than a review after the fact.
The paper measures all of it with a short list of numbers: interrupt rate, interrupt category breakdown, interrupt recurrence, autonomous completion rate, MTTR after an interruption, and feedback-to-demo cycle time.
One change in framing I’ve made lately is that a rising interrupt rate is not automatically a failure. It can be the system correctly refusing to outrun its own evidence. I wrote about interrupts as the core metric earlier this year in Toward Zero Interrupts.
Key Takeaways
- Decision quality is cheap now. Absorbing decisions at agent speed is the scarce part.
- Draw the decision topology first. If you can’t say where decisions are made, by whom, and on what authority, evaluation tooling has no grounding.
- Record what the agent believed, not just what it did. Provenance turns post-incident analysis from guesswork into reconstruction, and being able to review and replay agent trajectories is an emerging engineering requirement.
- Earn each rung of autonomy with evidence. Interrupt rate and autonomous completion rate turn “are we ready?” into a trend line, that we review like any other operational KPI.
- Treat cost as a design constraint. Routing to the cheapest model that meets the SLO is what makes higher autonomy affordable at production volume, something covered thoroughly by Leo Cui here.
Download the paper and tell me where you think we’re wrong.
P.S.: Jev / Decision Models
One thing the paper doesn’t cover showed up in September: Jev, an early-access model from TypeSafe AI. Jev doesn’t generate text. You define the choices for each request (a team, a model, a next action) and it returns a decision with a calibrated confidence score. TypeSafe calls this a “System One” model and points it at classification, routing, scoring, and guardrailing LLM output.
TypeSafe publishes speed and cost figures, up to 200x faster than frontier LLMs on equivalent tasks. Those are vendor numbers on an early-access, closed model, so treat them as a hypothesis.
The idea is what interests me, because I’ve been prototyping the same shape of problem where teams are given a fixed token budget. The money can run out before “valuable work” gets done. My value-based routing prototype is a Rust filter in an AI gateway that classifies each prompt against a weighted rule set (nine rules today, defined in YAML and changeable on the fly) and picks the model worth its cost. Users can still override the choice.
In the demo, traces land in MLflow review queues, and reviewed traces become candidates for a reinforcement learning pipeline that tunes the classifier.
One rule in the set I’m most interested in is self-hosted offload. I think it’s one of the keys to getting more capacity out of a budget.
The prototype doesn’t use Jev. Its classifier is hand-written rules, and it calls for exactly the component Jev represents: a model built to make the routing decision, improved from reviewed traces. The idea also fits the CDI loop, because a decision that arrives with a confidence score is easier to track and model.
Disclosure: the Fall 2026 Enterprise Technology Leadership Journal is sponsored by GitLab. The paper is licensed under CC BY-NC-SA 4.0.
Sources
- When Agents Decide: Continuous Decision Intelligence and the Substrate Beneath It – Cockcroft, Eder, Enochs, Spykerman, Wald. Enterprise Technology Leadership Journal, Fall 2026
- Introducing System One Models & Jev – TypeSafe AI, September 2026