The Execution Layer: Why AI Predictions Are Just Entertainment
Everyone agrees on where AI agents are going. The destination is not the differentiator. The route is. Constraint profiles, a task vocabulary, and democratized creation.
I wrote five predictions about AI in finance. APIs will go extinct as agents learn to read documents. CFOs will stop caring about accuracy and start obsessing over recovery speed. Business documents will wake up and start thinking. ERP will die death by a thousand agents. Finance jobs will split into two camps with no middle ground.
The piece got picked up by the press. People shared it. Conferences referenced it. And I’ve come to believe it is, like nearly every AI prediction piece published in the last two years, directionally correct and operationally useless.
Not wrong. Useless. There’s a difference.
Here is what I mean. Within weeks of publishing, I started seeing the same predictions show up everywhere, wearing different clothes. Tom Tunguz wrote about how skills, programs written in English, are replacing traditional software interfaces entirely. “The web required a URL and a browser. Mobile required a download and a homescreen slot. Skills require a sentence.” That’s my API extinction prediction, restated as a distribution thesis. Aaron Levie at Box declared that file systems are becoming a core primitive for agents, that agents need the ability to work with computers, execute code, store off work, and manage data. That documents need secure spaces where agents and humans collaborate, with data access controls, security, and governance baked in. That’s my documents-as-agents prediction, restated as an infrastructure thesis. Miles Deutscher cited the K-shaped economy, workers with AI skills earning fifty-six percent more than the same job without them, that premium doubling in a single year, while ninety percent of workers haven’t taken a single hour of AI training. Goldman Sachs estimates three hundred million jobs will be affected by AI by 2028. That’s my workforce bifurcation prediction, restated as an inequality thesis.
Meanwhile, survey data shows forty-seven percent of organizations plan to build AI agents for financial planning and analysis in the next twelve months, rising to fifty-one percent among enterprises. Fifty-six percent plan research and reporting agents. The breadth of planned use cases signals a shift toward treating AI agents as enterprise-wide infrastructure rather than department-specific tools. That’s my CFO prediction, restated as an adoption curve.
Everyone agrees on where we’re going. The consensus is overwhelming. And when everyone agrees on the destination, the destination is not the differentiator. The route is.
The route is what I want to talk about.
The Prediction Industry Has a Production Problem
Here is a number that should haunt every person who has ever published a prediction about AI: ninety-five percent of enterprise generative AI pilots have produced zero measurable return on the P&L. That’s not a guess. It’s from MIT’s NANDA initiative, based on 150 leadership interviews, 350 employee surveys, and 300 public AI deployment analyses. Five percent of pilots achieved rapid revenue acceleration. The other ninety-five percent produced nothing.
Gartner followed up by predicting that over forty percent of agentic AI projects will be cancelled by the end of 2027. Not paused. Cancelled. The reasons they cite are not technical. They are escalating costs, unclear business value, and inadequate risk controls. Gartner also estimates that only about 130 of the thousands of “agentic AI” vendors are real. The rest are engaging in what they call “agent washing,” rebranding chatbots and RPA as agentic without adding substantive new capabilities.
Deloitte’s 2026 State of AI report found that only twenty-one percent of companies have a mature governance model for autonomous agents. Only eleven percent have agents running in actual production. Seventy-four percent of organizations that invested in AI haven’t seen real value from those investments.
PwC quantified the expectations gap directly: seventy-nine percent of executives expect AI to significantly contribute to revenue by 2030. Only twenty-four percent can clearly see where that revenue will come from. That fifty-five-point spread between expectation and visibility is the prediction-to-execution chasm expressed as a single number.
And here’s the irony that should make every prediction-writer squirm: HBR published research in July 2025 showing that executives who used ChatGPT to make forecasts became significantly more confident in their predictions, and significantly more wrong, than those who simply discussed with their peers. The prediction-industrial complex is literally making decision-makers worse at predicting.
These are not technology failures. The models work. Tool calling has crossed ninety percent accuracy, up from below fifty percent two years ago. The capability threshold has been passed. What hasn’t been built is the layer between “this is possible” and “this runs in production.” The industry has spent years perfecting the what. Almost nobody has built the how.
The Thought Leadership Preparation Trap
There’s a pattern I’ve written about before that I call the Preparation Trap, the tendency for organizations to waste time perfecting their data, cleaning their processes, and building elaborate governance frameworks before deploying a single AI agent. It’s thinking borrowed from civil engineering, where change costs are high, applied to software, where iteration is cheap.
The thought leadership industry is caught in its own version of this trap. We endlessly refine predictions, the what, while avoiding the messy, iterative, unglamorous work of building execution primitives, the how. Predictions are high-status work. They get keynotes. They get press coverage. They get shared on LinkedIn. Execution primitives get Jira tickets.
HBR published a piece in August 2025 called “Beware the AI Experimentation Trap” that nailed this dynamic. They argued that leaders are repeating the mistakes of the digital transformation era, funding scattered pilots with a “let 10,000 flowers bloom” approach, hoping a few experiments produce outsized returns. The result, then and now, was a morass of unfocused, under-resourced teams that produced few scalable results.
MIT Technology Review declared 2025 “The Great AI Hype Correction.” They wrote that AI leaders made promises they couldn’t keep, that generative AI would replace white-collar work, usher in an age of abundance, make scientific discoveries. The correction isn’t that AI doesn’t work. It’s that predictions without execution are entertainment.
Every technology wave follows this arc. Cloud computing had “everything will move to the cloud” predictions for years before AWS published the Well-Architected Framework that gave teams an actual execution playbook. DevOps had “ship faster, break less” as a mantra long before the DORA metrics created shared accountability for deployment frequency, lead time, change failure rate, and recovery time. Digital transformation had a decade of “digital will disrupt everything” predictions before platform strategies and maturity models turned the rhetoric into operational discipline.
AI agents are deep in the prediction phase. What’s missing is the equivalent of what DORA gave DevOps and the Well-Architected Framework gave cloud, a set of execution primitives that turn vision into operating discipline.
I believe there are exactly three.
Primitive One: Formal Constraint Profiles
Every prediction about autonomous agents implicitly assumes boundaries. “Invoices will negotiate their own terms” assumes we know what terms the invoice is allowed to negotiate. “AI agents will handle collections end-to-end” assumes someone has defined what “end-to-end” means and where the agent stops. But almost nobody is building the specification layer that makes these boundaries explicit, auditable, and machine-readable.
Aaron Levie gets closest to naming this when he writes that agents need “effective data access controls, security, and governance” as they collaborate with humans and other agents. He’s right. But he describes these as properties of the file system, the container. I’d argue they’re properties of the agent itself. The governance doesn’t live in the platform. It travels with the agent.
Right now, agent governance exists primarily as prose. Policy documents written in natural language. Slide decks describing “human-in-the-loop” without defining what that means for a specific agent executing a specific task at a specific confidence level. Only ten percent of organizations have a strategy for managing autonomous systems, according to recent research, even as non-human identities are projected to exceed forty-five billion by the end of 2026.
What’s needed is a constraint profile, a structured specification that travels with each agent and defines its operational boundaries. Not a governance platform. Not a centralized control plane. A portable, auditable card that answers a set of non-negotiable questions. What is this agent’s autonomy level? What confidence threshold triggers human escalation? What data can it access? What actions can it take without approval? What is its cost budget per execution? What audit trail does it produce?
The industry is starting to grope toward this. Singapore launched the world’s first state-backed governance framework for agentic AI in January 2026 at Davos. It defines four dimensions: bounding risks upfront, making humans accountable at defined checkpoints, implementing technical controls throughout the lifecycle, and enabling end-user responsibility through transparency. The financial services industry has the FINOS AI Governance Framework v2.0, built by institutions like Morgan Stanley, which catalogs eleven operational risks, nine security risks, and a dedicated agentic AI risk catalogue covering multi-agent trust boundary violations, agent action authorization bypass, and agent state persistence poisoning.
These are good starts. But they’re still frameworks, guidance for humans to interpret and apply. A constraint profile is different. It’s a configuration, not a conversation. Without constraint profiles, agent governance is a meeting. With them, it’s a deployment artifact.
This connects to something I’ve argued before about the Autonomy Tax, the idea that reversibility matters more than accuracy when evaluating AI systems. A system with higher error rates but faster recovery can be more cost-effective than a highly accurate system with slow error correction. The constraint profile operationalizes this. It defines what “reversible” means for a specific agent: which actions it can undo, which require approval before execution, and which are prohibited entirely. You don’t grant autonomy by writing a policy. You grant it by modifying the constraint profile.
Consider the data point: fifty-one percent of enterprises plan to build agents for financial planning and analysis in the next twelve months. FP&A. Where forecasts directly influence capital allocation, headcount decisions, and board reporting. These aren’t chatbot experiments. These are agents operating in high-consequence territory where a wrong forecast can cascade into wrong hiring decisions and wrong investment theses. Deploying FP&A agents without constraint profiles, without explicit specifications for what the agent can modify, what confidence level triggers human review, and what decisions remain human-only, is not innovation. It’s negligence.
Consider the progression from assistant to autonomous agent, what I call the four-stage evolution: Assist, Automate, Advise, Configure. An agent doesn’t jump from Assist to Configure overnight. It earns autonomy. And the mechanism for earning that autonomy is the constraint profile. At the Assist stage, human_loop is set to “approve every action.” At Automate, it shifts to “approve exceptions.” At Advise, the agent recommends actions and the profile defines which recommendations can self-execute. At Configure, the agent modifies its own workflows within bounds defined by, you guessed it, the constraint profile.
The CIO recently published a piece describing something they called “Service Passports,” identity credentials that define an agent’s scope, budget, human manager, and probation status. That’s directionally right. But a passport tells you who someone is. A constraint profile tells you what they’re allowed to do, how much they can spend doing it, and who gets called when something goes wrong. The distinction matters when you have a thousand agents.
Primitive Two: A Task Vocabulary for Agents
Here is a sentence I hear constantly: “Build an agent that handles collections.”
That is not a requirement. It’s a prediction disguised as a requirement. It describes an outcome without decomposing it into buildable, testable, auditable units. And it’s the dominant way that business stakeholders communicate agent requirements to technical teams.
The problem is that “handles” is doing an enormous amount of work in that sentence. Does the agent retrieve outstanding balances? Classify accounts by risk? Rank them by likelihood of payment? Generate communications? Verify that the communication complies with regulatory requirements? Each of these is a discrete task with its own inputs, outputs, accuracy requirements, and failure modes. Lumping them into “handles collections” guarantees that the resulting agent will be untestable, ungovernable, and, based on the industry statistics, cancelled within eighteen months.
Tom Tunguz’s “Can You Fly That Thing?” essay illuminates what a real execution vocabulary looks like, even if he doesn’t frame it that way. He describes skills, programs written in English that encode institutional knowledge in executable form. An FP&A team doesn’t get “an AI tool.” They get an optimized budget variance skill that pulls NetSuite data and formats reports to CFO specifications. That’s decomposed. That’s specific. That’s testable. The skill isn’t “handle financial planning.” The skill is “retrieve budget data from NetSuite, calculate variance by department, format to the CFO’s report template, flag variances exceeding threshold.” Tunguz is describing a task vocabulary without calling it one.
And his insight about platform compression applies directly. Each technology cycle reduces friction: the web required a URL and a browser, mobile required a download and a homescreen slot, skills require a sentence. But that compression only works if the sentence is precise. “Handle my collections” is a sentence. It’s not a skill. “Retrieve outstanding balances, classify by risk tier, rank by payment probability, generate tiered communications, verify compliance, execute delivery,” that’s a skill. The vocabulary is what makes the compression useful rather than dangerous.
What’s needed is a standard vocabulary for describing what agents do. Not prose. A grammar. Something like: retrieve, classify, rank, simulate, generate, verify. Each verb has a defined meaning. Each agent gets a task sequence built from these verbs. The collections agent becomes an explicit chain: retrieve outstanding balances, classify accounts by risk tier, rank by payment probability, generate communication per tier template, verify regulatory compliance, execute delivery. That’s not a prediction. It’s an executable specification.
This matters for three reasons.
First, you cannot test what you cannot name. If the agent’s job is to “handle collections,” how do you write a test? If the agent’s job is to “classify accounts by risk tier using payment history, contract terms, and dispute patterns,” you can measure precision and recall. You can set thresholds. You can improve.
Second, you cannot govern what you cannot decompose. A constraint profile needs to specify what the agent does. If the vocabulary for what agents do is ad hoc and vendor-specific, as it is today, every constraint profile speaks a different language. Every registry defines capabilities differently. Every governance review starts from scratch.
Third, you cannot democratize what you cannot describe. If only developers understand what an agent does because they wrote the code, then only developers can build, modify, and evaluate agents. A shared task vocabulary lets a finance operations manager say “I need an agent that retrieves, classifies, and generates, but a human verifies before execution.” That’s a constraint profile written in a task vocabulary by someone who has never written a line of code. That’s democratization.
The closest thing to this today is the 12-Factor Agents framework, principles for production-ready LLM applications inspired by Heroku’s 12-Factor App methodology. Its core philosophy is right: “Most AI agents that actually succeed in production aren’t magical autonomous beings, they’re mostly well-engineered traditional software, with LLM capabilities carefully sprinkled in at key points.” But 12-Factor Agents is developer-facing. It codifies ownership, own your prompts, own your context window, own your control flow. It doesn’t provide a vocabulary that non-technical stakeholders can use to describe what an agent should do.
Anthropic published a guide on context engineering in September 2025 that identified strategies for managing agent context, write it, select it, compress it, isolate it. That’s useful for engineers building agents. It’s not useful for CFOs specifying what agents should do in their operations.
The gap is the vocabulary. And the vocabulary is what makes everything else, constraint profiles, registries, governance, democratization, composable.
Primitive Three: Democratized Agent Creation
Every prediction about agent proliferation contains an unstated assumption about who builds the agents. When I wrote that ERP would die death by a thousand agents, I was describing an end state with hundreds or thousands of specialized agents replacing monolithic systems. Someone has to build those agents. If the answer is “the same developers who built the last generation of software,” the prediction fails on supply-side economics alone.
This is not theoretical. Eighty percent of Fortune 500 companies already have employees building AI agents using low-code and no-code tools. Microsoft Copilot Studio, ChatGPT Enterprise, and open-source platforms have made agent creation accessible to anyone with a browser. The barrier to entry has collapsed.
And it’s created a governance nightmare. Ninety percent of IT directors and executives are concerned about shadow AI from a privacy and security standpoint. Nearly eighty percent have already experienced negative AI-related data incidents. Thirteen percent report those incidents caused financial, customer, or reputational harm. Employees are deploying agents faster than organizations can govern them, agents that respond to customer inquiries, approve transactions, and initiate workflow changes without centralized oversight.
This is shadow IT from a decade ago, except the shadow actors are autonomous.
Tunguz flagged the dark side of this acceleration. A recent analysis of 4,784 AI skill repositories uncovered embedded malware, credential harvesting, monitoring backdoors. When anyone can build and distribute a skill, anyone includes bad actors. The execution layer’s job is not to slow this down. It’s to make the fast path safe.
The instinct is to lock it down. Build a centralized control plane. Restrict agent creation to approved teams using approved tools. This is the response that feels safe and is structurally wrong. Centralized agent creation recreates the same bottleneck that centralized IT created, a queue of business needs waiting for technical capacity. When the prediction says “a thousand agents,” the question isn’t whether a centralized team can build a thousand agents. It’s whether you can create the conditions for a thousand agents to be built safely by the people closest to the problems they solve.
This is where the three primitives work as a system. Democratization without constraint profiles produces shadow AI, autonomous agents operating without boundaries, creating risk nobody can see. Constraint profiles without a task vocabulary produce governance that nobody can apply consistently, every team describing agents differently, every review starting from zero. And constraint profiles plus task vocabulary without democratization produces an elegant framework that only two people in the organization can use.
The execution layer is all three, working together. The task vocabulary gives everyone a shared language for describing what agents do. The constraint profile gives every agent auditable boundaries. And democratization gives everyone the ability to build agents within those boundaries using that language. Remove any one primitive and the system fails.
The Execution Layer Is the Moat
There’s a question hiding inside all of this that matters more than the operational details: where does competitive advantage live in an agent-driven world?
It’s not in the models. Models are commoditizing. Frontier models become orchestrators, routing to specialized agents. The model layer is a race to the bottom on price and a race to the top on capability that benefits everyone equally.
It’s not in the predictions. Everyone agrees on the destination. The prediction consensus is public knowledge.
It’s in the execution layer. A constraint profile encodes organizational knowledge, risk tolerance, regulatory requirements, escalation patterns, cost thresholds. That’s institutional context made machine-readable. A task vocabulary encodes domain expertise, what work actually consists of, decomposed into testable units. That’s operational knowledge formalized. Both are context artifacts. And as SaaS commoditizes under AI pressure, value migrates to whoever owns context.
The organizations that build execution layer primitives early will have something their competitors cannot easily replicate: not a model advantage, not a prediction advantage, but a structural advantage in how they design, deploy, and govern autonomous systems. The playbook exists before the agents do. The constraints are defined before the capabilities are tested. The vocabulary is shared before the first line of agent code is written.
Every previous technology wave followed this pattern. Cloud’s competitive advantage didn’t go to the companies that predicted cloud adoption. It went to the companies that adopted the Well-Architected Framework early and built operational discipline around it. DevOps advantage didn’t go to the companies that talked about shipping faster. It went to the companies that measured DORA metrics and improved them systematically. The advantage went to the executors, not the predictors.
The Inversion
I started by saying my own predictions were directionally correct and operationally useless. That’s not false modesty. It’s a structural observation about where value lives.
The thought leadership industry, and I’m including myself in this, has spent three years publishing increasingly convergent predictions about AI agents. Documents will become agents. APIs will dissolve into skills. CFOs will manage agent fleets instead of spreadsheet jockeys. The workforce will bifurcate into machine teachers and strategic thinkers. These predictions are correct. They are converging from every direction, from Tunguz’s distribution thesis, from Levie’s infrastructure thesis, from the survey data showing forty-seven percent of organizations already planning FP&A agents, from the economic data showing AI-skilled workers earning fifty-six percent premiums while ninety percent of the workforce hasn’t taken a single hour of training. The K-shaped economy isn’t a forecast anymore. It’s a measurement.
But the predictions are also the easy part. The hard part is building the execution primitives that make the predicted future actually arrive, and arrive in a way that doesn’t cancel itself through governance failures, untestable agents, and shadow AI chaos.
Constraint profiles that make governance a configuration instead of a conversation. A task vocabulary that makes agent capabilities describable, testable, and governable. Democratized creation that lets the people closest to problems build agents within auditable boundaries. These are not implementation details beneath the dignity of thought leadership. They are the thought leadership. They are what separates a prediction from a product, a keynote from an operating system, and entertainment from execution.
Whoever builds the primitives defines the industry. Not whoever predicts the destination first.
The predictions are the deck. The execution layer is the code.