New report: The AI Ship Date Dilemma. 132 enterprise AI teams on the gap between a working demo and production.New report: The AI Ship Date Dilemma. Get the report

Skip to main content

Autonomy as Infrastructure

The hardest part of running an AI agent is the infrastructure. It’s time to get it right the first time.

Angela McNeal - CEO & Co-Founder at Thread AI

September 29, 2026

The models are the easy part. Any company can sign an enterprise contract and have frontier AI running by the end of the week. The harder question comes right after: you have the models, now what?

At Thread AI, we ask this question at the start of every conversation with a new client. We often learn how companies have called model APIs, bought vertical-specific tools, and experimented with how to integrate AI into their work, yet most struggle to connect any of it to real business impact. That’s typically because a company’s most valuable asset, the expertise and IP it has built over years, can’t be captured in a simple prompt.

Moving beyond prompts, teams have turned to agents to deliver a more sophisticated solution, but challenges have appeared when they try to encode their expertise and complex logic and run it as a real workflow, feature, or product. We have watched this go wrong in two primary ways.

In the first, a team wires a model into an existing process and calls the result an agent. The agent performs well in a static environment with the constraints and guidelines it was prompted for, but if a sudden shift occurs, an agent does nothing, or worse, takes liberties and makes a bad call with confidence. Without embedding a protocol for deviations, an agent can go rogue.

In the second way we’ve seen, a team gives an agent the run of its systems because of a remarkable, controlled demo. Then a regulator or a customer’s security team asks a simple question: what did the agent do last Thursday, and why? Nobody can say. The logs show that a model was called, but not what it saw, what it decided, or who signed off. The team spends time troubleshooting or the agent is shut down, both leaving them without the productivity boost they originally sought.

Pure agent autonomy is a security incident waiting to happen, but pure control is just standard software. The companies that have gotten agents into trustworthy production at scale found a middle path, which we call Controlled Autonomy.

  1. Pure control

    Scripted

    Deterministic, capped by what you wrote down.

  2. Where Lemma sits

    Controlled Autonomy

    Agents decide inside bounds you can inspect.

  3. Pure autonomy

    Unbounded

    Fast to demo, impossible to defend in an audit.

Controlled Autonomy

Although control is often seen in opposition to autonomy, in our experience they’re intertwined.

The companies that have deployed agents successfully and safely did so by balancing the two extremes. The agent handles the routine, measurable share of a process on its own. At critical inflection points, the workflow stops and waits for a person. The person approves, corrects, or sends it back, and the run continues.

This is the core of everything we build at Thread AI. Lemma, our platform, has Controlled Autonomy baked into its architecture.

Lemma executes your agents, powered by whichever models you choose, gives each agent only the credentials it needs for only as long as it needs them, and when something fails at step fourteen, it resumes at step fourteen. You, the builders, define the boundaries, including when you want subject matter expert input, the actions the agent should and shouldn’t take, and the failsafe logic that brings additional assurance.

The platform keeps a record of every step, every human and non-human action, and every piece of data: what the agent saw, what it proposed, what a person changed, and who signed off. Over time, the judgment captured across runs becomes a record of how the company operates, which is then used to further refine the agent’s behavior.

Controlled Autonomy is an essential aspect of infrastructure and something that every company putting agents to work needs.

Where control has to live

Controlled Autonomy depends on boundaries the agent cannot avoid, hack, or override. Those boundaries are only as strong as the layer that enforces them, which is why they have to live in the runtime rather than in the prompt or the application code.

Most teams don’t build them that way. Their controls can be bypassed and won’t hold up to scrutiny. Take isolation, the guarantee that one customer’s data can never be reached from another customer’s session. If that guarantee depends on the model having been told not to look, it won’t survive a security review.

Agent frameworks are a common approach, but they struggle with the end-to-end controls. They supply a loop of prompts, tools, and model calls that make up a single agent execution for as long as the program running it stays alive. The builder is left to figure out what happens when that process dies mid-run, or when two hundred customers call the same agent at once, or when step nine needs a person’s approval and the person is out for three days.

Alternatively, teams may spend several quarters assembling the different components needed, and then have to manage them in perpetuity. Hyperscalers make this approach possible, but operability becomes the core challenge. Amazon’s Bedrock and AgentCore, Google’s Vertex, and Microsoft’s Copilot Studio supply every capability an agent needs, but companies still have to operate the distributed system underneath. As a result, teams often end up maintaining infrastructure instead of guaranteeing the capabilities of the end system.

We’ve researched the issues with assembled approaches in depth: in our new report, The AI Ship Date Dilemma, drawn from buying conversations with 132 companies, about half had weighed building the production layer themselves, and two-thirds of those ended up assembling it from multiple vendors instead.

In both approaches, teams are left operating pieces they weren’t hired to build, and all of it ends up on the on-call rotation. What teams actually need is the production layer itself, built once, enforced, and run by someone else.

Stop rebuilding the rails

We have seen this pattern before. Until about 2011, accepting a credit card online meant assembling the machinery yourself: a merchant account with a bank, a gateway to reach the card networks, and processes for security compliance, fraud checks, retries, and reconciliation. Then, a single layer took all of that behind one interface, and companies were able to keep their checkout pages without having to manage the plumbing underneath.

Identity, messaging, and cloud storage went the same way. Each time, the infrastructure every company needed, yet no one differentiated on, got pulled into a single layer everyone could call. Each time, that layer was built to be configured, not just consumed. Now, autonomy is becoming infrastructure, and Lemma is the layer that fills in the gaps.

Lemma is isolated like banking, metered like payments, and auditable like a ledger. Isolation, in particular, is a place where agents differ significantly from ordinary software. An agent does not follow a fixed path, so one customer’s information can leak into another’s session through the model’s context window, the vector store, caches, or memory. The same problem shows up in auditability and metering: when the path changes every run, you can’t anticipate what a run will cost or what the record will need to capture.

Lemma isolates each run, with tenancy layered from Thread AI to your company to each of your customers. Usage is metered to the token at every level, so you can price and bill for an agent inside your own product, and every run carries its own record of what happened and why.

None of this means giving up control. Lemma provides powerful defaults, but the rails are yours to shape: you decide where tenancy boundaries sit, what systems of record your audit trail flows into, and which models, tools, and data stores plug in. Lemma handles the infrastructure that every team would otherwise need to rebuild, and you customize what makes your agent yours.

One runtime, three kinds of production

At Thread AI, we see the same runtime do three kinds of work. What changes is whether the output is an internal process, a product customers buy, or a system deployed inside a secure boundary.

Operations is where we started. Teams connect Lemma to the systems they already run, put the routine share of a core process under their own protocol, route exceptions to a specialist, and write results back to their system of record. These processes focus on making businesses more effective and efficient at scale for whole teams, not for individual knowledge workers. The first workflow is typically in production in about a month, and each one after that builds on that foundation. Agents handle the workflow logic and volume, and specialists make the calls that require expertise.

Product is for companies whose customers are evaluating AI-native alternatives. For them, shipping an agent that customers can trust means building multi-tenancy, metering, credential handling, model routing, audit, and approval flows before the first feature ships. Instead of spending a year on the platform work, they run their products on Lemma under their own brand. Their customers decide how much autonomy to grant, and where their own people sign off.

Mission is for federal and defense programs that need to deploy agents on their own infrastructure, across complicated topologies. Lemma deploys inside those boundaries, in GovCloud, or in a disconnected enclave with self-hosted models, on the same build that serves commercial customers.

This breadth is deliberate. Our belief is that these pillars of Controlled Autonomy benefit every industry, no matter how regulated or complex. Lemma was built to the standard of the strictest customer, and every customer gets it.

Your judgment stays yours

When your IP runs as software on someone else’s platform, whose asset does it become?

On Lemma, that expert knowledge stays yours, whether you’re building a workflow or product. A Worker, our name for a complete workflow, is a versioned spec built on open-source standards, so it’s yours to keep even if you leave. And because Lemma routes across every major provider, the models underneath are always yours to change.

Over time, Workers capture a record of how a company operates: every time a person overruled a recommendation, added a condition, or approved something the model had flagged. Today, that knowledge lives in the heads of the people who made the calls, and it leaves when they do. On Lemma, these choices natively accumulate and can be fed back into agent logic and written out into your systems of record, not trapped in the platform. Even as the model, process, or product evolves, the records will remain.

Your judgment becomes more valuable and more capable of running autonomously the longer it runs on Lemma, not harder to take with you.

Now ship

We started with the question: you have the models, now what? For many of the companies we talk to, the answer so far has been months of work on the rails before an agent does anything real. Our report, The AI Ship Date Dilemma, found that only 5% of the companies we spoke with had committed to a ship date anyone outside the company could hold them to, and the work underneath the product is part of why.

The companies that pull ahead over the next few years will not be the ones with the best models. They will be the ones whose judgment is already running in production, under Controlled Autonomy, adding every day to a record no competitor can copy. Time spent building rails is time your expertise isn’t in production and compounding, and every quarter counts.

The models are bought. You bring the IP, and it stays yours.

You have the models. Now ship.