Why I Built an LLM Orchestration Layer in Rust Instead of Python
October 6, 2026 · 9 min read
Why I Built an LLM Orchestration Layer in Rust Instead of Python
By Matt Busel
I used to think the hard part of AI infrastructure was the model.
It is not.
The model is increasingly becoming the easiest part of the stack to replace.
OpenAI today. Anthropic tomorrow. A local model next week. A completely different architecture six months from now.
The real problem starts when you try to build everything around the model.
Queues. Retries. Rate limits. Tool calls. Memory. Streaming. Backpressure. Provider failures. Duplicate requests. Agent coordination. Observability. Persistence. Distributed execution.
Once you start running more than a few agents at the same time, the model stops being the system.
The infrastructure becomes the system.
That realization is why I started building my AI infrastructure in Rust.
More specifically, it is why I built Tokio Prompt Orchestrator.
Tokio Prompt Orchestrator on GitHub
The project started from a simple question:
What happens when AI systems stop being single prompts and start becoming persistent, concurrent software systems?
Most AI applications begin with something like this:
response = client.messages.create(...)
There is nothing wrong with that.
Python is incredible for experimentation. It lowered the barrier to machine learning, data science and AI development more than almost any language ever created.
But that simplicity hides the infrastructure problem.
One request becomes ten.
Ten requests become ten agents.
Those agents call tools.
The tools call other services.
One agent retries.
Another agent sends the same request.
A provider starts returning 429s.
Another provider has a temporary outage.
Memory consumption starts climbing because a queue somewhere is not bounded.
One slow downstream service backs up everything behind it.
Suddenly your AI application is not an API call anymore.
It is a distributed system.
And distributed systems have very different requirements from notebooks.
AI infrastructure needs backpressure
One of the first things I wanted was a pipeline that physically could not grow forever.
Tokio gives me a very natural model for this.
Each processing stage can exist as its own asynchronous task connected through bounded channels.
A request enters the pipeline. It moves through controlled stages. Each stage has a finite amount of capacity.
If the system is overloaded, the architecture has to make a decision. It cannot quietly consume more and more memory while pretending everything is fine.
That sounds obvious. It is also one of the easiest mistakes to make in real AI systems.
A producer can often generate work faster than a model provider, database, vector store or downstream service can consume it. Unbounded queues turn temporary load into permanent instability.
So Tokio Prompt Orchestrator treats backpressure as part of the architecture rather than something you bolt on after production starts failing.
When capacity is exhausted, the system can shed work deliberately and record why it happened.
That is a much better failure mode than watching memory usage climb until the operating system makes the decision for you.
The same prompt should not cost twice
Then there is deduplication.
This sounds almost trivial until you start running agent systems.
Agents repeat themselves constantly. Users repeat requests. Retries duplicate work. Separate workers arrive at the same question. Multiple parts of a system request the same information at almost the same time.
Every one of those can become another model call. That means more latency, more tokens, more money, and more pressure on provider rate limits.
The orchestrator has deduplication directly in the request path.
In one of the basic examples, twelve requests from four users collapse into three model calls because nine of the requests are repeats.
Twelve answers. Three model calls.
That is infrastructure doing useful work before the model ever gets involved.
I eventually pushed the same idea further with semantic deduplication.
Exact matches are easy. Real users do not always ask exact matches.
'What is the capital of France?' and 'Which city is the capital of France?' are different strings describing the same request.
So the orchestrator can optionally use local embeddings to identify semantically similar prompts and reuse a previous result.
But there is a dangerous side to semantic similarity.
Embeddings can tell you that two sentences are extremely similar even when a tiny difference completely changes the answer.
'Convert 10 miles to kilometers.' and 'Convert 10 kilometers to miles.' can have extremely high embedding similarity. They absolutely should not share an answer.
That means AI infrastructure cannot just be clever. It has to be conservative.
The system checks more than vector similarity before it reuses a semantic result. Numbers and word relationships matter too.
I would rather pay for one unnecessary model call than confidently return the wrong cached answer.
That principle runs through most of the system. Reliability beats cleverness.
Provider failures should remain provider failures
Another lesson becomes obvious when you run agents for long enough.
External APIs fail. That is not exceptional behavior. It is normal behavior.
Connections drop. Providers return 429. Services return 503. Keys expire. Latency spikes. Models become unavailable.
A temporary provider problem should not be allowed to become a system-wide problem.
This is why the orchestrator has circuit breakers.
If a provider begins failing consistently, the system can stop throwing more traffic at it. Requests fail quickly and clearly instead of sitting in an ever-growing pile waiting for a service that is already unhealthy.
After a recovery period, the breaker can allow a probe through and determine whether the provider is healthy again.
Retries exist for the opposite case. Some failures really are temporary.
A brief connection failure or rate limit might succeed moments later. So retry policies can use randomized exponential backoff and respect provider retry instructions.
But not every error should be retried.
A bad API key does not become a good API key if you send it five more times.
The infrastructure needs to understand the difference.
This is the kind of work that becomes increasingly important as AI agents move from demos into systems that are expected to stay running.
I wanted the infrastructure to be model agnostic
One thing I strongly believe is that model providers will continue changing rapidly.
Infrastructure should survive that.
Tokio Prompt Orchestrator can sit in front of Anthropic, OpenAI, llama.cpp, vLLM and compatible backends.
It can also act as a drop-in proxy.
An application already using an OpenAI client can point its base URL at the orchestrator and gain infrastructure around the request without rewriting the entire application.
The same concept works with Anthropic-compatible traffic.
That separation matters.
Your application should not have to understand how circuit breakers work. Your agent should not care which process is managing deduplication. Your frontend should not know how the dead-letter queue works.
Those are infrastructure concerns.
The model call should remain a model call. The system around it should make that model call reliable.
Local models changed how I thought about the architecture
Local inference makes this even more interesting.
Once you can route requests between hosted APIs and local models, AI architecture starts looking less like a web application and more like compute infrastructure.
You can have expensive models handle high-value reasoning. Smaller local models can handle simpler requests. Different workers can have different latency and cost profiles. Some workloads may never need to leave the machine.
The infrastructure should not care.
That is another reason Rust and Tokio make sense to me here.
I want the orchestration layer close to the metal. I want explicit concurrency. I want predictable memory behavior. I want thousands of asynchronous operations without designing the entire system around threads. I want the compiler helping enforce assumptions that would otherwise become production incidents.
Rust is not valuable here because a benchmark says one loop executes faster.
That misses the point.
Rust is valuable because infrastructure eventually becomes a problem of controlling complexity.
The next layer is not prompts. It is agents.
Once the orchestration layer was working, the obvious next question was: What sits on top of it?
That led to my second project, Agent Runtime.
Agent Runtime is a Tokio-native runtime for building actual LLM agents rather than treating every interaction as an isolated completion.
And that distinction matters.
An agent is not just a prompt plus a tool.
A useful agent needs state. It needs memory. It needs failure handling. It needs tools. It needs a lifecycle. It may need to communicate with other agents.
It needs to remember enough to remain useful without allowing memory to grow forever. It needs to recover from errors without turning one broken tool into a broken system. It needs observability because eventually something will behave differently from what you expected.
This is where I think a lot of current AI development is heading.
The model is becoming one component inside a much larger runtime.
Memory should be infrastructure too
Agent Runtime includes multiple forms of memory.
There is working memory for bounded short-term state. There is episodic memory for experiences associated with an agent. There is semantic memory that can retrieve related information using vector similarity. There is long-term memory with decay and consolidation.
There is memory compression because eventually every persistent agent runs into the same problem: You cannot keep everything forever.
This is a much deeper problem than stuffing the last fifty messages into a context window.
Persistent agents need memory policies.
What should be remembered? What should decay? What should be consolidated? What deserves to consume context? What can be safely forgotten?
Those are runtime questions. Not prompt-engineering questions.
Knowledge is not always a list
I also wanted agents to reason over relationships rather than treating everything as isolated chunks of text.
So the runtime includes a graph store.
That means agent state can represent connections between entities and traverse them using graph operations.
BFS. DFS. Shortest paths. Transitive relationships. Centrality. Communities. Cycles. Subgraphs.
The point is not to turn every agent into a graph database.
The point is that intelligence often depends on relationships.
A flat list of memories loses structure. Graphs give the runtime another way to represent what an agent knows.
Tool calling needs guardrails
Tool use is where agents become genuinely useful. It is also where they become dangerous from a software reliability perspective.
The runtime supports ReAct-style loops, but it can also use native structured tool calling.
Tools have schemas. Arguments can be validated before execution.
If the model generates an invalid call, that error can be returned to the model rather than executing malformed input and hoping for the best.
Individual tools can have their own failure handling.
MCP tools can also be pulled into the runtime.
That means an MCP server can expose capabilities and those capabilities become available to an agent through the same runtime model.
This is the direction I think agent infrastructure has to move.
Tools should not be random functions scattered throughout an application. They should be governed resources.
Multi-agent systems need actual coordination
Then you get to multiple agents. This is where the architecture becomes much more interesting.
If one agent is useful, the obvious temptation is to launch twenty.
But twenty independent agents are not automatically a system.
Without coordination, you just created twenty sources of race conditions, duplicate work and conflicting decisions.
Agent Runtime includes an asynchronous message bus so agents can communicate.
Teams can use different topologies. Star. Mesh. Ring.
Different coordination strategies make sense for different problems.
Some tasks want a central coordinator. Some want peer interaction. Some want parallel work followed by consensus. Some want pipeline execution where one agent's output becomes another agent's responsibility.
Again, the important part is that these become explicit infrastructure concepts. Not prompt conventions.
The compiler should catch what it can
One small design decision represents a much larger reason I like Rust.
The Agent Runtime builder uses typestate.
If required configuration is missing, I would rather have the program fail to compile than discover the mistake halfway through an agent session.
That is the Rust philosophy I want in AI infrastructure.
Move failures left. Make invalid states harder to represent. Use the type system where possible. Use bounded structures where possible. Fail clearly when something cannot be recovered.
AI itself is probabilistic.
That does not mean every layer surrounding AI also needs to be probabilistic.
In fact, I think the opposite is true.
The less predictable the model becomes, the more predictable the infrastructure around it needs to become.
AI models are getting smarter. Infrastructure has to catch up.
The industry is obsessed with model intelligence.
Every few months we get better reasoning. Larger context windows. Faster inference. Better multimodality. More capable coding models.
That progress is real.
But every increase in model capability makes the infrastructure problem larger.
A model that can do more will be asked to do more.
More tools. Longer sessions. More autonomous work. More concurrent agents. More memory. More external systems. More decisions without a human sitting inside every loop.
At that point, reliability stops being an optimization.
It becomes the product.
I do not think the future of AI infrastructure is thousands of Python scripts wrapped around increasingly intelligent models.
Python will remain enormously important for research, experimentation, training and data science.
But there is another layer forming underneath production AI.
A layer that looks much more like traditional systems engineering.
Schedulers. Queues. Runtimes. Memory systems. Routers. Streaming systems. Observability. Distributed coordination. Resource governance. Persistent execution.
And that is where I think Rust becomes extremely difficult to ignore.
Not because Rust is fashionable. Not because Python is bad. Because the problem is changing.
AI started as model research. Then it became APIs. Now it is becoming infrastructure.
The next generation of AI systems will not simply ask models better questions.
They will coordinate fleets of models, agents, tools, memory systems and compute resources continuously.
Someone has to build the layer that keeps all of that from collapsing under its own complexity.
That is the layer I am interested in.
The model can think. The infrastructure has to make sure everything around it survives.