OrchestrAI
Give it a goal and parallel AI research agents turn it into a plan, drawn as a live knowledge graph you can inspect, approve and steer.





Finished agents report cost, tokens and time, and their findings attach with confidence bars.
I built an agent orchestration platform to answer a design question: how do you stay in control of a system that does a lot of work while you watch? You give OrchestrAI a goal. It splits the goal into research questions, sends agents after them in parallel, synthesizes what they find and drafts a phased plan. All of that shows up as a knowledge graph you can inspect, approve and redirect. The engine has 529 passing tests. The screens below are the real web app replaying a real run of that engine, with the model's text canned so it runs without API keys.
The Problem
Turning a big idea into a plan takes a lot of research, and most AI tools make you pick between two bad options. A single agent is easy to follow but shallow. A swarm of agents goes deep but turns into a black box that spends money while you wait. I wanted the depth of many agents with the legibility of one, where you can see what each agent is doing, why, and at what cost, and stop it at any point.
The mental model: give the system a North Star like "Launch a design system for a fintech startup". It decomposes the goal into questions, researches them in parallel, synthesizes the findings, spots conflicts and gaps, and drafts a roadmap. Every step of that is visible and steerable.
From a Goal to a Researched Plan

Here's the run from start to finish, including the steering covered below:
Goal to graph to steering: agents run, a finding is inspected, a pattern switch approved and a follow-up rejected with a reason (44s)
Why a Graph
I started with a list. It didn't work, because research isn't linear. A market insight informs an architecture decision, which then runs into a compliance requirement. Those links are the most important part of the plan, and a list buries them.
The graph uses React Flow with 10 node types (North Star, research question, finding, decision, epic, task, file, agent, error and user note) and 7 edge types (for example informs, blocks and derived from). Every task in the final plan traces back through the findings that justified it. Agent nodes carry their model, cost, tokens, tool calls and time, so cost is visible right where the work happens.
Human in the Loop Is a Design Problem
Building approval gates is easy. Deciding when to interrupt someone is the hard part. Interrupt too often and the system is useless. Interrupt too rarely and nobody trusts it. I designed three layers:
- Plan and Act modes. Every run starts in Plan, which means read-only research. Nothing changes files or runs commands until you switch to Act, so you can explore freely.
- Approval gates. Plan starts, phase starts, budget increases, pattern switches and external actions can each require a person. Each request shows what it wants to do, why, what it costs and how risky it is. Low-risk ones can auto-approve.
- Stop points. The engine can pause on tool calls, agent spawns, budget thresholds, low confidence, errors or checkpoints, and wait to be told to continue, cancel, modify or retry.
The detail I'm proudest of is that a rejection takes a reason, and the reason steers the run. Rejecting a follow-up agent with "defer ROI research to the adoption-dashboard task" doesn't just stop it. The reason becomes a note node in the graph, so the next agent and the next person can both see why the plan changed.

Five Ways to Organize Agents
Different problems need different team shapes, so the engine has five orchestration patterns:
- Supervisor: one lead agent breaks down, delegates and synthesizes. This is the default.
- Map-Reduce: parallel work with a final aggregation, for independent questions.
- Pipeline: sequential stages that hand work down the line.
- Consensus: several agents vote on a decision.
- Hierarchical: a multi-level delegation tree.
The engine can switch patterns mid-run when budget, failure rate, complexity or scale changes. A switch can be gated behind an approval, so a person signs off first, as with the Supervisor to Map-Reduce request above.
Cost You Can See
Multi-agent systems burn tokens fast, so budget is a first-class screen, not a setting. Spend is tracked per call, per agent and per model. Alerts fire at 50%, 80% and 95% of the limit, the run can pause at the limit, and the router steps down to cheaper model tiers at 80% and again at 95%. Providers are listed in priority order with automatic failover: Anthropic, Bedrock, Vertex, OpenRouter, Ollama and LiteLLM.
Architecture
| Layer | Technology | Purpose |
|---|---|---|
| Frontend | Next.js 16, React Flow, 8 Zustand stores | Dashboard, graph, approvals, budget |
| Real-time | Socket.io | Streams graph and agent events |
| Engine | TypeScript, 17 modules, ~21k lines | Orchestration, research, planning, budget |
State is deliberately low-infrastructure. Plans, roadmaps and memory are Markdown files committed to Git, so they're human-readable and diffable and need no database.
What I Learned
Before the graph I thought of orchestration as a pipeline. Seeing findings form a web of dependencies and contradictions changed the design of the engine, not just the UI.
My first run without limits cost $47 in six minutes. Tiered routing and automatic downgrades turned it into a $4 experiment, and made cost a screen instead of a surprise.
Walking every screen for this write-up found what's next: the dashboard's Active Agents count stays at zero while agents run, graph text disappears in dark mode, and Pause, Stop, Filter and Layout aren't wired yet. The web UI also needs to listen to the engine's live run events directly, instead of a replay.