Agentic Pilot.
Four weeks. One production use case.
A demo proves an agent can do the job once. Production asks the harder questions, which are what the agent may touch, what it must prove before it acts, and who answers when it goes wrong at three in the morning. Four weeks settles them. One agent live, wrapped in tool boundaries, eval gates, tracing, cost caps and an audit trail, on the deployment target your constraints allow.
The four weeks, in order.
Scope the authority, draw the boundary
- Map the use case end to end, covering users, data, decisions and the latency budget
- Write the decision authority document and the tool access boundary, recording read, write and reversibility for every tool the agent gets
- Choose the deployment target against your residency rules, from managed cloud through private cloud to on-premise and air-gapped
Stand up the harness
- Sandbox the runtime with its own service account, egress restricted to the hosts it needs, and a filesystem it cannot escape
- Wire the approval gates, so every write touching money, customer records or production infrastructure stops for a human
- Turn tracing on before the first real run, capturing every prompt, tool call, argument, result and retry
Build the agent and prove it
- Ship the agent against your real source of truth, whether that is internal APIs, a database or your document store
- Stand up the eval suite on a golden dataset of at least 100 real tasks, grading tool calls and output schema deterministically and using LLM-as-judge on the subjective parts
- Wire the escalation path to a named queue with an owner, handing the human the full run context rather than a bare alert
Cap it, ship it, hand it over
- Set cost and step caps per run, per user and per tenant, so a runaway loop fails fast instead of degrading quietly
- Stream the audit trail to the SIEM your security team already operates, and keep prompts, model versions and tool schemas in Git so any one of them rolls back in a single command
- Launch to production with a 90-minute walkthrough, a recorded primer, and a 30-day support window
What's in the box.
- One agent running in production against your real systems
- Tool access boundary with approval gates on every irreversible write
- Sandboxed runtime with its own service account and restricted egress
- Retrieval with a citation surface where the agent grounds on your documents
- Eval suite on a golden dataset of at least 100 real tasks, wired into CI
- End-to-end tracing of every prompt, tool call, argument and result
- Cost and step caps per run, per user and per tenant
- Audit trail streamed to your SIEM, with prompts, models and tool schemas versioned in Git
- Architecture document, runbook, recorded walkthrough, and 30 days of post-launch support
What the pilot does not cover.
A fixed price only holds if the edges are drawn, so here is where this one stops. None of these are refusals. Each is a separate scope, and a discovery call will tell you whether you need any of them yet.
- A second agent, since the pilot ships one and multi-agent orchestration is separate scope
- A fine-tuned or newly trained model, since the pilot routes to existing models and measures them on your tasks
- Hardware procurement, if the deployment target lands on-premise or air-gapped
- A rewrite of the systems the agent talks to, which the pilot integrates with as they stand
- An open-ended research phase, since the pilot starts from a use case you have already chosen
Fit check.
The pilot ships best when these four things are true. If only some are true, a discovery call will sort whether the pilot or a different scope fits.
- You have a task an agent should own, not a research question about whether agents work
- You have systems the agent will act on, with an API or a database behind them
- You have someone who will accept or reject the agent's output the day it ships
- You have a security or compliance reviewer who will want to see the audit trail
One agent, four weeks, answered
What the pilot builds, what it costs, where it runs, and why the harness matters more than the model.
An agentic AI pilot is a fixed-scope engagement that puts one agent into production with the operational controls it needs to stay there. Pooya Golchian runs his as a four-week build covering one production use case. Week one draws the tool access boundary and picks the deployment target. Week two stands up the sandbox, the approval gates and tracing. Week three ships the agent and its eval suite. Week four sets cost and step caps, wires the audit trail, and launches with 30 days of post-launch support.
Agentic pilots with Pooya Golchian start at $15,000. That fixed price covers four weeks and one production use case, including the tool access boundary, a sandboxed runtime, an eval suite on a golden dataset of at least 100 real tasks, end-to-end tracing, cost and step caps, an audit trail, and 30 days of post-launch support. Agent harness engagements typically run four to six weeks in total, and larger rollouts move to fixed scope per phase.
An agent harness is everything around the model that decides what the agent may touch and what happens when it is wrong. It covers tool access boundaries, a sandbox that limits blast radius, permission and approval models, eval gates, tracing and observability, human escalation paths, cost and step caps, an audit trail, rollback of prompts, models and tool definitions, and the deployment target. Pooya Golchian builds the harness in the same four weeks as the agent, because an agent without one cannot pass a security review no matter how well it demos.
Pooya Golchian ships a single production agent in four weeks under a fixed-scope Agentic Pilot. Most agent harness engagements run four to six weeks end to end. The four weeks run in order, moving from decision authority and tool boundaries, through the sandbox, approval gates and tracing, to the agent and its eval suite, and finally to cost caps, the audit trail and launch.
Yes. The deployment target is a week-one decision on every Agentic Pilot, and it follows your residency rules rather than a default. Managed cloud, private cloud, on-premise and air-gapped all work, and the harness around the agent stays the same in each. Pooya Golchian settles the residency question with your compliance reviewer before anyone picks a model, because the target constrains model choice and never the other way round.
A demo proves an agent can complete one task, on a clean input, with a human watching. Production asks what the agent may touch, what it must prove before acting, who answers the page when it fails, and what a runaway loop costs before anyone notices. Those answers live in the harness rather than the model. Pooya Golchian publishes the fourteen checks that decide it in the Agent Readiness Checklist, and the four-week Agentic Pilot is the engagement that installs them around one agent.
Reserve a pilot slot.
I run two pilots in flight at any time. A discovery call locks the scope, the deployment target and a start date.