Production AI engineering.
Built for real traffic.
I build production AI into real products. Shipping software since 2009, the last three years in production AI.
See the case studies
What I build
Agent architecture
Multi-agent orchestration, where one agent plans, one acts and one checks. Built with retries for production traffic, not demos.
MCP server design
Custom Model Context Protocol servers exposing your internal systems to AI clients. Auth, rate limiting, observability, and tested capability surface.
RAG architecture
Search that grounds every answer in your own documents, cites its source, and keeps working as those documents change.
Eval frameworks
Tests that catch a worse answer before it ships, run against a golden dataset of known answers.
Durable workflows
AI jobs that run for minutes or hours and survive restarts. Each step retries without doing the work twice.
Token economics & FinOps
Spending caps per tenant or role. Caching, shorter prompts and model routing protect your margins as usage grows.
Production integration
AI added to your Python or TypeScript stack, with streamed answers and AI kept apart from business logic.
Model selection & routing
The right model for each task, from OpenAI, Anthropic, Google or open-weight models you can host yourself. No vendor lock-in.
Observability & audit
Trace every prompt, every token, every tool call. Linked to user sessions and stored where compliance can find it.
Where this fits
The work falls into one of these shapes. Each one runs under an approval-gated method I publish as open source.
AI inside an existing app
AI features added to a Python or TypeScript product without slowing the team. Its engineers own the code afterwards.
Agent platform for revenue ops
Agents that enrich leads, research accounts or draft proposals, using your CRM, documents and APIs.
MCP server for an internal data platform
Let Claude Desktop, code editors or custom AI clients use your data, through a tested server that survives model changes.
Migration off a stalled pilot
A notebook pilot breaks under production traffic. A rewrite with retries, evals and cost controls gets it shipped.
Stack I work in
Production AI, answered
Agents, MCP servers, platforms and evals.
Production AI agent development wires planner, executor and critic patterns to real tools, with evals that catch regressions before release. It adds structured outputs, retries and idempotency. Pooya Golchian builds these agents with MCP tools and human approval on customer-facing actions.
A custom Model Context Protocol server exposes internal systems to Claude Desktop, IDEs, and custom AI clients. Pooya Golchian ships each one with authentication, rate limiting, observability, and a tested capability surface that survives model upgrades.
Pooya Golchian builds agents on Amazon Bedrock, Microsoft Foundry or self-hosted models, on whichever cloud the product already runs. Deploys run through GitHub Actions or Azure DevOps Pipelines.
Pooya Golchian builds eval suites that catch regressions before they reach users. They combine golden datasets, LLM-as-judge scoring with calibration, and A/B harnesses for prompt and model changes. He holds a Master of Science in Software Engineering and has spent three years on production AI. His evals map to real failure modes rather than benchmark scores.