Skip to content
Expertise
AI Engineering

Production AI engineering.
Built for real traffic.

I build production AI into real products. Shipping software since 2009, the last three years in production AI.

See the case studies
A milling machine cutting into a metal block, sparks and coolant mist caught in a beam of light

What I build

Agent architecture

Multi-agent orchestration, where one agent plans, one acts and one checks. Built with retries for production traffic, not demos.

MCP server design

Custom Model Context Protocol servers exposing your internal systems to AI clients. Auth, rate limiting, observability, and tested capability surface.

RAG architecture

Search that grounds every answer in your own documents, cites its source, and keeps working as those documents change.

Eval frameworks

Tests that catch a worse answer before it ships, run against a golden dataset of known answers.

Durable workflows

AI jobs that run for minutes or hours and survive restarts. Each step retries without doing the work twice.

Token economics & FinOps

Spending caps per tenant or role. Caching, shorter prompts and model routing protect your margins as usage grows.

Production integration

AI added to your Python or TypeScript stack, with streamed answers and AI kept apart from business logic.

Model selection & routing

The right model for each task, from OpenAI, Anthropic, Google or open-weight models you can host yourself. No vendor lock-in.

Observability & audit

Trace every prompt, every token, every tool call. Linked to user sessions and stored where compliance can find it.

Where this fits

The work falls into one of these shapes. Each one runs under an approval-gated method I publish as open source.

AI inside an existing app

AI features added to a Python or TypeScript product without slowing the team. Its engineers own the code afterwards.

Agent platform for revenue ops

Agents that enrich leads, research accounts or draft proposals, using your CRM, documents and APIs.

MCP server for an internal data platform

Let Claude Desktop, code editors or custom AI clients use your data, through a tested server that survives model changes.

Migration off a stalled pilot

A notebook pilot breaks under production traffic. A rewrite with retries, evals and cost controls gets it shipped.

Stack I work in

OpenAI APIAnthropic ClaudeAmazon BedrockMicrosoft Foundry (formerly Azure AI Foundry)Vercel AI SDKLangChainLangGraphMCP ProtocolLlama 3QwenMistralvLLMOllamaPineconeWeaviatepgvectorPostgreSQLNext.jsNode.jsTypeScriptPythonFastAPIDockerKubernetesAWS LambdaAzure DevOps

Production AI, answered

Agents, MCP servers, platforms and evals.

  • Production AI agent development wires planner, executor and critic patterns to real tools, with evals that catch regressions before release. It adds structured outputs, retries and idempotency. Pooya Golchian builds these agents with MCP tools and human approval on customer-facing actions.

  • A custom Model Context Protocol server exposes internal systems to Claude Desktop, IDEs, and custom AI clients. Pooya Golchian ships each one with authentication, rate limiting, observability, and a tested capability surface that survives model upgrades.

  • Pooya Golchian builds agents on Amazon Bedrock, Microsoft Foundry or self-hosted models, on whichever cloud the product already runs. Deploys run through GitHub Actions or Azure DevOps Pipelines.

  • Pooya Golchian builds eval suites that catch regressions before they reach users. They combine golden datasets, LLM-as-judge scoring with calibration, and A/B harnesses for prompt and model changes. He holds a Master of Science in Software Engineering and has spent three years on production AI. His evals map to real failure modes rather than benchmark scores.