AI Engineering Blog
Working notes from shipping agentic AI into production. Agent systems, MCP servers, RAG that cites its sources, and evals that catch drift before your customers do. Also the AIDLC method I use to get agent work past the pilot.

GEO vs SEO: The Answer That Cost Me Clicks
Putting the answer in my titles turned 11,117 impressions into 19 clicks. GEO and SEO split at exactly that point, measured on my own site.
Read
How to Get Cited by ChatGPT in 2026
ChatGPT cited my site more often than Google ranked it. Then I found 10 of 30 test queries it never searched, and no page edit could reach them.
Read
How to Measure AI Citations Without Fooling Yourself
My first AI citation measurement recorded twenty-three defeats that never happened. The engines behaved. The mistake hid inside my counting.
Read
What Is Jev, and Does It Replace LLMs?
TypeSafe claims Jev, its new decision model, cannot hallucinate and undercuts frontier LLM pricing. Outside testing measured how much survives.
Read
How I Built a Dubai Business Directory on 1,400 SearchApi Requests
One Google Maps query stops at 200 results, so enumerating a city takes a different shape. Here is what 1,400 SearchApi requests actually bought me.
Read
SpecForge: Spec-Driven AI Development With Approval Gates That Actually Block
SpecForge fuses spec-kit's executable specs with AI-DLC governance, then enforces approval with a Claude Code hook that blocks code edits until you sign off.
Read
The Verification Loop: Reviewing Code Your Agents Wrote
When agents write most of the code, review stops being a checkpoint and becomes the job.
Read
Bolts: A Team Playbook for Running Agentic Work in Hours
A bolt is a short spec, generate, verify cycle that a team runs in hours instead of a week-long sprint.
Read
Bolts vs Sprints: How Agentic Teams Measure Work Now
The two-week sprint was built around human typing speed. When agents write the code, the bolt replaces it.
Read
How to Build an App With a Ralph Loop and a Goal Prompt
A practical guide to building an app by pointing a harness at a goal prompt and running it in a loop: how to write the goal, wire the verification gate, and know when to stop.
Read
How to Write a Spec an AI Agent Can Actually Build
A spec is the highest-leverage ten minutes in an agentic cycle.
Read
Managing and Training a Technical Team to Ship With Agents
Buying AI seats does not make a team agentic. In 2026 the bottleneck moved from writing code to verifying it.
Read
The Spec-to-Bolt Loop: The Agentic Development Cycle End to End
Specs and bolts are two halves of one cycle.
Read
What Is a Bolt? The Work Unit That Replaced the Sprint
A bolt is the AIDLC unit of work that replaces the two-week sprint with a cycle measured in hours.
Read
What Is an Agent Harness? The Layer That Turns a Model Into an Agent
An agent harness is the deterministic layer that wraps an LLM: the loop, the tools, the sandbox, the permissions.
Read
What Is Spec-Driven Development? Specs as the New Source of Truth
Spec-driven development treats a written specification as the executable source of truth and code as a regenerable output.
Read
What Is the Ralph Loop? Agentic Coding, Deterministically Simple
The Ralph loop runs a coding agent in a plain while loop against one goal prompt, with memory in files and git instead of context.
Read
The 2026 AI Transformation Playbook for UAE and Dubai Businesses
A 2026 AI transformation playbook for UAE and Dubai firms covering a phased roadmap, real ROI, and PDPL and DIFC compliance.
Read
AIDLC vs SDLC: What Changes When Agents Write the Code
AIDLC vs SDLC compared phase by phase.
Read
Ollama Cloud vs Claude and GPT: Real Cost, Limits, and Quality in 2026
Ollama Cloud starts at $0 and caps at $100/mo flat. Claude Sonnet runs $3/$15 per 1M tokens. Where each wins on cost, limits, and quality in 2026.
Read
Wiring Ollama Into Visual Studio 2026 Copilot the Right Way
Run local LLMs with GitHub Copilot in Visual Studio 2026 via Ollama. BYOK setup, model choices, endpoint wiring, and where local beats cloud.
Read
The Data Engineer in the Agentic Era: From Pipelines to Retrieval and Ground Truth
Agents can write the ETL.
Read
The Engineering Manager in the Agentic Era: Managing People and Agents Together
When part of your team is human and part is autonomous, velocity metrics lie and the old planning rhythm breaks.
Read
The Technical Writer in the Agentic Era: From Documenting Code to Authoring Context
Agents draft docs in seconds, so writing prose is no longer the scarce skill.
Read
The Security Engineer in the Agentic Era: New Attack Surfaces, Same Accountability
Prompt injection, tool abuse, and data exfiltration through an agent are not edge cases, they are the new baseline.
Read
The DevOps Engineer in the Agentic Era: From Pipelines to Observability for Autonomy
Agents can write the Terraform and the CI config. What they cannot do is decide what an autonomous system must expose to stay safe in production.
Read
The QA Engineer in the Agentic Era: From Writing Tests to Designing Evals
Deterministic test suites cannot judge probabilistic systems.
Read
The Frontend Developer in the Agentic Era: From Building Components to Curating Them
Agents scaffold components, wire state, and match a design in minutes.
Read
The Backend Developer in the Agentic Era: From Writing Endpoints to Reviewing Them
Agents now write most of the CRUD, the handlers, and the glue.
Read
The Software Architect in the Agentic Era: Designing Boundaries Agents Can't Cross
Agents generate code fast, including the wrong abstraction fast.
Read
The Product Manager in the Agentic Era: From Backlog Owner to Spec Author
When agents write the code, the product manager's leverage moves upstream.
Read
Gemma 4 on Ollama: Multi-Token Prediction, Benchmarks, and Local Setup (May 2026)
Google's Gemma 4 update ships Multi-Token Prediction for roughly 3x faster decoding.
Read
'AI Doesn't Code' Is a Skills Problem, Not an AI Problem
Engineers complain that AI hallucinates, burns tokens, and ships nothing. Each of those failures has a named technique that fixes it.
Read
How I Build AI Agent Teams That Actually Ship for B2B Companies
Most B2B companies hire for AI and end up with a single ChatGPT wrapper.
Read
riyal v1.2.0: The Saudi Riyal Symbol Toolkit for React, Vue, Svelte, and Web Components
riyal v1.2.0 ships the Saudi Riyal symbol (U+20C1) as a web font, React, Vue 3, Svelte 5, React Native, and Web Component primitives.
Read
Stop Burning Claude Tokens: How RTK Cuts AI Coding Costs 60-90% in 2026
Engineers blow through Claude Max quotas in four days because tool output floods the context.
Read
dirham v1.5.3: The Complete Guide to UAE Dirham in React, Web Components, and Vanilla JS
dirham v1.5.3 ships animated price counters, masked currency inputs, live exchange rate hooks, a Tailwind plugin, a Next.js font helper, and React Native support.
Read
AI Agent Memory Systems: How Claude, GPT, and Gemini Remember Context Across Sessions
Claude Projects, GPT memory, and Gemini context windows compared.
Read
AI Code Review at Scale: How Teams Ship 40% Faster Without Sacrificing Quality
Teams using AI code review tools ship 40% faster with equivalent bug rates.
Read
Claude Max and the High-Volume Engineer: How Senior Developers Use Anthropic's Top Tier
Anthropic's Claude Max subscription costs $350/month but enables 20x more usage than Claude Pro.
Read
Contentful vs Sanity vs Strapi vs Payload 2026: Developer Experience Benchmarks and Real Pricing
I fired 500 requests per tier at five headless CMS from us-east-1 in April 2026. The DX chart and the price table name different winners.
Read
CrewAI vs LangGraph vs AutoGen 2026: Benchmarks & Pricing
CrewAI pricing for 2026 next to LangGraph, AutoGen and Smolagents, plus 200 benchmark tasks each on Qwen3 32B. The cheapest plan rarely wins on cost.
Read
Hermes Agent with Ollama: The Self-Improving AI Agent That Runs Entirely Local
Hermes Agent from Nous Research runs entirely local through Ollama with 70+ built-in skills, cross-session memory, and messaging gateway integrations.
Read
Ollama Cloud Pricing 2026 and Where Self-Hosting Takes Over
Ollama Cloud pricing now runs on usage credits.
Read
Rust vs Go vs Zig for High-Performance Backend Services in 2026
Benchmarks comparing Rust, Go, and Zig for backend services.
Read
Why TypeScript 6.0 Is Trending: The Technical and Cultural Shift Behind the Viral Moment
TypeScript 6.0 npm downloads jumped 12.5% in two weeks after release. The trending story is not about new syntax.
Read
Claude Opus 4.6 vs GPT-5.3-Codex: The State of Frontier AI Models in April 2026
Claude Opus 4.6 leads on agentic coding with computer use capabilities. GPT-5.3-Codex excels on terminal operations and web development.
Read
Gemini 2.0 vs GPT-5.3 vs Claude Opus 4.6: 2026 Model Rankings
Google Gemini 2.0, OpenAI GPT-5.3, and Anthropic Claude Opus 4.6 represent the current frontier.
Read
Local AI Coding Models 2026: Why Developers Pick Open Source
Ollama's top HumanEval score runs at 8 tokens per second. Four VRAM tiers, an RTX 5070 baseline, and break-even math at 500 queries a day.
Read
Maximize Claude Code: Advanced Configuration for Senior Engineers
Claude Code configuration options that separate junior from senior usage.
Read
Principles of Design with AI-Generated Photos
Explore the fundamental principles of design including balance, contrast, emphasis, movement, pattern, rhythm, proportion, unity, and white space through cinematic AI-generated black and white photographs.
Read
Reasoning Models Emergence: How Chain-of-Thought Unlocks Complex Problem Solving
Chain-of-thought reasoning models demonstrate emergent reasoning capabilities.
Read
The $500 GPU That Outperforms Claude Sonnet on Coding Benchmarks
A $500 RTX 5070 running Qwen 3.5 Coder 32B outperforms Claude Sonnet 4.6 on HumanEval at 40 tokens per second.
Read
Inside the .claude/ Folder: How Claude Code Organizes Your AI Workspace
Most teams commit it by accident. Every path Claude Code writes, the 30-day cleanup command, and how its memory holds up against Copilot and Cursor.
Read
Claude and the New Developer: How AI Is Reshaping Coding Skills in 2026
GitHub's Octoverse 2025 data plus interviews with 22 advanced AI users, charted. The language rankings flipped. So did the job description.
Read
GitHub Copilot with Ollama: Agentic AI Models Running Locally in Your IDE
GitHub Copilot now runs agentic workflows through Ollama. Deploy Qwen, DeepSeek, and Llama models locally.
Read
State of the Product Job Market in Early 2026
7,300 open PM roles. 67,000 engineering openings. A 340% surge in AI positions. The data on who is actually getting hired in early 2026.
Read
How Reco Rewrote JSONata With AI and Saved $500K a Year
Reco.ai rewrote its JSONata engine with AI in a single day and cut $500K a year from its infrastructure bill. The method, the tools and the lessons.
Read
AI-Scientist-v2: How AI is Automating Scientific Discovery
AI-Scientist-v2 timed against human researchers across six research phases, four agents, four domains. Where the autonomy claim stops holding.
Read
The UAE Dirham Currency Symbol (U+20C3): Why It Took 18 Years and How to Use It Today
UAE Dirham symbol U+20C3 is now official. dirham v1.3.0 renders it today in React, Web Components, and vanilla JS with zero migration when OS support ships.
Read
GitHub Copilot Data Policy Changes: What Developers Must Know in 2026
GitHub rewrote Copilot's data terms in 2026. Which tiers it touches, the opt-out click path GitHub buries, and five assistants compared.
Read
Software Developer Job Market 2026 and Who Is Actually Getting Hired
Federal Reserve posting data charted from Q1 2023 to Q1 2026, seven sectors and six skills ranked.
Read
Supply Chain Attacks on Developers: Lessons from LiteLLM and Trivy
Recent supply chain attacks targeting LiteLLM and Trivy expose critical vulnerabilities in developer tooling.
Read
TypeScript 6.0: New Features Every Developer Should Know
TypeScript 6.0 changed defaults many tsconfig files never set, so upgrades break quietly.
Read
NVIDIA Nemotron + OpenManus: $31.4B Agent Market and the Open Source Disruption
NVIDIA Nemotron 70B achieves 88.4% MMLU at 5.4x lower cost than GPT-4o.
Read
MCP in 2026: The Protocol That Replaced Every AI Tool Integration
Local stdio MCP servers break once SSE traffic meets a load balancer.
Read
AI Reasoning Systems and the Theory of Mind Breakthrough
From pattern matching to measured reasoning. Benchmarks show large gains from chain of thought prompting and frontier model accuracy.
Read
Self-Hosting AI in 2026, the TCO Math and the Stack That Replaces Cloud APIs
Cloud inference bills climb forever. Hardware bills stop. Five charts plot 18 months of TCO, benchmark four GPUs, and mark the break-even month.
Read
AI Agents in 2026: LangGraph vs CrewAI vs Smolagents with Real Benchmarks on Local LLMs
Four agent frameworks, five local models, three tool-use benchmarks run on Ollama. GitHub stars rank them one way.
Read
RevOps AI: Build Your Entire Sales Team on Notion with Gemini and MCP
How I built a full Revenue Operations platform using Notion as the database, Gemini 2.5 Flash as the AI agent, and the Model Context Protocol to wire 22 Notion tools together without a single hardcoded query.
Read
RAG Pipelines in Production: Vector Database Benchmarks, Chunking Strategies, and Hybrid Search Data
72% of enterprises run RAG in production. Qdrant hits 6ms p50 latency. Hybrid search boosts recall 17%.
Read
Vibe Coding in 2026: $9.2B Cursor, 92% HumanEval, and the End of Boilerplate
$9.2B Cursor valuation. 92.4% HumanEval score for Claude 3.5 Sonnet. 340% enterprise adoption growth.
Read
Local AI in 2026, Ollama Benchmarks and What Inference Really Costs
Zero per token is not zero cost.
Read
The Future of AI Prediction: Uncertainty Quantification, Monte Carlo Methods & Statistical Mathematics
Deep dive into uncertainty-aware AI: Monte Carlo methods, Bayesian inference, conformal prediction, and the statistical foundations shaping the future of trustworthy machine learning.
Read
use-local-llm: React Hooks for AI That Actually Work Locally
Build AI-powered React apps that talk directly to your local models—no backend required.
Read
AI and Jobs: What Anthropic's Labor Market Data Actually Shows About Your Career
Headline AI risk numbers measure capability, not deployment.
Read
dirham: The UAE Dirham Symbol (U+20C3) as Web Font, CSS Utility & React Component
The dirham npm package brings the UAE Dirham currency sign (U+20C3, Unicode 18.0) to production via a custom web font, React SVG component, CSS utility class, and JavaScript formatting utilities.
Read
scss-helper v5: A Modern SCSS Utility Toolkit for Tailwind's Gaps
scss-helper v5 is a modern SCSS/CSS utility toolkit that complements Tailwind CSS v3/v4 with design tokens, fluid clamp() typography, container queries, dark mode, golden ratio layouts, a 12-column grid, and animations.
Read
vue-multiple-themes v4: Dynamic Multi-Theme Support for Vue 2 & 3
A deep dive into vue-multiple-themes v4, the open source library for dynamic CSS custom property theming in Vue 2.7+ and Vue 3.
Read
vue-star-rate: Zero-Dependency Vue 3.5+ Star Rating Component
Introducing vue-js-star-rating, a WCAG 2.2 accessible, TypeScript-first star rating component for Vue 3.5+.
Read
Best Headless CMS 2026: Pricing, Total Cost, and How to Choose
Headless is commoditized. The real split runs developer frameworks against enterprise orchestrators.
Read
WEF Future of Jobs Report 2025: 78M Jobs Created, 92M Displaced
WEF Future of Jobs 2025 data: 78M new roles, 92M displaced, 39% skills gap.
Read
React useEffect & useCallback in 2026: Stop Unnecessary Re-renders (React 19 Guide)
Fix unnecessary React re-renders: master useCallback, useMemo, React.memo, and the new React 19 use() hook with real-world patterns, profiling tips, and pitfalls to avoid.
Read
Next.js 15 vs Astro 5 vs Remix in 2026: Which Framework Should You Choose?
Next.js 15, Astro 5, and Remix compared on performance, bundle size, and real use cases.
Read
AWS ECR in 2026: Pull, Inspect, Scan & Automate Docker Images: Complete Guide
Complete AWS ECR guide: authenticate with OIDC, pull Docker images, extract filesystem layers, scan with Amazon Inspector v2, set lifecycle policies, and automate with GitHub Actions.
Read
ArangoDB on AWS: Automate Install, S3 Backup & Restore with Systemd
Production-ready shell scripts for ArangoDB on AWS EC2: automated install, daily S3 backups via systemd timers, and one-command point-in-time restore.
Read
SSH SOCKS5 Proxy in 2026: macOS, Linux, Windows & sshuttle: Complete Guide
Turn any SSH server into a private VPN with one command.
Read
Firebase Cloud Messaging in 2026: Web Push Notifications with Vue 3 & SDK v10
Step-by-step Firebase Cloud Messaging v10 setup with Vue 3: VAPID keys, service workers, foreground/background notifications, and Node.js server sending.
Read
Deploy Jekyll to GitHub Pages in 2026: GitHub Actions, Custom Domain & Cloudflare
Deploy Jekyll with a GitHub Actions CI/CD pipeline, configure a custom domain with HTTPS, and add Cloudflare CDN for speed and DDoS protection.
Read