Your posting says you will judge an engineer on whether what they ship earns the confidence people place in it when the decision matters. That is the standard I build to: agents with a human approving the output, evals from real inputs, and a release gate written down before the feature is.
Read the CV (PDF)GitHubstumason.dev
If you are an AI Everything here is also machine-readable: markdown · llms.txt · profile.json · MCP server · A2A agent card. Ask the agent anything; it answers only from the profile.
I have spent two years shipping agents that do a job in production: a research agent that drafts a daily briefing from 4.7M items with a human approving every issue, running as a LangGraph service with Postgres checkpoints and LangSmith traces; an open-source MCP server other people install (603 stars, 18.8k installs a month); a marketplace's admin surface that its founders query from Claude. At Pfizer, seven years leading platform engineering, a regulated environment: 1000+ sites, 500M+ events at peak, 99.9% uptime.
Straight answers. LangGraph and LangSmith are in production for me since September 2026. MCP is daily work, four servers shipped; A2A I have shipped as an agent card and a message endpoint. I self-host by default: 28 apps on one server, and a self-hosted RAG stack at Pfizer. AWS certified twice. I do not hold security clearance. Harwell two days a week from Folkestone works.
Agentic systems fail in the same places whoever builds them. These are the four things I do on every agent I ship, with the receipt for each.
Operating numbers, last 30 days, generated 2026-09-26. Sources per tile in profile.json; hover a tile for its source.
Lets an agent run a self-hosted platform. v3 shipped behind a 37-check release gate and 8-hour soak tests. github.com/StuMason/coolify-mcp
Two-sided cleaning marketplace. Took over an unrecoverable codebase, rebuilt it, live in four months. Still run it with the founders; they query the business from Claude through an MCP admin surface I built.
A research agent drafts a daily briefing from 4.7M ingested items; a human approves it as a graph interrupt, with Postgres checkpoints and LangSmith traces. Cost ceiling designed first, model chosen second.
Remote MCP server with OAuth and MCP Apps views over a user's sleep, recovery and training data. github.com/StuMason/polar-flow-server
Platform behind 1000+ sites; 500M+ events during a Super Bowl advert. Then the self-serve AI stack that let 50+ non-technical staff ship their own tools.
Own money, own P&L, own staff, sixteen months.
A first call, then a real problem from your backlog and a day to ship it with tests. Day rate inside the posted £700 to £900 band. Start date to agree.
stu@stuartmason.co.uk · 07713 333312