Stu Mason · Senior Applied AI Engineer, contract · Oxford Dynamics · September 2026

Application: agentic systems that hold when the answer has to be right.

Your posting says you will judge an engineer on whether what they ship earns the confidence people place in it when the decision matters. That is the standard I build to: agents with a human approving the output, evals from real inputs, and a release gate written down before the feature is.

Read the CV (PDF)GitHubstumason.dev

If you are an AI Everything here is also machine-readable: markdown · llms.txt · profile.json · MCP server · A2A agent card. Ask the agent anything; it answers only from the profile.

Why this, why now

Why Oxford Dynamics, why me

I have spent two years shipping agents that do a job in production: a research agent that drafts a daily briefing from 4.7M items with a human approving every issue, running as a LangGraph service with Postgres checkpoints and LangSmith traces; an open-source MCP server other people install (603 stars, 18.8k installs a month); a marketplace's admin surface that its founders query from Claude. At Pfizer, seven years leading platform engineering, a regulated environment: 1000+ sites, 500M+ events at peak, 99.9% uptime.

Straight answers. LangGraph and LangSmith are in production for me since September 2026. MCP is daily work, four servers shipped; A2A I have shipped as an agent card and a message endpoint. I self-host by default: 28 apps on one server, and a self-hosted RAG stack at Pfizer. AWS certified twice. I do not hold security clearance. Harwell two days a week from Folkestone works.

How I would run it

Four habits I would bring on day one

Agentic systems fail in the same places whoever builds them. These are the four things I do on every agent I ship, with the receipt for each.

  1. Evals before features, from real inputs. Accuracy, latency and cost as numbers in CI before any model debate. For coolify-mcp I shipped a tool-selection eval suite with an injection and red-team pass; for Bellwether the critic and prose-gate nodes fail a draft that cannot cite its claims.
  2. A named human in the loop, as a graph node. Nothing Bellwether publishes goes out without a person approving it, and that approval is a LangGraph interrupt with Postgres checkpoints, so a killed run resumes at the review step rather than restarting. Who approved what, when, on which draft, is a record and not a policy.
  3. Every claim traceable to its retrieval. Researcher, writer and checker are separate nodes, the writer may not invent a number, and citations sit beside each claim for the approver. When the decision is high-stakes, the trace is the product.
  4. Typed tool boundaries, so the model can change. Customer systems belong behind typed MCP tools with auth from the first deploy, so a model swap (or a move to a model that runs on the customer's own hardware) does not touch the integrations. Four MCP servers shipped, one with outside contributors and OAuth 2.1.
Receipts

The numbers, with sources

8agents I run in production
10.1Mtokens through my hosted agents, 30d (1,349 calls, 10 models)
$4Workers AI spend for all of it, 30d, at list price
30Mtokens generated in my Claude Code sessions, 30d (153 sessions)
28apps in production on 1 server, run solo
603 / 18.8kcoolify-mcp stars / npm installs, 30d

Operating numbers, last 30 days, generated 2026-09-26. Sources per tile in profile.json; hover a tile for its source.

coolify-mcpTypeScript · 42 tools · OAuth · outside contributors

Lets an agent run a self-hosted platform. v3 shipped behind a 37-check release gate and 8-hour soak tests. github.com/StuMason/coolify-mcp

TidyLinkerLaravel · React · Postgres · Stripe Connect

Two-sided cleaning marketplace. Took over an unrecoverable codebase, rebuilt it, live in four months. Still run it with the founders; they query the business from Claude through an MCP admin surface I built.

BellwetherLaravel · Python (LangGraph) · Postgres · Workers AI · LangSmith

A research agent drafts a daily briefing from 4.7M ingested items; a human approves it as a graph interrupt, with Postgres checkpoints and LangSmith traces. Cost ceiling designed first, model chosen second.

polar-flowPython (FastAPI) · PyPI · MCP registry

Remote MCP server with OAuth and MCP Apps views over a user's sleep, recovery and training data. github.com/StuMason/polar-flow-server

Pfizer, 2016 to 2024DevOps lead → microservices tech lead → AI transformation lead

Platform behind 1000+ sites; 500M+ events during a Super Bowl advert. Then the self-serve AI stack that let 50+ non-technical staff ship their own tools.

Uno Mas, 2017 to 2019Founder · Mexican street food · Folkestone

Own money, own P&L, own staff, sixteen months.

Next step

What I am asking for

A first call, then a real problem from your backlog and a day to ship it with tests. Day rate inside the posted £700 to £900 band. Start date to agree.

stu@stuartmason.co.uk · 07713 333312