I assess whether AI systems do what their builders say they do.
- reported
- held out
I've spent twelve years building reinforcement learning and optimisation systems, including Gran Turismo Sophy, the agent that beat the world's top Gran Turismo Sport drivers, published in Nature in 2022. I do this work independently of my employer, for investors deciding whether to fund a technical claim and for teams deciding whether to build on one.
01 — offer Technical due diligence
- Problem
- Technical claims in AI companies are hard to falsify from a pitch deck, and the failure modes surface after the money is committed.
- Method
- Architecture and code review, evaluation-harness audit, prior-art check, interviews with the technical team, reproduction of headline results where feasible.
- Deliverable
- Written report: claim-by-claim assessment, scaling risk, defensibility, team capability, and the questions to ask before signing.
- Turnaround
- Three weeks from data-room access. Two weeks at +40%.
- Price
- €12,000. €20,000 where scope includes reproducing results or reviewing a training pipeline.
What I check
- Whether reported evaluation numbers survive a held-out setting, or whether the benchmark was selected after the results were known.
- Whether the reward, environment, or data pipeline already contains the result it appears to be discovering.
- What the training and inference cost curve does at ten times current scale, and whether the architecture survives it.
- Whether the moat is the model, the data, the distribution, or nothing.
- Whether the team has shipped a system of this class before or is describing one.
For: funds at seed to Series B evaluating an AI-first company; corporate development assessing an acquisition or partnership; boards where a technical bet has gone quiet.
Not for: companies who want help building the thing. I assess systems, I don't staff them.
Conflicts: I decline any engagement overlapping my employer's domain or portfolio, and sign your NDA before receiving materials. If a conflict emerges mid-engagement I stop and refund the unused portion.
02 — offer Fractional research lead
- Problem
- Teams building RL, agentic, or optimisation systems make architecture decisions once and live with them for years, usually with nobody senior to argue against.
- Method
- Two half-days a month of direct work, plus asynchronous review of designs, evaluation plans, and candidates.
- Deliverable
- Written recommendations after each session. Direct access between sessions for decisions that can't wait.
- Term
- Three-month minimum, then monthly.
- Price
- From €6,000/month. Equity considered in addition to cash, never instead of it.
The ground this covers, stated plainly because these are the words people search and ask with: reinforcement learning, imitation learning, reward and objective design, explainable AI and model interpretability, robot learning and control, agentic system architecture, evaluation and benchmark design, optimisation. Twelve years of published research across them, listed on the publications page.
Typical sessions: evaluation design before a model is trained rather than after; whether an agentic architecture needs the orchestration layer it's about to get; reward and objective specification; making a model's behaviour explainable when a regulator or a customer asks why; reading research candidates accurately in hiring; killing a research direction early enough that it's cheap.
Limit: three retainer clients at a time.
My publication record is on the publications page.
If you're evaluating a system and want to know whether the claims hold, book 30 minutes. If it isn't something I can help with, I'll say so on the call.
Book 30 minutes