Offers
I take on four kinds of engagement. Each is scoped in writing before it starts and priced by the day or as a fixed fee, and each ends with a one-page memo for the person who has to make the decision and a repository for whoever inherits the model after me.
Two days a week from start month TBD. Remote from the UK, with on-site days by arrangement.
£500 a day. Fixed fees against a written scope on request. Inside or outside IR35, as the engagement requires.
Fedelm Advisory invoices directly. If your organisation buys through a framework or an intermediary, tell me and we will find the route.
In every engagement
Whatever the size of the work, the same things are true of what I hand over.
- The claim is written down, with the result that would refute it, before I test it.
- Baselines are fitted and reported before any candidate model: seasonal naive, ETS and ARIMA for forecasting; unpooled and fully pooled models for hierarchical work.
- Uncertainty is delivered as quantiles or posterior samples, and the coverage of the intervals is measured on held-out data rather than assumed.
- Backtests use rolling origins and strict time splits, with features built only from what was known at each origin. Where the source data is revised after publication, I archive the vintages and evaluate on what existed at the time.
- The answer is expressed in your units, beds, staff, stock, pounds or experiments, and scored on the cost of getting it wrong.
- The code installs and tests in two commands, and the README gets a stranger to the baseline table.
- Where a simpler model beat a more complex one, the write-up says which and by how much.
Forecast audit
2–4 weeks, typically 4 to 8 days.
You already have a forecast: a spreadsheet, a vendor’s tool, or a model built by someone who has since moved on. It feeds a plan, and nobody has recently checked whether it deserves to.
I check five things.
- Does it beat a seasonal naive baseline on a proper backtest with rolling origins, and by how much?
- Do its ranges hold? A nominal 90% interval should contain about 90% of outcomes. Most do not, and whether the miss is too wide or too narrow changes what you should do about it.
- Is there leakage: features or revisions that were not available when each forecast was made?
- How much does its accuracy fall on the data you actually had at each origin, as opposed to the revised file you have now?
- What would you need to watch each month to know when it has stopped working?
You get a written audit with a baseline table and calibration plots; a one-page memo saying whether to rely on the forecast, for which horizons and with what range; and a monitoring plan your own team can run.
This is not a rebuild. If the audit finds the forecast should be replaced, the harness I built to test it carries straight into the next engagement.
Probabilistic demand forecasting
8–12 weeks, two days a week, typically 16 to 24 days.
For a planning decision that repeats: winter beds, call-centre staffing, stock levels, energy, or anything else counted weekly or monthly across many sites. The decision currently rests on a point forecast, last year’s number or someone’s judgement, and the cost of being wrong falls mostly on one side.
How it runs:
- Weeks 1–2: data audit; a vintage archive if the source is revised; hypotheses written down with refutation rules and frozen before evaluation.
- Weeks 3–6: baselines (seasonal naive, ETS, ARIMA), then the candidates, all scored on rolling-origin backtests with proper scoring rules for the whole predictive distribution.
- Weeks 7–8: the decision table in your units, scored on the cost of getting it wrong; a scheduled pipeline; the handover pack.
- Weeks 9–12, if scoped: hierarchy and reconciliation across sites or regions; cold-start handling for new units; a scorecard that updates itself after each data release.
You get a repository that installs and tests in two commands; a scheduled pipeline (GitHub Actions or your equivalent) that refits, forecasts and scores itself; a scorecard; a handover pack covering backfill, retraining, drift triggers and what to do when the feed breaks; a technical write-up; and a one-page memo giving the range, the drivers, the assumptions and how much to trust it.
The NHS A&E project is this engagement run in public, on open data, with every result posted.
Bayesian measurement
Fixed scope or multi-month, quoted against the scope.
For questions where a single number would mislead and the data arrives from many units at once: which channels or components drive the response and where they saturate; how demand responds to price across stores; what a newly launched unit can borrow from the ones that came before it.
I build these as hierarchical models in PyMC and deliver the workflow that makes them defensible when someone senior asks how you know:
- prior predictive checks, and a record of any prior change that mattered;
- an identifiability audit by simulation-based calibration, so you know which parameters your data can pin down before you commission the full model;
- posterior predictive checks and calibration against any experiments or holdouts you have;
- a production wrapper if the model will be refit: diagnostic gates in CI, versioned priors, drift alerts on the predictive checks.
You get the model and its tests in a repository; the identifiability report; a decision output in your units with its posterior uncertainty; a write-up positioned against the published method you would otherwise have used; and the one-page memo.
The screening-campaign replay project carries the Bayesian workflow this engagement uses: pre-registered comparisons, sampler health gates in CI, proper scoring of the predictive distribution, and a decision log with every rejected alternative.
Campaign replay audit
1–2 weeks, fixed fee.
For an R&D team that runs screening campaigns in batches, in media, formulation, fermentation conditions or assay development, and wants to know before committing to a model whether it would have chosen better than their current practice.
Send the CSV of a finished or running campaign: design columns, a response, and a batch or plate column if you have one. I replay it on its own measurements: a policy sees a small seed of designs, picks a batch, is shown the real outcomes, refits and repeats, against random selection on the same pool. Nothing is simulated.
You get a one-page replay report giving the experiments each policy needed to reach a top design, with its spread and the share of runs that reached the pool’s best; a planning table for the next campaign, seed size against refit interval; a recommendation on what to run in every batch so that drift can be removed; and the caveats printed beside the numbers, starting with how enriched your pool already is.
The screening-campaign replay project is this engagement on three published campaigns, with the tool that runs it public.
Calls and workshops
- Expert calls, one hour, through expert networks or directly, on forecasting practice and evaluation; Bayesian workflow and when it is worth the cost; bioprocess modelling, digital twins and scale-up for alternative proteins. I discuss public knowledge and my own judgement, and nothing confidential from any party.
- Half-day workshops for analytics teams: how to backtest without fooling yourself; ranges instead of points, and how to check them; a Bayesian model from prior to decision.
How to start
Send a paragraph on the decision, the data and the deadline to email address TBD. Within a week you will have a written scope, a day count and a fee, or a short note saying it is not a fit and who might be.