Offers

Three fixed-scope engagements in probabilistic forecasting and Bayesian modelling, priced at £500 a day or as a fixed fee against a written scope.

I take on four kinds of engagement. Each is scoped in writing before it starts and priced by the day or as a fixed fee, and each ends with a one-page memo for the person who has to make the decision and a repository for whoever inherits the model after me.

Two days a week from start month TBD. Remote from the UK, with on-site days by arrangement.

£500 a day. Fixed fees against a written scope on request. Inside or outside IR35, as the engagement requires.

Fedelm Advisory invoices directly. If your organisation buys through a framework or an intermediary, tell me and we will find the route.

In every engagement

Whatever the size of the work, the same things are true of what I hand over.

  • The claim is written down, with the result that would refute it, before I test it.
  • Baselines are fitted and reported before any candidate model: seasonal naive, ETS and ARIMA for forecasting; unpooled and fully pooled models for hierarchical work.
  • Uncertainty is delivered as quantiles or posterior samples, and the coverage of the intervals is measured on held-out data rather than assumed.
  • Backtests use rolling origins and strict time splits, with features built only from what was known at each origin. Where the source data is revised after publication, I archive the vintages and evaluate on what existed at the time.
  • The answer is expressed in your units, beds, staff, stock, pounds or experiments, and scored on the cost of getting it wrong.
  • The code installs and tests in two commands, and the README gets a stranger to the baseline table.
  • Where a simpler model beat a more complex one, the write-up says which and by how much.

Forecast audit

2–4 weeks, typically 4 to 8 days.

You already have a forecast: a spreadsheet, a vendor’s tool, or a model built by someone who has since moved on. It feeds a plan, and nobody has recently checked whether it deserves to.

I check five things.

  • Does it beat a seasonal naive baseline on a proper backtest with rolling origins, and by how much?
  • Do its ranges hold? A nominal 90% interval should contain about 90% of outcomes. Most do not, and whether the miss is too wide or too narrow changes what you should do about it.
  • Is there leakage: features or revisions that were not available when each forecast was made?
  • How much does its accuracy fall on the data you actually had at each origin, as opposed to the revised file you have now?
  • What would you need to watch each month to know when it has stopped working?

You get a written audit with a baseline table and calibration plots; a one-page memo saying whether to rely on the forecast, for which horizons and with what range; and a monitoring plan your own team can run.

This is not a rebuild. If the audit finds the forecast should be replaced, the harness I built to test it carries straight into the next engagement.

Probabilistic demand forecasting

8–12 weeks, two days a week, typically 16 to 24 days.

For a planning decision that repeats: winter beds, call-centre staffing, stock levels, energy, or anything else counted weekly or monthly across many sites. The decision currently rests on a point forecast, last year’s number or someone’s judgement, and the cost of being wrong falls mostly on one side.

How it runs:

  • Weeks 1–2: data audit; a vintage archive if the source is revised; hypotheses written down with refutation rules and frozen before evaluation.
  • Weeks 3–6: baselines (seasonal naive, ETS, ARIMA), then the candidates, all scored on rolling-origin backtests with proper scoring rules for the whole predictive distribution.
  • Weeks 7–8: the decision table in your units, scored on the cost of getting it wrong; a scheduled pipeline; the handover pack.
  • Weeks 9–12, if scoped: hierarchy and reconciliation across sites or regions; cold-start handling for new units; a scorecard that updates itself after each data release.

You get a repository that installs and tests in two commands; a scheduled pipeline (GitHub Actions or your equivalent) that refits, forecasts and scores itself; a scorecard; a handover pack covering backfill, retraining, drift triggers and what to do when the feed breaks; a technical write-up; and a one-page memo giving the range, the drivers, the assumptions and how much to trust it.

The NHS A&E project is this engagement run in public, on open data, with every result posted.

Bayesian measurement

Fixed scope or multi-month, quoted against the scope.

For questions where a single number would mislead and the data arrives from many units at once: which channels or components drive the response and where they saturate; how demand responds to price across stores; what a newly launched unit can borrow from the ones that came before it.

I build these as hierarchical models in PyMC and deliver the workflow that makes them defensible when someone senior asks how you know:

  • prior predictive checks, and a record of any prior change that mattered;
  • an identifiability audit by simulation-based calibration, so you know which parameters your data can pin down before you commission the full model;
  • posterior predictive checks and calibration against any experiments or holdouts you have;
  • a production wrapper if the model will be refit: diagnostic gates in CI, versioned priors, drift alerts on the predictive checks.

You get the model and its tests in a repository; the identifiability report; a decision output in your units with its posterior uncertainty; a write-up positioned against the published method you would otherwise have used; and the one-page memo.

The screening-campaign replay project carries the Bayesian workflow this engagement uses: pre-registered comparisons, sampler health gates in CI, proper scoring of the predictive distribution, and a decision log with every rejected alternative.

Campaign replay audit

1–2 weeks, fixed fee.

For an R&D team that runs screening campaigns in batches, in media, formulation, fermentation conditions or assay development, and wants to know before committing to a model whether it would have chosen better than their current practice.

Send the CSV of a finished or running campaign: design columns, a response, and a batch or plate column if you have one. I replay it on its own measurements: a policy sees a small seed of designs, picks a batch, is shown the real outcomes, refits and repeats, against random selection on the same pool. Nothing is simulated.

You get a one-page replay report giving the experiments each policy needed to reach a top design, with its spread and the share of runs that reached the pool’s best; a planning table for the next campaign, seed size against refit interval; a recommendation on what to run in every batch so that drift can be removed; and the caveats printed beside the numbers, starting with how enriched your pool already is.

The screening-campaign replay project is this engagement on three published campaigns, with the tool that runs it public.

Calls and workshops

  • Expert calls, one hour, through expert networks or directly, on forecasting practice and evaluation; Bayesian workflow and when it is worth the cost; bioprocess modelling, digital twins and scale-up for alternative proteins. I discuss public knowledge and my own judgement, and nothing confidential from any party.
  • Half-day workshops for analytics teams: how to backtest without fooling yourself; ranges instead of points, and how to check them; a Bayesian model from prior to decision.

How to start

Send a paragraph on the decision, the data and the deadline to email address TBD. Within a week you will have a written scope, a day count and a fee, or a short note saying it is not a fit and who might be.