Research

Pre-registration

The hypothesis, the universe, the evaluation window and the stop conditions are written down and frozen before the first backtest runs.

A specification that cannot fail is not a specification. Every programme states in advance what result would end it.

Parameters are fixed before evaluation, so a disappointing run cannot be quietly re-specified into a good one.

When a programme hits its stop condition we close it and record why, including the variants we agreed not to try next.

Out-of-sample discipline

In-sample performance is a description of the past. Everything is evaluated walk-forward, across ten horizons from one month to five years.

A result that only exists on one lookback window is a property of that window.

Holdout periods are reserved at the outset and touched once.

Regime changes are treated as information about robustness, not as an excuse to exclude a period.

Multiple-testing correction

Test enough ideas and some will look excellent by accident. We deflate for the number of trials rather than reporting the survivor and calling it an edge.

Deflated Sharpe ratios with an explicit floor agreed before the search begins.

Candidates are grouped into families so a hundred variations of one idea count as one idea.

The size of the search is recorded alongside the result, because a Sharpe ratio without a trial count means nothing.

Realistic cost modelling

Costs are modelled from actual order-book depth at the size being traded, not assumed as a flat spread. It is the most common way a paper edge disappears.

Fill prices derived from depth at the moment of execution.

Fee schedules modelled exactly as the venue charges them, including where the formula is non-linear in size.

Latency and partial fills treated as costs rather than as rounding.

Paper before capital

Every candidate runs as a live paper soak on real market data before any conversation about capital. Most do not survive that stage.

Soaks run against live feeds with the same execution path production would use.

Instrumentation is added before the soak starts, so a silent failure is visible rather than inferred afterwards.

A soak that produces the expected null is a successful soak. It has told us something true.

Publishing the nulls

Internally, a failed programme is documented as thoroughly as a successful one — the mechanism, the evidence, and the reason not to revisit it.

Closed programmes carry a written finding so the same idea is not rediscovered in six months.

Where a whole venue is ruled out on cost grounds, that conclusion is recorded and enforced against new proposals.

The archive of what does not work is the most valuable thing the research function owns.