Combinatorial Purged Cross-Validation

SkillDev tools

Combinatorial Purged CV generates a distribution of backtest paths instead of a single estimate. Use when quantifying strategy robustness and overfitting probability.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Combinatorial Purged Cross-Validation skill

What this skill tells your AI

The instructions your AI receives, as published by ml4t/skills in validation/cpcv/SKILL.md and read by ahel’s review.

Standard k-fold CV on time series produces one biased performance estimate. CPCV generates C(N,k) train/test combinations with purging and embargo, yielding a distribution of results that reveals overfitting.

The Problem

A single train/test split gives one Sharpe ratio - you cannot tell if it is skill or luck. Standard k-fold shuffles temporal order, leaking future information. Even TimeSeriesSplit produces only a handful of sequential folds, each with different train sizes, making comparison unreliable. You need many unbiased performance samples to build a distribution.

The Pattern

Partition data into N groups, select k as test sets, train on the rest. Purge samples whose labels overlap the test boundary, add an embargo buffer. Repeat for all C(N,k) combinations.

WRONG

from sklearn.model_selection import KFold

# Shuffled k-fold on time series - future leaks into training
cv = KFold(n_splits=5, shuffle=True, random_state=42)
scores = []
for train_idx, test_idx in cv.split(X):
    model.fit(X[train_idx], y[train_idx])
    scores.append(model.score(X[test_idx], y[test_idx]))
print(f"Mean score: {np.mean(scores):.3f}")  # Overly optimistic

CORRECT

import numpy as np
from itertools import combinations

# Manual CPCV with purging using standard tools
n_groups, n_test, horizon, embargo = 8, 2, 5, 2
n_samples = len(X)
# array_split, not n_samples // n_groups: fixed-width groups leave the
# remainder outside every test group, so those rows are never tested.
groups = np.array_split(np.arange(n_samples), n_groups)
scores = []

for test_groups in combinations(range(n_groups), n_test):
    test_mask = np.zeros(n_samples, dtype=bool)
    for g in test_groups:
        test_mask[groups[g]] = True

    # Purge: remove training samples within horizon of test boundaries
    train_mask = ~test_mask.copy()
    for i in np.where(np.diff(test_mask.astype(int)) != 0)[0]:
        purge_start = max(0, i + 1 - horizon)
        purge_end = min(n_samples, i + 1 + embargo)
        train_mask[purge_start:purge_end] = False

    model.fit(X[train_mask], y[train_mask])
    scores.append(model.score(X[test_mask], y[test_mask]))  # one split, not a path

# C(8,2) = 28 split scores, which assemble into N-1 = 7 backtest paths
print(f"Mean: {np.mean(scores):.3f}, Std: {np.std(scores):.3f}")

Parameter Selection

n_groupsn_test_groupsCombinationsUse case
6215Small datasets
8228Standard
103120Deep analysis
  • Heuristic: set n_groups = desired paths + 1, n_test_groups = 2 (de Prado)
  • label_horizon: must match label construction (5-day returns = 5)
  • embargo_size: ~10-20% of label_horizon (prevents serial correlation post-test)

Guardrails

  • More paths → lower variance of the mean Sharpe estimate (var ∝ 1/φ when paths are uncorrelated), directly reducing false discovery
  • Verify training set size after purging is still sufficient (>60% of data)
  • Combine with PBO / Deflated Sharpe Ratio (see deflated-sharpe skill) for statistical significance
  • Never report the best fold - report the full distribution (mean, std, worst-fold)

Production Implementation

ml4t-diagnostic provides a validated, sklearn-compatible splitter:

from ml4t.diagnostic.splitters import CombinatorialCV

cv = CombinatorialCV(
    n_groups=8,
    n_test_groups=2,
    label_horizon=5,
    embargo_size=2,
    max_combinations=28,
    random_state=42,
)
for train_idx, test_idx in cv.split(X):
    model.fit(X[train_idx], y[train_idx])
    scores.append(model.score(X[test_idx], y[test_idx]))

Checklist

  • Using CPCV (not KFold or single split) for strategy evaluation
  • label_horizon matches actual label construction
  • embargo_size > 0 for autocorrelated features
  • Reporting distribution statistics (mean, std, min), not single score
  • Training set size after purging verified as sufficient

Signals

GitHub stars
20
Forks
11
Last commit
Sep 2026
Advanced
Item type
skill
Key
ml4t-cpcv
Source
github.com/ml4t/skills