Strategy Development Workflow

SkillCommerce & finance

End-to-end strategy development lifecycle from hypothesis to live trading. Use when starting a new strategy project or onboarding to the ML4T workflow.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Strategy Development Workflow skill

What this skill tells your AI

The instructions your AI receives, as published by ml4t/skills in workflows/strategy-workflow/SKILL.md and read by ahel’s review.

Strategies fail because developers skip straight to modeling. The correct process spends most time on hypothesis and data, with modeling as a small fraction.

The Problem

A researcher downloads data, fits a model, runs a backtest, and sees a 2.5 Sharpe ratio. They deploy, then lose money because there was no documented hypothesis, no feature validation, no cost model, and no holdout.

The Pattern

WRONG

# Jump straight to modeling - no hypothesis, no validation gates
import lightgbm as lgb

data = load_data()
features = data[["momentum", "volatility", "volume"]]
labels = data["next_day_return"]

model = lgb.LGBMRegressor().fit(features, labels)
predictions = model.predict(features)  # Predicting on training data!

# "Looks great, ship it"
print(f"R2: {r2_score(labels, predictions):.3f}")  # 0.95 - overfitting

CORRECT

# Seven stages with quality gates between each
hypothesis = {
    "mechanism": "Momentum persists due to slow institutional rebalancing",
    "signal": "12-month risk-adjusted return predicts 1-month forward return",
    "kill_criteria": "IC < 0.02 or non-monotonic quintiles",
    "capacity": "ETFs with >$50M daily volume",
}
# Commit term sheet to version control before proceeding

# Stage 2: Data  → fetch + validate
# Stage 3: Features → compute + factor research
# Stage 4: Model → CPCV + model validation
# Stage 5: Backtest → realistic costs
# Stage 6: Paper trade → live data, simulated fills
# Stage 7: Live → small size, slow ramp

Quality Gates

GateConditionFail Action
HypothesisDocumented mechanism, kill criteria, capacity estimateDo not start coding
DataNo gaps > 2 days, point-in-time correct, survivorship-freeFix data pipeline
FeaturesIC > 0.02 (HAC-adjusted), stable across subperiodsDrop factor or redesign
ModelLoss rate < 50%, best config clears the selection boundSimplify model or revisit features
BacktestSharpe > 0.5 net of costs, max DD < 20%Revise sizing or cost assumptions
Paper tradeFills within expected slippage, no execution anomaliesFix execution logic

Guardrails

  • If Sharpe > 2.0 on daily equity data, assume lookahead bias or selection bias until proven otherwise - inspect with ml4t-lookahead-bias and ml4t-deflated-sharpe
  • If in-sample and out-of-sample performance match closely, suspect data leakage
  • If the strategy requires > 20% annual turnover to work, verify cost assumptions with ml4t-transaction-costs
  • If no documented hypothesis exists, stop and write one before any other work

Production Implementation

import asyncio

from ml4t.backtest import Strategy, run_backtest, BacktestConfig  # MyStrategy subclasses Strategy
from ml4t.live import LiveEngine, AlpacaBroker, AlpacaDataFeed

results = run_backtest(
    prices=prices, signals=signals, strategy=MyStrategy(), config=BacktestConfig()
)

async def trade_live():
    broker = AlpacaBroker(api_key, secret_key, paper=True)
    feed = AlpacaDataFeed(api_key, secret_key, symbols=["SPY"], experimental=True)
    engine = LiveEngine(MyStrategy(), broker, feed)
    await engine.connect()
    await engine.run()

asyncio.run(trade_live())

Checklist

  • Hypothesis documented in term sheet before any code
  • Data validated for gaps, survivorship bias, point-in-time correctness
  • Features pass IC significance test with HAC standard errors
  • Model validated via CPCV; loss rate < 50% (PBO: ml4t-backtest-overfitting)
  • Backtest includes realistic transaction costs
  • Deflated Sharpe computed across all trials (ml4t-deflated-sharpe)
  • Paper trading completed for minimum 4 weeks
  • Kill criteria defined and monitoring configured before going live

Signals

GitHub stars
20
Forks
11
Last commit
Sep 2026
Advanced
Item type
skill
Key
ml4t-strategy-workflow
Source
github.com/ml4t/skills