Forecast Accuracy Review
SkillAI & modelsForecast-accuracy-review is a skill that evaluates demand-forecast quality using error metrics such as WMAPE, bias, and Forecast Value Added measured against a naive benchmark. It runs a rolling-origin backtest to show whether a forecasting process or tool is actually performing well.
Available today. Use it from your connected AI after setup.
No other account needed.
Have demand-forecast data available for the period you want to evaluate.
Then ask your AI: use the Forecast Accuracy Review skill
What your AI can do with it
- Compute WMAPE, bias, and Forecast Value Added for demand forecasts
- Compare forecast accuracy against a naive benchmark
- Run a rolling-origin backtest to evaluate forecast quality
- Assess whether a forecasting process or tool performs well
- Report demand planning performance in plain terms
Getting started
- Have demand-forecast data available for the period you want to evaluate.
- Add the forecast-accuracy-review skill to your agent setup.
- Ask the agent to review forecast accuracy, mentioning the metrics or questions you care about.
- Review the reported WMAPE, bias, and Forecast Value Added results against the naive benchmark.
What this skill tells your AI
The instructions your AI receives, as published by davila7/claude-code-templates in cli-tool/components/skills/operations/forecast-accuracy-review/SKILL.md and read by ahel’s review.
A forecast is only worth what it adds over the free alternative: shipping last period's number. Every review must answer "how many points does this process add over naive?" before any model discussion.
Required data
Per-SKU demand history at the planning bucket (usually monthly): sku, period, qty. If evaluating an existing forecast, also the forecast values with their creation dates (to avoid hindsight leakage). 18+ periods per SKU for a meaningful backtest; flag SKUs with less.
Workflow
- Profile the demand first. Per SKU compute mean, CV and zero-period share; classify smooth / erratic / intermittent / lumpy (defaults: CV 0.5 and 1.0 boundaries, intermittency at >25% zero periods - state them, adjust to natural breaks). Accuracy expectations differ by class; never report one blended number alone.
- Set the benchmarks. Naive (last period) always; seasonal naive when 2+ full seasons exist. These are non-negotiable controls.
- Backtest rolling-origin. One-step-ahead forecasts for each of the last 6+ periods, expanding window, using only data before each origin. A single train/test split is one lucky draw - do not accept it.
- Score with honest metrics:
- WMAPE = sum(|error|) / sum(actual) - the volume-weighted headline
- Bias = sum(error) / sum(actual) - direction; a fine WMAPE with persistent bias is quietly building excess stock or stockouts
- MAPE only as a footnote, and always disclose how many zero-actual periods it dropped
- Deliver the FVA verdict. FVA = WMAPE(naive) - WMAPE(candidate), per segment and overall. Negative FVA means the process destroys value - say it plainly.
- Validate. Recompute WMAPE for one model directly from the raw backtest rows and confirm it matches the table before presenting.
Pitfalls to check explicitly
- MAPE with zeros: undefined on zero-actual periods; silently dropping them fakes precision on intermittent SKUs.
- MAPE asymmetry rewards under-forecasting (errors capped at 100% below, unbounded above).
- Aggregation mix: a good total can hide terrible A-item accuracy; always show the value-weighted cut.
- Lumpy segments: if WMAPE > ~100%, the honest recommendation is an inventory-policy answer (buffers, MTO), not a better model.
- Hindsight leakage: forecasts must predate actuals; check timestamps when auditing an existing process.
Output format
- Scoreboard table: model x (WMAPE, bias, MAPE-footnote), sorted by WMAPE
- FVA statement: "the process adds/destroys X points vs naive" - overall and per segment
- Segment table (pattern x best approach)
- Two or three recommendation sentences tied to segments, not globals
Worked example with five baseline models and charts: https://github.com/gulmezeren2-byte/forecast-accuracy-lab
Source: industrial-engineering-ai-skills by Eren Gulmez (MIT). The full method pack - entry skill, role agents, data-hygiene rules and artifact templates - lives there.
Signals
- GitHub stars
- 32k
- Forks
- 4k
- Last commit
- Sep 2026
Questions
- What does it measure?
- It measures forecast accuracy using WMAPE, bias, and Forecast Value Added, comparing results against a naive benchmark over a rolling-origin backtest.
- When should it be used?
- Use it when forecast accuracy, MAPE, or demand planning performance comes up, or when you want to know whether a forecasting process or tool is any good.
- Why compare against a naive benchmark?
- A naive benchmark shows whether forecasts add value over a simple baseline, which is what Forecast Value Added captures.
Advanced
- Item type
- skill
- Key
forecast-accuracy-review- Source
- github.com/davila7/claude-code-templates