# Phase 3 Design Review: HermesUSstock Signal Research Pipeline

**Reviewer:** Senior ML / Quant Researcher  
**Scope:** Read-Only Architectural & Empirical Design Review (No File Modifications)  
**Status:** Backtest / Research Evaluation  

---

## Executive Summary

The **HermesUSstock Phase 3** pipeline is a well-structured, modular, and disciplined walk-forward ML research framework. The code exhibits high engineering hygiene: strict data freeze discipline (`2026-04-30`), explicit temporal purging/label guards, reproducible random seeds, and an automated registry audit trail.

However, from a quantitative research and production-readiness perspective, several critical methodology issues, statistical evaluation flaws, and code reporting bugs were identified during this review. These include nonsensical portfolio drawdown metrics, out-of-sample dynamic percentile selection during confirmation, severe test-window truncation for long-horizon scenarios, and unmodeled gap-down execution slippage.

Below are detailed answers to your 8 specific questions, followed by formal numbered findings with file/line citations, and the top 5 recommended improvements before publishing.

---

## Section I: Answers to Specific Questions

### 1. Selection Bias & Fold-6 Confirmation Evaluation
* **Is Fold-6 a trustworthy honest estimate?**  
  Fold-6 serves as a single temporal holdout window (`2026-03-15` to `2026-04-30`), evaluated after candidate model selection on Folds 1–5 out-of-fold (OOF) predictions. Within a single scenario, Fold-6 is relatively clean from candidate selection bias. However:
  1. **Cross-Scenario Selection Bias:** If scenarios are selected for deployment based on their Fold-6 metrics (e.g., favoring Scenario C because Fold-6 hit 71.75% precision / +3.10% net vs. Scenario E's +1.096%), residual multiple-testing inflation is introduced across the 5 scenario trials.
  2. **Evaluation Flaw in Fold-6 Execution (`[phase3_train.py:209](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_train.py#L209)`):** In `phase3_train.py`, Fold-6 confirmation evaluates signal performance using `coverage_metrics()`, which dynamically computes top-5% percentile rank on Fold-6 probabilities (`np.argsort(-p)[:n]`). It **fails to enforce the fixed decision score threshold** (`thr`) derived from Folds 1–5 (`[phase3_reg.py:180](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_reg.py#L180)`). In live trading, future test probability distributions cannot be sorted in hindsight; signals must be generated via a static cutoff (`P(WIN) >= thr`).

### 2. Ensemble Design: Simple Average vs. Stacking vs. Rank Average
* **Is stacking worth it?**  
  No. With only 5 temporal folds and up to 5 model candidates, training a stacking meta-model (e.g., Logistic Regression or Ridge on OOF probabilities) risks overfitting to the specific volatility regimes of 2021–2026.
* **Simple Average vs. Rank-Based Average:**  
  The pipeline currently uses a simple average (or log-AUROC weighted average) of isotonic-calibrated probabilities (`[phase3_reg.py:161-170](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_reg.py#L161-L170)`). Isotonic regression produces non-parametric step-functions, creating score ties and plateau artifacts near the 95th percentile threshold.  
  *Recommendation:* Replace calibrated probability averaging with **Percentile Rank Averaging** (converting each classifier's predictions to $[0, 1]$ fractional ranks before taking the mean). Rank averaging is non-parametric, immune to monotonic scaling differences across algorithm families, and avoids step-function calibration distortions.

### 3. Scenario E Profile (+15% / -10% / 30 Sessions)
* **Robust Edge or Lucky Trades?**  
  Scenario E achieves a 42.30% precision @ top-5% vs. a 20.66% base rate, with a positive net return (+1.391% validation, +1.096% Fold-6). While the asymmetric payoff (+15% vs. -10%) maintains positive expectancy, its statistical robustness is fragile due to **effective sample size truncation**:
  * For 30 trading sessions (~6 calendar weeks), adjacent 5-minute bars share overlapping forward outcome windows.
  * In Fold-6, because the freeze date is `2026-04-30`, the 30-session forward label window truncates valid signal bars to **only 4 trading days** (`2026-03-15` to `2026-03-19`).
  * The $n = 1,508$ signals in Fold-6 represent **less than 1 independent 30-day market window**. Thus, Fold-6 confirmation for Scenario E is effectively a single market regime snapshot.
* **How to test robustness:**  
  1. Perform **Combinatorial Purged Cross-Validation (CPCV)** or temporal block-bootstrapping to calculate block-adjusted confidence intervals for precision.
  2. Perform **Barrier Sensitivity Analysis** by perturbing target/stop thresholds (e.g., +14%/-9%, +16%/-11%, 25 vs. 35 sessions).

### 4. Labeler Implementation & Boundary Leakage
* **Lookahead / Leakage Risk:**  
  The vectorized barrier labeler in `[labeler.py](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/labeler.py)` correctly enforces next-bar-open entry (`entry_idx = np.arange(n) + 1`). However, two subtle risks exist:
  1. **Guard Sessions Leak (`[purge.py:60](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/purge.py#L60)`):** `label_guard_mask()` evaluates `entries + horizon_sessions - 1 <= train_end_session + guard_sessions` with `guard_sessions = 2`. This permits training trades whose forward label windows extend up to 2 trading sessions *into* the test fold.
  2. **Horizon Truncation at Data Freeze:** For long horizons, signals occurring within $H$ sessions of `2026-04-30` are dropped (`labeler.py:70-72`). This correctly prevents lookahead past the freeze line, but severely starves Fold-6 of test data for Scenarios D (20d) and E (30d).

### 5. Diversity Gate (Max Pairwise Correlation < 0.9)
* **Is the 0.9 threshold sound?**  
  A threshold of 0.9 correlation between predicted probabilities is extremely lenient. Classifiers with $r = 0.88$ share ~77% of their variance ($R^2$). While it prevents identical model clones, it allows highly collinear models into the ensemble.
* **Scenario D with 4 models:**  
  Ending with 4 models in Scenario D (`xgb_f2, cat_f1, lr_f2, lr_f1`) is sound and preferable to forcing a 5th model that is 0.95 correlated. However, notice that Scenario D selected **two Logistic Regression models** (`lr_f1` and `lr_f2`), making 50% of the team linear models. A family cap of 1 model per family when $N < 5$ would provide better algorithmic diversity.

### 6. Cost Model (0.10% Round-Trip Fee)
* **Realism for US Equities:**  
  A flat 0.10% (10 bps) round-trip fee is reasonable for liquid mega-caps (e.g., AAPL, MSFT, SPY) under interactive broker fee schedules (~0.5–2 bps commission + 1–2 bps spread). However, for mid/small-cap tickers in the 97-ticker universe (e.g., AXTI, BE, CRDO, IREN), 5-minute bar executions incur wider spreads (5–15 bps one-way), market impact, and overnight gap slippage.
* **Missing Execution Frictions:**  
  * **Gap-Down Slippage on Stops:** `[labeler.py:129-137](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/labeler.py#L129-L137)` assumes execution at *exact* stop price `prices * (1 - stop_pct)`. If a stock opens 14% down overnight on a 10% stop, the labeler credits a -10% return instead of the actual -14% exit price.
  * **Capital Holding Cost / Borrow:** Scenario E locks capital for up to 30 sessions (~42 calendar days). Holding costs and opportunity costs are unmodeled.

### 7. Public Website Strategy (quanttradingweb.pages.dev)
* **What to Publish:** High-level win-rates, target/stop definitions, cumulative return curves (normalized), binned calibration plots, per-scenario precision, and risk metrics.
* **What to Keep Internal:** Model joblib binaries, raw hyperparameter grids, 27 proprietary F2 feature definitions, exact real-time feature extraction code.
* **Additional Metrics Required Before Going Public:**  
  1. Annualized Sharpe and Sortino Ratios.
  2. Corrected Portfolio-Level Max Drawdown.
  3. Per-Ticker Contribution & Concentration (ensuring performance isn't driven by 2–3 volatile tech stocks like NVDA/TSLA).
  4. Feature Importance / SHAP summary plots.
  5. Empirical Reliability / Probability Calibration plots.

---

## Section II: Detailed Technical Findings

### 1. Nonsensical Cumulative Max Drawdown Calculation
* **[Severity: Critical]**  
* **Finding:** In `[phase3_common.py:215-224](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_common.py#L215-L224)`, `max_drawdown()` computes `net = s["ret"].to_numpy() - COST_RT` and `cum = np.cumsum(net)`.
* **Why It Matters:** `cum` sums uncompounded return percentages sequentially across 70,000+ non-sequential, overlapping trades across 97 tickers. In `[10pct_20d.md:20](file:///F:/GITHUB/HermesUSstock/MachineLearning/evaluations/scenario_matrix_phase1/phase3_gates/10pct_20d.md#L20)`, the gate report publishes `Max drawdown (cum net): -30745.52%`. This is a major mathematical and reporting bug that destroys the credibility of the research metrics.
* **Concrete Fix:** Reconstruct a proper daily time-series equity curve by aggregating daily realized returns weighted by active portfolio capital (or equal-weighted across active signals per day), then compute peak-to-trough percentage drawdown:
  ```python
  daily_ret = oof_signals.groupby(oof_signals["ts"].dt.date)["net_ret"].mean()
  equity_curve = (1 + daily_ret).cumprod()
  mdd = ((equity_curve - equity_curve.cummax()) / equity_curve.cummax()).min()
  ```

---

### 2. In-Sample Dynamic Percentile Selection in Fold-6 Evaluation
* **[Severity: Important]**  
* **Finding:** In `[phase3_train.py:209](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_train.py#L209)`, Fold-6 confirmation evaluates signals via `conf_rec = coverage_metrics(te_df["win"], cp, te_df["realized_return"])`. Inside `coverage_metrics()` (`[phase3_common.py:166](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_common.py#L166)`), `order = np.argsort(-p)[:n]` selects the top 5% highest probability scores *within Fold-6*.
* **Why It Matters:** The probability threshold `thr` locked on Folds 1–5 (`[phase3_train.py:194](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_train.py#L194)`) is ignored. In live trading, future test set probability distributions cannot be sorted in hindsight; trades must be triggered by `p >= thr`. Dynamic percentile sorting in Fold-6 overstates confirmation performance if the Fold-6 score distribution shifted.
* **Concrete Fix:** Update Fold-6 evaluation in `[phase3_train.py](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_train.py)` to apply the static threshold `thr` determined during validation:
  ```python
  sig_mask = cp >= ens["threshold"]
  conf_prec = te_df.loc[sig_mask, "win"].mean() if sig_mask.any() else 0.0
  conf_net = te_df.loc[sig_mask, "realized_return"].mean() - COST_RT if sig_mask.any() else 0.0
  ```

---

### 3. Fold-6 Sample Size Truncation for Long-Horizon Scenarios
* **[Severity: Important]**  
* **Finding:** For Scenario E (30-session horizon), `[labeler.py:70-72](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/labeler.py#L70-L72)` drops bars whose forward window extends past `2026-04-30`. Consequently, the last labeled bar is `2026-03-19`.
* **Why It Matters:** Fold-6 test start is `2026-03-15`. Thus, Fold-6 for Scenario E contains **only 4 trading days of labeled data** (`2026-03-15` to `2026-03-19`, $n=1,508$ signals). Evaluating a 30-day holding horizon strategy on 4 days of test signals provides less than 1 independent trade horizon, rendering Fold-6 confirmation uninformative for Scenarios D and E.
* **Concrete Fix:** Extend the dataset history further into the past or adjust Fold-6 test start boundaries proportional to the scenario horizon $H$ (e.g., $test\_window\_length \ge 3 \times H$).

---

### 4. Unmodeled Gap-Down Execution Slippage on Stop-Losses
* **[Severity: Important]**  
* **Finding:** In `[labeler.py:128-137](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/labeler.py#L128-L137)`, when a stop loss is triggered (`defeats`), `exit_price` is set to `prices * (1 - stop_pct)` and `realized` return is set to `-stop_pct`.
* **Why It Matters:** If a stock opens -15% lower overnight on earnings, a -10% stop loss will actually execute at -15% (or worse). The labeler artificially truncates losses at exactly -10.0%, overstating strategy net expectancy in volatile market regimes.
* **Concrete Fix:** Update `[labeler.py](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/labeler.py)` so that when a stop is hit on bar $k$, the exit return uses $\min(\text{stop\_pct}, 1 - \text{low}_k / \text{entry\_price})$ or the Open price of the gap bar if the bar opened below the stop.

---

### 5. Label Guard Allowance of 2-Session Test Period Overlap
* **[Severity: Minor]**  
* **Finding:** `[purge.py:60](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/purge.py#L60)` evaluates `entries + horizon_sessions - 1 <= train_end_session + guard_sessions` with `guard_sessions = 2`.
* **Why It Matters:** This allows training rows whose label windows extend up to 2 sessions into the test period, introducing a slight 2-session lookahead leak from test fold prices into training set binary labels.
* **Concrete Fix:** Set `guard_sessions = 0` in `[phase3_common.py:146](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_common.py#L146)` and `[purge.py:45](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/purge.py#L45)` to guarantee complete label separation.

---

### 6. Isotonic Step-Function Calibration Artifacts
* **[Severity: Minor]**  
* **Finding:** `[phase3_reg.py:94-96](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_reg.py#L94-L96)` fits `IsotonicRegression(out_of_bounds="clip")` on pooled OOF predictions.
* **Why It Matters:** Isotonic regression produces piecewise constant step-functions. When multiple model probabilities are averaged, step-function plateaus generate identical probability clusters, causing arbitrary tie-breaking when ranking signals at the top-5% threshold.
* **Concrete Fix:** Use `CalibratedClassifierCV` with `method="sigmoid"` (Platt scaling) or apply Gaussian kernel smoothing to the isotonic mapping curves.

---

## Section III: Top 5 Improvements Before Going Public

1. **Fix Portfolio Drawdown & Equity Curve Calculation (`[phase3_common.py:215](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_common.py#L215)`):** Replace the uncompounded cumulative sum of trade returns with a true daily portfolio simulation (allocating capital across active signals per day) and publish realistic Max Drawdown and Sharpe Ratios.
2. **Enforce Static Score Thresholds in Holdout Confirmation (`[phase3_train.py:209](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_train.py#L209)`):** Evaluate Fold-6 using the fixed probability threshold (`P(WIN) >= thr`) derived from validation folds rather than re-sorting Fold-6 predictions in hindsight.
3. **Transition to Percentile Rank Ensemble Averaging (`[phase3_reg.py:161](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/phase3_reg.py#L161)`):** Convert candidate output probabilities to percentile ranks before averaging to eliminate calibration step-function artifacts and handle algorithm score scaling differences gracefully.
4. **Incorporate Gap-Slippage into Label Realized Returns (`[labeler.py:128](file:///F:/GITHUB/HermesUSstock/MachineLearning/src/labels/labeler.py#L128)`):** Account for overnight gaps and fast-market slippage on stop-loss exits rather than crediting exact barrier return values.
5. **Publish Per-Ticker Risk Concentration & SHAP Diagnostics:** Include ticker concentration analysis (proving edge is not dependent on 2–3 mega-cap tech stocks) and binned calibration reliability diagrams on the public website.
AGY_REVIEW_EXIT=0
