loading…Scoreboard …Right now — current conditions (click to expand)
Right now — current conditions
Field
Base forecast (L1)
Production
Correction
Field
Base forecast (L1)
Production
Correction
Loading current tick…
Confidence: —· — stations reporting
Briefing source: —· —
Base forecast (L1) = the selector's chosen raw source fed into our correction stack (HRRR raw or NBM raw, per field per lead). Production = final Wyman Cove forecast (after L2/L3/L4/gates). Correction = Production − L1 for this tick.
Total Lift
Prod vs NBM raw (the field your default weather app already shows you). The referee's number.
7 Day
median
—
mean
—
24 Hour
median
—
mean
—
1 winning
ch
2 flat
cl · cm
8 losing
t · h · dp · ws · wg · wd · sr · cc
Pipeline Lift
Local correction stack value: Prod vs L1 selector's pick. Answers "are the pipelines pulling their weight on top of raw?"
7 Day
median
—
mean
—
24 Hour
median
—
mean
—
3 winning
cl · cm · ch
0 flat
—
8 losing
t · h · dp · ws · wg · wd · sr · cc
Selector Skill
Value Captured: of the routing gain a perfect picker could have captured, how much did we actually get? Magnitude-aware and n-weighted.
7 Day · Value Captured
median
—
mean
—
24 Hour · Value Captured
median
—
mean
—
9 winning
t · h · dp · ws · wg · wd · sr · cc · ch
0 flat
—
0 losing
—
Attribution — Routing · Cascade · Total (median, pp of user default)
Median across fields per column — the typical field's split. Three independent stats: medians don't add, so Total is its own median, not Routing + Cascade.
7 Day
routing
—
cascade
—
total
—
24 Hour
routing
—
cascade
—
total
—
Attribution — Total = Routing + Cascade (mean, pp of user default)
Additive split of Total Lift. Routing = pick vs user default (0 when the pick IS the default — most rows today). Cascade = local stack on top of that pick.
7 Day
routing
—
cascade
—
= total
—
24 Hour
routing
—
cascade
—
= total
—
Prod Trend
Change vs the prior equal window. Are we improving over time?
7 Day
median
—
mean
—
24 Hour
median
—
mean
—
4 improving
t · cm · ch · cc
3 flat
cl · pr · pp
4 regressing
h · dp · ws · sr
Notable Calls (7D)
Biggest wins and losses at the (field, lead-band) level. Referee highlights.
Best
loading…
Worst
loading…
2nd worst
—
3rd worst
—
Top-1 win + top-3 losses at (field, lead-band). Scored on current-config counterfactual Total Lift (v0.6.494) so killed layers don't drag the rankings. dp + cc excluded (derived); pa + pp excluded (non-MAE); n < 200 filtered.
National Source (7D · lower MAE wins)
Which national raw source is currently scoring better per field. Diagnostic only — Total Lift already scores against NBM raw (or HRRR raw for the 5 HRRR-only fields), so National Source doesn't set the baseline; it's context for interpreting where the pipeline's lift is coming from.
—
loading…
Health & Reliability (7D)
Trust indicators. Independent of who's winning.
Confidenceloading…
High-conf cellsloading…
Halves agreeloading…
Halves agree below ~70% is a noisy week — read the score tiles above with caution. Halves agree = both halves of the 7d window show the same sign of lift.
Current state
Stack health trajectory — aggregate Prod vs 90d-ref raw, per day (12 fields, pa/pp excluded)
Per-obs-day cross-field Median + Mean + P25–P75 band of (1 − prod_MAE / raw_MAE_90d_ref) × 100. Denominator is each field's fixed 90-day-reference raw MAE — daily weather difficulty is absorbed by the fixed baseline, so the ratio only moves when Prod moves. Positive = Prod beating the 90d baseline; a ship that improves the stack pushes the 7d rolling median UP over the following week. Full per-field trajectory + drill-down: Accuracy over time ↓.
Per-field diagnostic — 7 day
Coach's table: each field's total, the pipeline lift on the selector's chosen stack (mirrors the tile), the two cascade skills (independent of routing), and the two selector-quality views. If Total Lift is red, this table shows which lever to pull. Live from per_field_scoring.json (refits hourly). Win Rate + Value Captured are diagnostic on the paired chosen/alt Prod pool — NBM-scope fields only. dp and cc omitted — both derived, no independent skill to score (see note below).
Field
Total LiftProd vs Public Baseline
Difficulty7d raw MAE ÷ 90d ref · >1 = harder than usual
Pipeline Lift(L1 pick − Prod) / L1 pick
HRRR Pipeline Skill(raw − Prod) / raw
NBM Pipeline Skill(raw − Prod) / raw
Win Ratewins / (wins + losses); ties excluded
Value Captured% of oracle
n
loading per_field_scoring.json…
dp and cc omitted — both derived at prod-time (dp = Magnus(prod_t, prod_h); cc = Ccd max(prod_cl, prod_cm, prod_ch)) with no independent skill chain. Future dp bias layer or cc composition tuner would restore their rows.
🪦 cfg X% annotation on a Total Lift cell = current-config counterfactual (v0.6.494). Computed by walking the cascade stamps and skipping layers currently ENABLED=False. Big delta between the as-shipped number and cfg = a recent kill is still poisoning the rolling window; expect it to fade over ~7-30 days as new rows accumulate. Snapshot 2026-08-26: skips error_l5_nbm (killed 08-25) + error_l6_nbm (scaffold). HRRR-side counterfactual deferred (needs per-field disable logic because HRRR layer slots are field-dependent).
Windows: all six lift/skill columns computed on the same paired 7d/24h pool (v0.6.492) — rows where every relevant residual (HRRR raw, NBM raw, selector output, Prod, both cascade Prods) exists. Total Lift, Pipeline Lift, and per-cascade Pipeline Skill columns share denominators (each vs its own reference: user default, L1 pick, own raw), so causal chaining is meaningful: if Total Lift moves, at least one of Pipeline Lift, a Pipeline Skill, or the selector's Win Rate moved on the same rows to cause it.
Per-field diagnostic — 24 hour
Same columns as the 7d table above, computed over the last 24 hours of paired residuals. Read this table for "did anything break today" — a field whose 24h row goes red while its 7d row stays green flagged a regression the daily view is still averaging away. Live from per_field_scoring.json's windows.24h.per_field block. Same aggregation rules as the 7d table (dp/cc derived-omitted; pa/pp non-MAE excluded from Median/Mean).
Field
Total LiftProd vs Public Baseline
Difficulty7d raw MAE ÷ 90d ref · >1 = harder than usual
Same columns as the 24h table, computed over the last 12 hours of paired residuals. Read this AFTER a ship — the 24h and 7d windows lag any override/selector-fit change by up to 24h/7d respectively, so a fresh routing change or L2 fix reads as noise there. 12h reads first. Thin-window caveat: pair-log typically has an 8h backstamp lag, so the 12h window is really "recent 4h of ripe pair-log rows"; fields with n≤50 read as noise, not signal. Live from per_field_scoring.json's windows.12h.per_field block. Same aggregation rules as 7d/24h.
Field
Total LiftProd vs Public Baseline
Difficulty7d raw MAE ÷ 90d ref · >1 = harder than usual
Pipeline Lift(L1 pick − Prod) / L1 pick
HRRR Pipeline Skill(raw − Prod) / raw
NBM Pipeline Skill(raw − Prod) / raw
Win Ratewins / (wins + losses); ties excluded
Value Captured% of oracle
n
loading per_field_scoring.json…
Per-field pipeline architecture + status
Companion to the per-field scoring table above (numeric selection/correction/total lifts). This section shows the pipeline shape for each field — HRRR-side cascade, NBM-side cascade (if in scope), which side the selector picks — plus hand-curated narrative context (open work, recent ships, current state). Numeric cells live in the scoring table; this table is the architecture + journal.
Positive vs raw on 7d. Numeric values live in the per-field scoring table above (7d/24h/12h). Structural state (09-29):L3_NBM dropped 09-05 v0.6.551 (sentry HOT + 14/15 cells help_fresh negative), so NBM cascade is raw_nbm → l2_nbm only. HRRR side runs the full L1→L2 additive → L4 → chp/dpbp/lsb/C1 stack. Routing has been carrying the field vs cascade at various times; v0.6.588 escalation cell h/calm/24-47 live since 09-11. Current shadow: v0.7.6 L1 static blender covers h with universal ω=0.44 on 10 cells (SHADOW ENABLED=False); 09-29 retro shows h/nw_flow/24-47 +39.4% SHIP-READY (n=336, below gate 400 — narrow-flip decision ~10-02).
Best-performing field vs raw. Full stack L1→L2→L3→L4→Lc→chp. 10-cell _CELL_SKIP forces mid/long-lead cells back to L4 (v0.6.405/409 emergency demotes).
raw_nbm → l2_nbm (l3_nbm dropped 08-25 v0.6.472; l5_nbm KILLED 08-25 v0.6.471 — sentry +238% MAE + walkforward −126%/−145% agreed; sr has no HRRR L2 so l2_nbm ≈ raw_nbm)
loading…
Lsr LIVE, skip ne_flow + calm. Lsb sea_breeze cc<25 override LIVE. NBM cascade minimal — l3_nbm/l5_nbm both killed for sr. Selector pick per band shown live in the Selector column above; direction can flip with recent-window MAE.
L2 wind_blend + wdp both LIVE. NBM cascade full (l2/l3/wdp_nbm).
Note: the narrative snippet below is the LEGACY per-layer MAE narrative — it reads mae_over_time.json's last_7d block (with an unweighted-across-leads fallback when that block is empty). It is NOT the top scoreboard's source. The scoreboard's Total Lift / Pipeline Lift / Notable Calls / High-Conf tiles all read per_field_scoring.json's paired-pool numbers (v0.6.492+), which don't touch this fallback. Kept here as a per-field cross-check; retire when the paired-pool numbers are trusted for the same purpose.
Loading 7-day per-layer narrative from mae_over_time.json…
Loading provenance…
What's running · improving · being evaluated next
What's running
Stack
L2 mesonet blend: t · h · dp · cc · cl · cm · ch · ws · wg (pr on nw_flow 0-11h since 08-10)
L3 lead-decay: wg · ch (pp dropped 2026-07-04 v0.6.304; ws dropped 2026-08-08 v0.6.397; cm dropped 2026-09-15 v0.6.625 — walkforward L3 fc −1.6% / obs −1.2% pooled, calm/24-47h LOSS −17.2%; wg L3 skip cells: calm 0-5 + 12-23 + 24-47, ne_flow 6-11, sea_breeze 6-11 + 12-23 + 24-47, frontal 12-23)
L4 diurnal: {ch}· cc dropped 2026-09-08 v0.6.563 (30d marginal lift +0.79% pooled, best skip-table only +1.38% — not worth the maintenance)
Lsr synoptic (sr): skip ne_flow + calm · Lt reclassified as telemetry experiment 07-27 (Fix B held-out +0.29%)
NBM parallel cascade — AT ARCHITECTURAL PARITY (v0.6.499, 2026-08-26): raw_nbm for 9 fields; native L2 for all 8 NBM-scope fields (t/h/dp additive Kalman+decay, ws/wg/wd wind_blend, cc/ch cloud_obs_blend hourly[0]; sr identity — no L2 either side); l3_nbm for wg/ch/sr (scope trimmed 08-25 v0.6.472; sr re-added 09-04 v0.6.549 then per-cell skipped 09-29 v0.7.11 on nor_easter/12-23h + /24-47h; h dropped 09-05 v0.6.551; cm dropped 09-15 v0.6.625; cc dropped 09-10 v0.6.577); l4_nbm for ch only (cc dropped 09-08 v0.6.563); l5_nbm for sr KILLED 08-25 v0.6.471; l6_nbm for t scaffold (ENABLED=False); chp_nbm for ch; wdp_nbm for wd. · Field-level pick distribution changes daily — see the per-field Selector column above and the National Source tile for real-time per-band picks. The static summary sentence that used to live here drifted from reality within days of every ship. · skip_table_nbm_curated.json 17 cells post-v0.7.11 (was 15 post-v0.6.622; +2 for sr × nor_easter on 09-29). Symmetric ADD + REMOVE audit loop runs daily in the digest.
Lc cloud saturation-unbiasing (cm · ch only) · FLIPPED 07-17 v0.6.355 · cl+cc both OFF via _FIELD_SKIP after 07-30 v0.6.389d-390 intervention
Ccd cc from-derivation · LIVE ENABLED=True 07-30 v0.6.390 (flipped 6d ahead of gate) · cc = max(cl_l6, cm_l6, ch_l6) except SKIP_REGIMES {se_flow, unknown}
15 active Stage 1+3 candidates · all auto-run in daily digest · shipped items live under Post-ship watches below
L1 blender (v0.7.3 · shadow-only)⚠ v0.7.2 APPLY flip rolled back 2026-09-24 · BLENDER_APPLIED_FIELDS = frozenset() · curated table stale-window fit · re-curate needed· sibling: v0.7.6 L1 static blender (2026-09-26, separate mechanism — universal ω per field, L1 seat, cascade BYPASS, curated for h/dp only, ENABLED=False shadow). Both mechanisms are wired in parallel; both currently shadow. v0.7.6 09-29 shadow retro (post-v0.7.8):2 SHIP-READY (h/nw_flow/24-47 +39.4%, dp/nw_flow/24-47 +8.6%, both halves-stable, n=336 each — below gate min_n=400, wait ~3d) / 13 HOLD / 7 KILL / 7 THIN on 20 curated cells. Universal apply-flip 10-03 still not on track — pre_frontal cells KILL, and nor_easter (629 dp + 471 h rows, 41% of stamps) is off-curated because its best-ω is HRRR-favoring (1.00 dp, 0.65 h) vs universal 0.27/0.44. Off-curated stamp rate dropped 84%→41% post-v0.7.8, residual is entirely nor_easter — a curation gap, not a plumbing bug. Narrow-flip candidate: nw_flow/24-47 h+dp only, decision around 10-02 when n crosses 400.
First architectural change at the L1 seat since the selector. Replaces binary picking with continuous per-obs blend weight ω ∈ [0,1] from an L2-regularized ridge on 12 features (ims, xr_spread, lead, hour circular, cc_inter_sigma, pressure_trend, wd circular, ws, cloud_low, solar). At runtime: forecast = ω · HRRR_terminal + (1−ω) · NBM_terminal, per (field, regime, band) cell. Fields in BLENDER_APPLIED_FIELDS (currently empty) apply the blend to the served forecast; every other field gets shadow telemetry only. Un-curated cells fall through to the selector unchanged.
→ Stale-window fit discovered 2026-09-24. The GCS forecast_error_log_backstamped.jsonl had been frozen since 2026-08-21 (uploaded manually once, then no auto-refresh). l1_blender_stage1.py reads that URL via cached_path(); every fit since Aug 21 saw data ending 2026-08-20T20:07. The 13 STABLE cells shipped v0.7.0 and today's v0.7.2 dp APPLY flip were both fit on that frozen window. Fresh-data refit results: 11 of 13 cells fail halves-stable. Survivors: h/pre_frontal/24-47 (A+25.7% / B+8.1%, ω̄ 0.44-0.55) and t/se_flow/24-47 (A+11.8% / B+14.9%, ω̄ 0.37-0.46). One new candidate not in the shipped set: wg/pre_frontal/12-23. v0.7.3 (this session) reverts BLENDER_APPLIED_FIELDS to empty and wires analysis/nbm_backstamp_append.py into the publisher CF (runs first each hour, appends new pair-log bytes to the GCS backstamped file via bucket.compose(), HWM tracked in gs://myweather-data/backstamp_hwm.json) so this staleness class can't recur silently. l1_blender_curated.json intentionally not re-curated in v0.7.3 — since applied-fields is empty, no cell fires and shadow telemetry now accumulates against fresh data. Re-curation is a follow-up ship. Related audit: per-obs classifier v2 (v0.6.644-646) fit on the same stale window; fresh-data refit shows STAGE 1 HOLD on both h and t with every cell marked degenerate (fNBM 67-99% — the fit rubber-stamps NBM rather than selecting per-obs).
NBM parallel cascade✓ AT ARCHITECTURAL PARITY 2026-08-26 v0.6.499 · native L2 shipped for all 8 NBM-scope fields · all layer slots + specialists mirrored
Full mirror of the HRRR cascade on the NBM side, layer-for-layer: raw_nbm → l2_nbm → l3_nbm → l4_nbm (ch) → l5_nbm (sr) → l6_nbm (t, ENABLED=False) → chp_nbm (ch) / wdp_nbm (wd). Each layer has its own fitter (analysis/l{3-6}_nbm_*.py) writing a curated JSON that the collector loads at import; per-hour stamps land in the snapshot log and per-row error_l{n}_nbm columns land in the pair log. L1 selector picks argmin(MAE(HRRR-deepest-Prod), MAE(NBM-deepest-Prod)) per (field, lead-band). Bins reached L3 maturity via analysis/nbm_backstamp.py (walks pair log + backfilled NBM point extracts, reconstructs l2_nbm exactly from snapshotted HRRR values); deeper layers warm naturally over the next ~30 days as error_l3_nbm/error_l4_nbm/error_l5_nbm rows accumulate. See per-layer fit-status tiles below.
→ Live per-field / per-band picks in the National Source tile above. Router-scope aggregate lift last refit at v0.6.500 (2026-08-26) was +53.5% on n=73,023; refreshed daily by the pool table. v0.7.5 (09-26) added learned_gbm + ims_threshold precedence on top of the pool for ch and sr; v0.7.11 (09-29) added a per-cell L3 skip for sr × nor_easter that runs after the selector picks NBM. Static "which side wins" summaries drift within days of every ship — read the live tile.
15 SHIP cells (all WIDEN) live-stamping on confidence.cells[field][band].c1h. Per-cell co-axis ortho gate (v0.6.321) suppresses cells non-orthogonal to the currently-firing co-axis — cl fires freely, ch 24-47h + t × 3 never fire.
→ Already live-stamping (confidence_layer reads the curated table each tick). The narrow-promote counter is auto-suppressed in the digest per KNOWN_LIVE_PIPELINES — the ship gate for user-visible band widening is C1 Stage 4 (last HOLD 65.00% — moved 57.14% → 65.00% on 08-15 re-cure).
dp depression regimefrontal branch closed · nor_easter watch opened
Frontal decayed −2.19 → −1.98 → −1.51 → −0.87°F — below the 1.5°F action floor as of 07-09. Branch retired.
New: nor_easter +3.79°F ★ but n=279 — magnitude passes floor, sample thin. Gate: 3 consecutive reads with n growing.
C1d cloud disagreement✓ Stage 3 wired 07-08 · gated OFF
13 SHIP + 1 MARGINAL cells stamping on confidence.cells[field][band].c1d. Fires when live cloud_inter_source_sigma ≥ Q3 (44.55). cc 24-47h MARGINAL NARROW −7.06% is a documented outlier — flagged for Stage 4 review.
→ Already live-stamping (curated table read per tick). Narrow-promote counter auto-suppressed in digest per KNOWN_LIVE_PIPELINES — ship gate for user-visible band widening is C1 Stage 4 (last HOLD 65.00% — moved 57.14% → 65.00% on 08-15 re-cure).
Pre-frontal cloud widening⚙ Stage 2 curated 07-12 · counter wired · blocked on population (n=8 passages, THIN)
Latest ortho read: 7 SHIP cells (cell is SHIP iff ORTHOGONAL vs C1a AND vs C1e). Written to weather_collector/data/pre_frontal_curated.json. v0.6.372b matched-regime fix HOLDS PROMOTE with the corrected baseline; SHIP set expanded 5→7. Digest currently at streak 1/7 due to cell-set drift + population still THIN.
→ Real 7-day narrow-promote counter wired 2026-07-12 v0.6.328d. Blocked on population: only n=8 frontal passages in the current window, 17% pair-log join rate — both matched-regime KILL (h_hsf) and matched-regime PROMOTE (h_pre_front) share this THIN caveat. Do NOT act on either verdict until passage count ≥ 15 (probably 08-05 to 08-15). See [[project_c1e_hsf_kill_investigation]].
cl persistence gate (clp)⚙ Stage 3 wired 07-24 v0.6.379 · ENABLED=False · streak walker tracked continuously via auto-rolling windows (v0.6.398, 08-09); flip when 7-day Jaccard clears — no fixed date target
Successor to the retired cl_persistence_short_lead (all-9-regimes design gate not met). Broader regime × lead_band gate: 12 SHIP / 8 MARGIN / 16 SKIP / 1 THIN. Whitelist: persistence on calm/se_flow/unknown all-leads + nw_flow 24-47h + short-lead in sw_flow/ne_flow/nw_flow/pre_frontal + pre_frontal 6-11h. Wired after Lc. 7-day flip gate: SHIP cell-set Jaccard ≥ 0.8 across 7 daily reads. See [[project_cl_persistence_investigation]].
wg residual persistence gate⚙ Stage 3 wired 07-14 v0.6.351 · 07-27 flip HELD · awaiting Stage 1 recovery
Long-lead-only: adds per-clock-hour L2-residual mean (prior 14d) on top of fc_l2, replaces L3-corrected wg on gate-fired cells. Short leads always SKIP (L2's Kalman handles close-in). 07-27 flip HELD — h_wg_residual_persistence_stage1 flipped PROMOTE → MARGINAL 07-26 (+20.24%, was +17.74%). Persistence hypothesis's aggregate weakened; hold until Stage 1 recovers PROMOTE. (Divergence report drop-{wg,ws} was stale-tool artifact per [[feedback_walkforward_skip_table_blind]]; fix v0.6.381 landed 07-26.)
07-27 Stage 1 SHIP (Brier lift −28.56%/−18.73%, n=292) re-flipped HOLD on 08-02 halves-strict re-cut — recalibration params don't transfer across time at this sample size. Root: Reliability ≈ 5% of Brier (0.005 of 0.086); Uncertainty + Resolution dominate. All Platt-family scripts (h_pp_platt_calibration, h_pp_platt_by_regime, h_pp_frontal_platt_stage1, h_pp_bin_calibration, h_pp_bias_persistence_stage0) retired to .skip.py on 08-10 with the daily-digest retirement pass (10 dead-verdict scripts total). Structural next step if pp ever reopens: physical-feature gating (CAPE, RH-500mb, front proximity), not calibration parameter fitting.
LC_RECENT_BIAS_GATE_ENABLED = True live in cloud_saturation_correction.py. Applies existing lc_correction_table.json shift only where recent 3-day observed bias still agrees with the historical fit (sign match + magnitude ≥ 0.5×|hist|). Smaller surface than the EMA/Kalman rewrite alternative — no new lookup table, just a per-cell gate on the live shift. 08-18 refresh (v0.6.430): ch per-field streak now 11 days; cl and cm remain CHURN (in and out of promoted set day-over-day). Deploy carries refreshed lc_correction_table.json + lc_recent_bias_gate.json into the collector image; runtime toggle unchanged. Live cells suppressed today = 0/3 — all three ch per-bin cells (20-50, 50-80, 80-95) have gate_apply=True because recent 3d bias tracks historical (or is thin) on every bin. cc excluded (derived via Ccd). Ship is a live no-op today but the self-healing mechanism is in place for any future per-bin drift — the anti-scar-tissue counterpart to hand-curated skip frozensets. Runtime telemetry verified 10:48 UTC tick: weather_data["cloud_saturation_correction"]["recent_bias_gate"] = {enabled: true, fields_cleared: ["ch"]}. See [[project_lc_regime_conditional]]. Follow-on: cl remains in _FIELD_SKIP (Lc architecture can't rescue cl regime) — needs EMA/Kalman fallback, separate workstream [[project_lc_cl_unskip_investigation]].
chp full-shape refinement (adopt L6-baseline 5-cell SHIP set)⚙ 6-cell emergency demote LIVE 07-27 v0.6.382t · full-shape adoption 7-day gate closed 08-03 · verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09)
Preview curated JSON at weather_collector/data/ch_persistence_gate_curated_vs_l6.json. L6-baseline Stage 2 rebuild (h_ch_persistence_blend_stage2_vs_l6.py) says chp's honest SHIP set is 5 cells: calm/0-5 (−70%), pre_frontal/0-5 (−41%), sw_flow/0-5 (−30%), se_flow/0-5 (−26%), ne_flow/6-11 (−17%). Currently-live chp is at 14 SHIP + 8 MARGIN post-demote; 5 more halves-disagreement/small-magnitude cells still fire but should also fall away. 7-day live-layer change gate closed 08-03; verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09). Adopt condition: 7 daily reads of the L6-baseline Stage 2 continue to agree on the 5-cell SHIP set + no new frontal-only-recalibration interaction.
What's being evaluated next
Upcoming
Forward-looking only. For today/yesterday narrative see Recent activity below.
Fri 10-03
KEY DATE — v0.7.5 + v0.7.6 fresh-data verdicts, scope narrowed.(a) v0.7.5 router-as-authority — ch-only now. 09-28 finding: sr's learned_gbm path never fires because l1_learned_selector_curated.json has no sr cells; v0.7.5 verdict is really about ch via ims_threshold. Attribute per-mechanism using v0.7.7's selector_mechanism stamp. Rollback = one flag flip. (b) v0.7.6 L1 static blender shadow retro — universal apply-flip NOT on track, but nw_flow/24-47 narrow-flip candidate emerging. 09-29 retro: 2 SHIP-READY (h/nw_flow/24-47 +39.4%, dp/nw_flow/24-47 +8.6%, both halves-stable, n=336) / 13 HOLD / 7 KILL / 7 THIN on 20 curated cells. Off-curated ratio dropped 84%→41% (residual all nor_easter, not in curated). Held on narrow flip: n=336 < gate min_n_rows=400 AND most of 7d predates v0.7.8's clean-plumbing. Wait ~3 days for n≥400 on strictly post-v0.7.8 rows; if halves hold, ship narrow (nw_flow/24-47 only, h+dp). Universal apply-flip still blocked — pre_frontal cells KILL in shadow, nor_easter needs per-regime ω (best-ω is HRRR-favoring, opposite of universal 0.44/0.27).
~10-02 (n gate)
v0.7.6 nw_flow/24-47 narrow-flip decision. Re-run analysis/l1_static_blend_shadow_verify.py when h/nw_flow/24-47 and dp/nw_flow/24-47 both cross n≥400 on strictly post-v0.7.8 rows. If both halves still positive ≥5%, ship narrow flip: modify l1_static_blend_curated.json to keep only nw_flow/24-47 cells per field, set ENABLED=True. Rollback = flag back to False.
~10-04 (n gate)
nor_easter static-blend re-evaluation. Today's nor_easter fit shows h/12-23 best-ω=0.65 with +28% lift halves-stable +23.5/+38.0, n=381 (below 400 gate). If n crosses 400 next 3-5 days: decide schema — extend l1_static_blend_curated.json to support per-regime ω override (nor_easter isn't compatible with the field's universal ω=0.44). Also revisit dp/nor_easter (best-ω=1.00, halves unstable). If no cells clear, hold and revisit when regime accumulates more sample.
Mon 10-05
v0.7.9 chp dynamic gate 7d verify.h_ch_persistence_blend_stage2_vs_l6 WATCH count should drop from 4 live losing cells to ≤1. If not, one of the 5 dynamic-only cells has a Simpson-paradox artifact and the flag should be reverted; investigate before shipping any additional cells.
next session
dp NWS coherence — SUPERSEDED by v0.7.6 blender.Original plan (09-13): route dp through NWS at the hourly-array level via decay_apply. Superseded 09-26 by the L1 static blender's coverage of dp with universal ω=0.27 on 10 (regime, band) cells; the analysis showed L1-blend-no-cascade beats every routing/blending option that keeps the cascade. If v0.7.6 apply-flip lands 10-03 clean, this workstream closes. If not, re-scope NWS coherence as fallback.
rolling
L1 by-regime walker — near-miss watch. First wire landed 09-11 v0.6.581 (ws/calm/12-23, ws/nw_flow/12-23). Under the fixed gate (sum(n_today) ≥ 60 across 3d), three PPP cells sit below threshold and will clear as sum_dn accumulates: wg/sw_flow/0-5 (sum_dn=32, lift +7.4%), ws/calm/0-5 (sum_dn=6, lift +21.4%), ws/calm/6-11 (sum_dn=2, lift +28.7%). Digest watch — no action.
rolling
Prune HRRR PBL morning-overshoot gate. v0.6.606 hardcoded gate is a stop-gap until walker's stagnant_high × t × 0-5 cell clears the wire naturally. Once the walker routes t at that regime/band, the named gate becomes redundant. Remove HRRR_PBL_MORNING_OVERSHOOT_* block from l1_selector.py; drop the hour_local kwarg call from forecast_snapshot.py. Was "4-7d" (09-13); regime rarely active recently.
Backlog
NBM sea-breeze specialist — dedicated workstream. Only if 09-11 walker read + skip curation loop don't close enough of the aggregate gap. Stack-health trajectory (top of page) is the arbiter — if 7d median plateaus after this week's ships, structural NBM specialist work is next; if it keeps trending up, keep tuning.
clp deferred
clp Stage 3 flip gate — 08-16 walker FAIL (min-J 0.250; fire set churned 7 → 3 cells over 7d window). No ENABLED flip. Walker continues; re-check when digest reports PASS.
ongoing
C1d narrow-promote — auto-suppressed (KNOWN_LIVE_PIPELINES). Already live-stamping via confidence_layer; user-visible band-widening gate is C1 Stage 4.
held
pre-frontal C1e narrow-promote — HELD 08-11 fresh re-run. Down from 16 ortho cells (June) to 2 SHIP (ch 24-47h, cl 24-47h). Only 3 frontal passages in the 45d window; most cells THIN. Re-run when autumn front cadence returns.
held
frontal-t bias Stage 0 — Stage 1 gate day 1/7 (08-12 v0.6.401g). 3/4 bands HIT (0-5h +1.94, 6-11h +2.74, 12-23h +1.06); 24-47h SIGN_FLIPS. Half A concentrated in one 07-13/16 event. Rolling 7-day gate added to h_frontal_t_bias_stage0.py mirroring l6_fix_b_refit: requires ≥7 distinct days, no HOLD days, ship_bands STABLE, ≥1 band. Verdict escalates to STAGE 1 CLEAR — the trigger to scope frontal_t_bias.py specialist. Earliest clear 2026-08-19.
held
sr obs-recent override Stage 0 — 08-11 near-miss. Fired-subset lift +23.5% but pooled test lift 4.58% under 5% gate (77 fires, 91% non-Lsb). Two paths to Stage 1: loosen pooled gate OR tighten trigger to 300 W/m². Backlog #8.
blocked
wsbp — flip gate still HELD as of 08-08 (v0.6.388): calm regime shadow still below MIN_N_ANTECEDENT=20 in the 24h antecedent window. Predict: one more calm overnight pushes it over.
Post-ship watches — active
v0.7.15 sr learned_gbm cells finally live (opened 09-29). 5 sr STABLE GBM cells (nw_flow/12-23, nw_flow/24-47, se_flow/12-23, se_flow/24-47, sw_flow/6-11) populated in l1_learned_selector_curated.json. Completes v0.7.5's sr side after 3 days as a silent no-op. Deploy 15:50 UTC clean. Watch: (a) overnight — first sr rows in the 5 covered cells stamp selector_mechanism=learned_gbm (was band_pool); (b) ~7d — sr Value Captured 7d trends positive as covered cells accumulate; (c) if any covered cell shows negative lift, that cell needs to be dropped and refit. See feedback_shipped_flag_verify_effect.
v0.7.16 v5 candidate writer schema fix (opened 09-29). One-line preventive fix in analysis/l1_selector_per_obs_classifier_stage1_v5.py: candidate JSON now emits band "12-23" instead of "12-23h" to match runtime _band_for_lead(). Prevents recurrence of the v0.7.5→v0.7.15 silent-no-op bug on future ships from this pipeline. Watch: next ship attempt from this script produces cells that fire without band-key mismatch.
v0.7.14 selector terminology rule (opened 09-29). "Selector" is the primary name in prose. Code entities keep their names (l1_selector.py, pick_source(), selector_source/_mechanism). "Router-as-authority" retained only as the v0.7.5 pivot's framing name. Motivated by v0.6.432 "L1 router" archive collision. See feedback_selector_is_primary_name. Watch: new prose in future sessions holds the line (no drift back to "router" as a synonym for selector).
v0.7.13 operator narrative structural cleanup (opened 09-29). 22 old post-ship watches (opened 08-30 through 09-15, ≥14d concluded) archived via display:none. Recent Activity today entry compressed 5k→1.7k chars. CLOSED CLEAN summary line below expanded to catalog every archived entry. Watch: next session's Recent Activity stays landmark-only (memory owns full detail).
v0.7.11 sr × nor_easter L3 bypass — TEMPORARY (opened 09-29). Added sr × nor_easter × 12-23h and 24-47h to l3_nbm in skip_table_nbm_curated.json. Circuit-breaker per feedback_fresh_fire_vs_circuit_breaker_frames: known layer/regime/direction, two-day worsening, n=220 above threshold. Watch: (a) tomorrow's pair-log 12-23h backstamps show applied_layer=l2_nbm (not l3_nbm); (b) nbm_regression_sentry sr.l3_nbm exits HOT; (c) 24h/12h sr Total Lift recovers toward raw NBM MAE. Re-review 2026-10-13: either remove (regime faded, walkforward proposes cleanly) or keep as formal skip. Reversal = delete the two entries in the JSON.
v0.7.12 debug narrative de-staled (opened 09-29). Humidity row narrative rewritten event-based (no drifting numbers); two NBM cascade "which side wins" summaries replaced with pointers to live tiles; cm-dropped note added to L3_NBM sentence. Watch: whether future ship notes stay landmark-only (no live numbers in prose) — recurring drift class.
v0.7.10 h_cc_derivation format guard (opened 09-29). Two pct()=None format crashes fixed in per-day and per-regime tables. Substantive verdict unchanged (long-standing PROMOTE: derived-random beats prod cc +47.9%). Watch: next digest shows h_cc_derivation OK (was FAIL(1)).
v0.7.9 chp dynamic gate flipped (opened 09-28).CHP_CELL_GATE_ENABLED = True in ch_persistence_gate.py. 9 cells cleared 7-day gate (5 dynamic-only new suppressions: ne_flow/6-11, ne_flow/12-23, ne_flow/24-47, nw_flow/24-47, se_flow/12-23). Watch: (a) h_ch_persistence_blend_stage2_vs_l6 WATCH count drops from 4 live losing cells to ≤1 within a week; (b) any of the 5 new dynamic-only cells flipping OUT of gate = churn signal, revisit threshold; (c) ch pooled Prod-vs-L6 delta should improve ~5-10% MAE per 09-21 66-day anchor. Reversal: flag back to False; static _CELL_SKIP retained.
v0.7.8 regime-source reconciliation (opened 09-28).forecast_snapshot.py field loop now classifies regime inline from entry values with cloud_cover (same signature as forecast_error_log). Watch: (a) 09-29 VERIFY PARTIAL PASS — off-curated ratio dropped 84%→41%; residual (1,126 rows) is entirely nor_easter, a regime not in the curated table at all. Fix confirmed working; the residual is a curation gap, not divergence. (b) Non-nor_easter off-curated: 26 rows. (c) New workstream: nor_easter needs per-regime ω (universal 0.44/0.27 hurt −13 to −152%; best-ω is HRRR-favoring). Schema extension pending n≥400 on h/nor_easter/12-23 (~10-04). (d) v0.7.6 apply-flip partially unblocked — 2 SHIP-READY cells emerging.
v0.7.7 mechanism attribution stamp (opened 09-27).{f}_selector_mechanism pair-log tag written by pick_source_with_mechanism(). Watch: (a) 10-03 v0.7.5 verdict now attributable per-mechanism — filter pair-log by selector_mechanism=='ims_threshold' for ch. Sr has zero learned_gbm rows because curated table has no sr cells — verdict scope narrowed to ch-only. (b) analysis/l1_static_blend_shadow_verify.py retro scorer available in digest.
sr regression — L3_nbm / nor_easter watch (opened 09-28, worsening 09-29). Post-v0.7.7 sr rows losing to raw_nbm on nor_easter cells. 09-29: nbm_regression_sentry sr.l3_nbm HOT +8.1%→−46.7% (Δ +54.8pp), worsened vs 09-28's +5.8%→−1.9%. NBM skip-add proposal l3_nbm sr nor_easter 12-23h now n=220 lift −47.2% (up from n=39 −74% on 09-28), above skip-threshold, but 14d+50d two-window verdict not yet cleared (nor_easter regime too new). Router correctly picks NBM (band_pool); downstream l3_nbm correction inflates error on low-solar. NOT a v0.7.5 issue — GBM never fires on sr. Still held on emergency skip per fresh-fire lucky-baseline discipline. Watch: 14d+50d gate should promote cleanly within ~1 week if regime persists.
L3_FIELDS drop cm (opened 09-15 v0.6.625). Walkforward gate cleared 7/7 on the drop proposal (L3 fc -1.6% / obs -1.2% pooled, no regime × band WIN, calm/24-47h LOSS -17.2%). Same fix pattern as ws v0.6.397, cc L4 v0.6.515. Watch: (a) next digest divergence report should show L3_FIELDS AGREE (was READY, drop cm); (b) pair-log rows for cm at calm/24-47h stop stamping applied_layer=l3; (c) cm pooled prod-vs-raw should not regress vs prior 30d — the L3 correction was net-negative, so removing it is a small positive expected.
Digest streak = same-proposal (opened 09-15 v0.6.625).build_executive_summary.py streak walkback now requires matching normalized verdict text. Watch: next digest walkforward_l3l4_validator shows 7/7 (or drops out of ship-eligible once drop-cm shipped verdict text changes). Any other ship-resolution script's streak count should stay reasonable — sudden resets on a stable proposal would indicate the normalizer is too aggressive (strip too much) or too lax (miss meta-clauses).
KILLED_LAYERS 09-05 pruned (opened 09-15 v0.6.626). Removed (ch, chp_nbm) and (h, l3_nbm) from nbm_regression_sentry.py registry. Both windows post-date the kill so sentry reads natural nominal/THIN. No runtime change. Watch: sentry verdict lines for ch.chp_nbm and h.l3_nbm should read CLEAN/THIN, never HOT/WATCH. (cc, l4_nbm) 09-08 and (cc, l3_nbm) 09-10 remain — prune 09-16 and 09-18 respectively when sustained window post-dates.
nbm_skip_add_audit filters shipped (opened 09-15 v0.6.627). Audit now drops proposals already in skip_table_nbm_curated.json before evaluating. Watch: next digest's CONFIRMED list should show only genuinely new proposals. If a shipped cell somehow appears, filter key format may not match — check the audit against a known-shipped cell that would previously have surfaced.
Publisher CF redeployed (opened 09-15).make deploy-publisher at 14:00:58 UTC picked up v0.6.615's 12h WINDOWS addition to analysis/per_field_scoring.py and analysis/scoreboard_v2.py. Publisher had been running 09-13 code for 2 days, blanking the 12h per-field diagnostic table. Watch: (a) hourly GCS publish continues to write a 3-window per_field_scoring.json; (b) any future analysis edit that touches a publisher-hosted script triggers a paired publisher redeploy — [[feedback_deploy_hygiene_publisher_pairs_analysis]].
pr L2 unwire nw_flow/6-11h (opened 09-12 v0.6.590). Three-tool ship: retro Δ-13.6% over 1,447 pairs since 08-13 (both halves negative -8.5/-17.7), layer-shape sentry pr/production@6-11h +10.6%, yesterday's Notable Calls pr 6-11h -8.7%. Watch: pair-log MAE at nw_flow/6-11h should return toward raw as new obs stamp applied_layer=l1; layer-shape sentry for pr 6-11h should clear within days. Stability re-read scheduled 09-19 — surviving nw_flow/0-5h should stay halves-positive.
inter_model_spread Stage 2 7-day gate (opened 09-12 v0.6.595). 33 SHIP cells cleared |premium| ≥ 30% halves-stable in 14d test window. Gate: SHIP set must stay stable across 7 daily reads before wire. Daily verification via h_inter_model_spread_c1_stage2.py output. Earliest wire flip to axis_6: 2026-09-19. First new C1-axis candidate to reach Stage 2 since cross_run_spread in June.
L1 3-way walker + runtime NWS routing (opened 09-13 v0.6.601, extends 09-12 v0.6.597). Walker adds NWS as 3rd direction; runtime pick_source() returns "nws" with `entry[f"{f}_nws"]` swap in forecast_snapshot.py. dp gated at wire via _NWS_FIELDS_WIRE_ELIGIBLE = {t, ws, wd, pp} pending option B design (Magnus consistency at hourly-array level cascades to h/AH/feels-like — see [[project_nws_dp_coherence_wire]]). Today: 3 dp cells cleared escalation but fall through to pool; runtime is a no-op today by design. Watch: first non-dp NWS cell (t/ws/wd/pp) to clear the 3-way fitter — the runtime wire fires immediately. Cell-stability across daily reads for the 5 dp cells (informational only until option B lands).
l3_nbm.wg.nw_flow/12-23h skip (opened 09-13 v0.6.602). Two-window CONFIRMED: 14d n=472 lift -3.8%, 50d n=1,545 lift -5.0% halves -8.87/-1.86. Watch: pair-log MAE at wg.nw_flow/12-23h should trend back toward raw as new obs land with the l3_nbm skip in place.
l3_nbm.wd.se_flow/12-23h skip (opened 09-14 v0.6.609). Two-window CONFIRMED: 14d n=953 lift -4.60%, 50d n=2,581 lift -9.79% halves -14.24/-5.26. Passes v0.6.574 two-window gate. Watch: pair-log MAE at wd.se_flow/12-23h should trend back toward raw as new obs stamp applied_layer=l2_nbm instead of l3_nbm.
l3_nbm.wd.se_flow/0-5h skip (opened 09-14 pm v0.6.622). Two-window CONFIRMED afternoon: 14d n=209 lift -3.6%, 50d n=808 lift -8.9% halves -13.4/-3.0. Completes wd.se_flow short-lead pattern — all three bands (0-5h, 6-11h, 12-23h) now skipped. Watch: pair-log MAE at wd.se_flow/0-5h should trend toward raw as new obs stamp applied_layer=l2_nbm. Same shape as v0.6.609.
h L2 soft_ramp retune (opened 09-14 v0.6.618).H_SOFT_RAMP_FLOOR 0.1→0.4, H_SOFT_RAMP_END 10→24 in corrected_hourly.py. Addresses digest top-alert τ-suspect (h/production helps 0-5h -43.7% but hurts 24-47h +6.6% under old shape). Source: h_l2_shape_sweep STAGE 1 PROMOTE 7/7 rolling days, halves-stable at 24-47h. Watch: (a) 09-21 the 7d walker window fully rotates post-ship — τ-suspect should clear; (b) per-band vs raw stays positive across all 4 bands (0-5h +48.22%, 6-11h +11.10%, 12-23h +2.72%, 24-47h +1.46% projected).
frontal_detector_health daily verdict (opened 09-14 v0.6.619).analysis/frontal_detector_health.py runs in daily digest, emits Verdict: HOLD/CLEAN line based on threshold-vs-observed distribution + type='cold' reachability + miss-rate. Currently HOLD — 4 events / 336h, type='cold' never classified, runtime missed 9/13 candidates (pre-v0.6.620 threshold data). Watch: verdict auto-flips HOLD→CLEAN as new obs accumulate under 4.0°F threshold; first CLEAN expected within days of first cold-front-shaped signal firing.
Frontal detector calibration + diagnostic logging (opened 09-14 v0.6.620/v0.6.621).DP_DROP_THRESHOLD 8.0→4.0°F (was above p99.9=6.3°F, unreachable). Diagnostic print(" frontal: score=…") when score≥1 for miss-rate traceability (v0.6.620 used logging.info which Cloud Run drops; v0.6.621 fixed to print(..., flush=True)). Watches: (a) first frontal: score=N sigs=... log line via gcloud functions logs read myweather-collector --region=us-east1 --gen2 | grep frontal whenever any signal fires; (b) first type='cold' event tagged in frontal_events_log.json; (c) if any logged score≥2 line does NOT produce a corresponding events-log write, the miss-rate root cause becomes traceable; (d) C1e "post-front" pool re-fits over ~4 weeks as c1_confidence_calibration_v2.py rolls its window — unblocks [[project_ch_24_47h_c1d_c1e_split]].
HRRR-wire cells first-gated live (opened 09-14 v0.6.609 deploy, walker generated 10:57 UTC, deploy 11:06 UTC). 3 cells cleared 3-day gate: cc/ne_flow/12-23, wd/se_flow/24-47, ws/sea_breeze/24-47. The last was already firing via 09-11 escalation clause — today formalizes gate-clearance. Watches: (a) pair-log MAE at each cell should trend toward the HRRR raw MAE as new obs stamp selector_source=hrrr; (b) any of the 3 flipping OUT of the walker within days = churn signal, revisit MIN_LIFT threshold; (c) 5 flipped_in_window cells (walker-noted) not eligible for wire without operator review.
stagnant_high regime label (opened 09-13 v0.6.605). New synoptic bin — ws<5 mph AND cc<0.40 AND |pt_3h|<0.5 hPa, checked before frontal/calm in classify_synoptic_regime(). Retroactive pair-log sweep at these thresholds: ws +16.6% NBM-vs-HRRR lift under stag (vs +3.7% non-stag); sr +35.4% (vs +13.8%); ch reverse anti-signal −46.1% vs +34.6%; t not clean. Stamp-only ship — walker consumes automatically. Watches: (a) first labeled rows on next joiner write (~16:07 UTC); (b) stag×ws + stag×sr cells escalation-wire earliest (magnitudes above 20% clause); (c) stag×ch cell should route to HRRR (reverse direction) when it clears the 3-way walker's hrrr-wire gate; (d) stag rate expected around 2% of pair-log rows in typical regimes, higher this week during current stagnant-high pattern.
HRRR PBL morning-overshoot routing gate (opened 09-13 v0.6.606). Named hardcoded gate: field=="t" AND regime=="stagnant_high" AND hour_local ∈ {4,5,6,7,8} (EDT) → NBM. Highest precedence in l1_selector.pick_source(). Diagnosis: pair-log dig showed HRRR MAE 2.74-3.40 with bias +2.48 to +3.13°F at UTC 09-11 (EDT 05-07); NBM MAE 0.59-0.67 same hours. Boundary-layer scheme mixes down aloft warm air too aggressively during morning heating under stagnant clear-air. Watches: (a) t 24h Total Lift trends toward zero as morning-hour rows use NBM; (b) selector_picks for t stagnant_high show NBM dominant in EDT 04-08 window; (c) walker stagnant_high × t × 0-5 cell clears wire (4-7d) → prune this named gate; (d) reversibility via HRRR_PBL_MORNING_OVERSHOOT_KILL.
h Stage 1 halves watch (opened 08-30 v0.6.520) — CLEARED same-session v0.6.522. Morning MARGINAL (halves −0.93 / +3.12) was methodology artifact. /code-review high exposed 3 real load-path bugs (fc_prod L2 contamination + test-window off-by-one + halves-crash-on-n=0); shared harness fix moved h to +24.04% STAGE 1 PROMOTE (halves +1.28 / +0.88 BOTH WIN). Sibling numbers moved: dp +21.25 → +28.83, wg → +36.53. Stage 2 harness extracted v0.6.524; h Stage 2 written same session. See [[project_h_residual_persistence_attribution_08_30]].
h Stage 2 walker watch (opened 08-30 v0.6.524; automated 08-31 v0.6.529). Day-1 rollup 9 SHIP / 8 MARGIN / 9 SKIP / 7 THIN. SHIP cells today: nw_flow all 4 bands, sea_breeze 6-11/12-23/24-47, sw_flow 12-23/24-47. Walker (h_h_residual_persistence_walker.py) accumulates per-cell verdicts to .cache_h_residual_persistence_walker_history.json and gates the Stage 3 flip on 7/7 consecutive days in {SHIP, MARGIN} per cell. Earliest clear ~2026-09-06.
Stage 1 harness halves-preference watch (opened 08-31 v0.6.528). Bug fixed: grid picked max-test-MAE (window=5d for h today, halves ✗) over halves-stable window=14d (halves ✓ +0.56/+4.10) and downgraded h verdict to MARGINAL. Now iterates every combo, computes halves, picks halves-passing max. h corrected to STAGE 1 PROMOTE at 14d (+24.51%); dp/wg unaffected. Wrote [[feedback_grid_select_halves_stable]]. Watch: does daily verdict now stay PROMOTE across the 7 walker days? Any override-line firings on wg/dp? Analysis-only fix — no runtime change.
h_residual_persistence Stage 3 shadow watch (opened 08-31 v0.6.530 collector deploy). Processor pre-staged (cloned from wg template with FIELD=h, HOURLY_KEY=corrected_humidity, [0,100] both-ends clamp for RH). Wired into collector.py ENABLED=False; all 4 v0.6.525 bugfixes baked in. Deploy-verified 06:33 UTC tick: nw_flow fires {0-5:5, 6-11:6, 12-23:12, 24-47:24} (47/47), clamped_out_by_band present with zeros, corrected_humidity_post_l3_pre_hrp absent (ENABLED=False semantics correct). 7-day shadow accumulation pairs with the walker cell watch. Ship-day (~09-06) is now a one-line ENABLED=False→True flip + redeploy.
Stage 3 processor bug-fix watch (opened 08-30 v0.6.525 collector deploy). 4 real bugs fixed in wg_residual_persistence.py + dp_residual_persistence.py: (1) telemetry pre/post-clamp mismatch (wg only — dp doesn't clamp), (2) _TABLE_CACHE never invalidated (mtime check + MYWEATHER_REFRESH now respected), (3) clamp-out silently pooled with normal skips (new clamped_out_by_band counter), (4) no_hourly_array early return skipped gate_firing_log.record_firing. Post-deploy verify 18:17 UTC tick on new revision myweather-collector-00544-fuj: clamped_out_by_band key live with expected zeros on nw_flow regime; table_generated_at reflects fresh v0.6.524 Stage 2 refit (mtime cache invalidation working). Watch 7 daily digest reads for gate_firing_rollup regressions vs pre-fix baseline.
Lc emergency intervention (v0.6.389d-g + v0.6.390, 07-30) — active state, no close date. cl + cc BOTH off Lc via _FIELD_SKIP. cl held out (walk-forward: pool_vs_raw −22.15%, reg_vs_raw −30.37%). cc retired v0.6.390 architecturally → Ccd derives max(cl_l6, cm_l6, ch_l6). cm + ch untouched, +34-50% vs raw held-out. Next step for cl: EMA/Kalman shift tracker OR recent-bias gate. See [[project_lc_regime_conditional]] + [[project_lc_regime_stage1_pool_prereq]].
wsbp — HELD (v0.6.388, 07-28). Sibling of dpbp, calm-only, sign-inverted. ENABLED=False. Preflight 08-04: calm regime n=0 in shadow window. Wait for calm regime accumulation. Cost of waiting: nothing (dpbp covers dp side).
l6_fix_b_refit — HOLD-GATE (rolling gate added 08-08). Was single-day SHIP +2.15% today, HOLD at +0.29% 26 days prior; no consistency check between. Rolling gate needs 7 distinct days + same ship_bins before verdict can promote. Lt stays on do-not-reopen. See [[project_l6_fix_b_rolling_gate]].
wg persistence-skill thin margin — pooled L4 skill +0.17 today (was +0.19 08-29; < +0.20 hold margin). BUT per-cell check 08-30: Prod-vs-persistence skill +0.215 pooled (above +0.20 margin, healthy at 6-11h/12-23h/24-47h; 0-5h at parity is structural — wind_blend uses recent obs). The "at-risk" line measures L4-vs-persistence; users-see production is fine. Not a real live-product regression. Re-flag only if Prod skill drops < +0.20 for 2+ days OR any of 6-11h/12-23h/24-47h halves specialist contribution.
🧪 Architectural backlog — longer-horizon items, no date target
MLC — marine-layer correction sandbox. Stamps weather_data["marine_layer_correction"] every tick, ENABLED=False. In-bin cc bias collapsed 06-30 (real break, not the 07-07 "cliff" the anomaly detector reported — that was cumulative-window artifact). Split at 07-04 (cm HRRR-anomaly onset): stratum-local, pre-HRRR, rules out HRRR-anomaly causation. Companion cl signal points seasonal. Hold OFF indefinitely. Will not re-arm when cm HRRR anomaly clears — different event. Redesign candidate: time-of-year gating on the MLC bin.
L2-as-observation-only. Remove L2 from forecast pipeline, keep as training target for L3/L4/Lsr only. Architectural exploration.
Tight-τ cloud bias propagation across leads 1–3h. Pearson at lead 1h is +0.50/+0.57/+0.55/+0.66 for cc/cl/cm/ch — strong, clean; degrades by lead 3h. Architectural slot: short-lead τ-decayed bias propagation (~2h half-life) on top of the existing hourly[0] override. Needs Stage 0 MAE-impact estimate (probably small — first 3 leads only). Not blocking; logged for backlog. Spun out from resolved KBOS+KBVY cloud blend (2026-06-30, ρ ≤ 0 at lead 6h — hourly[0]-only formalized as intentional).
Recent activity — today + 2 prior days (older entries in docs/CHANGELOG.md, trimmed on next curation)
Badges:
PIPELINE ·
DISCOVERY ·
INFRA ·
DASHBOARD ·
PREFLIGHT
2026-09-29 (Tue) — today· 7 ships (v0.7.10–v0.7.16). 2 collector deploys. Biggest: v0.7.15 finally wires sr's `learned_gbm` cells (v0.7.5 pivot had shipped as silent no-op for 3 days). Also v0.7.11 sr × nor_easter L3 bypass (circuit-breaker), v0.7.13 archived 22 stale post-ship watches, v0.7.14 nailed down "selector" as primary name, v0.7.16 preventive schema-mismatch fix in v5 classifier + wd dig.v0.7.15 SHIP + deploy (15:50 UTC clean) — weather_collector/data/l1_learned_selector_curated.json populated with 5 sr STABLE GBM cells (nw_flow/12-23, nw_flow/24-47, se_flow/12-23, se_flow/24-47, sw_flow/6-11). Halves-stable A/B +14 to +37% on current pair-log incl. 4d nor'easter. Two independent bugs found: (1) shipped JSON had been left at 0 cells since v0.7.5 despite `LEARNED_SELECTOR_SHADOW_ENABLED = True`; (2) v5 candidate emitted band "12-23h" (h suffix) but runtime `_band_for_lead()` returns "12-23" — cells would never have keyed correctly even if copied verbatim. Stripped suffix on the shipped file; v0.7.16 fixes the writer script. See feedback_shipped_flag_verify_effect. v0.7.16 SHIP — one-line fix in analysis/l1_selector_per_obs_classifier_stage1_v5.py candidate writer to strip "h" from band. Future ships from this pipeline won't recur the bug. v0.7.14 SHIP — terminology consistency: "selector" is the primary name (prose, section headings, new writing). Kept: code entities (`l1_selector.py`, `pick_source()`, `selector_source`/`_mechanism`), UI ("Selector Skill"). "Router-as-authority" retained ONLY as the v0.7.5 pivot's framing name. Historical v0.6.432 "L1 router" archive collision motivates the rule. See feedback_selector_is_primary_name. v0.7.13 SHIP — 22 old post-ship watches archived (display:none, opened 08-30 through 09-15); today's Recent Activity compressed 5k→1.7k chars; landmark-only, memory owns full detail. v0.7.11 SHIP + deploy (13:10 UTC clean) — sr × nor_easter × 12-23h + 24-47h added to l3_nbm skip. Circuit-breaker per feedback_fresh_fire_vs_circuit_breaker_frames: known layer/regime/direction, two-day worsening (Δ +7.6pp → +54.8pp), n=39→220. Re-review 2026-10-13. v0.7.12 SHIP — humidity row narrative rewritten event-based; NBM cascade "which side wins" summaries pointed at live tiles. v0.7.10 SHIP — h_cc_derivation.py pct=None format guard. Report totals reconciled in l1_static_blend_shadow_verify.py (was 2+13+7+7 labeled "of 20"; now splits curated=20 / off-curated=9). Diagnostics (no ships): v0.7.8 verify off-curated 84%→41% (residual nor_easter, curation gap); l1_selector_fit_3way [promote→hold] is real nor_easter signal decay; nor_easter static-blend fit shows universal ω doesn't fit (best-ω HRRR-favoring); cc FRESH FIRE = lucky-baseline (3 all-clear days). Wd dig — held on ship. 7d Total Lift −2.67% is entirely cascade (routing 0%, cascade −2.67%). NBM L3 sentry WATCH on wd (help −3.4% → −16.1% fresh). Same pattern as sr today but smaller magnitude (Δ +12.8pp vs sr's +54.8pp) and WATCH not HOT. Held to verify v0.7.11 sr ship first; revisit tomorrow with data. Watches: overnight 09-30 v0.7.11 sr L2 stamps + v0.7.15 sr `selector_mechanism=learned_gbm` firing · ~10-02 v0.7.6 nw_flow/24-47 narrow flip · 10-03 v0.7.5 + v0.7.6 verdicts · ~10-04 nor_easter blender schema · 10-05 v0.7.9 chp verify · ~10-06 v0.7.15 sr Value Captured 7d · 10-13 v0.7.11 re-review. Memory: new project_09_29_session, feedback_shipped_flag_verify_effect, feedback_fresh_fire_vs_circuit_breaker_frames, feedback_selector_is_primary_name; updated project_router_as_authority_pivot with silent-no-op story; MEMORY.md READ FIRST refreshed.
2026-09-28 (Mon) — 1 day ago· 2 ships (v0.7.8 regime-source fix + v0.7.9 chp dynamic gate flip) · 2 collector deploys, both first ticks clean · sr audit-trigger cleared. Morning digest triage on the 09-27 scheduled watch: 24h scorecard showed sr −77% / wg −39% / t −17% yesterday; today's read was sr −77% (WORSE), wg +10% (recovered), t +4% (recovered). Only sr persisted → v0.7.4/v0.7.5 audit triggered per schedule. Deep dig on sr revealed the router is NOT the cause. Post-v0.7.7-deploy sr rows (filter by run_time, not valid_time): 100% stamped band_pool, zero learned_gbm — the v0.7.5 GBM path never fires on sr because l1_learned_selector_curated.json has no sr cells (v0.7.5 shipped ims_threshold for ch and GBM for sr, but the sr GBM curated set was never populated). Per-layer breakdown on the losing regime (nor_easter, new since ~09-26): raw_nbm 5.6-15 W/m², L2_nbm=raw_nbm (clean), L1(HRRR) 15-40, prod tracks HRRR-side because band_pool table routes sr→NBM correctly BUT downstream l3_nbm correction inflates error on low-solar. Corroborated by nbm_regression_sentry: sr.l3_nbm HOT — layer help +5.8% → −1.9%, Δ +7.6pp. The v0.7.5 router audit is CLEARED — sr regression is l3_nbm sr / nor_easter, downstream correction miscalibration on new regime. Held on emergency skip (n=39 on 0-5h too thin; nor_easter transient; halves-stable impossible). Waiting for regime to accumulate n≥200/band so walkforward proposes the skip cleanly. v0.7.8 SHIP + collector deploy (08:02:57 EDT, first tick 08:07 EDT confirmed clean) — regime-source reconciliation in forecast_snapshot.py. Previously the field loop used _wdp_state_fc_by_lead[i] (state_stamp's regime, built from raw hourly[] arrays WITHOUT cloud_cover) for its shadow-stamp cell decisions, while forecast_error_log.py rebuilds regime from the snapshot's per-hour dict WITH cloud_cover. Divergence: stagnant_high can only fire in the pair-log path (cloud_cover gated), and L2-corrected wind/pressure can cross the nor_easter threshold when raw values don't. Per 09-27 finding: 78/93 v0.7.6 shadow rows in first 12h landed on nor_easter cells not in curated JSON — the stamp used state_stamp's regime, the pair-log recorded the state-stamp's DIFFERENT regime. Fix: inline classifier call in the field loop, same signature forecast_error_log uses (entry.get("wd"/"ws"/"pr"/"cc"/"t") + snap-level pressure_trend + local valid-hour). Flows into _learned_predict, _blender_omega, _l1_static_blend.blend_l1, _selector_pick_source_with_mech. _wdp_state_fc_by_lead still consumed by wd_persistence_gate — different concern, not touched. Removes pre-flip blocker for v0.7.6. Shadow-only, no user-visible change. Verify in ~6h: analysis/l1_static_blend_shadow_verify.py off-curated stamp ratio should drop from ~84% to near-zero. v0.7.9 SHIP + collector deploy (~11:2X EDT, first tick 11:27 EDT confirmed) — CHP_CELL_GATE_ENABLED = True in ch_persistence_gate.py. Digest 09-28 h_chp_cell_gate cleared 9 cells with days_lose=7 / days_win=0 / days_thin=0: ne_flow/6-11, ne_flow/12-23, ne_flow/24-47, nw_flow/12-23, nw_flow/24-47, pre_frontal/12-23, pre_frontal/24-47, se_flow/12-23, se_flow/24-47. 4 overlap the hand-curated _CELL_SKIP (already suppressed); 5 net-new dynamic-only suppressions. Per project_chp_narrow_to_0_5h_watch memory (09-21 66-day anchor): 6h+ regime cells lose to L6 by +10-37% MAE. Killing 5 new losers should net ~5-10% MAE reduction on ch. Static _CELL_SKIP retained (still holds calm/12-23, calm/24-47, sea_breeze/12-23, sea_breeze/24-47, sw_flow/12-23, sw_flow/24-47 that dynamic hasn't cleared yet). Reversal: flag back to False. Watch: h_ch_persistence_blend_stage2_vs_l6 WATCH count should drop from 4 live losing cells to ≤1 within a week. Both ships committed 70cf2cda. Sr investigation memory: project_09_28_session. MEMORY.md READ FIRST reordered — sr framing corrected. Session lessons: (1) filter pair-log by run_time not valid_time when isolating post-deploy effects — legacy snapshots close against fresh obs and confound the mechanism attribution. (2) When a router audit fires, verify the router ACTUALLY made the pick before rolling back — the mechanism stamp is the source of truth; zero learned_gbm rows means the router never even entered the picture. (3) Two callers computing the same regime independently is a divergence trap. Any shared classifier call should read a single stamped value or share a helper. Non-goals held: did not roll back v0.7.5 (evidence pointed elsewhere); did not ship an emergency nor_easter skip on n=39 (fresh-fire lucky-baseline discipline). Clock-watches held: 10-03 v0.7.5 verdict now scoped to ch-only (sr never routed via GBM); v0.7.6 shadow retro first read (0 SHIP-READY / 11 HOLD / 4 KILL / 13 THIN) — apply-flip target not on track, cells may need refit on fresh data. Files this session:weather_collector/processors/forecast_snapshot.py, weather_collector/processors/ch_persistence_gate.py, index.html v0.7.7→v0.7.9, docs/CHANGELOG.md. Memory: new project_09_28_session.
2026-09-27 (Sun) — 2 days ago· v0.7.7 shipped and committed (commit e403e2ad) — telemetry stack for the 10-03 verdicts.v0.7.7 SHIP + collector deploy (07:17 EDT commit, deploy same session). Three-part telemetry ship. (a) {f}_selector_mechanism pair-log tag alongside {f}_selector_source. Written by new pick_source_with_mechanism() in l1_selector.py; passed through by forecast_error_log.py. Values: pbl_morning_kill / learned_gbm / ims_threshold / regime_override / band_pool / default_hrrr. Backwards-compat wrapper preserves pick_source(). Unblocks the v0.7.5 10-03 verdict — without this we couldn't separate router-driven picks from precedence-chain picks that happened to land on the same source. (b) New analysis/l1_static_blend_shadow_verify.py — v0.7.6-specific retro scorer (distinct from v0.7.0-era l1_blender_shadow_verify.py). Reads l1_blend_shadow stamps on covered cells, computes served/blend/L1/raw_NBM MAE per cell over 7d + 30d, halves-stable A/B, verdict SHIP-READY/HOLD/KILL/THIN. Digest driver picks it up automatically. First read on 09-28: 0 SHIP-READY / 11 HOLD / 4 KILL / 13 THIN. (c) DISCOVERY — off-curated stamping bug surfaced by the new verifier: 78/93 v0.7.6 shadow rows in first 12h landed on nor_easter cells NOT in curated JSON. Root cause: forecast_snapshot passes its own per-lead _fc_regime_i (from state_stamp's array, built without cloud_cover) to l1_static_blend.blend_l1, but the pair-log records a DIFFERENT regime later. Shadow-only impact today; blocks the v0.7.6 apply-flip. Fix landed 09-28 v0.7.8. Also: 24h scorecard slice showed sr −68% / wg −39% / t −17% vs raw_nbm — flagged as audit trigger for 09-28 (resolved next session — sr was L3_nbm/nor_easter, not v0.7.5). Memory:project_09_27_session (superseded 09-28 in framing).
2026-09-26 (Sat) — trimmed· 2 architectural ships (v0.7.5 + v0.7.6) · 2 collector deploys, both first ticks clean. Morning session started as digest triage on Selector Skill 24h Value Captured showing −33% median / −100% mean while trajectory chart held +36% median 7d. Argued the tile was noise, user pushed back correctly — trajectory tail was diving too, with ws Prod worse than Raw for two straight days (raw HRRR 3.86 vs served 6.18 on 09-26). Attribution reframing: Routing 7d −6.3% / 24h −8.3% (selector subtracts value), Cascade +18.9% / +23.5% (carries the total). Reframed the whole day. v0.7.5 SHIP + collector deploy (rev deployed 10:53 UTC, first tick 10:57 UTC kq96blFRm2nY, MEMPROBE 48.7→465.7 mib clean) — router-as-authority pivot live in one commit. weather_collector/processors/l1_selector.py: (a) _IMS_SELECTOR_CELLS replaced with 10 refit cells from analysis/l1_selector_ims_threshold_refit.py (halves-A/B +37 to +76% lift on error_prod_real, ablation-cleared); IMS_SELECTOR_SHADOW_ENABLED = True. (b) weather_collector/data/l1_learned_selector_curated.json populated with 5 sr STABLE GBM cells from v5 sweep (ch cells stripped — non-overlapping mechanism split); LEARNED_SELECTOR_SHADOW_ENABLED = True. Flag names preserved for wire compat; docstrings already treated True as live apply. Killed the 10-02 fresh-corpus gate as over-caution — halves-stable A/B was the safety net that caught 09-24's stale-fit, and the v5 sweep passed it on 09-25 post backstamp appender. Shipped 6 days early. Precedence in pick_source(): HRRR-PBL → LEARNED → IMS → by-regime walker → band pool → HRRR fall-through, so non-covered fields fall through cleanly (zero blast radius outside ch/sr). v0.7.6 SHIP + collector deploy (rev deployed 16:05 UTC, first tick 16:07 UTC IHi3BOYQZmI9, MEMPROBE 50.7→470.2 mib clean) — L1 static blender, shadow-only. New processor weather_collector/processors/l1_static_blend.py + curated table weather_collector/data/l1_static_blend_curated.json (h ω=0.44 on 10 cells, dp ω=0.27 on 10 cells). Wired into forecast_snapshot.py next to v0.7.0 blender pattern — stamps {f}_l1_blend_shadow unconditionally on every covered row; apply gated by ENABLED = False. Verified: 39/48 hours in first snapshot have shadow stamps for h and dp. Architectural mechanic: unlike v0.7.0 blender which blends TERMINAL forecasts, v0.7.6 blends at the L1 seat (raw HRRR L1 + raw NBM L1) and BYPASSES the cascade for covered rows. On h/dp, cascade applied to a blended L1 hurts more than it helps — L2/L3/L4 residual models are source-specific and miscorrect blended input; terminal blending also loses to L1-blend-no-cascade. Backed by 90d pair-log analysis (6 scratchpad scripts, ~11 hypotheses): between-vs-outside classification (h 51% obs-between, dp 47%); 9-mode seat comparison (L1-no-cascade +20% h / +30% dp vs live); halves-A/B stability (real half-B floor +9% h / +23% dp); universal-ω-per-field beat per-cell (h per-cell 29.1% / universal 29.0% identical; dp universal actually WINS +1.5pp) — one ω per field, not per cell; tail compression p95 dp 7.10→3.99 (44% catastrophic-error reduction); 79-90% row-win rate in strong cells. Rejected during analysis: feature-based ω (ridge overfits −3 to −18pp vs static), cross-field ω sharing (h and dp want different weights), universal-regime rule (cell heterogeneity too high), rolling refit (30d vs 90d STABLE cells nearly identical — ω is time-stable), blender for cc (0 real STABLE cells) or wind fields (7-10% blend-over-pick ceiling only). Watch 10-03 for v0.7.5 first 7d Value Captured verdict AND v0.7.6 shadow retro. If v0.7.6 halves-stable holds on fresh live data, v0.7.7 flips ENABLED = True AND extends _SELECTOR_WRITEBACK at forecast_snapshot.py:1211 to also cover "l1_blend" source (found post-deploy — less invasive than modifying corrected_hourly.py). Session lessons: (1) Trajectory chart tail is not always small-n noise — check whether raw sources are genuinely diverging before trusting the +weekly-trend line. ws Prod > ws Raw for two consecutive days was real, not variance. (2) Selector work has been polishing a lever that averages negative-value; instead of another selector cell, the between-fraction analysis found that h and dp are blender-territory fields where a naive constant-ω blend beats the entire cascade. Architectural shift, not another cell-by-cell grind. (3) Static ω per cell → universal ω per field simplification is real when halves-stable A/B measurement can carry it. Ridge features on top of ω are overfitting bait. Non-goals held: did not touch ws routing despite the trajectory tail (2 days is not enough to whipsaw a table fit on 30d evidence); did not open cc/cm/pr/cl DIVERGE cells (short-lead cascade genuinely wins there). Clock-watches: 10-03 v0.7.5 7d + v0.7.6 shadow retro (both). Files this session: 09-26 memory files project_09_26_session, project_router_as_authority_pivot (updated as SHIPPED), project_l1_static_blend_v076 (new). MEMORY.md READ FIRST reordered.
2026-09-15 (Tue) — trimmed· 9 ships (v0.6.625 collector · v0.6.626 registry+debug · v0.6.627 audit filter · v0.6.628 sweep · v0.6.629 layout · v0.6.630 attribution reconcile · v0.6.631 12h default-open · v0.6.632 09-15 entry fill · v0.6.633 debug page consistency sweep) · 1 collector deploy · 2 publisher deploys. Morning digest triage: dp + h "FRESH FIRE" regression (3d prod_real +30/+35% vs raw) diagnosed as a v0.6.620 frontal-detector shadow — prod-replay stays clean but 09-13/14 pairs were served under the old regime stamps; no ship needed. Expect HEALING on 09-17 digest as pre-fix pairs roll off the 3d window. Walkforward's "48/7 days confirmed" clarified as bucket-continuity, not proposal-continuity — the actual "drop cm" proposal age is 7 days. Also fixed the publisher CF being 2 days stale, which was blanking the 12h per-field diagnostic table.v0.6.625 SHIP + collector deploy (rev myweather-collector-00574-guj, 13:07 UTC) — two changes bundled: (a) L3_FIELDS = {"wg", "ch"} in decay_apply.py:78, dropped cm. Walkforward cleared its 7-day gate on this proposal (L3 fc -1.6% / obs -1.2% pooled; only calm/24-47h shows LOSS -17.2%, no regime × band cell reaches WIN). Same fix pattern as ws (v0.6.397) and cc L4 (v0.6.515). (b) build_executive_summary.py streak walkback now requires matching normalized verdict text, not just "promote" bucket. Strips bracket clauses ([entangled: N], [N thin]) so meta-noise doesn't reset; keeps ship-count digits so a real proposal change ("L3 ship 3" → "L3 ship 2") does reset. Motivating case: walkforward L3L4's counter has been ticking since 2026-06-25 through three intermediate ships (ws v0.6.397, cc L4 v0.6.515, pp v0.6.304), reading "48/7 days confirmed" while the current "drop cm" proposal was actually only 7 days old — user pushed back on the streak, verifying via history file confirmed it. Effect on next digest: walkforward L3L4 shows 7/7 (real age), or drops out of ship-eligible entirely after the drop lands. v0.6.626 SHIP — analysis/nbm_regression_sentry.py KILLED_LAYERS registry pruned (ch, chp_nbm) and (h, l3_nbm), both kill-dated 2026-09-05 (v0.6.551). Sustained (7d) and fresh (3d) windows now post-date the kill so cells naturally read THIN/nominal — registry entries no longer suppress anything. (cc, l4_nbm) 09-08 and (cc, l3_nbm) 09-10 kept — sustained window still overlaps. Also removed the two closed 09-15 clock-watches (L3 DROP cm shipped v0.6.625; KILLED_LAYERS prune shipped here) from the Upcoming grid. v0.6.627 SHIP — analysis/nbm_skip_add_audit.py now filters already-shipped cells. Walkforward re-emits every proposal that clears its 14d gate, including cells shipped weeks ago. Today's digest CONFIRMED list had 3 cells and all 3 were already live: wd/se_flow/0-5h (v0.6.622), wd/se_flow/12-23h (v0.6.609), wg/nw_flow/12-23h (v0.6.602). New _load_shipped_cells() reads skip_table_nbm_curated.json into a {(cand_layer, field, regime, lo, hi)} set; _accumulate_50d drops any proposal already in the set before building the target dict. Mirrors the KILLED_LAYERS suppression pattern (silent no-op if the JSON is missing). Verified: 17 shipped cells load, all 3 known-shipped CONFIRMED entries match, filter drops them from the 21 walkforward proposals. Next digest's CONFIRMED list should show only genuinely new proposals. DISCOVERY — 12h per-field diagnostic blank: user pointed at the debug page's 12h table showing "—" everywhere. Traced to the publisher CF (myweather-publisher) last deployed 2026-09-13, before v0.6.615 (09-14) added the 12h window to analysis/per_field_scoring.py. Publisher was running old code with 2 windows; local file had all 3. GCS-published JSON only had 7d + 24h. Fix: make deploy-publisher at 14:00:58 UTC, then manual scheduler trigger produced fresh JSON at 14:11:43 UTC with all 3 windows. Table populated on next page refresh. Root cause pattern: [[feedback_deploy_hygiene_publisher_pairs_analysis]] — analysis edits that touch publisher-hosted scripts need a paired publisher redeploy. v0.6.615 shipped the code + a claim in its own changelog note that "GCS-published files now carry 12h", but the publisher redeploy was never run. Two days of user-visible blank table before it was noticed. v0.6.628 SHIP — this Recent Activity entry + Upcoming grid 09-17 dp/h HEALING watch + Post-ship watches for v0.6.625/626/627 + publisher redeploy. Text-only, no runtime change. v0.6.629 SHIP — What's running section outer wrapper flipped from 3-column auto-fit grid to flex-column. The three panels (running · improving · being evaluated next) are dense text blocks; narrow columns caused excessive vertical growth. Full-width stack gives content the horizontal room it deserves. One-line CSS change. v0.6.630 SHIP + publisher redeploy (rev 14:46:15 UTC) — Attribution + Total Lift reconcile via paired decomposition fields. per_field_scoring.py now emits routing_paired_pct and cascade_paired_pct alongside total_vs_best_raw_pct, all rounded to 2dp on the same paired pool (pool_ok row set where best_raw + l1_selected + prod all exist). Frontend Attribution reads them directly instead of recomputing from *_mae. Previously the tile recomputed 100*(b-s)/b etc. from best_raw_mae / l1_selected_mae / prod_mae, each rounded to 3dp separately — the compounded rounding drifted the mean/median aggregate ~0.5pp from Total Lift's headline, reading as a mysterious mismatch on the scoreboard even though the identity holds algebraically. Now Routing + Cascade = Total per field, and Attribution Total = Total Lift by construction. Verified end-to-end: 16:00 UTC GCS tick, t 7d = 7.40 + 0.72 = 8.12 = total_vs_best_raw_pct exactly. v0.6.631 SHIP — 12h per-field diagnostic <details> now default-open, matching 7d and 24h above it. One-char edit. v0.6.633 SHIP — debug page consistency sweep (5 items flagged on end-of-day read of the v0.6.628–632 debug page). (a) version.json bumped 624 → 633 — the header meta-line reads from it and had been showing v0.6.624 while Recent Activity documented through v0.6.632 (index.html was already at 632; version.json was the lagging file). (b) sr row status narrative dropped the hardcoded "selector picks HRRR" clause and now defers to the live-loaded Selector column above the row; the architecture-summary picks list also corrected (sr moved out of NBM picks into HRRR — 30d NBM lift −34% to −43%, 7d recent lift −0.8% to −6.9%, all bands source=hrrr in l1_selector_table_curated.json). (c) h row narrative refreshed to current values: 7d Total Lift +3.7% (was +1.0% stale), cascade component still −3.9%, routing +7.6%, HRRR pipeline skill +10.3% / NBM +5.4%; added a live fresh-fire regression note pointing at the 09-17 clock-watch. (d) dp+h regression clock-watch language softened from "v0.6.620 caused" to a tentative attribution: the prod_real vs prod-replay divergence isolates the cause class to stateful regime handling, and v0.6.620 is the most likely specific driver, but a second regime-stamp bug in the same window would produce the same signature — so the diagnosis is tentative until 09-17 HEALING confirms or a second dig turns up a peer cause. (e) 8°F "unreachable" wording in v0.6.619 recap clarified: the 9 candidates the sweep found fired via the remaining wd_shift + press_bounce signals with the dp arm dead; they are not "9 candidates at 8°F" but 2-of-3 fires that never used the dp signal. Text-only sweep, no runtime change, no publisher redeploy needed. This-session lessons: (1) diagnostic-blank triage — check GCS first, then live vs local, then publisher-deploy state. Fast path from blank to root cause. (2) "48/7 confirmed" is bucket-continuity, not proposal-continuity — if user pushes back on a long streak, verify what proposal has been standing by digest_history.jsonl. Prior ships in the same bucket reset the proposal but not the streak counter. (3) prod-real vs prod-replay divergence isolates stateful/runtime-only regressions cleanly — fitter says shape is fine, prod is hot → suspect regime-stamp or state-plumbing change, not the corrections themselves.
2026-09-14 (Mon) — 1 day ago· 14 ships (v0.6.609 collector · v0.6.610 debug sweep · v0.6.611 registry prune · v0.6.612 clock-watch clearance · v0.6.613 residual-walker gate off-by-one · v0.6.614 chp-cell-gate same off-by-one · v0.6.615 12h short-window · v0.6.616 C1d KILL scope-artifact fix · v0.6.617 tool audit + stale-comment cleanup · v0.6.618 h L2 soft_ramp retune · v0.6.619 frontal-detector health check · v0.6.620 DP_DROP_THRESHOLD 8→4°F · v0.6.621 logging.info→print fix · v0.6.622 wd.se_flow/0-5h skip) · 5 collector deploys.v0.6.609 SHIP + collector deploy (rev myweather-collector-00568-wot, 11:06 UTC) — l3_nbm.wd.se_flow/[12,24) skip cell added to skip_table_nbm_curated.json. Two-window CONFIRMED on today's nbm_skip_add_audit: 14d n=953 lift −4.60%, 50d n=2,581 lift −9.79% halves −14.24/−5.26. Passes v0.6.574 two-window gate. Third net-new CONFIRMED proposal from the audit; the other two lines (wd/se_flow/6-11h, wg/nw_flow/12-23h) were re-surfaces of already-shipped cells (v0.6.500 and v0.6.608). Post-ship dig — wg.l3_nbm sentry HOT investigated: sentry flagged pooled layer help +1.06% sust → −2.36% fresh (Δ −3.43pp). Per-regime × band across 5 windows: nw_flow/24-47h is the un-skipped negative driver (fresh −14.02% n=307, sust +6.64% n=1,225) but 50d +1.75% n=4,387 halves +0.70/+2.98 both positive — 3-day dip against a durably-earning cell; the v0.6.574 two-window gate exists exactly for this and correctly excludes. Real watch target is frontal/24-47h: fresh −11.18% / sust −5.93% / 14d −4.54% / 30d −1.61% / 50d −1.26%, halves 50d +0.38 / −2.74 (unstable). Under ADD threshold today; if the slide continues, promotes to CONFIRMED on next 3-4 days of audit. pre_frontal/24-47h small fresh dip, 14d/30d/50d all positive — not skip. Sentry doing its job on the aggregate; audit doing its job on the per-cell filter. v0.6.610 SHIP — this Recent Activity entry + Upcoming grid Wed 09-17/18 sentry watch + Post-ship watches v0.6.609 entry. Text-only debug page sweep, no runtime change. v0.6.611 SHIP — ADDED_LAYERS registry prune: removed sr.l5_nbm + sr.l3_nbm (both added 2026-09-04 v0.6.548) from analysis/nbm_regression_sentry.py. Both sr sentries CLEAN today (l3_nbm sust +16.2% → fresh +24.7%; l5_nbm CLEAN with THIN counts). Suppression was already no-op (add_date < sustained_start today) — clean-up only. Registry now empty until the next NBM-layer add. Clock-watch clearance — HRRR-wire day 3/3 done: walker's HRRR-wire direction cleared 3 cells on the 3-day gate today — cc/ne_flow/12-23, wd/se_flow/24-47, ws/sea_breeze/24-47 (the last was already firing via 09-11 escalation; today formally gate-clears too). 5 flipped_in_window (not ship-eligible until operator review): documented per walker output. Runtime: walker JSON regenerated 10:57 UTC before this morning's v0.6.609 deploy at 11:06 UTC, so the 2 net-new HRRR routings (cc/ne_flow/12-23, wd/se_flow/24-47) are live in production this tick. NWS-wire direction cleared 3 dp cells (dp/nw_flow/0-5, dp/nw_flow/12-23, dp/pre_frontal/12-23) — walker-diagnostic only, dp still gated at wire via _NWS_FIELDS_WIRE_ELIGIBLE pending option B design. v0.6.613 SHIP — RESIDUAL-WALKER GATE OFF-BY-ONE: investigated why h_residual_persistence Stage 3 flip has been stuck at 0/8 cells cleared for 8 days. Root cause: shared analysis/_residual_persistence_walker.py cutoff was now - GATE_WINDOW_DAYS which produces an 8-day window for a 7-day gate. Line 149 checks n_seen == GATE_WINDOW_DAYS (7). With 8 days seen, the check never fires — even for cells SHIP every single day. Fingerprint that surfaced it: h/sw_flow/12-23 with days_seen=8, days_positive=8, days_ship=8, flipped=false, cleared=false. Applies to all three residual walkers (wg/dp/h) — zero cells cleared since inception across the entire mechanism. Memory index had wg listed as "live via shared harness" — that was aspiration, wg has been shadow ENABLED=False the entire time. Fix: cutoff = now - (GATE_WINDOW_DAYS - 1). Post-fix walker output: wg 16 cells cleared, h 7 cells cleared, dp 6 cells cleared.No runtime change today — all three processors ENABLED=False. Follow-up ships (not today): consider ENABLED=True on wg + h after one more week of stability. dp stays ENABLED=False per derived-field rule. See [[project_residual_walker_gate_off_by_one_09_14]]. Lesson: off-by-one class of bug. If ever writing "last N days" filter, sanity-check window contains exactly N (not N+1) dates. And digest "0 cells cleared" was normal-looking output nobody investigated — same class of "mechanism silently mis-firing" trap [[project_h_residual_persistence_attribution_08_30]] fell into with the Stage 1 harness. v0.6.614 SHIP — SAME BUG in h_chp_cell_gate: after v0.6.613 audit-swept other walkers for the same off-by-one pattern. Two candidates matched (h_cc_blend_formula_stage1, h_frontal_t_bias_stage0) but both use the tolerant len(by_day) >= GATE_WINDOW_DAYS check — unaffected. h_chp_cell_gate.py:137 had the identical bug: cutoff now - GATE_WINDOW_DAYS + clearance check n_seen == GATE_WINDOW_DAYS. Yesterday's digest said "HOLD — walker at day 8/7 distinct dates" — same fingerprint. Same fix applied. Post-fix: 11 cells cleared the dynamic chp cell gate. 4 overlap _CELL_SKIP (dynamic gate catches manual skips). 7 net-new dynamic suppressions: ne_flow all 3 bands, nw_flow/6-11, nw_flow/24-47, pre_frontal/6-11, se_flow/12-23. No runtime change today — CHP_CELL_GATE_ENABLED = False. Follow-up: fresh 7-day post-fix window closes 2026-09-21; if the 7 net-new cells stay in the clear-list, consider flipping the dynamic gate to True (would suppress the 7 net-new cells at production; chp there falls to L6 baseline). Correlates with h_chp_midlead_regression ESCALATE verdict on lead 11 (+24.3%) — the two tools are converging on the same mid-lead chp regression. v0.6.615 SHIP — 12h short-window on scoreboard + per-field diagnostic: from the stale TODO backlog. Added "12h" (days=0.5) to WINDOWS in both analysis/scoreboard_v2.py and analysis/per_field_scoring.py. Tried 6h too but dropped it — pair-log lags real-time by 8+ hours (backstamp cadence) so a 6h window has zero rows most of the time; 12h reliably has data. Also added a Per-field diagnostic 12h table on the debug page below the 24h one (uses parameterized renderPerFieldDiagnostic("12h", ...) from v0.6.604). Purpose: post-ship read. The 24h and 7d windows lag any routing change by up to 24h/7d, so a fresh L2 fix or selector flip reads as noise there. The 12h table catches the effect first. Thin-window caveat noted in the table subhead: n≤50 fields read as noise. Sample from today's run: ch 24h Total Lift −10.8% but 12h +11.1% (flipped positive fresh); wg 12h +16.2% vs 24h +8.0% (improving); ws 12h −14.6% vs 24h −1.8% (fresh regression). GCS-published scoreboard_v2.json + per_field_scoring.json now carry the 12h block. v0.6.616 SHIP — C1d KILL scope-artifact fix: investigated today's digest verdict "→ KILL C1d: signal captured by C1a/C1e (6/8 redundant)". Finding: KILL is a scope artifact of the tool's default MIN_N_PER_CELL=100, which limits it to 24-47h cells only. At that threshold the tool tests 8 of the 32 possible cells and is blind to 3 of the 5 live C1d SHIP cells (all at 12-23h). Lowered MIN_N_PER_CELL to 50 in analysis/h_cloud_disagreement_orthogonality.py — now tests 28 cells across all 4 bands. New verdict: MIXED — 3 orthogonal / 17 redundant / 8 other. Narrow promote on the orthogonal cells. The ORTHOGONAL cells are cc/0-5h (2.57× no-trans / 2.82× trans on C1a; 2.57× / 3.38× on C1e) and ch/12-23h × C1a (1.20× / 1.37×) — real signals a KILL would have erased. Do NOT unwire C1d. Updated feedback_digest_triage_discipline memory with a new step 6: for KILL verdicts on live axes, verify the tool's test scope covers all live SHIP cells before acting — same class of "tool with a blindspot" trap as the walker off-by-ones today. See [[project_c1d_kill_scope_artifact_09_14]]. Follow-ups queued: cc/0-5h narrow-promote candidate; ch/24-47h re-calibration flag (post-front σH/σL=3.20× is OPPOSITE direction to live NARROW premium). Post-ship watches: v0.6.609 — pair-log MAE at wd.se_flow/12-23h should trend back toward raw as new obs stamp applied_layer=l2_nbm. Clock-watch added: Wed 09-17 → Thu 09-18 wg.l3_nbm sentry HOT self-clear OR frontal/24-47h CONFIRMED promote (see Upcoming). Memory: [[project_wg_l3_nbm_sentry_09_14]] documents the 5-window per-cell dig with numbers for the next audit read. Lesson: the two-window gate is doing real work — the sentry alone would have proposed shipping nw_flow/24-47h today; the 50d halves-stable check caught it as noise. Same shape of premature-ship trap v0.6.573 fell into pre-v0.6.574.
Afternoon session — diagnostic-driven chain from the digest:v0.6.618 SHIP + collector deploy — h L2 soft_ramp retune (H_SOFT_RAMP_FLOOR 0.1→0.4, H_SOFT_RAMP_END 10→24 in corrected_hourly.py). Resolves the top-alert τ-suspect from morning digest (h/production helps 0-5h −43.7% but hurts 24-47h +6.6% under old shape). Source: h_l2_shape_sweep STAGE 1 PROMOTE, 7/7 rolling days, halves-stable A +6.99% / B +10.39% at 24-47h. Per-band vs raw new: 0-5h +48.22%, 6-11h +11.10%, 12-23h +2.72%, 24-47h +1.46% — every band positive. v0.6.619 SHIP — analysis/frontal_detector_health.py daily digest calibration audit of the frontal-passage detector. Sweeps frontal_obs_log.json with the runtime's 2-of-3 signal logic and emits a Verdict: line. Chronic HOLD until the underlying miscalibration is fixed. Source of investigation: ch/24-47h C1d × C1e split (next-session brief) traced back to the frontal detector as the root blocker. DISCOVERY — DP_DROP_THRESHOLD unreachable: observed 60-min dp drops over 14 days had max=6.8°F, p99.9=6.3°F. Live threshold of 8.0°F sat above the 99.9th percentile. The dp arm of the 2-of-3 detector was effectively dead — no observed 60-min drop ever reached 8°F in the 14d window — so passages could only fire via the remaining two signals (wd_shift + press_bounce) and the _classify_type='cold' branch (which requires the dp arm) was unreachable (0 events ever tagged cold in the log). The 9 candidates the sweep found in 14d are exactly that: 2-of-3 fires on wd_shift + press_bounce without the dp signal. 4 of the 9 landed in the events log — 5/9 miss-rate with no obvious cause (all missed candidates had ≥6 obs entries in their 60-min windows). v0.6.620 SHIP + collector deploy — DP_DROP_THRESHOLD 8.0→4.0°F in frontal_detection.py. Lands at p99.5 (top 0.5% of ticks). Simulated event rate at 4.0°F: 13 / 14d vs 9 at 8.0°F. Restores dp as a real signal and makes type='cold' reachable. Also added diagnostic print(..., flush=True) when score≥1 so the miss-rate mystery is traceable via gcloud functions logs on next occurrence. C1e "post-front" pool will re-fit automatically over ~4 weeks as c1_confidence_calibration_v2.py rolls window forward — unblocks [[project_ch_24_47h_c1d_c1e_split]]. v0.6.621 SHIP + collector deploy — DIAGNOSTIC BUG: v0.6.620's log line used logging.info(...) which is silently dropped in Cloud Run (no basicConfig; root logger at WARNING). Verified after deploy — three runs (17:27/17:37/17:47 UTC) produced zero frontal: lines despite the detector running. Same trap the 2026-08-15 MEMPROBE fix hit; comment in collector.py:495 explicitly warned about it. Should have caught before shipping. Swapped to print(..., flush=True). Lesson: Cloud Run logging behavior is documented in-repo — read the comment before adding new log calls. v0.6.622 SHIP + collector deploy — l3_nbm.wd.se_flow/[0,6) skip cell added. Two-window CONFIRMED on this-afternoon's digest re-read: 14d n=209 lift −3.6%, 50d n=808 lift −8.9% (halves −13.4/−3.0). Completes the wd.se_flow short-lead pattern: all three of 0-5h (today), 6-11h (v0.6.584, 2026-09-04), 12-23h (v0.6.609, this morning) now skipped. 24-47h is FRESH not CONFIRMED (50d only −2.9%) — hold. Afternoon post-ship watches: (a) frontal_detector_health daily verdict flips HOLD→CLEAN as new obs accumulate under 4.0°F threshold (rolling 14d window; expect within ~2-3 days of first cold-front-shaped signal); (b) first frontal: score=N sigs=... log line in Cloud Run whenever any signal fires; (c) first type='cold' event tagged in frontal_events_log.json; (d) miss-rate: if any new frontal: score≥2 log line does NOT produce a corresponding events-log write, the miss-rate root cause becomes traceable. Afternoon lessons: (1) the C1d × C1e split investigation traced back through 4 layers of "why is that number wrong" before landing at the actual root (frontal detector threshold miscalibration). Chase the chain don't stop at the first plausible cause. (2) Deploying a diagnostic is not the same as verifying the diagnostic works. Always confirm output landed before trusting the diagnostic on the next signal. (3) In-repo comments carry hard-won operational knowledge; a code sweep before adding a similar pattern is worth 2 minutes.
2026-09-13 (Sun) — trimmed· 7 ships (v0.6.601 → v0.6.608) · 4 collector deploys · 1 publisher deploy. Morning: L1 3-way walker+runtime (dp gated) + wg.nw_flow/12-23h skip + dp NWS coherence scouts. Afternoon: diagnostic-driven pipeline — 24h per-field table surfaced t 24h Total Lift −43.8%, dig identified HRRR PBL morning-heating overshoot under stagnant clear-air (HRRR MAE 2.74-3.40 at UTC 09-11 vs NBM 0.59-0.67), shipped stagnant_high regime label + a named routing gate to close it before the walker escalation clause catches up (~4-7d). Evening: L1 recency override Simpson-guard shadow — designed a per-regime veto guard on the pooled 7d flip, shadow-measured over 7d, DATA REJECTED THE SHIP (guard would regress pooled prod-MAE −14.2% on the 5 flagged cells). Recency mechanism vindicated on Simpson-shaped flips; h 24h regression diagnosed as short-lived HRRR-favorable pattern, not a routing bug.v0.6.601 SHIP + collector deploy (rev myweather-collector-00564-kex, 11:18 UTC) — extended L1 selector to 3-way HRRR/NBM/NWS. Three files: analysis/l1_selector_fit_3way.py adds n_today per-cell so walker's WINDOW_SUM_N_MIN works; analysis/l1_selector_fit_by_regime_walker.py adds NWS-wire as 3rd direction alongside NBM/HRRR-wire (same 3d gate + escalation clause); weather_collector/processors/l1_selector.pypick_source() returns "hrrr" | "nbm" | "nws"; forecast_snapshot.py NWS branch swaps entry[f] = entry[f"{f}_nws"]. dp gated at wire (_NWS_FIELDS_WIRE_ELIGIBLE = {t, ws, wd, pp}). Yesterday's v0.6.600 flagged the Magnus-consistency concern for dp; gating at wire-time lets the infrastructure ship without inheriting the block. Walker still tracks dp — today 3 dp cells cleared escalation (dp/nw_flow/0-5 +25.0% n=1,238; dp/nw_flow/12-23 +29.3% n=1,721; dp/pre_frontal/12-23 +23.6% n=1,078) but all fall through to pool (nbm). Zero non-dp NWS cells cleared, so runtime wire is a no-op today by design — fires the moment any t/ws/wd/pp cell clears. v0.6.602 SHIP + collector deploy — l3_nbm.wg.nw_flow/12-23h skip cell added (14d n=472 lift -3.8%, 50d n=1,545 lift -5.0% halves -8.87/-1.86). Cleared both windows of the two-window audit (v0.6.584 gate). l3_nbm.wd.se_flow/6-11h also CONFIRMED today but already in an earlier curated JSON — audit re-surfaces every passing proposal. Afternoon scouts: (1) analysis/scout_nws_dp_coherence.py — empirical scout of option A (NWS-t agreement gate) on the 3 cleared dp cells. Tight gates DESTROY lift on the biggest cell (dp/nw_flow/12-23 baseline +45.5% vs live-prod → tol=1°F +37.2%, tol=0.5°F +33.2%). NWS advantage lives ON rows where NWS-t disagrees with selected t — Option A dead. Deeper audit: option D-as-conceived is a NO-OP for users — frontend reads hourly.corrected_dew_point from decay_apply, not forecast_snapshot; snapshot-level routing only fixes pair-log attribution. Real user impact requires option B at the hourly-array level (touches _ensure_derived_moisture_consistency, cascades to inverse-Magnus h + AH + feels-like + cc=Ccd). Deferred pending option B design. Cell dp/nw_flow/0-5 is a fitter-vs-raw artifact — flat vs live-prod, would not ship. (2) analysis/h_state_fc_obs_disagreement_orthogonality.py — Stage 1 orthogonality on cloud_delta + solar_delta (mirrors inter_model_spread Stage 1 structure from v0.6.594). Result: neither is a general C1 axis — both HOLD-MIXED (redundant cells slightly outnumber orthogonal on 5 of 6 axis combos). But strong per-cell orthogonality in narrow pockets: cloud_delta on cc/cl/cm/dp/wd at 0-5h; solar_delta on sr/cm/h at 0-5h. Belongs in C1 marginal-axis narrow-ship pool alongside pre-frontal/hsf, not full-axis pipeline. Curation deferred. Post-ship watches: v0.6.601 — first t/ws/wd/pp NWS-wire cell clearance (runtime fires immediately). v0.6.602 — wg pair-log MAE at nw_flow/12-23h should trend back toward raw. Clock-watches unchanged: 09-14 first HRRR-wire GATED read + sr ADDED_LAYERS prunable; 09-15 L3 DROP cm streak + KILLED_LAYERS 09-05 prunable; 09-19 inter_model_spread wire day + pr L2 gate re-read. Lessons: yesterday's v0.6.600 blocker note saved this session — I nearly wired dp NWS without checking the derivation cascade. Reading the previous day's changelog before touching related code is a first-read discipline. Also: option-D "downstream audit clean" claim I made mid-session was WRONG — I audited entry["dp"] consumers but missed that entry["dp"] and hourly.corrected_dew_point are two SURFACES; frontend reads the latter. Reversed my own recommendation mid-session before shipping.
Afternoon session — diagnostic-driven ship chain: user pointed at the Stack Health 24h daily/mean-median gap; per-field query identified dp -6.6% + h -1.7% dragging the mean. User asked whether a 24h version of the per-field diagnostic would surface it → yes → shipped. v0.6.603 SHIP — this Recent Activity entry rewrite (was 2 ships → now 6, morning + afternoon narrative). v0.6.604 SHIP — 24h Per-field diagnostic companion table below the existing 7d table. renderPerFieldDiagnostic() parameterized on (windowKey, tbodyId, tfootId); called twice on page load. Difficulty column always sources from 7d block (raw_difficulty_ratio is 7d-computed by definition). Immediately surfaced t 24h Total Lift −43.8% — biggest per-field regression of the day. Both cascades individually fine (HRRR skill +3.1%, NBM skill −1.2%), selector routing wrong: HRRR raw 1.16 vs NBM raw 0.515 in the last 24h — 2.25× inversion of the 30d pattern where they're within 3%. DISCOVERY — regime axis dig: checked by-regime walker (t has zero cleared cells across all 36 regime×band combos — all lift is NEGATIVE historically). Sketched a "stagnant_high" regime axis (weak synoptic forcing) as a peer label to existing wind-flow bins. Retroactive pair-log sweep at ws<5 & cc<0.4 & |pt_3h|<0.5 threshold: ws +16.6% NBM lift under stag vs +3.7% non-stag; sr +35.4% vs +13.8%; ch a reverse anti-signal −46.1% vs +34.6% (HRRR wins in stagnant clear-air on cirrus). t not a clean signal across threshold sweep — non-monotonic, no ship candidate at this axis alone. v0.6.605 SHIP + collector deploy — stagnant_high regime label added to classify_synoptic_regime(), checked before frontal/calm so light-wind + clear + steady-pressure states get their own bin. Wire-through in forecast_error_log.py (both fc + obs) and solar_correction.py. Stamp-only — zero routing change; walker consumes the new label transparently. First stagnant_high rows expected on next joiner write (~16:07 UTC). DISCOVERY — t regression by-hour dig: broke 24h t errors down by obs-hour × regime × lead-band. Signature is unambiguous — HRRR MAE at UTC 09/10/11 is 2.74 / 3.40 / 2.49 with bias +2.48 / +3.13 / +2.28 (consistently ~3°F warm). NBM MAE at same hours 0.67 / 0.62 / 0.59. Classic HRRR boundary-layer overshoot: under clear skies + light wind, overnight radiational cooling makes surface very cold; HRRR's PBL scheme mixes down aloft warm air too aggressively as the sun rises. NBM's climatological smoothing sidesteps it. Prior 6d se_flow data shows NBM was actually WORSE than HRRR at short-lead t (1.49 vs 1.06); today is a complete flip on the specific morning-heating hours only. v0.6.606 SHIP + collector deploy — l1_selector.pick_source() gains optional hour_local kwarg + a named routing gate: field=="t" AND regime=="stagnant_high" AND hour_local ∈ {4,5,6,7,8} (EDT) → return "nbm". Highest precedence (fires before walker overrides and pool table). Reversible via HRRR_PBL_MORNING_OVERSHOOT_KILL. Wire-through in forecast_snapshot.py passes _chp_valid_hour_local(times, i). Why hardcoded not walker-driven: escalation clause needs ~500 stag t rows in one cell (~4-7d accumulation at 2% stag rate). Named gate closes the gap immediately today. Prune when walker's stagnant_high × t × 0-5 cell wires. Afternoon reversals: (1) I over-scoped the "24H chart" ask — built a 24h Stack Health aggregate chart when the user wanted a 24h per-field table. Ripped it out including the producer schema change (last_48h_hourly bucket) and re-shipped correctly as the 24h table. (2) I said "would a 24h version have surfaced dp?" YES → was WRONG (dp is derived-omitted from the diagnostic table by design). Corrected mid-session before user acted on the claim. Afternoon post-ship watches: (a) first stagnant_high pair-log rows next joiner write; (b) t 24h Total Lift trends toward zero as morning-hour rows use NBM; (c) selector_picks for t stagnant_high should show NBM dominant in EDT 04-08 window; (d) walker stagnant_high × t × 0-5 cell — remove the named gate when it clears the wire; (e) stagnant_high × ws and stagnant_high × sr cells expected to escalation-wire earliest given retroactive lift magnitude. Afternoon lessons: (1) build the diagnostic FIRST, then the fix — the 24h table surfaced the regression, the by-hour dig identified the mechanism, and both landed same-session with real understanding not a guess. (2) "Would X have surfaced Y" claims need code-check not intuition — I got the dp caveat wrong initially. (3) Stagnant-high axis was correct as a general mechanism (ws/sr/wg/ch all real signal) but wrong for t specifically — the t failure is hour-of-day × stagnancy, not stagnancy alone. Named gate is the right shape when the axis is real but sample thin per specific cell.
Evening session — L1 recency-override audit + Simpson-guard shadow: user asked whether the selector's "still sucking" 24h read on h (sel_h −66% at scan-start) was a routing bug or a mechanism-working-correctly result. Dig sequence: (1) per_field_scoring 24h read — cascade heavily positive (ch +66%, cm +14%, sr +15%, wg +16%, wd +12%), selector red on h (−66%), dp (−30%), ws NBM-oracle (−62%). (2) sel_n metric verified as symmetric alt-baseline artifact, not evidence of a bad NWS wire. (3) h/*/0-5 by-regime picture — 8/8 regimes show HRRR winning 25-127% on 30d, all halves_stable_hrrr=True on 7 of 8; yet current routing sends h/0-5 to NBM. (4) Traced to L1 recency override — v0.6.546 mechanism pooled at (field, band) flipped h/0-5 hrrr→nbm on a +5.4% recent 7d lift while every regime individually favored HRRR on 30d. Textbook Simpson's-paradox shape. (5) Re-audited all 9 current overrides via per-regime lens: 6 of 9 are Simpson-shaped (h/0-5 8/8, sr all 4 bands 7-8/8, dp/0-5 8/8), 3 are legit (t/24-47, wg/0-5, wd/0-5 all had marginal 30d + regime disagreement). v0.6.608 SHIP (analysis-only, SHADOW ONLY, no collector/PWA change) — analysis/l1_selector_fit.py gains a per-regime 30d accumulator + Simpson-guard shadow annotations. New per-cell fields: simpson_guard_would_veto, source_under_simpson_guard, simpson_guard_note. New top-level simpson_guard_shadow block with vetoed-cell list. Guard params: MIN_N_PER_REGIME=100, MIN_MEASURED_REGIMES=6, MIN_AGREE=7, MIN_MEDIAN_LIFT_PCT=20.0. Runtime pick_source unchanged. Companion analysis/simpson_guard_shadow.py re-scans pair-log over 7d, computes pooled prod-MAE under source (current NBM) vs source_under_simpson_guard (counterfactual HRRR), publishes to GCS simpson_guard_shadow.json. DECISIVE SHADOW READ: today's fresh fit fires 5 vetoes (sr all 4 bands + dp/0-5; h/0-5's recency override quietly stopped firing today at +0.3% recent lift). 7d prod-MAE deltas: sr/0-5 −6.2%, sr/6-11 −16.2%, sr/12-23 −10.5%, sr/24-47 −8.6%, dp/0-5 −8.9%. Pooled −14.2%. Guard would regress prod on every flagged cell. Shipping the guard is REJECTED by data. Verdict: the recency override IS earning its keep even on cells that look Simpson-shaped by per-regime 30d comparison — 30d per-regime data is a stale/wrong baseline for judging recent flips. Joe's [[project_selector_recency_override_watch]] Checkpoint 2 verdict ("DO NOT tighten threshold — mechanism working") is reinforced hard. h 24h regression follow-up: post-fit sel_h moved −66% → −41% because h/0-5 flipped hrrr in this run. Layer-by-layer decomp on the h 24h pool showed 100% of rows delivered by l2_nbm (bands 6-11 and 12-23 still route NBM in fresh fit); HRRR-side L2 station bias delivers 3.6-5.8 MAE across bands vs L2_NBM 6-10 in the 24h window — but 30d by-regime prod-vs-prod shows NBM cascade wins at h/6-11 and h/12-23 in 7/8 regimes with strong effect sizes. Short-lived HRRR-favorable pattern, not a routing bug. If HRRR keeps winning, 7d recency will flip h/6-11 in 3-4 more days automatically. Watch item logged: L2_NBM MAE on h/6-11 (9.02) is barely different from NBM raw (9.06) — station-bias Kalman doing nothing on that cell in the 24h window; compare h/0-5 where L2_NBM=6.02 vs raw 9.28 (−35%). Likely stale/thin Kalman state. Log-only, not urgent. Evening lesson: shadow-first was the right call. Designed guard looked mathematically defensible against pooling artifacts; measured it against real 7d prod and got the opposite answer. The [[feedback_shadow_write_applied_layer_trap]] discipline paid off — no runtime touched, no bad ship. See [[project_simpson_guard_shadow]].
2026-09-12 (Sat) — 2 days ago· 10 commits (v0.6.590 → v0.6.599) · 1 production change · 1 collector deploy · 8 analysis/documentation commits. Morning read on 09-11 escalation-clause bet vindicated: h -522% → +7.2% total lift, ws to +12.5%, sr +21.4%. v0.6.590 pr L2 unwire `nw_flow/6-11h` on 3-tool agreement (retro Δ-13.6% halves-negative both, layer-shape sentry +10.6%, yesterday's scoreboard -8.7%). Collector deployed 10:19 UTC rev `00563-few`. v0.6.591 session-end sweep for v0.6.590. Strategic pivot — I said the digest was steady-state incapable of producing new correction layers. Joe pushed: "come up with new shit to try." What followed: v0.6.592 digest pruning — 13 CLOSED-MISS/retired-mechanism scripts retired to .skip.py, 4 new (1 real rewrite of pre_front ortho with matched-regime baseline surfacing 4 ORTHOGONAL cells first-run + 3 Stage 0 scaffolds). v0.6.593 4 real Stage 0 hypothesis-tests: h_inter_model_spread_stage0.py → PROMOTE 35/36 cells (|forecast_l1 − forecast_raw_nbm| per row, Q4/Q1 MAE ratios 1.4×-5× halves-stable every field); h_prior_day_error_c1_stage0.py → PROMOTE 14/40 (strong 0-5h short-lead); h_diurnal_l2_tau_stage0.py → MARGINAL (sr 4.69pp + wg 3.37pp HOD spread); h_buoy_sst_gradient_stage0.py → SCAFFOLDING (plumbing-blocked). v0.6.594 Stage 1 orthogonality on both PROMOTEs vs 3 per-row-available C1 axes (C1a, cluster, pt_mag). inter_model_spread cleared all three (C1a 33/36, cluster 36/36, pt_mag 34/36 — 0 REDUNDANT on cluster). prior_day_err mixed 19/13/8 — narrow-ship 0-5h only. v0.6.595 Stage 2 preview h_inter_model_spread_c1_stage2.py — STAGE 2 PROMOTE, 33 SHIP cells (halves-stable premiums 40-500% for t/h/ws/wg/wd/cc/ch/dp; sr's night-Q1 pathology needs floor filter). 7-day gate armed. Earliest wire 2026-09-19 as C1 axis_6. First new C1-axis candidate to reach Stage 2 since cross_run_spread in June. v0.6.596 3 more Stage 0s exploiting under-used pair-log features (found forecast_nws/error_nws/selector_source per-row): 3-way spread PROMOTE 15/16 (not superior to 2-way, reduced coverage); NWS-optimal cell detection PROMOTE — 7 halves-stable cells where NWS beats best-of-HRRR-NBM: dp/nw_flow at 3 leads (25-30% better), dp/pre_frontal, dp/calm, ws/sw_flow 2 leads; state_fc-vs-state_obs disagreement PROMOTE both cloud_delta (12) + solar_delta (13). v0.6.597analysis/l1_selector_fit_3way.py — 3-way fitter (HRRR/NBM/NWS) analysis-only, mirrors by-regime fitter's 30d window + halves-stability. NWS covers 5 fields (t/wd/ws/dp 100%, pp 76%). First-run 6 NWS-wire cells clear: dp/nw_flow/12-23 +29.3% h1/h2 13.4/33.7 n=1,721 (escalation-clause eligible); dp/nw_flow/0-5 +26.1%; dp/pre_frontal/12-23 +24.8%; dp/nw_flow/24-47 +13.6%; dp/sw_flow/24-47 +9.8%; t/frontal/6-11 +9.7% (thin, masked). Runtime NOT touched — walker+pick_source extension is next session's opener. v0.6.598 session-end sweep (Recent Activity + memory). v0.6.599 follow-up sweep completion — Upcoming grid + Post-ship watches (the v0.6.598 sweep missed forward-looking sections; Joe caught it). v0.6.600 reviewer-driven fixes — h humidity row unstuck to "Recovered", h + cc NBM topology corrected to end at l2_nbm (both l3_nbm dropped 09-05 + 09-10 but text still showed them), scope reframed to "10 commits · 1 production change · 1 deploy". Session pipeline: 5 hypothesis-tests written, 4 PROMOTE, 1 MARGINAL, 1 HOLD; 1 candidate (inter_model_spread) fully through Stage 0 → Stage 1 → Stage 2 with 7-day gate armed; 1 concrete routing miss identified (NWS on dp/nw_flow, expected +975°F-days of dp error saved per month on top cell alone). Lesson: the digest as pipeline works when we feed it fresh hypotheses. What was "steady-state incapable" 6 hours ago is now a candidate 7 days from wire + a routing extension queued for the next session. Pipeline gates (Stage 0 → Stage 1 orthogonality → Stage 2 preview → 7-day stability) all fired correctly on real signal.
2026-09-12 (Sat) — trimmed-old· 1 ship (v0.6.590 pr L2 unwire) + 1 collector deploy. Escalation clause vindicated + 4 candidates reviewed and held with reason. Morning first-read on the 09-11 escalation-clause bet: h moved from -522% to +7.2% total lift, ws to +12.5%, sr +21.4%. 7d value-add mean +10.98% (6g/3a/0r). 24h still -67% aggregate but the fields targeted (h/ws) recovered. Two 09-11 wires flipped OUT of the walker today (ws/calm/0-5, ws/nw_flow/12-23) — walker correctly withdrew. HRRR-wire escalation cells (h/calm/24-47, ws/sea_breeze/24-47) both holding. v0.6.590 SHIP + collector deploy (10:19 UTC, rev myweather-collector-00563-few) — pr L2 gate unwired nw_flow/6-11h. Three-tool agreement: pr_l2_regime_lead_retro pooled Δ -13.6% over 1,447 pairs since 08-13, both halves negative (A -8.5% n=658 / B -17.7% n=789); layer-shape sentry pr/production@6-11h +10.6% vs raw; yesterday's Notable Calls flagged pr 6-11h -8.7% n=942. The 08-10 Stage 1 both-halves that justified the ship (A +10.3% / B +13.1%) has inverted — persistent regression across a month of live data, not noise. Fix: removed ("nw_flow", "6-11") from _PR_L2_FIRE_CELLS in weather_collector/processors/corrected_hourly.py:38. nw_flow/0-5h retained — retro confirms HEALTHY (pooled +6.4%, halves +11.7%/+1.9%). Shadow-wire stays unconditional. Holds (with reason): (1) dp -18% corr regression — standing rule per [[project_dp_is_derived_no_dp_work]]: dp = Magnus(t, h), both t (+14.4%) and h (+7.2%) winning at prod → the regression is either small-window artifact (like cm -654% today, CLEAN on anomaly detector) or dpbp misfiring on today's regime mix; not opening a dp workstream. (2) ch -17.7% / chp L6-baseline gate — h_ch_persistence_blend_stage2_vs_l6 shows 7 live chp cells losing to L6 (worst se_flow/6-11 Δ+29.91%), but escalation playbook requires 7 daily reads / 2-tool / per-cell / no-ENT — today is day 1. Playbook discipline over urgency. (3) wg.l3_nbm sentry HOT (+5.3% → -4.6%) — single-tool signal; walkforward disagrees (aggregate +3.1% over 30d); all 3 wg skip proposals STALE on today's 50d two-window audit — exactly what v0.6.584 exists to catch. (4) walkforward wdp_nbm DROP wd — on n=107 THIN, below any ship floor; l3_nbm ADD cc,wd — cc killed 09-10 (walkforward hasn't caught up), wd resolved via yesterday's CONFIRMED ship. Post-ship watches: pr pair-log MAE at nw_flow/6-11h should return to raw as new obs stamp applied_layer=l1; layer-shape sentry for pr 6-11h clears within days; ~09-19 pr L2 gate stability re-read. Clock-watches: 09-13 day 2/7 chp L6-baseline gate. 09-14 first HRRR-wire GATED read (7 candidates); sr ADDED_LAYERS prunable. 09-15 L3 DROP cm walkforward streak (day 5/7 today); KILLED_LAYERS 09-05 entries prunable. Lessons: standing rules first — the derived-field rule closed dp diagnostic in one memory read; playbook discipline over co-owner urgency — chp gate at day 1/7 is a hold, not a ship; two-window audit paying dividends — 3 wg skip proposals + 1 sentry HOT would have shipped 3 skip cells under the old 14d-only gate, all STALE on 50d; escalation-clause bet paid off — h VC -522% → +7.2% in one day validates the "halves-stable × large × plenty-of-n cells shouldn't wait 3 days" hypothesis.
2026-09-09 (Wed) — trimmed· 4 ships (v0.6.571 → v0.6.574). NBM skip-table symmetric REMOVE curation loop built + first REMOVE curation shipped + two-window verdict cross-check + revert. Morning digest 180/180 clean; 3 HOT/WATCH sentry rows all logged false-positives (cc.l3_nbm watch through 09-11, sr.l5_nbm/sr.l3_nbm through 09-14). v0.6.571 SHIP — debug page cleanup pass on Upcoming grid + Post-ship watches: 5 ✓-completed rows pruned from Upcoming, 3 CLOSED-CLEAN items dropped from active Post-ship watches (still in footer summary), L1 walker milestone rewritten from Mon-09-07 + ~Mon-09-14 to ~Fri 09-11 reflecting v0.6.566 gate loosen. Also backfilled CHANGELOG.md for v0.6.569 (attribution decomp) + v0.6.570 (debug text sweep) which had shipped 09-08 but were missing from the changelog. v0.6.572 SHIP — analysis/nbm_skip_earning_audit.py. Symmetric REMOVE mechanism on the NBM skip table. For every cell in skip_table_nbm_curated.json, evaluate on recent pair-log data: what would error_l3_nbm have been if the cell were not skipped, using the current pooled l3_nbm_curated.json bias. Compare MAE(counterfactual) vs MAE(input). If lift ≥ 3% with n ≥ 50 AND both halves positive → REMOVE candidate. Runs in daily digest via run_digest.sh's auto-discovery of analysis/*.py; wired into build_executive_summary.py's new "NBM stale-skip proposals" section directly below the existing ADD-side "NBM skip-table proposals." Analysis-only. First-run finding: 7 of 19 skip cells flagged REMOVE (37% of table apparently stale). Motivation: the ADD-side skip curation has been running daily since v0.6.462 (08-21); the REMOVE side had never existed — cells added weeks ago on evidence were never re-tested against fresh data. Asymmetric process → skip table drifts toward over-skipping as regimes shift. v0.6.573 SHIP + collector deploy — first-ever REMOVE curation on the NBM skip table. All 7 REMOVE candidates shipped: wg se_flow 0-5h/6-11h/12-23h, wg nw_flow 6-11h/12-23h, wg pre_frontal 12-23h, h se_flow 24-47h. Table: 19 → 12 cells. Deploy 15:19 UTC, revision myweather-collector-00555-hiz, first tick 15:27 UTC clean. v0.6.574 SHIP + collector deploy — 50d cross-check on the 7 shipped removes surfaced 2 premature: h se_flow 24-47h (50d lift −2.22%, halves −8.60/+8.38 — skip still earns overall) and wg nw_flow 6-11h (50d halves-unstable −4.55/+12.76). Both cleared 14d only. Reverted both — table 12 → 14 cells. Audit refactored to require BOTH a 14d fresh AND 50d long window to clear before emitting REMOVE; new WATCH verdict for 14d-only signals (flagged in digest, no action). build_executive_summary.py emits separate REMOVE + WATCH blocks. Deploy 16:23 UTC revision myweather-collector-00556-yoy, first tick 16:27 UTC clean. Post-v0.6.574 audit state on the 14 remaining cells: 0 REMOVE, 3 WATCH, 10 HOLD, 1 THIN — the 3 WATCH are the 2 reverts plus wg sea_breeze 6-11h. The 5 net removes that stayed (all wg L3, all halves-stable on both 14d and 50d): se_flow 0-5h (+5.01%), se_flow 6-11h (+5.92%), se_flow 12-23h (+4.62%), nw_flow 12-23h (+6.64%), pre_frontal 12-23h (+7.76%). NBM cascade now applies pooled L3 bias in these cells. Selector already routes NBM for most (wg at 6-47h in most regimes). NBM specialist workstream reframed: 2 mis-framings during the session's exploratory reasoning had to be retracted before shipping — first was "NBM has no specialists" (wrong — v0.6.437-499 fully ported layer-for-layer); second was "L3_NBM is regime-blind" (wrong — the skip table is the regime layer; 19 cells sat on top of pooled bias). The real "why NBM specialists aren't earning" isn't missing gates — it's the ADD/REMOVE asymmetry. Also caught mid-session: my first diagnostic (analysis/l3_nbm_fit_by_regime.py) reported L3_NBM ch/sr as net-negative in production. Protocol bug — my strict train/test split used 7-14d-stale training data, but production refits daily. Actual stamped error_l3_nbm shows L3_NBM helping +19% on ch, +16% on sr in production. Sentry was right; my diagnostic was wrong. Corrected before recommending kills. Followups queued: (1) 12 days of accumulated ADD-side walkforward SKIP proposals since 08-28 — same rescore protocol applies, next-session workstream. (2) ADDED_LAYERS registry mirror of KILLED_LAYERS (small ship candidate, defers false-positive sentry HOTs on recently-added layers). (3) NBM sea-breeze specialist as a native workstream if 09-11 walker read + skip curation loop doesn't close enough of the gap. Clock-watches: 09-11 L1 by-regime walker first cell-wire read (day 3/3). 09-15 L3 DROP cm walkforward streak potentially clears (day 1/7 today). 09-15 KILLED_LAYERS registry prunable. Lessons: whenever we add a curation mechanism (ADD), verify a REMOVE mirror exists — asymmetric process drifts. Two-window verdict is the right shape for time-window-sensitive gates: 14d catches freshness, 50d catches robustness; neither alone. When re-fitting a diagnostic to compare against production behavior, mirror production's fit protocol (daily rolling refit), not a static one-shot. Read the codebase before making architectural claims — I mis-framed NBM state twice.
2026-09-08 (Tue) — trimmed· 3 ships (v0.6.561 → v0.6.563) + 1 collector deploy. Sentry plumbing + l4_nbm cc DROP. Morning digest triage on 2 sentry-HOT rows (ch.chp_nbm, h.l3_nbm) + t/production τ-suspect top alert. Both HOTs traced to legacy pair-log rows from v0.6.551 kills (2026-09-05) whose fresh 3d window still contained ~9-13h of pre-kill fires. Kills already live via v0.6.552 collector 09-06; sentry had no way to know. v0.6.561 SHIP — analysis/nbm_regression_sentry.py gained KILLED_LAYERS registry {(field, layer): kill_date_iso} seeded with (ch, chp_nbm) and (h, l3_nbm), both 2026-09-05. When sustained-window start ≤ kill_date and cell has any rows, verdict is KILLED with a "pre-kill rows aging out" note instead of HOT/WATCH. Digest exec-summary greps for HOT/WATCH → alerts stop firing automatically; row still prints for visibility. Verified: verdict flipped from "2 HOT" to "CLEAN — 6 NBM cells nominal (7 THIN)". Registry prunes naturally when both windows post-date the kill (rough clearance: 09-15 for today's entries). Analysis-only — no publisher redeploy (nbm_regression_sentry not in publisher/main.py:PUBLISHERS). v0.6.562 SHIP — τ-suspect flag in analysis/runlog/build_executive_summary.py:layer_shape_sentry now requires L2 MAE to differ from L1 MAE by ≥1% at the long-hurt band before labeling the shape as decay-τ. Driver: today's t/production τ-suspect (helps 0-5h -7.7%, hurts 12-23h +5.3%) was mislabeled — per-lead diagnostic from time_series_diagnostic.json showed L2 (τ=4h) fully decayed by lead 6, L2 MAE equalled L1 MAE from lead 6 onward. Real driver was L1 selector routing to NBM at leads 12-27 where NBM raw is ~5-14% worse than HRRR. decay_tau_tuning already stable on t at τ=4 (τ=42 shift +1.0%, noise) so shortening was a no-op; the alert misdirected the fix. Verified all HRRR-side layers (l2/l3/l4/l6/l1r) fully converged to l1 by lead 12; only nws (NBM-derived) column materially different at 12-23h. Production tracks nws row-for-row (lead 15: nws-l1 +0.31, prod-l1 +0.15 → ~48% NBM-routing fraction). Selector-routing issue is a known transient — L1 recency-override aftershock post-v0.6.540; by-regime walker armed 09-06 (v0.6.552) is the fix, earliest 7/7 clear 09-14. Not shipped: the routing fix itself, only the sentry mislabel guard. v0.6.563 SHIP + collector deploy — weather_collector/processors/l4_nbm.py: L4_NBM_FIELDS = ("ch",). cc dropped. Pulled from 09-09 scheduled slot ([[l4-nbm-cc-drop-prep]]) — walkforward has proposed DROP cc for multiple runs; 30d prep (n=27,313) confirmed layer's total marginal lift over L3 is +0.79% pooled, best-possible skip-table +1.38%, cells flip verdict week-to-week (se_flow 24-47h flipped -4.8% → +1.2% helping between windows). ≤60bp on a field with pool-wide MAE 19.5 = not worth skip-table maintenance. Scope-checked all L4_NBM consumers iterate the tuple dynamically (apply loop, gate telemetry, chp_nbm's ch_l4_nbm reference); no dangling refs. Deploy landed 12:14 UTC, first tick 12:17:02 clean (MEMPROBE 50.4 mib cold start, no NameErrors, no missing-attr errors). Selector's deepest NBM layer for cc now falls through to L3_NBM. Aside — Selector Skill tile color question: user asked why 24h median 54.2% is orange while 7d 57.4% is green. Thresholds at corrections_debug.html:4413 are {win: 55, lose: 50} for the selector_skill tile — ≥55 green (winning), 50-55 orange (flat), <50 red (losing). 54.2 is 0.8pp shy of winning; working as designed. Flags left open: h_hsf KILL from digest is not passive — hsf is wired as C1e axis in confidence_layer.py:464 (5-tuple spread_q::pt::trans::c1f::hsf_group keys). Retiring means confidence-layer refactor + curator changes + walkforward gate — dedicated session, not housekeeping. Deferred. Walkforward wants ADD t/wd/ws to l3_nbm; sentry shows THIN (n=0-28) on those fields and t is a historical kill (v0.6.472), so hold pending its own investigation. Followups queued (unshipped): (1) prune KILLED_LAYERS entries once both windows post-date the kill (~09-15). (2) deploy-hygiene enforcement in CLAUDE.md §8 or pre-push hook (carried from 09-07). (3) NBM-side specialist workstream (carried from 09-07). (4) h_hsf C1e-axis retirement. Clock-watches: 09-14 earliest by-regime walker wire (day 3/7 today). 09-15 KILLED_LAYERS registry prunable. Lessons: τ-suspect heuristic was checking "helps short, hurts long" without checking which layer is actually applied at the long band — same class as the sentry flip-magnitude fix (v0.6.558). Symptoms-vs-mechanism gap in the alert. Also: legacy pair-log data outlives kill commits by up to a week (fresh 3d + sustained 7d = 10d); alerts on killed layers are a known false-positive class the KILLED registry now handles. Evening session (v0.6.564–569): scoreboard trust-check + selector-and-attribution reframing. v0.6.566 L1 by-regime walker gate 7d→3d + per-day n_today ≥ 20 floor; earliest cell wire 09-11. v0.6.567 Selector Skill tile primary Win Rate → Value Captured. v0.6.568 Selector Skill tile: drop Win Rate secondary + VC-tuned WFL thresholds ≥+33%/≤0%. v0.6.569 Attribution panels (mean + median) shipped — additive decomposition Total = Routing + Cascade in pp of user default; per-field Pipeline Lift column added to diagnostic; Selector Skill VC thresholds retuned to ≥+70%/≤+30% (prior +33%/0% painted +48% VC green even though the selector was leaving over half the oracle gap on the table). Also traced cc 24h REGRESS to l3_nbm hurting cc by ~1.5 MAE (sentry still CLEAN via marginal help vs l2 — 3-day watch, revisit 09-11); noted sr.l5_nbm HOT / sr.l3_nbm WATCH as false positives from the 09-04 sr add (sustained window pre-ramp; ignore through 09-14).
2026-09-07 (Mon) — trimmed· 5 ships (v0.6.554 → v0.6.558) + full scoring audit. Publisher CF was 12 days stale — make deploy-publisher shipped it forward; every hourly tick had been clobbering fresh output with pre-fix code. v0.6.555 _selected_l1_error NBM walk simplified (L3_NBM moved to correction side, not reference); Selector Skill primary swapped chooser_vs_prod → hit_rate. v0.6.556 Hit Rate → Win Rate (ties dropped from denominator). v0.6.557 scoreboard_v2 baseline reframed argmin → user-default (matches per_field_scoring's best_raw). v0.6.558 nbm_regression_sentry.py:220 flip-to-hurt clause requires abs(help_s) + abs(help_f) ≥ 3.0. Full context in [[project_09_07_session]].
2026-09-06 (Sun) — trimmed· 1 ship + 1 collector deploy + session-end sweep. Wired the by-regime walker into the L1 selector. Digest 219/219 pass, all OK. Morning scoreboard read: 7d value-add mean +0.72%, 24h -0.66%. h flipped GOOD +22.2% on 24h (v0.6.546 override + v0.6.551 kill both paying out), ch GOOD +39% (chp_nbm kill paying out). Ship gate router-scope +39.7% (was +48.4% 09-03) — continues to erode as 30d anchor loses volume. Recency override still 11 flips one-directional (HRRR→NBM), v0.6.540 aftershock bleeding out. v0.6.552 SHIP + collector deploy — weather_collector/processors/l1_selector.py extended: pick_source(field, lead_h, regime=None) gains optional regime param and loads l1_selector_by_regime_walker.json at module init. When a cell has cleared_for_wire=True AND flipped_in_window=False, routes NBM — precedence over band-pool pick. forecast_snapshot.py passes _wdp_state_fc_by_lead[i] (fc-time regime for that lead). Wire contract matches the walker's docstring exactly. Passive today — walker suppression ends 09-07, earliest 7/7 clear 09-14. Rev 00552-vut ACTIVE 11:15 UTC; first run 11:17 UTC cold-start clean (MEMPROBE 47.7→430.7 mib, no import errors, no NameErrors). Drove by cell-level Value Captured audit (option-4-first per user). User asked "what can we work on to make the selector better." Offered 4 options: (1) wire by-regime walker, (2) shorten 30d anchor to 14d, (3) drop ≥5% floor on recency override, (4) diagnostic first to decide. User chose 4-then-1. Scratchpad selector_cell_audit.py computed per (field, band, regime) cell — n_paired, hit_rate, chosen/alt/oracle prod MAE, value_captured_pct. 7d median VC -51% has two distinct drivers: (a) transitional plumbing — sr + wind cells the recency-override flipped to NBM this week still have 7d rows mostly pre-flip with no error_l3_nbm stamped → hrrr_fallback pick, HRRR shipped when NBM would have won (9 of top-20 worst cells are sr × hrrr_fallback, sr 12-23 calm VC -310%, ws 24-47 se_flow -360%). Evaporates on its own. (b) real per-regime routing errors band pool cannot see — ne_flow is the outlier: dp 24-47 ne_flow VC -818% (n=393), cc 12-23 ne_flow -150% (n=233), cc 6-11 ne_flow -64% (n=151), wd 12-23 ne_flow -55% (n=233), h 6-11 ne_flow -137% (n=124). By-regime walker was built for exactly (b). Positive-VC cells confirm flip potential (wd × nw_flow 12-23 VC +82%, ch × sw_flow 24-47 VC +83%, cc × pre_frontal +63-71%). Diagnostic conversation — pipeline vs selector reframing: user pushed back twice. (1) I called the 7d regression "an old wound rolling off the window." User pointed at the 24h Selector Skill card at median -19.9%. Four days after v0.6.540, that IS live. Retracted "wait" framing for the 24h card. (2) I over-committed to "kill L3_NBM for t/ws/dp" as same-medicine-as-yesterday. User pointed at the Accuracy scoreboard — HRRR Pipeline Skill +2.9/+11.1%, NBM Pipeline Skill +4.8/+10.0% (median/mean) all positive across every field. Both cascades are working. What's red is Hit Rate 47.5% + Value Captured -51.4% = pure selector metrics. Correct answer: selector issue, not pipeline. My kill-recommendation was wrong mechanism (L3_NBM barely runs on t/ws/dp — sentry columns THIN n=0-28). Retracted before shipping. Sentry triage — wd.l3_nbm HOT NOT disease. Sentry logged HOT via flipped_to_hurt = help_s > 0 AND help_f < 0 clause (help_s +5.35% → help_f -0.18% crosses zero). But move is only 5.5pp. Cell-level breakdown on paired raw+l3 rows: raw wd MAE spiked +31% by itself (nw_flow 24-47 raw +60%, pre_frontal 24-47 raw +34%, sw_flow 24-47 raw +121%, sea_breeze 24-47 raw +155%). L3 tracked raw. Paired-total Δhelp only +0.88pp, well below WATCH's 8pp. Not disease. Same class as 09-04 raw-drift false positive. Followup queued, NOT shipped: flip clause needs abs(help_s) + abs(help_f) ≥ 3pp min-magnitude gate at nbm_regression_sentry.py:220. Same class as v0.6.548's marginal-help refactor. Scope creep to ship today. Sentry stale HOTs: h.l3_nbm and ch.chp_nbm both still HOT because sustained/fresh windows include pre-kill data — expected to clear ~09-08. v0.6.553 session-end sweep — this Recent-Activity entry + 09-05 entry + memory MEMORY.md READ FIRST/SECOND/THIRD refreshed + new [[project_09_06_session]] file. Day-labels shifted: 09-05 → 1 day ago, 09-04 → 2 days ago, 09-03 dropped from recent. Lessons: answer the actual question, not yesterday's narrative — Joe caught two "fit today's data into yesterday's story" errors. [[feedback_refresh_current_state_before_defending]] applies. Also: diagnostic-before-ship pays when audit could invalidate the assumed lever — Joe's #4-first ordering was correct even though the audit confirmed the ship (it might not have). [[feedback_stop_after_minimum_ship]] held — flagged wd sentry followup, did not ship it. Clock-watches: L1 by-regime walker suppression ends 09-07 (tomorrow) → day 1/7 → earliest wire 09-14. l4_nbm cc DROP scheduled 09-09. h_cc_blend Stage 1 day 4 or 5/7 (holding). NBM skip-table curation 09-09.
2026-09-05 (Sat) — trimmed· 1 ship + 1 collector deploy. Killed two NBM correction layers found sentry-HOT for a second straight day.v0.6.551 SHIP + collector deploy — two module-level flags kill h from L3_NBM_FIELDS in weather_collector/processors/l3_nbm.py and set CHP_NBM_CH_KILL = True in weather_collector/processors/forecast_snapshot.py. Both cells had been sentry-HOT under the marginal-help metric for the 2nd consecutive day. Per-cell breakdown (regime × band, n≥40 per window): h.l3_nbm shows 14 of 15 cells with help_fresh negative — only 1 improver (se_flow 24-47 flipped +8 → +42). Systemic degradation, not weather-mix. ch.chp_nbm had only 4 cells clear n floor; ALL 4 (se_flow 12-23/6-11/24-47, pre_frontal 6-11) flipped help→hurt. Pattern: L2_NBM MAE dropped fresh (input got easier), L3/chp static bias tables kept applying. Both reversible one-liners; flip back when sustained-3d sentry clears. Rev 00551-xob serving 100% traffic ~15:00 EDT, subsequent runs clean (MEMPROBE 985 mib steady). Same disease identified on t/ws/dp/sr — NBM corrections dragging NBM prod below NBM raw — but sentry can't diagnose yet because L3_NBM error columns still THIN for those fields (n_sust 0-28). ~2-3 days out until diagnosable. Diagnostic conversation reframed scoreboard columns: user pushback landed twice. (1) Selector Skill picks min-PROD not min-raw (per per_field_scoring.py:596, value_captured_pct = (alt_prod − chosen_prod) / (alt_prod − oracle_prod) where oracle is per-row min-prod). [[feedback_selector_prod_vs_prod]]. (2) Total Lift baseline is NBM raw (user default), not min(HRRR raw, NBM raw). Per per_field_scoring.py:522best_raw = NBM for any NBM-scope field. [[feedback_baseline_is_user_default]]. Value-Add negative means one of {routing stale, corrections regress on routed path, both stacks under water} — read the sel_h / sel_n / corr / total decomposition to know which. h Value-Add already recovered +21% → +4.0% GOOD today (v0.6.546 override + v0.6.551 kill both paying out on 24h). Debug page sweep held for 09-06 pending pair-log rotation — swept today.
2026-09-04 (Fri) and earlier— trimmed to rolling 3-day window. Older narrative in docs/CHANGELOG.md and git log.
Applicability map — what corrections trigger, why, and when they actually fire
Two lenses on the same object. Applicability (top) describes what's configured to fire and under what gates — built each tick from describe_applicability() in each correction module, so the page can never drift from the code. Runtime firing frequency (bottom, 7-day rolling) shows what's actually firing per operator × field × regime, so silent dormancy (config says "enabled" but the code path never mutates a value — the class of bug that hid ws L3 for 4 days after v0.6.279) surfaces as an ★ on cells with 0 fires despite ≥5 ticks in that regime.
How to read this
Three categories: General-purpose layers (L1–L4) can apply to any field; the per-field applicability set controls which actually trigger. Specialists (Lsr, Lt, Lc) are domain-scoped by construction — the physics of the correction binds them to a field type. Confidence (C1) doesn't modify forecast values; it widens or narrows uncertainty bands along orthogonal axes (transition, pressure tendency, mesonet spread, pre-frontal proximity) that cross multiple (field, band) cells.
HRRR + NBM cascades render together below. HRRR-side layers (L2 mesonet blend, L3 decay, L4 diurnal, Lsr, Lc, chp/clp/wdp/dpbp/wsbp/lsb, C1) plus NBM-side layers (l3_nbm, l4_nbm, l5_nbm, l6_nbm, skip_table_nbm) are all unioned into the same block. Selector (l1_selector) picks between HRRR-Prod and NBM-Prod per (field, lead-band) after both cascades finish. L2_NBM as of v0.6.499 is a NATIVE fit (not a delta reconstruction) — described in the hand-curated block below.
applies when is the predicate in plain text. gated by names the module-level constant the gate reads (omitted if always-on). current state resolves the constant for the reader — what the gate is doing this very tick.
Terminology key:applies = correction applied to a (field, row) this tick · enabled = feature switch (the module-level constant is on) · active = runtime branch (which code path actually fired).
Source: weather_data.applicability_map.layers. Schema: weather_collector/data/applicability_map_schema.json. If this section says "not available," the collector hasn't shipped the block yet — it landed v0.6.260.
L2 — Aggregate bias (mesonet blend)general-purposehand-curated · promote to describe_applicability()
field
applies when
gated by
current state
t
Always — additive bias from station network, scaled by Kalman gain K
always on
applies
dp
Always — additive bias from station network
always on
applies
h
Always — additive bias scaled by Kalman gain K
always on (K-taper 1.0 → 0.4 by lead 24h)
applies
pr
RE-ENABLED 2026-08-10 v0.6.401 with regime gate — K=1 additive, τ=8h fitted, applied on (regime, lead_band) ∈ {(nw_flow, 0-5h), (nw_flow, 6-11h)}; every other cell stays raw. Cell selection from analysis/pr_l2_regime_lead_retro.py 08-10 (Jaccard 0.50 STAGE 1 SHIP, both-halves winners A +21.8%/B +41.6% and A +10.3%/B +13.1% on 8,596 shadow rows). History: DISABLED 2026-07-01 v0.6.276 after pooled Production +2.4% BAD — but that pooled read was regime-blind. The nw_flow short-lead win was real all along; needed regime-conditional cross-cut + halves verification to surface. Shadow-wire since 2026-07-29 v0.6.389 supplied the fresh data.
regime-gated (current-tick regime × lead band)
applies on nw_flow 0-11h; other cells raw
cc
Always — Kalman-blended KBOS + KBVY METAR override on hourly[0] (current-hour value used by cards). Not propagated across forecast leads — promotion to a true per-lead L2 bias correction is queued (see Open architectural questions).
always on (needs KBOS or KBVY at current hour)
applies
cl / cm / ch
Derived — same Kalman K applied to L/M/H splits on hourly[0] to keep them self-consistent with cc. Not propagated across leads, same as cc.
always on (derives from cc)
applies
ws / wg
Always — direct selection from per-octant median (no per-station bias track); KBOS+KBVY authoritative-source floor when both agree at >1.4× the octant median
always on (direct-selection, not additive)
applies (direct)
sr / pp / pa
N/A — no station network for this field; L2 row is structurally empty
n/a — no obs network
n/a
wd
Always — SHIPPED 2026-07-20 v0.6.368a. Circular unit-vector blend in wind_blend.py: obs wd + fc wd → (sin, cos), weighted mean by linear decay over 24 leads, atan2 back to degrees. Solves the wrap-around that made "N/A — L2's linear math doesn't apply" true prior to this ship. Calm-floor guard WIND_DIR_MIN_SPEED = 3.0 mph skips cells where both obs and fc speed are below floor. raw_wind_direction preserved for baseline. Complementary wd persistence gate (Stage 2 07-20 v0.6.365) targets long-lead regime-transition cells where L2 decay has expired.
always on (needs both obs+fc speed ≥3 mph)
applies (circular)
L2_NBM — NATIVE fit (v0.6.499)general-purpose · NBM cascadehand-curated · not a standalone module
v0.6.499 (2026-08-26) — L2_NBM is now NATIVE for all 8 NBM-scope fields. Full delta-transfer retirement. Every family mirrors its HRRR counterpart, re-anchored against NBM raw:
t/h/dp — station-consensus additive bias + Kalman gain + decay curve from corrected_hourly.py, anchored to raw_nbm at h=0. dp = Magnus(l2_nbm_t, l2_nbm_h).
ws/wg/wd — wind_blend.py's linear time-decay blend of the current observed value into [0, BLEND_HOURS), applied to raw_nbm arrays. wd uses circular unit-vector blend with WIND_DIR_MIN_SPEED calm-floor guard.
cc/ch — cloud_obs_blend.py's K·(obs_mean − raw) at hourly[0] only, reusing HRRR's K + obs means from cloud_l2_meta.
sr — identity. No station network for solar → no L2 either side.
All native math shares the same station observations / Kalman gain / decay curves the HRRR side uses. The only substitution is which raw forecast the correction anchors against. This is exactly the fix the 08-26 nbm_l2_delta_audit prescribed: the station-network bias IS model-independent, but the anchor (raw at h=0) is not, and the delta transfer conflated the two.
Always — direct-selection delta from HRRR L2 wind blend rides across
rides on HRRR L2
applies (direct)
wd
Always — circular unit-vector delta from HRRR L2 wind_blend rides across (calm-floor 3 mph inherited)
rides on HRRR L2
applies (circular)
cc / ch
hourly[0] only — inherits HRRR L2 KBOS+KBVY METAR delta on current hour; not propagated across leads (same limitation as HRRR L2 cloud fields)
rides on HRRR L2
applies on hourly[0]; other leads pass through
sr
Passes through raw — HRRR L2 has no sr correction (no station network), so delta is 0.
n/a — HRRR L2 sr is n/a
l2_nbm = raw_nbm
cl / cm / pp / pa / pr
Not emitted by NBM CO grib — permanent HRRR-only. Selector never picks NBM for these; no L2_NBM row exists.
n/a
n/a
🔍 L2_NBM soundness — 30d field × leadmeasurement
Grades the HRRR-delta transfer per (field, lead-band). Positive lift = delta helps NBM (assumption vindicated); near-zero = neutral (delta washes at long lead by construction — K-taper); negative = delta HURTS raw NBM for this cell (Wyman Cove bias is model-dependent). ★ = both halves of window agree on sign. Source: analysis/nbm_l2_delta_audit.py → gs://myweather-data/nbm_l2_delta_audit.json.
loading nbm_l2_delta_audit.json…
L1 blender shadow statusv0.7.3 · shadow only · fresh-data since 2026-09-24 16:00 UTC
Per-cell verdict on whether the blender's shadow forecast is beating what the selector actually served. SHIP-READY = 7d n≥50, lift ≥ +3% vs served, ω drift ≤ 0.20 · HOLD = insufficient fires or below threshold · KILL = 7d lift ≤ −3% (blend hurts) · THIN = zero fires (regime hasn't matched). Add cells to BLENDER_APPLIED_FIELDS in weather_collector/processors/l1_selector.py only after SHIP-READY on fresh data. Caveat: the 13 rows in the table reflect the stale-window l1_blender_curated.json shipped v0.7.0. Table rebuild after re-curation on fresh data will trim it to the 2-3 cells that survive halves-stable. Source: analysis/l1_blender_shadow_verify.py → gs://myweather-data/l1_blender_shadow_verify.json.
loading l1_blender_shadow_verify.json…
Runtime firing frequency — 7-day rolling
Fires = the correction actually mutated a value at that lead (not "would have applied"). Skips = would-fire cells suppressed by the skip table (L3 ws in ne_flow / short-lead sea_breeze; Lsr sr in ne_flow / calm) or by a gate being OFF (MLC currently ENABLED=False → all matching cells count as skips). Rate/tick = fires ÷ ticks_in_regime; for L3/L4 with 48 leads, healthy is close to 48. ★ marks (operator × field × regime) cells with 0 fires despite ≥5 ticks in that regime — silent-dormancy candidates.
Source: gate_firing_rollup.json (nightly digest via analysis/gate_firing_rollup.py) over per-tick gate_firing_log.jsonl (written by weather_collector/processors/gate_firing_log.py, v0.6.318). NBM operators added v0.6.458 (F4, 2026-08-21): L3_NBM, L4_NBM, L5_NBM, L6_NBM, CHP_NBM, WDP_NBM now emit per-tick fires/skips from forecast_snapshot.py; expect them to appear in the 7-day rollup once the rollup script's next digest cycle picks up the newly-accumulating log rows.
L1 — Raw model input (parallel HRRR + NBM cascades; selector picks per cell)
Two forecast sources run in parallel for every hour: Open-Meteo HRRR/GFS on one side, direct NBM (National Blend of Models) on the other. Each side runs its own full correction cascade. The 🎯 selector below picks per (field, lead-band) which cascade's output becomes user-visible. Multi-kilometer grid resolution on both — still knows nothing specifically about Wyman Cove; that's what the layers below add.
About the L1 sources (post-selector-arm v0.6.445)
HRRR-side (Open-Meteo). Open-Meteo HRRR (next 48h) + GFS (days 3-7), fetched every 10-min tick. Feeds the full HRRR cascade: L1 → L2 mesonet blend → L3 lead-decay → L4 diurnal → L5 Lsr (sr) / L6 Lc (cm/ch) / specialists (chp, wdp, clp, dpbp) / Ccd (cc derivation). Every field starts from a HRRR value at L1.
NBM-side (direct grib ingester). NBM CO 2.5km grib byte-range fetched hourly at :45 UTC by the nbm-ingester CF (v0.6.433). Extracts 9 fields at Wyman Cove: t / dp / ws / wd / wg / sr / cc / ch / h. Full parallel cascade shipped 2026-08-21: raw_nbm → l2_nbm (HRRR-delta reconstruction for all 9) → l3_nbm (per-lead bias, LIVE v0.6.445, scope trimmed 08-25 v0.6.472 to {wg, h, ch, cc}; sr re-added 09-04 v0.6.549) → l4_nbm for ch only (cc dropped 09-08 v0.6.563; diurnal hour-of-day) → l5_nbm for srKILLED 08-25 v0.6.471 (sentry +238% MAE + walkforward −126% to −147% agreed) → l6_nbm for t (v0.6.453 scaffold, ENABLED=False mirroring HRRR L6) → chp_nbm for ch (v0.6.454) / wdp_nbm for wd (v0.6.437). clp_nbm N/A (NBM doesn't emit cl). Path B discipline (owner call 08-25): any NBM cascade layer that can't beat its own raw gets killed, not routed around — a broken cascade should be diagnosed loudly, not silently masked. Adding raw as a selector fallback was considered and rejected on those grounds.
🎯 L1 selector (Prod-vs-Prod, v0.6.440 fit / v0.6.445 armed). For each (field, lead-band), argmin of recent-window MAE(HRRR-Prod, deepest applied HRRR layer) vs MAE(NBM-Prod, deepest applied NBM layer — l3_nbm → l4_nbm → l5_nbm → l6_nbm → chp_nbm/wdp_nbm as each fires). Fires per l1_selector_table_curated.json. In-scope fields: t / ws / wg / wd / h / ch / sr / dp / cc (all 9 NBM emits). Out-of-scope (HRRR-only): cl / cm / pr / pa / pp — NBM doesn't publish them. Currently 10 cells pick NBM: wg 12-47h, wd 12-23h (added 08-26 v0.6.500 after skip-table expansion + native L2 refit), dp 6-47h, cc all leads.
v0.6.432 L1 router retired 2026-08-19 v0.6.437. Superseded by the selector on all counts: wider scope (9 fields vs 3), stronger evidence base (30-day scoreboard vs 14-day, plus 108-day backfill for L3 fit), post-cascade application (all corrections apply first, selector picks last).
Curves below show the current L1 forecast per field for the HRRR-side only. NBM raw is available on the per-field accuracy charts as a dashed line (raw_nbm) once the Fitter has a day of pair-log coverage.
🔀 L1 router — current tick's per-hour source per field (runs first: what enters the pipeline as L1)
v0.6.432, live-tick view. Reads hourly.corrected_temperature_router_source / wind_speed_router_source / wind_direction_router_source from weather_data.json. Cascade output (*_pre_router) shown side-by-side with the router's live value at every hour the router fired, so you can see the router's per-hour delta as raw numbers before the next Fitter run rolls the L1r line onto the accuracy chart above. The raw-model curves below reflect the router's per-hour selection: cyan segments = NBM (via NWS-gridpoint), gray segments = Open-Meteo HRRR/GFS.
For each (field, lead-band), argmin of recent-window MAE(HRRR-Prod, deepest applied HRRR layer) vs MAE(NBM-Prod, deepest applied NBM layer — l3_nbm → l4_nbm → l5_nbm → l6_nbm → chp_nbm/wdp_nbm as each fires) picks the cascade output that becomes the user-facing forecast. Table refit by analysis/l1_selector_fit.py; stored at weather_collector/data/l1_selector_table_curated.json. Selector runs after all NBM stamping in forecast_snapshot; when picked "nbm", replaces top-level {field} with the deepest applied NBM layer's value (e.g. {field}_chp_nbm for ch when the specialist fires, else {field}_l4_nbm, else {field}_l3_nbm, ... falling through to {field}_raw_nbm). Fall-through to HRRR is safe.
loading l1_selector_table_curated.json…
Raw L1 forecast per field (post-router where the router fired)
L2 — Aggregate-bias correction (local station network)
What our 66 nearby weather stations say the model is getting wrong right now — and how much of that signal we trust.
How the network correction is built
Two parallel aggregation paths under this layer, each suited to the noise behavior of its metric:
Temp / humidity / pressure (additive bias): per-station Kalman offset (rolling 48h) → per-octant 1/distance² × exp(-|elev_diff|/30)-weighted (station − model) bias → unweighted mean across non-empty octants. t and h scale by per-field network Kalman K (sec 2c); pr applies full strength.
Lead-decay:bias_applied(lead) = current_bias × exp(-lead/τ) per field. Fitter (03/15 local) refits τ on train/test split; loader adopts fitted τ only if it beat hardcoded default on held-out RMSE AND stays within 0.25×–4× default. Defaults: τ_t=4h, τ_h=240h, τ_pr=12h. Fields without τ (wind/clouds/solar) get flat L2. See sec 2d for live curves.
Wind / gust (direct selection): no per-station calibration, no additive bias. Per-octant MAX gust → MEDIAN across octants; linear-decay blend into next 24h (100% h0 → 0% h24). KBOS+KBVY floor: if both airports agree on speed >1.4× octant-median, defer to their median. Direction rejected if >60° off airport+buoy+Tempest consensus. Gust override allows single-source (METAR omits when steady). Physical floor: gust ≥ wind.
Cloud cover (Kalman METAR blend): KBOS+KBVY sky obs (BKN/SCT/FEW/OVC → percent + L/M/H) blended against HRRR hourly[0] via _kalman_gain_cloud. K=0.90 if airports agree within 20pp, 0.70 at 20-40pp, 0.50 at >40pp, 0.35 for single-source. Same K on total + L/M/H for self-consistency. Live values in cloud_l2_meta.
2a. Octant coverage — where this tick's stations came from
2e. Post-aggregate-bias forecast — what gets passed to L3 · engineering view (pre-clamp)
⚠ These are internal pipeline values, not user forecasts. After the aggregate-bias correction is applied but before downstream layers and physical-bounds clamping (FIELD_BOUNDS in decay_apply.py). Values can legitimately fall outside physical ranges here — cloud_cover 121%, precip_probability −6%, precip_amount −0.025 in — because the bias offset is additive and the clamp comes downstream. If any of these look wrong for user display, check the L3 / L4 / cloud-clamp path, not this section.
L3 — Lead-decay correction
How the model's error tends to grow with each hour into the future — and the per-field nudge we apply to push back against that drift.
How decay correction works
For each field and each lead hour, we learn from millions of historical pairs how far off the model usually is — then bend the forecast back toward truth by that amount. Some fields (wind, gusts, high & mid cloud, POP) genuinely benefit; others net-negative on held-out data and are paused (see Applicability below). Tracked over a 30-day window with an exponential recency weighting — default τ=14 days, with per-field overrides for fields where analysis/decay_tau_tuning.py shows ≥5% MAE improvement vs the default. Current overrides: pp (POP) at τ=28d (+11.1% held-out, 2026-06-21 v0.6.167; note: today's read shows only +3.2% vs τ=14, now below the 5% floor — flagged for next τ-audit day). pa (precip amount) reverted 2026-07-19 v0.6.358 after today's read landed +0.9% vs τ=14 (was +5.9% yesterday); pa's best-τ swung 28→42→7 across three IMPLEMENT reads. See [[feedback_tau_streak_gate_limits]].
Applicability
Held-out MAE audit picks the fields where L3 actually beats L2 — currently wind speed, gusts, high cloud, mid cloud. Other fields stay off because the correction at best ties (temperature, pressure) or actively hurts (humidity, dew point, solar, low cloud, precip). POP is the special case: it's evaluated by Brier score, not MAE, so the audit's MAE-based ⚠ rule is suppressed for it — the v0.6.20 calibration analysis showed flat-additive correction cuts Brier 5%. Currently applied: —. Brier-evaluated: —. See Applicability map for live triggers and per-field gates.
Live state — with vs without decay correction
Historical calibration — fitted correction curves per lead hour
Calibration history — decay curves over time
L4 — Diurnal correction
A separate correction for the part of model error that follows the sun — e.g., bias that's different at 3 AM than at 3 PM. Currently applied to high cloud and cloud cover.
How diurnal correction works
Bins historical errors by hour-of-day (0–23) and fits the persistent pattern at each hour. Most fields don't have a clean diurnal signal once L2/L3 have run, so this layer's applicability set is narrow — currently {ch, cc} (cc added 2026-06-24 v0.6.214). Fits for fields outside the applicability set are still computed below for diagnostic purposes. See the Applicability map for the live trigger state.
Applicability
L4 corrects the portion of forecast error that repeats with time of day. It is the hardest layer to earn because the same hour-of-day bias must recur consistently across many days. Most fields fail that test because their dominant errors are driven by changing weather regimes (air mass, cloud regime, frontal timing, etc.) rather than the clock. Currently applied to {ch, cc} — the two fields showing a stable enough diurnal signal to pass the held-out audit. Fits for the remaining fields are retained in the calibration sections below as diagnostics. See Applicability map for live gates.
Live state — what L4 is doing this tick
No dedicated live-state panel yet. Current per-tick L4 deltas land in weather_data.hourly[*].corrected_<field>; the applied-vs-unapplied comparison is visible via the L4 lines on the per-field accuracy charts at the top of the page. Future build-out — placeholder for a dedicated live snapshot mirroring L3 §3.
Historical calibration — diurnal correction curves over time
Calibration history
L4 doesn't currently maintain a separate fit-history view — the curves-over-time panel above serves both "current fit" and "how it's changed" roles. Split out if/when a dedicated current-fit snapshot ships.
A per-regime W/m² delta applied to direct solar radiation. The classifier reads current wind direction, speed, pressure trend, hour, and temp; the regime label keys into a calibrated per-regime delta. L1–L4 are trained on general bias correction; Lsr is the first layer trained on synoptic-state stratification — different "kinds of weather" get different corrections. Skip regimes (v0.6.280): ne_flow and calm return 0.0 — l5_solar_analysis showed Lsr hurt sr in those regimes (+32% and +10% worse) while helping everywhere else. Skip regimes are a targeted patch until Fix B / refit against L4 baseline lands.
How Lsr works — Classify → Lookup → Apply
1. Classify.regime_classifier.classify_synoptic_regime(wind_dir, wind_speed, pressure_in, pressure_trend_hpa_3h, hour_local, temp) returns one of nine labels — nw_flow, ne_flow, sw_flow, se_flow, sea_breeze, frontal, pre_frontal, nor_easter, calm (plus unknown when inputs are missing).
2. Lookup. Per-(regime, hour) Δ W/m² from _BIAS_BY_REGIME_HOUR[regime][hour_local] in solar_correction.py; falls back to _BIAS_FALLBACK_BY_REGIME[regime] when the hour cell has too few samples.
3. Apply. Added to every lead's direct_radiation where the lead's raw value is above SUN_UP_THRESHOLD (50 W/m²). Pre-sunrise / fully overcast leads get Δ = 0 regardless of regime.
Applicability
Lsr is a specialist — it only applies to direct solar radiation (sr). Other fields have no Lsr line on their accuracy charts by construction. See Applicability map for the live trigger predicate (ENABLED in solar_correction.py + the sun-up gate + skip regimes ne_flow and calm since v0.6.280).
Live state — what Lsr is doing this tick
Loading…
Read directly via curl -s https://data.wymancove.com/weather_data.json | jq .solar_correction. The per-lead Δ is also baked into the sr card on the Forecast Accuracy chart above.
Engineering status
Where we are (2026-07-28):
Lsr lives in production for sr. Skip regimes: ne_flow + calm. Structural raw-baseline verifier (v0.6.291) catches raw-column mutation drift automatically. Unit-mismatch resolution — Lsb Stage 3 wired 07-17 v0.6.354, ENABLED=False. Overrides Lsr's direct-beam output with (shortwave − bias(hod)) on sea_breeze rows where cc < 25 — narrowed 2026-07-28 v0.6.383b from the original two-sided (cc < 25) OR (cc >= 75) after 07-24 halves re-run showed the overcast half (75-100 cc-bin) actively regressed (Δ=−17.6%, halves +6%/−23%) while the clear-sky half held robustly (Δ=+35.8%, halves +33%/+40%). The "thick attenuation missed" hypothesis for overcast was wrong; data says model attenuation is fine or over-corrected. Middle-and-overcast cc bins (≥25) now all fall back to Lsr's original direct-beam correction. Narrowed-shape Stage 2 re-run (07-28) PROMOTES cleanly: pooled +31.50%, halves +43.78% / +22.69% (both above +10% ship gate), 4/4 lead-bands SHIP. Landed ENABLED=False with fresh 7-day live-layer gate 07-28 → 08-04. Broader unit-mismatch across other regimes (pre_frontal +12, unknown +367 W/m²) still open. Shortwave shadow-log infrastructure stays alive (underpins future sr L2 work once unit mismatch resolves per project_sr_unit_mismatch). Divergence-report LSR_ENABLED source bug fixed 07-20 v0.6.365 (was reading l5_solar_analysis, now routes through .cache_l5_gate_history.json) — status now correctly AGREE, live gate 100% SHIP for 30+ days.
Full Lsr shipped-history log (v0.6.248 → v0.6.291)
Shipped 2026-06-28 v0.6.248 — ENABLED=True after the Lsr gate cleared 7/7 ship days (12-cycle SHIP streak).
Attribution fix v0.6.249 — initial ship was silently absorbing Lsr into the L4 column (same bug shape as the Lt-absorbed-into-L2 issue). Fixed by preserving direct_radiation_post_l4 before mutation and adding sr_l5 to the snapshot writer.
Chart wiring v0.6.250 — sr card now shows an amber Lsr line + column + "Lsr ✓ synoptic" badge. _layersFor() filters the Lsr entry to the sr card only.
Skip regimes v0.6.280 (2026-07-02) — ne_flow (+32.3% worse) and calm (+10.7% worse) now return 0.0 in compute_solar_correction. Per today's l5_solar_analysis: overall Lsr net −7.5% MAE across 8 regimes, but only 3 clear the +3% improvement threshold (sea_breeze, nw_flow, se_flow) — verdict HOLD gates further refinement, not the shipped layer.
Raw-array pollution fix v0.6.285 (2026-07-02 evening) — raw_direct_radiation was being captured AFTER Lsr mutated direct_radiation, so the "raw" baseline in the debug page + Production accumulator was itself Lsr-corrected on all post-06-28 pair rows. Fixed by extracting raw preservation into preserve_raw_forecast_arrays() and calling it BEFORE stamp_solar_correction. Structural guard added v0.6.291 (2026-07-04):weather_collector/processors/raw_integrity.py deep-copies every raw-counterpart hourly array at pipeline start and asserts equality at pipeline end. Any future layer that mutates a source array before its raw_* copy exists will now fire a drift event on the next tick and land in the daily digest, instead of hiding for a week.
Per-lead delta fix 2026-07-03 — stamp_solar_correction was computing a single delta from the CURRENT tick's raw solar + hour, then applying that scalar to all 48 forecast leads. At pre-dawn / dawn ticks the current raw was below the 50 W/m² sun-up threshold, delta = 0, no Lsr applied to any lead. Fixed by iterating over the forecast array with each lead's own raw value + parsed local hour. Lsr now fires at every daytime lead in every non-skip regime.
Audit status: divergence-report row says AGREE (production matches script verdict). L5_VALID_FROM = 2026-06-28T07:05, but the two collector bugs above mean the pair-log record of Lsr in the wild before 07-03 is only a sliver of what Lsr will do post-fix. First actually clean 7-day audit window closes ~2026-07-10. Today's simulate_windows HOLD verdict + today's l5_solar_analysis HOLD verdict both read against the corrupted window — treat as informational until the 07-10 re-run.
Regime classifier.regime_classifier.classify_synoptic_regime(...) returns one of nine labels; same source feeds C1a's transition axis. Lsr reads via state.regime_synoptic. Live label surfaces in Production Stack box, R2.
Per-regime delta table — what the lookup contains. Lsr's bias tables live in weather_collector/processors/solar_correction.py as two Python dicts:
_BIAS_FALLBACK_BY_REGIME — single Δ W/m² per regime; used when the per-hour cell has too few samples. Latest fit values: frontal −81.1, sw_flow −79.3, pre_frontal −107.3, sea_breeze −44.6, nw_flow −40.8, calm −57.2, se_flow −114.9, ne_flow −206.3, unknown 0.0. Sign convention: positive = forecast over-predicts; correction = −bias.
_BIAS_BY_REGIME_HOUR — nested {regime: {hour_local: Δ}}; primary source per tick. Diagnoses the within-regime hour-of-day swing that the regime-only first pass missed (e.g. ne_flow goes from −238 W/m² at 10:00 to +247 W/m² at 14:00).
Refit cadence: regenerate via python3 analysis/l5_recompute_biases_hourly.py after at least 7 days of accumulated DAYTIME pair rows (raw_solar ≥ 50 W/m²). The daily digest prints a drop-in replacement table; mid-trajectory refits invalidate the audit window — wait for a clean break.
Lsr vs L4 audit — held-out MAE.Primary view: the Forecast Accuracy chart above. The sr card carries five lines (Raw → L2 → L3 → L4 → Lsr); the green box on the Lsr line is the user-visible MAE on rows where Lsr fired. Note: on rows where Lsr's skip regimes fire (ne_flow, calm since v0.6.280) Lsr returns 0.0, so those rows contribute L4-value MAE to the Lsr aggregate — the Lsr line converges toward L4 as skip-regime coverage grows.
Fitter audit: verdict logged to conditional_audits.l5 each Fitter cycle, recency-weighted since v0.6.178 (exp(−age_days/14)). The trailing 7-day rolling gate lives in l5_gate_history.json and surfaces beneath the Lsr row in the S1 Shadow Tuner section below. First actually clean 7-day window closes ~2026-07-10 (7 days from the 07-03 per-lead delta deploy — the earlier ~07-05 date was against the pre-fix deploy timeline).
RESOLVED 07-20 v0.6.365: the divergence report's LSR_ENABLED claim was sourced from l5_solar_analysis — a candidate script testing a simpler regime-only bias lookup, NOT the live hourly Lsr. Its HOLD verdict meant "don't ship the candidate refinement," but the divergence table was rendering it as "retire live Lsr." Live Fitter cycle gate history (l5_gate_history.json) shows 100% SHIP across 30+ days, 13-21% MAE improvement per cycle. Divergence report now routes LSR_ENABLED through the live gate history and correctly shows AGREE.
A per-(field, value_bin) percentage-point shift applied to the L4-corrected forecast for cc / cl / cm / ch. The bias pattern Lc unwinds is saturation: the models systematically overshoot at 0-5% cloud and undershoot at high fractions. The same shape holds across all four cloud fields but with different magnitudes. Lc reads each L4-corrected value, keys into a curated (field, value_bin) → shift table where the fitter's SHIP verdict is set, adds the shift, and clamps to [0, 100]. Emergency intervention 2026-07-30 v0.6.389d-g — pooled shift table diverged from recent regime-conditional truth, causing 8× cl MAE blow-up on 07-30. Live SHIP surface after intervention: cl fully off (_FIELD_SKIP); cc ships at 0-5 (all regimes), 50-80 + 80-95 (all regimes except ne_flow), 95-100 KILLED universally; cm unchanged (20-50 / 50-80 / 80-95 / 95-100, all regimes); ch unchanged (20-50 / 50-80 / 80-95, all regimes). SKIP cells (bins that stayed within the ±5 pp noise floor across the 30-day window) pass through unchanged. See Engineering status below for the walk-forward evidence that motivated the interventions.
How Lc works — Preserve → Lookup → Shift → Clamp
1. Preserve. Before mutation, stash the L4-corrected array as hourly.<field>_post_l4. Pair log + debug page read this to attribute L4 vs (L4+Lc) cleanly.
2. Lookup. For each lead 0-47, determine the value bin from [0, 5, 20, 50, 80, 95, 100]. If the (field, bin) cell has SHIP verdict in lc_correction_table.json, read its Δ pp; else Δ = 0 (SKIP — the L4 value passes through unchanged).
3. Apply. Add Δ to the value and clamp to [0, 100]. Writes back to hourly.<field>.
Sign convention: the fitter's stored bias is (forecast − observed); the applied shift is −bias, pulling the forecast toward the observation.
Applicability
Lc is a specialist — applies only to cc / cl / cm / ch. Fires when the L4-corrected forecast falls in a SHIP-verdict value bin (16 of 24 cells; see Applicability map for the full per-cell shift table + verdict list).
Live state — what Lc is doing this tick
Loading…
Read directly via curl -s https://data.wymancove.com/weather_data.json | jq .cloud_saturation_correction. The per-lead Δ is also baked into cc / cl / cm / ch cards on the Forecast Accuracy chart above.
Engineering status
Where we are (2026-07-30 — emergency intervention day):
Overall Prod-vs-Raw compressed from ~−13% a week ago to −5.2% today, driven by cl and cc l6 MAE blowing up. On 2026-07-30: raw cl MAE 7.29 (model + obs both mostly-clear) but Lc-corrected cl MAE 56.96 (8× worse). Same on cc (raw 7.21 → l6 44.23). Diagnosed via 4 new analysis scripts (Stage 0 regime × bin sweep, Stage 1 halves-strict fit, walk-forward validator, rolling-window sweep) as an architectural failure of the shift-table itself: the historical fit table subtracts 46-88pp at overcast bins because model USED to over-forecast overcast heavily, but the model no longer does — bias shrunk 4-38× or sign-flipped in the last 3-10 days. No rolling window length recovers cl on held-out (best W=3d still −3.7% vs raw). Regime-conditional slicing doesn't fix it either (walk-forward: cl reg_vs_raw −30.4%). Emergency bandages shipped v0.6.389d-g: cl fully off Lc, cc/95-100 universal bin-skip, cc/ne_flow/50-80 + cc/ne_flow/80-95 regime-conditional demote. cm and ch untouched (walk-forward +34% and +50% vs raw on held-out — still helping). Original 14-day post-ship watch (07-17 → 07-31) cannot close routinely. New tracking axis is the regime-conditional Lc Stage 1 gate (7-day walk-forward stability starting today) — but even that's superseded by the walk-forward validator's finding that the naive regime-conditional shape isn't the fix. Architectural next step (unshipped): EMA/Kalman shift tracker OR recent-bias gate on the existing lc_fit table. Both multi-day. Contingency: per-bandage reversibility is a single-line edit to _FIELD_SKIP or _CELL_SKIP frozenset in cloud_saturation_correction.py. See [[project_lc_regime_conditional]] for the full pipeline state.
Prior state (2026-07-17):
FLIPPED 2026-07-17 v0.6.355 after 8/7-day gate clear, 16 SHIP cells stable 7 consecutive days (07-11 → 07-17), LC_ENABLED READY on divergence report, and no cc/cl/cm/ch ANOMALY in the pair-log anomaly detector. First live tick (15:27 EDT): 113 cells fired — cc 46/48, ch 39/48, cm 19/48, cl 9/48; mean |Δ| in the expected 28-42 pp range per field. Predicted biggest Prod MAE lifts: cl 80-95 −55%, cl 95-100 −47%, ch 50-80 −37%. Two-gates-per-layer cross-check via [[feedback_two_gates_per_layer]]. 07-31 clean close superseded by the 07-30 intervention above.
Full Lc shipped-history log (v0.6.298 → v0.6.355)
Code shipped 2026-07-04 v0.6.298 — weather_collector/processors/cloud_saturation_correction.py at ENABLED=False. Fit script analysis/lc_fit.py emits lc_correction_table.json with per-(field, value_bin) shift + SHIP/SKIP verdict + n_samples.
Two-gates-per-layer codified 2026-07-15 v0.6.352a — divergence-report streak counter is not the same gate as lc_fit's own SHIP-set stability gate. Cross-check both before flipping. Applied here: divergence gate cleared 07-10 (7/7) while lc_fit gate remained False through 07-16 because cm 20-50 flipped into the SHIP set 07-11 (inside the 7-day stability window).
ENABLED=True 2026-07-17 v0.6.355 — 8/7-day divergence-report clear + lc_fit gate True (07-10 rolled out of the 7-day window today) + 16 SHIP cells identical 7 consecutive days (07-11 → 07-17) + no cc/cl/cm/ch ANOMALY. Deploy sequence: line 29 flip → make deploy-collector → :07 tick verified → GCS confirmed → frontend push. 14-day post-ship watch begins.
Developer notes — fit table, refit cadence
Fit table location.weather_collector/data/lc_correction_table.json. Structure: {field: {value_bin: {shift, verdict, n_samples, mae_pre, mae_post, delta_pct}}}. Loaded at import time in cloud_saturation_correction.py via _load_correction_table().
Refit cadence.analysis/lc_fit.py runs in the daily digest and rewrites the JSON in-place. SHIP-set stability (7 consecutive days with the same SHIP cells + no HOLD days) is the fitter's own gate; it lives in .cache_lc_gate_history.json.
Value-bin edges.[0, 5, 20, 50, 80, 95, 100] percentage-points — six bins per field × four fields = 24 cells. Bin 5-20 is currently SKIP for all four fields (mid-low cloud has near-zero systematic bias).
Research & Diagnostics — experimental signals + audit views (not applied to live forecast)
Diagnostics — audit live behavior
R0. Per-layer audit — is each layer + specialist earning its keep?
Held-out MAE per field per layer (leads 1–47), refit every Fitter cycle. Dim subtext = signed bias. Δ columns compare to layer below: green = beats + applied; amber = beats but NOT applied; red = loses; gray = tied. Banners fire when an enabled layer or specialist loses >3% (hidden regression) or a disabled specialist / layer wins >3% (missed opportunity / post-ship watch signal). Specialists column (v0.6.382n) chains Lsr → Lc → chp / clp / wdp for the fields each owns; each rung compares to the previous rung (or L4 if first). Production column shows the real per-row Production MAE (aggregated from applied_layer stamps). v0.6.456 (F3-A):applied_layer now stamps NBM layers (l3_nbm/l4_nbm/l5_nbm/l6_nbm/chp_nbm/wdp_nbm) on rows where the selector picked NBM — before v0.6.456 those rows carried the HRRR walker's stamp and Prod attribution silently misclassified. Historical rows carry the old stamp; only post-08-21 rows carry NBM-aware attribution. MAE-only caveat: RMSE tells a different story on L2-additive fields (wg Production improvement drops −33%→−26%, dp −17%→−13%, h −7%→−3%). Per-layer RMSE+bias column rewrite queued.
Forecast accuracy — Raw vs Production + per-field lead-band breakdown
Two lenses on the same question. First (Accuracy over time, below): the per-obs-day trajectory of Raw vs Production — is the stack drifting? Did a recent ship move the needle? Second (per-field tables, further down): the current-window breakdown by lead band per field — where in the 0-47h horizon is each correction layer helping or hurting? The over-time view is the drift detector; the per-band tables are the shipping-decision granularity. Production in both views IS a real per-row aggregate. Per-band tables: per_layer_mae_by_lead[field].production where the per-lead sample count clears n≥30, keyed on applied_layer stamps. Over-time chart: mae_over_time.json's "prod_real" series, bucketed per obs-day from the same applied_layer stamps (v0.6.371). Both surface the same "what users actually saw" quantity, just at different aggregation grains. Specialist lines (Lsr, Lc, chp, clp) remain visible on the over-time chart as intermediate series showing gate-fired-only MAE — smaller sample than Raw/L2/L3/Prod, useful for isolating each specialist's per-lead lift during 14-day watches.
How we measure whether the forecast is good — the metric framework
Core comparison: pipeline error vs. raw L1 error on the same forecast-observation pairs. L1 = HRRR-side (Open-Meteo HRRR/GFS) baseline. NBM cascade runs in parallel; the 🎯 L1 selector (Prod-vs-Prod as of v0.6.440, armed via backstamp v0.6.445) picks per (field, lead-band) which cascade's output the user sees. Every metric on this page averages those errors differently.
"Observed" per field (best local measurement, not absolute truth):
cc, cl, cm, ch: mean of KBOS + KBVY METAR sky reports (octa → percent). Coarser than mesonet.
sr: median across valid Tempest solar_radiation_wm2.
pa: max WU precip_rate_in across stations (patchy-rain-aware).
pp: binary — any measurable rain this hour (0 or 100).
wd: station wind_direction, circular error.
Three metrics per NWS/ECMWF conventions:
MAE — mean |error|. "Typical-day miss." Historical page text usually means this.
RMSE — root mean squared error. Weights big misses harder; if RMSE improvement lags MAE, pipeline has hidden blow-ups.
Bias — signed mean error. Positive = over-forecast; negative = under. Catches systematic drift MAE masks.
pp uses Brier score, not MAE (binary obs vs. probability forecast). Brier = mean of (fc/100 − obs/100)². Same "% erased" framing.
Measurement gaps:
Persistence skill — shipped 2026-07-11 (h_persistence_skill.py). Baseline: 5 ADD VALUE (t/h/pr/ws/sr), 4 MIXED (dp/wg/cc/pp), 3 NO SKILL — cl/cm/ch lose to persistence at every band despite L3+L4. Scorecard integration shipped 07-12 v0.6.328 — the "vs Persistence" line in the scorecard banner (above) is live-updated from persistence_skill.json. Response for ch: ch persistence gate FLIPPED LIVE 07-19 v0.6.358 (14 SHIP cells post-emergency-demote 07-27 v0.6.382t (was 27) on refreshed windows; 14-day watch CLOSED CLEAN 08-02). Landmark answered 07-14 v0.6.351a: persist_only ties gate pooled but gate wins halves-hedge; keep the gate. cl persistence gate wired Stage 3 07-24 v0.6.379 (replaces retired cl_persistence_short_lead — narrow 0-5h hypothesis disproven by halves-verified Stage 2). 08-09 Stage 2 rerun — 6 SHIP / 22 SKIP whitelist shape ready for flip decision. SHIP cells: calm/0-5 (−35.7%), ne_flow/0-5 (−54.4%), ne_flow/12-23 (−16.6%), nw_flow/0-5 (−18.7%), pre_frontal/0-5 (−17.8%), sea_breeze/0-5 (−32.9%). Flip = cl_persistence_gate.ENABLED=True. 08-09 c1_stage4_mixture_check: cl/6-11h [transition] b3 no longer in DEGRADED list (07-29 escalation window CLOSED CLEAN, never reached 3× ≥+200%); cm/12-23h [transition] IMPROVED; t/24-47h [transition] newly DEGRADED (unrelated). Joint cl/cm correction hypothesis rejected 07-29 (cl has no correction stack; cm corrections are stable improvers; drift is raw-HRRR-side on top-fc transition rows and fitter absorbs via c1 confidence bands).
Climatology skill — long-lead reference. Needs a climatology dataset.
pp Brier reliability decomposition — does "30% chance" actually happen 30% of the time? Aggregate Brier only, no decomposition.
Historical wording caveat: section text written pre-v0.6.325 (2026-07-10) treats MAE as the only measure. Read alongside RMSE + bias — disagreements usually mean occasional big misses or systematic drift.
Accuracy over time — Raw vs Production, per-day trajectory
Per-obs-day rollup. Chart shows the trajectory of Raw (Open-Meteo L1 seed — HRRR/GFS) and every correction layer applicable to the selected field (for t/ws/wd, "L1r (router)" line shows the v0.6.432 NBM router output at leads ≥6h) — legend adapts per field so noise-lines that don't do work for that field are hidden (e.g. sr shows Raw / Lsr; pa shows Raw / Prod only). Rolling 7-day mean overlays on Raw and the final applied layer smooth out day-to-day noise. Ship-date vertical annotations mark when live-layer changes landed. Source: mae_over_time.json — accumulating history, x-axis grows one day at a time (retention-independent; the pair log is capped at 30 days but this rollup persists prior days). Refreshed hourly by the myweather-publisher Cloud Function (v0.6.395f, hourly cron); prior to that it was tied to the daily digest. (v0.6.361: per-specialist attribution — Lsr, Lc, and post-Lc specialists (ch-persist, cl-persist) each plotted separately once ≥3 days of data have accumulated. New collector-side pair-log columns start accruing today; specialist lines will appear on the chart around 2026-07-22. v0.6.371: the 'Prod' line is now the real per-row aggregate keyed on applied_layer stamps — sample-comparable to Raw/L2/L3 on the same chart. Pre-v0.6.371 it was L4's daily MAE for non-specialist fields and a specialist line for cc/cl/cm/ch/sr (gate-fired-only, sample-mismatched).)
Detail view for selected field (Raw / L2 / L3 / Prod + rolling means + ship annotations):
Scan all fields (click any panel to focus the detail chart above):
📊 Per-field breakdown by lead band — MAE / RMSE / bias per correction layer
The bright white "Production" column is the forecast users actually see. Lower is better. For each field, one combined table shows how far off the forecast tends to be at each lead band — MAE (typical error), RMSE (occasional big misses that MAE averages away), bias (systematic drift, signed) — per correction layer. Compare a MAE row across columns to see where each layer helps or hurts.
How to read these tables
What you're looking at. One card per field. Each card shows a table with rows grouped by lead band (0-5h, 6-11h, 12-23h, 24-47h, ALL). Under each band, three metric rows: MAE (primary), RMSE (secondary), bias (secondary, signed). Columns are the correction layers actually applied to that field, plus the Production output on the right.
Only-applied layers. Layers that aren't applied to a field don't render as columns because their per-lead MAE array equals the previous applied layer's — noise. The badges above the table (L2 ✓, L3 ✓, L4 off, etc.) confirm the applied set. If Production isn't the lowest MAE, check the Applicability map for a gate that's wrong.
Per-metric usage:
MAE — average absolute error. The primary metric; drives L3/L4 whitelist decisions. Best cell in row highlighted green, worst red.
RMSE — squared-error root. Higher than MAE by a factor set by how tail-heavy the errors are. Watch for a band where RMSE jumps proportionally more than MAE — that's occasional big misses hiding behind an OK average.
bias — mean signed error (forecast − obs). Near-zero = calibrated on average. Persistent + or − by band signals systematic drift.
Layer labels:
Raw: bare Open-Meteo HRRR/GFS forecast for this coordinate. Knows nothing about Wyman Cove. (For t/ws/wd at leads ≥6h, the user-facing L1 is the NBM router output — see the L1r line, added v0.6.432.)
Aggregate bias (L2): what 40+ nearby weather stations say the model is currently getting wrong, blended in by distance.
Lead decay (L3): historical per-lead bias correction from the pair log.
Diurnal (L4): hour-of-day bias correction. Final line for t / dp / h / ws / wg / pp.
Synoptic-regime (Lsr): per-regime W/m² delta on direct solar radiation. Only present on sr.
Cloud saturation (Lc): per-(field, value_bin) percentage-point shift, clamped to [0, 100]. Final line for cm / ch (unchanged); cl and cc both fully off as of 07-30 (v0.6.389f cl · v0.6.390 cc). cc now derived downstream via Ccd = max(cl_l6, cm_l6, ch_l6). FLIPPED 2026-07-17 v0.6.355; emergency intervention 07-30 — see Lc layer section for details.
L2 badge variants:L2 ✓ additive = bias added (t, dp, h, pr). L2 ✓ direct = station median replaces model wind (ws, wg). L2 n/a = no station network reports this field (clouds, solar, precip). Note for clouds: L2 is n/a as a forecast correction, but obs truth for cc/cl/cm/ch comes from a KBOS + KBVY METAR blend (v0.6.134) — feeds the joiner so L3/L4 can be evaluated, not applied at L2.
Bands match the walk-forward validator buckets — this is exactly the granularity that drives shipping decisions. Source: time_series_diagnostic.json::per_layer_{mae,rmse,bias}_by_lead, 7-day window. Design note (v0.6.350): previously carried a per-card chart + 3 separate tables. The chart mostly restated what the MAE band table already showed; killed for signal density.
Live readout of detected frontal passages from frontal_events_log.json. Detector runs every tick in frontal_detection.py: passage fires when ≥2 of 3 signals hit (dp drop ≥4°F over 60min — lowered from 8°F on 2026-09-14 v0.6.620 after frontal_detector_health.py found observed 60-min dp drop max=6.8°F, p99.9=6.3°F; original 8°F threshold was unreachable, wd shift >60°, pressure inflection with ≥0.02″ rise). Confidence 67% with 2 signals, 100% with 3. Diagnostic print(" frontal: score=…") log line fires whenever score≥1 (2026-09-14 v0.6.620/v0.6.621) — traceable via gcloud functions logs read for miss-rate diagnosis. Health-check verdict runs in daily digest as analysis/frontal_detector_health.py. Live consumers: PWA front-passage card + C1e confidence axis + 10 analysis scripts joining on passage timestamps.
Loading detected passages...
R2. State-stratified accuracy — which regimes does the model fail in?
Per-field MAE + bias by regime. Big MAE spread across bins = regime-aware correction candidate. Addressed rows excluded from top-10 (shipped corrections absorb raw signal downstream); excluded rows in collapsible below. Temperature back in the top-10 since Lt retired 07-13. Refit twice daily; published to state_stratified_accuracy.json. MAE-only; RMSE would rerank L2-additive fields (dp/h/ws/wg).
Tools — evaluate candidates before promotion
S1. Shadow applicability tuner — what would auto-tuner have chosen?
Per Fitter cycle, log which fields a naive MAE auto-tuner would include in L3_FIELDS/L4_FIELDS, alongside production. Recommend ON if layer beats layer-below by ≥3% in any band AND bias no worse. Field-membership only — doesn't reason about per-field gates or skip tables (those in Applicability map). Precondition for automation is agreement after 90+ days; mismatches informative, not actionable. Also surfaces the live conditional_audits.r6 C1a-transition verdict per Fitter cycle.
Loading…
B1. Backtest sweep — alternative L3/L4 configs vs production
A/B comparison of candidate L3/L4 applicability configs vs production, computed by replaying the held-out pair log. Current production: L3 = {ws, wg, ch, cm, pp}, L4 = {ch, cc}. Run via python3 -m backtest.sweep --write-gcs (add --local-file ~/.cache/myweather/forecast_error_log.jsonl for fast iteration).
7 axes stamping every tick, applied=False until Stage 4 clears. Curated v3: 296 SHIP / 42 MARGINAL / 1048 SKIP across 39 axis-keys. C1a shipped as regime-transition axis (was R6; verdict logged under conditional_audits.r6 per Fitter cycle, surfaced via S1). Latest audit: HOLD 65.00% (13 CALIBRATED / 7 DRIFTED, 2 BRIER_EXEMPT — re-cured 08-15 from 57.14%). Prior cl "DEGRADED" reading was driven by the applied_layer poison fixed 08-02; next re-audit gated on window roll past cl-poisoned days (08-04+). Mixture-normalized view: only 2/112 cells are REAL DRIFT (per c1_stage4_difficulty_lens); legacy FAIL count is inflated by weather-mixture shift. Narrow-promote counters (today): pre-frontal 6/7 (2 SHIP, 1 to go). C1h + C1d auto-suppressed as KNOWN_LIVE_PIPELINES (already live-stamping). Ship gated on Stage 4. Per-cell co-axis ortho gate v0.6.321 in confidence_layer.py: cl fires freely, cc/cm/ch conditionally suppressed by the non-ortho co-axis, ch 24-47h + t × 3 never fire (REDUND both). Individual axes tracked in Backlog Group A below.
Stage 0-2 hypotheses (pre-wire). Promotion path: 0 exploration → 1 curated finding → 2 per-cell verification → 3 wired ENABLED=False → 4 shipped. Stage 3+ detail lives in Post-ship watches / layer sections, not here. Most surviving hypotheses measure forecast uncertainty (C1 axes), not forecast bias. Real bias candidates (marine layer, radiational cooling) overlap heavily with L2 and L4.
Group A — C1 multi-axis confidence extension
Individual axes. Five join the multi-axis Stage 4 calibration (C1a/b/c/f/e); C1h + C1d compose as marginal-premium tables (kept off the join to avoid cell-dilution). Stack-level status in C1 confidence stack card above.
C1e Hours-since-front — SHIPPED 2026-07-01 v0.6.272 (post <24h vs baseline ≥24h from frontal_events_log.json).
C1h Forecast trend-direction — Stage 3 wired 2026-07-08 v0.6.316, per-cell co-axis ortho gate wired 2026-07-10 v0.6.321. 15 SHIP cells; cl × 3 fire freely, cc/cm/ch conditionally suppressed, ch 24-47h + t × 3 never fire. Already live-stamping; narrow-promote counter auto-suppressed as KNOWN_LIVE_PIPELINES. User-visible band-widening ship gated on C1 Stage 4.
Group B — Bias candidates (paced, individually)
Cloud-ceiling regime correction. Dormant since 06-30 — per-(field, regime, lead_band) whitelist work supersedes the simpler ceiling slice for now. Revisit when capacity opens; may come out as a bias or as a confidence axis.
Marine-layer / harbor inversion. Stage 3 sandbox live 2026-06-24 v0.6.199, ENABLED=False. Diagnosis 2026-07-16: hold indefinitely — real break was 06-30 (seasonal marine-layer weakening), stratum-local, not HRRR-anomaly-coincident. Redesign candidate: time-of-year gating. See Archive for full history.
Group C — Lower priority (dominated by existing layers)
Wind-direction sector correction for gusts. L3 already does heavy lifting on ws/wg. Bar is "does directional structure survive L3?" — high bar. Revisit if Group A/B come up empty.
Sea-breeze onset/decay timing. Overlaps marine layer (onshore-flow hours) and L4 diurnal (timing). Folds into marine layer if that pilots out.
In-flight candidates + current stage counters live in Current state → What's improving at the top of the page (single source of truth).
Group D — Methodological refinements (modify existing layers, not new ones)
Joiner per-layer wd errors SHIPPED 2026-07-20 v0.6.367. Fixed silent wd branch that had skipped the per-layer loop; wd now first-class across every measurement surface.
Humidity K-taper SHIPPED 2026-06-24 v0.6.218. soft_ramp K curve for L2 humidity bias.
Cloud cover → L4 SHIPPED 2026-06-24 v0.6.214. L4_FIELDS = {ch} — cc dropped 2026-08-28 v0.6.515 after 7-day drop-gate cleared (cc is Ccd-overwritten downstream; L4 cc was code-hygiene deadweight).
Cloud saturation-unbiasing (Lc) SHIPPED 2026-07-04 v0.6.298 · FLIPPED 2026-07-17 v0.6.355. Per-(field, value_bin) shift after L4. See Lc section.
Regime-conditional dp depression correction — Stage 1 open. Frontal branch closed 07-09; nor_easter watch dormant; see Current state for latest.
Promotion rule: single-shot script in analysis/. Stage 2 verdict must hold across ≥2 reads spaced 3+ days apart. Group A → C1 axes (C1a/b/c/...); Group B → R-numbers if they clear the 7-window walk-forward gate.
Experiments — open design seeds + live telemetry probes
Only genuinely-open items live here. Promoted (→ Stage 1+), killed by orthogonality, and settled-null items live in Archive — single source of truth.
Asymmetric L3 (over- vs under-call). Wind L3 +57% on over-calls, −120 to −166% on under-calls. Real asymmetry but not directly actionable (can't predict which side a future pair falls on). Design seed for a future "L3-with-confidence-gate" hypothesis. h_asymmetric_l3.py
Run-time issuance bias. HRRR init hour shows clean sinusoidal pattern on humidity L1 MAE (8.0 → 10.5, 33% spread) but potentially confounded with valid-time-of-day. Needs a controlled follow-up holding valid_time fixed. h_run_time_bias.py
Front-type asymmetry (cold vs warm).PARTIALLY UNBLOCKED 2026-09-14 v0.6.620. Original detector had 0 type='cold' classifications since deployment because _classify_type's cold branch required dp_drop ≥ 8.0°F — threshold sat above observed p99.9 (6.3°F, max 6.8°F). DP_DROP_THRESHOLD lowered 8.0→4.0°F; cold branch now reachable. Warm-front separation still requires the reverse dp signature (dp rising) which no signal captures yet. First live type='cold' tag expected within a few days. Re-test h_front_type.py once the events log accumulates a mix of cold/sea_breeze/unknown.
precip_obs > 0 as obs-keyed confidence mirror. Fc-only false alarms (n=2,904): cl +283%, wg +57%. Architecturally tricky — we have precip_obs[now], not [future]. Different mechanism than C1f. Design seed. h_precip_obs.py
Cloud composition (single vs layered). wg +65% on 3-layer skies, ws +25% on 2-layer. Real but small; ch +106% partly tautological (composition includes the field). Weak design seed. h_cloud_composition.py
wind speed bin × wind direction error.SCRIPT BUG. Join between wd-row and ws-row returned empty. Per-pair key structure may not align across fields. h_ws_wd_error.py
Lightning proximity × MAE.INFRASTRUCTURE GAP. Pair log doesn't carry lightning data (Tempest strike fields live in station_history.json, not forecast_error_log.jsonl). Re-test after adding lightning_proximity_km to pair-log rows. h_lightning_proximity.py
Status: active telemetry, correction dormant. The processor code runs each tick (compute_cove_correction() returns 0.0 in both branches; telemetry stamps for retro-check). cove_gradient_log.json continues to append per-tick waterfront-vs-inland gradient data on GCS (14-day rolling window) — cheap, kept alive in case seasonal shift ever changes the microclimate signal and a new refit becomes worth trying. The Lt row in the divergence report now reads from l6_fix_b_refit (07-14 v0.6.351b), which reports HOLD and matches production, so the row stays AGREE forever unless a future refit crosses the +1% gate.
Original design. A small Δ°F added to the temperature forecast based on how much the waterfront stations (Willow Rd, Neptune Rd) typically diverge from the inland-network median under different wind / sea-breeze / hour combinations. Lt was the first correction trained on a spatial differential between station subgroups (L1–Lsr are trained on forecast-vs-aggregated-obs errors).
Shipped 2026-06-26 v0.6.238. Cleared a 2-read confirmation gate on r5_cove_analysis.py (06-25 SHIP + 06-26 SHIP). Per-lead projection fix 2026-06-26 v0.6.237 (initial ship applied current-tick Δ to all 48 leads — wrong by 3–5°F at distant leads when the table crossed zero). Cooling branch disabled 2026-06-30 v0.6.259 (paired-MAE on 19,975 t rows: cooling branch made cooling rows −74.9% worse). Warming branch disabled 2026-07-01 v0.6.276 after per-row Production data (v0.6.269 applied-layer stamping + v0.6.275 backfill on 47,301 T pair rows) showed warming-branch rows ran ~40% worse than L2 on the same rows. Hypothesis at that time: double-counting — L2's Kalman blend for T is dominated by waterfront Tempests (Willow Rd, Neptune Rd at ~0.1–0.2 mi from the cove) — L2 already carries "waterfront bias" by station weighting; adding a (waterfront − inland) Δ on top re-adds the same signal.
Fix B — tried 2026-07-13 v0.6.329, failed the +1% gate.analysis/l6_fix_b_refit.py refit both lookup tables against (cove obs − L2 forecast at cove) instead of the raw (waterfront − inland) gradient — the design the section header used to promise. Ran on 202,321 pair rows (154k train / 47,757 held-out). Panel B (sb_off × hour) looked like a real overnight cove-cooling signal on training: 00-06h +0.9 to +2.1°F, 7 SHIP bins. But held-out MAE improvement was +0.29% vs the +1.0% ship gate. Panel A (sb_on × octant) came back with 0 SHIP bins after refit — the warming-branch signal was largely a fitting-against-raw-L1 artifact. Mechanism confirmed: L2's Kalman blend re-fits per-tick based on obs-vs-model bias and absorbs the same microclimate signal dynamically. A static hourly cove table adds a delta L2 already added → wash on held-out.
Reactivation criterion. Not queued for any specific date. Would require the held-out MAE on a re-run of l6_fix_b_refit.py against fresh pair-log data to clear +1.0% for a sustained window (3+ reads). Possible triggers: seasonal shift (fall/winter offshore vs summer sea-breeze patterns), sensor swap changing L2 station weighting, or a materially different mechanism than the one that failed. Until then, Lt telemetry runs dry.
Things we built, ran, and stopped running. Two kinds live here: hypotheses the data answered no, and settled tunings where the sweep concluded "current value is fine." Kept as institutional memory. Scripts in analysis/ can be re-run if conditions shift. Retired layers (code intact but permanently no-op) stay in R&D — see Lt above.
wd L2 blend (v0.6.368a 07-20 → v0.6.384 07-28) — CLOSED CLEAN 2026-08-11. First circular field in L2 (sin/cos unit-vector mean, per-cell calm-floor 3 mph). Watch reset to 08-11 after v0.6.384 shrank BLEND_HOURS 24→4. Post-fix (obs_time ≥ 07-28) production: 0-5h L2 35.44 vs L1 46.16 (−23% MAE, 65% fire, fired-subset −36%); 6-11h L2 48.94 vs L1 48.79 (+0.3%, 6% fire); 12-23h L2 52.65 vs L1 52.82 (−0.3%, 8% fire); 24-47h never fires (band ≥ BLEND_HOURS). Predicted MIXED → ADDS VALUE flip landed on schedule as pre-fix fossil damage aged out. Design intent delivered.
C1h narrow-promote Stage 3 (v0.6.393, 08-05) — SUPERSEDED 2026-08-08 by v0.6.396 re-curate. Original ship: 9 SHIP cells cc/cl/cm/ch 12-23h + 24-47h + t/12-23h. Diff vs prior committed: cc/6-11h, ch/6-11h, cm/6-11h demoted SHIP WIDEN → SKIP (sample floor); t/12-23h flipped MARGINAL WIDEN → SHIP NARROW. Co-axis gate refreshed: t/12-23h + ch/24-47h from stale always_skip → require_c1f_off + require_c1e_off. First post-deploy tick fired 3 cells (cc/12-23h 1.83× widen, cl/24-47h 2.32×, ch/12-23h 3.78×) and coax_gated 2 correctly. Superseded when the 08-08 re-cut refreshed t band selection.
ws L3 SKIP_TABLE (v0.6.386, 07-28) — INERT 2026-08-08 (ws dropped from L3_FIELDS in v0.6.397). Cells wired at ship: frontal 24-47 (+5.8%, n=987), pre_frontal 24-47 (+8.0%, n=9,268). Also existing ne_flow 0-48 + sea_breeze 0-11. Full skip-table + asymmetric additive (v0.6.370) now dead code because ws is no longer in L3_FIELDS. Kept in decay_apply.py for historical documentation.
h L2 shape re-tune (v0.6.390g, 07-31) — CLOSED CLEAN 2026-08-07. H_SOFT_RAMP_FLOOR 0.4→0.1, H_SOFT_RAMP_END 24→10 after L2 station_bias flipped helping → hurting 07-25. Grid sweep: 7d best {0.0, 8}, 14d best {0.4, 12} — middle path shipped +2.8% vs raw 7d + +6.0% 14d. Layer-shape sentry green all bands, h pair-log ΔMAE −28.3%. Reversible one-line edit. See [[project_h_dp_tau_refit]].
Ccd 7-day live-shadow gate — CLEARED EARLY 07-30 v0.6.390. Rerun post-cl-field-kill: +8.48% pooled MAE, halves +11.4%/+5.8% both positive, 6/10 regimes win. Flipped ENABLED=True; cc _CELL_SKIP retired same commit. Verified live at 22:19 UTC: cells_fired=44, cc[0:3] = [81,100,19] matches max(cl,cm,ch), cloud_cover_pirate_raw preserved. Forward triggers: cc-derived MAE creeping worse than cm+ch-corrected trajectory; SKIP_REGIMES {se_flow, unknown} coverage % changing. 08-17 — first trigger FIRED. h_cc_derivation re-run shows random-overlap now beats max by +8.13% pooled, wins 9/9 regimes (halves A +0.86% / B +16.55%). Response: opened [[project_cc_combine_walker]] rather than hand-editing FORMULA — same anti-scar-tissue pattern as chp/Lc gates.
cc composition tuner (v0.6.390e, 07-31) — analysis-only, HOLD.h_cc_blend_formula.py Stage 0 grid on 123K quads: max wins in 8/10 regimes; frontal +4.3% for random but halves-unstable (A=−2%, B=+8%); nor_easter +30.9% but one-sided (A=8, B=233). Only clean per-cell candidate: pre_frontal/0-5h random +3.5% n=2,115. Ccd stays max-only. Re-check when frontal + nor_easter data accumulates. See [[project_cc_blend_tuner]].
4-verdict regression sentry (v0.6.390f, 07-31) — permanent infrastructure. Pair-window classifier (sustained 7d + fresh 3d) replaced flat 7d sentry that lagged interventions by 7 days. Emits SUSTAINED FIRE / HEALING / FRESH FIRE / clean per field. First live run 07-31: cl SUSTAINED FIRE, cc FRESH FIRE. 08-02 all fields clean — cl regression was stale applied_layer stamp artifact (v0.6.390j guard + one-shot backfill restored honest metrics). Not a watch — diagnostic used by all field-level watches.
Recently ruled out — 2026-06-22 to 06-29 Stage 0/1 kills
One-shot smoke tests that landed at "no signal" or "captured by an existing axis." Re-run in 2-3 months if the seasonal regime shifts.
[HYPOTHESIS] Regime-conditional L3 efficacy.Killed 2026-06-22. L3 wins in every regime for ws/wg/ch/cm; pp loses everywhere but that's the documented Brier exception. Current whitelist is correctly tuned per-regime. (analysis/h_regime_l3.py)
[HYPOTHESIS] L3 regime-mismatch gating.Killed 2026-06-22. When state_fc.regime ≠ state_obs.regime, ws L3 win drops from +50% (match) to +40% (mismatch) — still big wins on both sides. Gating L3 to regime-agreement would lose the +40% to clean up marginal noise. (analysis/h_l3_regime_mismatch.py)
[HYPOTHESIS] Lead-h × C1a transition interaction.Killed 2026-06-23. Hypothesis: C1a penalty grows with lead. Result: mostly flat. Only ch shows monotonic growth (+0.31 short→long); cm marginal. ch is already in C1e, so redundant. No need for lead-conditional C1a bands. (analysis/h_lead_c1a.py)
[HYPOTHESIS] Solar zenith × cloud MAE.Killed 2026-06-23 — duplicate. Strong cloud-MAE spread by solar bin (cl 60%, cm 43%, ch 36%) but this is the same day/night cloud bias the cc→L4 Stage 1 hypothesis already addresses. Same signal sliced differently. (analysis/h_solar_cloud_selfcheck.py)
[HYPOTHESIS] Weekday vs weekend anthropogenic bias.Killed 2026-06-22 — likely artifact. Apparent gaps (t +1.14°F, h -2.65%, ws +1.34 mph between weekday and weekend) but the 21-day window contains only 3 weekends — one unusual Saturday biases the whole number. Day-of-week × weather correlation high at small N. Revisit with ≥6 months data (earliest April 2027). (analysis/h_weekday_temp.py)
[TUNING] L4 fit-window size sweep (7/14/21/30d).Settled 2026-06-23. All window sizes essentially tie for every L4-applied field; 7d marginally worse for ch/cm. Current 30d is correct. (analysis/h_l4_window_size.py)
[TUNING] Lead-bin granularity (10 fine bands).Settled 2026-06-23. Finer bins reveal expected lead-decay shapes (wind L2 ramps from full at lead 1h to zero at lead 24h) but no hidden ship gains. Current 4-band [0-5/6-11/12-23/24-47] structure isn't hiding signal. (analysis/h_lead_granularity.py)
[HYPOTHESIS] Mesonet confidence regime × MAE.Killed 2026-06-24 — null. Used state_obs.regime_synoptic (sea_breeze + frontal = high-scatter; calm + nw_flow + sw_flow = low-scatter) as proxy for mesonet confidence label. High/low scatter MAE ratios across all 9 fields: 0.93×–1.22× (only cm hit ⚠ at 1.22×). No meaningful spread; regime classification doesn't carry through as a confidence axis on its own. (analysis/h_mesonet_conf.py)
[HYPOTHESIS] C1g — RH ≥95% (fog) as confidence axis.Killed 2026-06-24 — orthogonality check. Stage 0 (h_rh_saturation.py) showed cm +134%, ch +149%, cl saturating +67% MAE elevation. Promoted to Stage 1 Tier 2, then killed same week by h_c1g_orthogonality.py: 1 ortho / 69 redundant / 0 confounded / 2 ambiguous. The elevation was sampling-driven — fog (obs_humidity ≥95) heavily co-occurs with both rain-forecast (C1f) and cc_fc ≥95 (cc-saturation). In the F=False or S=False subsets, fog rows show ratios 0.02–0.25× (smaller MAE than non-fog) — i.e. C1g either reduces error (when conditions disagree with model) or merely tracks F/S (when conditions agree). No independent widening signal. (analysis/h_rh_saturation.py, analysis/h_c1g_orthogonality.py)
[HYPOTHESIS] Wind shift rate (|Δwd_3h|) as alt-transition axis.Killed 06-24 via orthogonality. Stage 0: rotating ≥80° wind elevates cloud MAE (ch +33%, cm +24%, cc +15%). Ortho vs C1a: 1 ortho / 22 redundant / 2 confounded / 11 ambiguous. Only ch/24-47h shows independent signal — too narrow.
[GATE] CALM_GATE_ENABLED — fc_ws<3 ws/wg L3 skip.Killed 06-30 — wrong intervention. Intended to skip ws/wg L3 in calm winds. 06-30 l3_regime_lead_analysis on 694k rows showed calm-regime L3 WINS every band (+15% to +44%). Real ws L3 losers: ne_flow all bands + short-lead sea_breeze. Gate would have skipped where L3 wins biggest. Removed v0.6.267. Real fix: per-(field, regime, lead_band) skip table.
[HYPOTHESIS] Lt cooling branch — "double-counting" diagnosis.Killed 06-30 — rejected by data. Initial: L2 mesonet blend + Lt cooling double-counts. l6_l2_double_counting.py on 19,975 pairs: L2 only erases 3.7% of L1's MAE on cove rows. Real cause: L1 cold ~2.25°F at cove (HRRR microclimate gap). Cooling branch on cold baseline → 74.9% MAE increase in Δ ≤ −2°F bucket. Cooling branch disabled v0.6.259; warming retained (later killed too).
[HYPOTHESIS] Naive persistence vs HRRR.Killed 06-24 — expected behavior. Persistence wins ws/wg at all leads +95% to +526%, cc at short leads. Investigated unit mismatch — ruled out (both fc + obs are mph). 2× ratio is real model over-prediction for sheltered coastal location; L2+L3 correct it (raw L1 4.17 → L3 2.44 mph). Frontend does persistence-blend via blend_observed_into_hourly. Nothing actionable.
[HYPOTHESIS] pp Brier recalibration (all approaches).Parked 2026-08-02 — non-stationary. 07-21 → 08-02 pp calibration investigation. Reliability component ≈ 5% of total Brier (0.005 of 0.086); Uncertainty + Resolution dominate the error budget. Six Stage 0 halves-tests all HOLD: h_pp_bin_calibration (per-decile lift), h_pp_platt_calibration (2-param logistic |Δb|=0.06 stable, |Δa|=0.30 unstable), h_pp_kalman_recalibrate (rolling a_t), h_pp_source_blend (HRRR+GFS near-collinear α→~1.0), h_pp_platt_by_regime (weighted agg WORSE +21%/+25%), h_pp_frontal_platt_stage1 — the one Stage 1 SHIP (frontal × 6-11h fixed-b=0.6, initial Brier lift −28.56%/−18.73% at n=292) has since re-flipped HOLD in 08-02 halves-strict re-cut. Root problem: recalibration parameters don't transfer across time at this sample size. pp is L1 in production; scoreboard 0.0% is legitimate. Re-check earliest 08-05 → 08-15 as more frontal passages accumulate; if frontal-narrow re-attempt fails too, structural (physical-feature gating: CAPE, RH-500mb, front proximity) may be the only path. Live section retired from Research & Diagnostics. (analysis/pp_brier_reliability.py, analysis/h_pp_*.py)
[HYPOTHESIS] Tide-phase corrections — does forecast error track the tide cycle? (RETIRED 2026-06-08)
Verdict: weak signal, mostly entangled with diurnal cycle. Per-field tide-phase curves were tracked across weeks; the signal that survived stratification was hard to distinguish from hour-of-day patterns we're already correcting in L4. Cost of keeping it running (NOAA fetch + 12-bin accumulator + GCS history per Fitter pass) wasn't justified. Analysis: analysis/tide_hypothesis.py.skip.py (renamed 07-18 — removed from digest run so it stops spurious-FAILing on NOAA fetch). Fitter module flag: RUN_TIDE_TRACKING. Frozen charts below show the final state at retirement; they will not update.
Companion view — error vs tide elevation over time (frozen)
Time-domain rendering of the same data as the bucketed chart above. Different angle on the same retired hypothesis. The frozen state below is the last fit before tide tracking was disabled.
Higher leads = forecast made further ahead. Switch to see if the tide pattern is lead-specific.
Verdict: equivalent. Tested whether deriving humidity from corrected temperature + corrected dew point via Magnus outperforms the L2 network-blended humidity. 27k triples, identical MAE within noise. We kept the derived path anyway because it keeps the (T, T_d, RH, AH) quadruple internally consistent — but the hypothesis "derivation is more accurate" is closed. Analysis: analysis/derived_humidity.py.
Verdict: τ=14 days is fine within noise. Not a hypothesis — a parameter sweep over the Fitter's recency-weighting τ (how much old pairs count when fitting decay curves). Not the L2 lead-decay τ added in v0.6.44, which controls how a current bias is spread across forecast leads (see sec 2d). Tested τ ∈ {7, 14, 21} days across four reports. Held-out MAE differences under 2%, well below run-to-run variance. τ=14d stays. With L3/L4 mostly disabled in v0.6.45, this knob barely matters anymore. Analysis: analysis/decay_tau_tuning.py.
[HYPOTHESIS] R4 — HRRR vs GFS spread as confidence signal (RETIRED 2026-06-17, verdict: CLOSE)
Hypothesis: when HRRR and GFS disagree at a given forecast hour, actual error magnitude tends to be higher — i.e. |HRRR − GFS| per (field, lead) predicts |forecast − obs|. If true, the spread becomes a free uncertainty number that can widen displayed intervals and feed Gemini hedge language ("models disagree on tomorrow's high"). Data collection: HRRR L1 already in forecast_log.json; gfs_l1_log.json captures GFS L1 per tick for the same 0-48h window. Decision rule was: ship if median Spearman ρ > 0.25 for ≥3 fields, consistent across lead bands.
CLOSE verdict (2026-06-17, 112,877 joined pairs):
0 of 6 fields above the 0.25 ρ threshold. Maximum observed |ρ| = 0.012 (wind speed at 1-6h) — essentially zero correlation. HRRR vs GFS spread does NOT predict forecast error magnitude. Retired without auto-wiring. Manual script:analysis/r4_spread_analysis.py — re-run quarterly or after a model release.
[HYPOTHESIS] R5 — Cove warming — sea breeze across the peninsula heats Wyman Cove vs inland (RETIRED 2026-06-17, verdict: HOLD — L2 already captures it)
Reframed hypothesis (2026-06-13): Wyman Cove sits in the lee of the Marblehead peninsula on a S/SE/SW sea breeze. Marine air crosses ~2 miles of sun-heated land before reaching the cove, picking up surface heat in transit. Expected pattern: delta_wf_inland = waterfront_median − inland_median goes positive (cove warmer) when wind is from the S half AND sea breeze is active, with magnitude scaling to solar input (peaking ~12-14 EDT). Should flatten to zero when wind is from N/NE (cove is windward of peninsula) or after sunset (no surface heating). Original hypothesis ("waterfront cools during sea breeze") was geographically backwards and is closed.
Day-12 refit (1,732 ticks through 2026-06-24): matches the reframed model; magnitudes tightened further as the sample grew. NW flipped from neutral to weakly negative; E cooled further.
Wind
Sea breeze
n
mean Δ°F
S
active
186
+1.5
SE
active
88
+2.0
SW
active
79
+1.1
N
inactive
378
−1.0
NE
inactive
103
−1.0
E
inactive
86
−1.3
NW
inactive
459
−0.9
Diurnal curve under offshore/calm conditions shows clean morning-marine-cooling: trough around −3.7°F at 12:00 EDT (refit 06-24, n=1,732 entries over 12 days; cool air pool over Salem Sound persists; inland warms fast with sun while cove stays anchored to marine boundary). Both signals are physically coherent with the lee-warming model.
Data collection:cove_gradient_log.json captures waterfront-tagged Tempest median (Willow Rd, Neptune Rd — both at cliff-edge elevations on the harbor, confirmed by Joe), inland Tempest median (~18 stations), ambient T, wind dir/speed, salem_water_temp_f, buoy_water_temp_f, sb_active, sb_likelihood per tick (14-day retention).
Two-step plan:
Step 1 — measurement is stable (analysis/r5_cove_analysis.py). Confirms the (wind × sb × hour) lookup table reflects a real, repeatable physical signal. Day-4 already passes both regime tests; 7-day re-run on 06-19 just confirms stability.
Step 2 — held-out MAE audit (analysis/r5_audit.py). The actual ship question: does APPLYING the correction improve cove temperature forecast accuracy? Joins the pair log against the cove log, computes error_l4 + R5_delta and error_l1 + R5_delta, compares MAE against the existing L4-corrected baseline.
Step 2 verdict: HOLD (run 2026-06-16, n=29,444 matched pairs)
Baseline (existing L4-corrected, which for temperature = L2-corrected since L3/L4 are off): 2.547°F MAE
R5 added on top of L4: 3.045°F MAE (−19.58% — significantly worse)
R5 replacing the entire stack: 3.066°F MAE (−20.39% — also worse)
The L2-overlap hypothesis was empirically confirmed. L2's 1/distance² × elevation station weighting for the cove is dominated by the two waterfront Tempests (Willow Rd, Neptune Rd at ~0.1–0.2 mi). L2's "cove bias" is effectively "waterfront bias" by construction. Layering R5's (waterfront − inland) delta on top double-counts the same signal — the cove obs is already waterfront-influenced via L2, so adding more waterfront delta pushes the forecast AWAY from the obs.
Decision (global R5, retired 2026-06-17):r5_audit.py's held-out test of R5 applied across the full pipeline showed it makes cove temp 20–22% worse — L2's waterfront-weighted station blend already captures the signal, so layering R5 on top double-counts. That global formulation stays retired.
Current status (Lt — microclimate correction, RETIRED 2026-07-13 v0.6.329): Both branches return 0.0. Cooling branch killed 2026-06-30 v0.6.259 (paired-MAE showed it materially increased cove temp MAE); warming branch killed 2026-07-01 v0.6.276 after real per-row Production exposed the same double-counting for the sea-breeze direction; Fix B refit tried 2026-07-13 (analysis/l6_fix_b_refit.py) and failed at held-out +0.29% vs +1.0% ship gate. Lookup tables + cove_gradient_log.json continue to log per-tick against a possible future reactivation (seasonal shift, sensor swap, or materially different mechanism). Full rationale in R&D → Lt.
One niche subtlety in the breakdown: long-lead (24-47h) sea-breeze forecasts get +7.85% MAE improvement with R5. At long leads, L2's τ=4h decay has long since faded, so R5 has something L2 doesn't. Not worth shipping a conditional correction for, but documented.