Forecast pipeline

loading… MAE data …
Right now — what the pipeline is doing
Field Raw model Production Correction Field Raw model Production Correction
Loading current tick…
Confidence: · — stations reporting
Briefing source: · —

Current state — what's running · improving · being evaluated

Current pipeline state — per-field snapshot
One-glance state of every field. Both columns read from mae_over_time.json7-day from last_7d (n-weighted rolling 7d), 24-hour from last_24h (always contains a full diurnal cycle, unlike calendar-day "today" which is a partial sample before end-of-day). Refreshed hourly by the myweather-publisher Cloud Function (v0.6.395f). 24-hour catches a bad ship on day 1; 7-day catches sustained drift. Watch both.
Field Applied layers 7d avg 24-hour Status
t temperature L1 → L2 additive Stable, no open work. L2 Kalman blend absorbs the microclimate signal on its own (persistence-skill +0.68 pooled); Lt reclassified as telemetry experiment 07-27 after Fix B refit came in +0.29% held-out (below +1% ship gate).
h humidity L1 → L2 additive → L4 diurnal h status loading… h L2 retune v0.6.390g (07-31) landed cleanly — 08-04 verdict: retune working, direction right, drawing board avoided. τ-suspect signature at 6-11h/12-23h to be watched as it compresses toward 0. h/l4 narrow-add HELD 07-26.
ws wind speed L1 → L2 direct (L3 dropped 08-08 v0.6.397) ws L3 dropped from L3_FIELDS 2026-08-08 v0.6.397 — walkforward cleared 7-day gate; pooled ws L3 fc +0.1% / obs +0.6% is noise. L3 asymmetric additive + SKIP_TABLE entries now inert (kept for docs). Root cause was L2 blend, not L3. L2 blend shrunk 24h → 4h (v0.6.384, 07-28) fixed the +30% Prod-vs-Raw regression that ran 2 weeks. Per-lead diagnostic showed L2 lost catastrophically at leads 3-16 (+40 to +85% vs L1). Sweep BLEND_HOURS ∈ {1..24}: 4 wins at +14.72% pooled. Physical: observed KBVY wind has more variance than the next-few-hour average — 24h bleed imported point-in-time noise. wsbp gate (v0.6.388, 07-28) — sibling of dpbp, calm-only, sign-inverted. ENABLED=False. Preflight failed 08-04: calm regime n=0 in shadow window. Wait for calm regime to accumulate.
wg wind gust L1 → L2 direct → L3 lead-decay → wg residual persistence (dormant) Stable win vs raw; existing asymmetric wg SKIP table is LIVE (07-20 v0.6.366, 48 cells). wg L3 flat SKIP_TABLE extended 07-28 v0.6.385 — 4 clean-halves cells wired (calm 0-5 +9.4%, calm 12-23 +40.0%, ne_flow 6-11 +11.8% NEW, sea_breeze 6-11 +5.6% NEW). Composition rotated from 07-14 Stage 1: 07-14 SHIP cells calm 6-11 fell to THIN, sea_breeze 0-5 to MARGIN, unknown 24-47 to PERSISTENCE_TERRITORY. wg L3 +3 cells 08-04 v0.6.391 — 08-04 re-cut cleared the two held cells (calm 24-47 +45.2% halves +4.3/+53.1, sea_breeze 24-47 +31.1% halves +4.0/+45.0) plus new frontal 12-23 (+8.8% halves +13.2/+4.5). Live table now 7 cells; demote-on-A-negative watch at next re-cut. wg residual persistence gate (Stage 3 wired 07-14 v0.6.351, ENABLED=False) still HELD 07-27 — blocked by h_wg_residual_persistence_stage1 flip PROMOTE → MARGINAL; wait for Stage 1 recovery.
dp dew point L1 → L2 additive → L3 skip-table → dp residual persistence (dormant) (derived from corrected t, h via Magnus) Stable win, derived. Rides t + h improvements. Frontal regime bias closed 07-09. dp residual persistence gate Stage 3 wired 2026-07-25 v0.6.380 ENABLED=False (cloned from wg template; 8 SHIP + 2 MARGIN long-lead cells). Gate closed 08-01 — verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09). dp_bias_persistence gate FLIPPED 2026-08-04 v0.6.391 (shipped 07-28 v0.6.387) — first antecedent-error-based specialist LIVE. Backlog #6 closed via full Stage 0 → Stage 3 arc: chunk-stability HOLD but per-day distribution + antecedent lag-1 r=+0.583 unlocked design. Gate: regime ∈ {pre_frontal, nw_flow, sw_flow} AND lead ≥ 6h AND prev_24h_dp_bias < -1.5°F → +2°F. Pooled dp MAE 3.108 → 2.808. 14-day post-flip watch through 08-18. Preflight gap: dpbp shadow key not stamped in pair log — same infra as chp/wdp — measurable pre-flip gate would require v0.6.382p-style shadow-key stamping (follow-up).
cc cloud cover (derived) L1 → L2 KBOS+KBVY hourly[0] blend → LcCcd = max(cl_l6, cm_l6, ch_l6) (SKIP_REGIMES: se_flow, unknown → Pirate fallback) cc retired from Lc as independent field 07-30 v0.6.390. cc is HRRR's own combine of cl/cm/ch — running Lc on it in parallel with the three components was double-correction from Lc's original ship. Ccd (cc_from_derivation.py) overwrites hourly.cloud_cover with max(cl_l6, cm_l6, ch_l6) after Lc. h_cc_derivation.py on 123K held-out quads: +8.48% pooled MAE vs pre-Ccd production, halves +11.4%/+5.8%, wins 6/10 regimes. cc excluded from Overall MAE mean in the scorecard above (derived field would double-weight cl/cm/ch). cc status loading… — 08-04 v0.6.391 Ccd saturation guard (raw cc ≥ 90 keeps raw) cleared the Last-24h loss to −9.9%; the 7d cell still carries pre-guard poison at Ccd's obs bin 95-100 (cm/ch Lc dragged max below METAR total-sky) and rolls off through ~08-11.
ch high cloud L1 → L2 KBOS+KBVY blend → L3 lead-decay → L4 diurnal → Lc (flipped 07-17) → chp (flipped 07-19) Best-performing field vs raw. Lc flipped 07-17 + chp LIVE 07-19 v0.6.358 stack for the fullest non-frontal ch coverage yet. Rollup post 07-27 v0.6.382t emergency demote: 14 SHIP / 8 MARGIN / 13 SKIP / 2 THIN (was 19/9/7/2 pre-demote; 6 mid-lead cells regressing vs Lc under L6 baseline flipped SHIP/MARGIN → SKIP). 14-day post-ship watch closed clean 08-02. Full-shape refinement to the 5-cell L6-baseline SHIP set still under normal 7-day live-layer gate. 08-10 diurnal gate (v0.6.401): chp suppressed on nw_flow + pre_frontal daytime valid hours [10, 18) local after 48h band-cut traced +7.48 bias at 12-23h to daytime nw_flow cells burning off overnight residuals. 08-13 emergency demote (v0.6.405): Prod-vs-L6 MAE gap widened +3.4→+12.7 over 08-06→08-13; h_ch_persistence_blend_stage2_vs_l6 flagged 9 live chp cells losing to L6 (worst pre_frontal/24-47 Δ +78%, sw_flow/24-47 +31%). Added _CELL_SKIP frozenset in ch_persistence_gate.py: calm/{12-23,24-47}, nw_flow/12-23, pre_frontal/12-23, se_flow/24-47, sea_breeze/{12-23,24-47}, sw_flow/{12-23,24-47} forced back to L4 regardless of curated verdict. Fitter regenerates the curated JSON daily so JSON edits would be transient. See Post-ship watches. See [[project_ch_chp_regression_watch_08_13]].
cl low cloud L1 → L2 KBOS+KBVY hourly[0] blend → Lc → clp (dormant) (cl in _FIELD_SKIP; cl_l6 = raw) cl status loading… (post applied_layer poison backfill 08-02; note auto-populated from same source as the 7-day cell above). cl fully off Lc as of 07-30 v0.6.389f — walk-forward held-out showed both shift-table shapes hurt cl (pool_vs_raw −22.15%, reg_vs_raw −30.37%). Shift-table can't distinguish "model over-forecasts overcast" from "correctly forecasts overcast" — cl currently in the latter regime. Reversibility: remove "cl" from _FIELD_SKIP. clp Stage 3 wired 07-24 v0.6.379 ENABLED=False (unaffected by cl Lc kill). Gate closed 08-03; 08-04 verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09).
cm mid cloud L1 → L2 KBOS+KBVY blend → L3 lead-decay → Lc (flipped 07-17) cm status loading… Stable win + Lc live. HRRR cm-distribution anomaly 07-04→07-11 recovered. C1 Stage 4 last audit: NOT READY (32.43% pass rate). Next re-audit gated on window roll.
pp precip probability L1 (L3 dropped 07-04 — hurts Brier despite MAE gain) L1 in production. All recalibration paths PARKED 2026-08-02 — non-stationary. L3 dropped 07-04 (hurt Brier despite MAE gain), so live Brier = raw Brier (matches the 0.0% headline tile). Separate offline v0.6.335 Phase 3 decomposition showed CALIBRATED +8.4% vs raw + BSS +0.126. The 07-27 h_pp_frontal_platt_stage1 SHIP at frontal × 6-11h re-flipped HOLD on 08-02 halves-strict re-cut — Reliability ≈ 5% of Brier (0.005 of 0.086); Uncertainty + Resolution dominate the error budget; recalibration params don't transfer across time at this sample size. All Platt-family scripts (h_pp_platt_calibration, h_pp_platt_by_regime, h_pp_frontal_platt_stage1, h_pp_bin_calibration, h_pp_bias_persistence_stage0) retired to .skip.py on 08-10 (dead-verdict HOLD/KILL, no reopen path from daily data). If pp_brier_reliability trips STAGE0_OPEN, h_pp_bin_calibration.skip.py runs on demand. Structural next step if this ever reopens: physical-feature gating (CAPE, RH-500mb, front proximity).
pr pressure L2 additive (gated on nw_flow/0-5h + 6-11h since 08-10) L2 SHIP with regime gate — v0.6.401 (08-10). pr_l2_regime_lead_retro cleared halves (Jaccard 0.50): nw_flow/0-5h A +21.8%/B +41.6%, nw_flow/6-11h A +10.3%/B +13.1% on 8,596 shadow rows. Only both-halves winners shipped; pooled-only WINs (nw_flow/12-23h, pre_frontal/{0-5,6-11}, sw_flow/{0-5,6-11}) held pending 7-day gate agreement. Shadow-write (corrected_pressure_in_post_l2) remains unconditional so retro keeps evaluating skip cells. Predecessor: L2-disabled-everywhere 07-01 → 08-10 (K=1 pooled noise dominated) — that was the pooled-blind read; regime-conditional found the win. Gate ordering bug fixed 08-13 v0.6.402: live gate had never fired since 08-10 ship — collector.py ran add_corrected_hourly_arrays before stamp_state, so derived.state.regime_synoptic was None when the gate read it and the guard short-circuited. Pair-log audit: 2,688 pr rows post-08-10 with 0 stamped applied_layer=l2. Backstamp early read (scratchpad/pr_l2_backstamp.py, simulating live-apply using state_obs.regime@lead=0): gated-vs-raw −5.19% on 2,688 rows; nw_flow/0-5h Δ −43.6% n=223, nw_flow/6-11h Δ −21.9% n=221 — matches Stage 1 A/B halves. 14-day post-fix watch begins 08-13.
sr solar radiation L1 → Lsr synoptic-regime (skip: ne_flow + calm) Lsr LIVE for sr. Skip regimes: ne_flow + calm; Lsr fires on the other 7. Bias table refreshed 08-04 v0.6.391 (fitter now writes lsr_bias_table_curated.json, module loads at import — retires 5+ weeks of manual-sync drift). sr sea_breeze Lsb FLIPPED 2026-08-05 v0.6.394 · ENABLED=True · gate cc < 25. Fresh 7-day gate cleared 7/7 daily PROMOTE reads (07-30 → 08-05) on the narrowed-shape Stage 2. Today's Stage 2: pooled Δ +11.34%, halves +23.9%/+4.0%, cc-bin 0-25 +44.6% (n=718, halves +49%/+34%). Post-deploy verified: sr_sea_breeze_correction.applied = True at 12:27 UTC tick. 14-day post-ship watch through 08-19. Watch points: hour 13 SKIP on test (bias +16.49 W/m² still applied — needs per-hour SHIP filter in curator if sr Last-24h drags); halves B +4% is recency-weaker than halves A +23.9%.
pa precip amount L1 (no L2/L3/L4 apply) L1 only, no open work. Field-specific τ swept 28 → 42 → 7 → reverted to global τ=14 on 07-19 v0.6.358 (+0.9% vs τ=14, below the 5% floor).
wd wind direction L1 → L2 wind_blend → wdp (flipped 07-27) L2 + wdp both LIVE — real trend visible post-08-02 metric fix. L1_ONLY routing bug had hidden L2/wdp contributions for 13 days (mae_over_time was reading L2-view error as raw); fixed 08-02, wd now honestly emits prod_real. Real picture: L2 −2 to −5% vs raw most days, wdp adds another 0-4%. 07-31 outlier: calm/24-47 wdp +72%, sea_breeze/0-5h wdp +109% — real gate misfires; if either recurs, narrow gate. L2 wind_blend shipped 07-20 v0.6.368a; wdp FLIPPED 2026-07-27 v0.6.382 (5 SHIP cells). wdp 14-day post-ship watch CLOSED CLEAN 08-11 — Δ −2.7% aggregate n=16,416; 07-31 outlier cells did not recur. Collector-side applied_layer stamp gap fixed 08-09 v0.6.400 (bug silent since v0.6.269) did not affect the debug-page metric — mae_over_time's L1_ONLY_FIELDS workaround bypassed the stamp; real impact was only in Fitter's decay_corrections.json fallback.
Loading 7-day snapshot from mae_over_time.json's last_7d block…
Loading provenance…
🟢 What's running
Stack
  • L2 mesonet blend: t · h · dp · cc · cl · cm · ch · ws · wg (pr on nw_flow 0-11h since 08-10)
  • L3 lead-decay: wg · ch · cm (pp dropped 2026-07-04 v0.6.304; ws dropped 2026-08-08 v0.6.397 — walkforward cleared 7-day gate; ws L3 fc +0.1% / obs +0.6% pooled is noise; wg L3 skip cells: calm 0-5 + 12-23 + 24-47, ne_flow 6-11, sea_breeze 6-11 + 12-23 + 24-47, frontal 12-23)
  • L4 diurnal: {ch, cc} · cc skip cells added 2026-08-08 v0.6.397: nw_flow 0-5 + 12-23, pre_frontal 0-5 + 6-11 + 12-23 (cc L4 only wins at longer leads in these regimes)
  • Lsr synoptic (sr): skip ne_flow + calm · Lt reclassified as telemetry experiment 07-27 (Fix B held-out +0.29%)
  • Lsb sr sea_breeze cc<25 Lsr override (sr): per-hour bias hours 12-19 · FLIPPED 2026-08-05 v0.6.394 · gate: regime == "sea_breeze" AND cc < 25 · 14-day watch through 08-19
  • Lc cloud saturation-unbiasing (cm · ch only) · FLIPPED 07-17 v0.6.355 · cl+cc both OFF via _FIELD_SKIP after 07-30 v0.6.389d-390 intervention
  • Ccd cc from-derivation · LIVE ENABLED=True 07-30 v0.6.390 (flipped 6d ahead of gate) · cc = max(cl_l6, cm_l6, ch_l6) except SKIP_REGIMES {se_flow, unknown}
  • chp high-cloud persistence-of-obs (ch): 14 SHIP cells post-07-27 demote, 9 additional cells force-skipped via _CELL_SKIP in processor 08-13 v0.6.405 · FLIPPED 07-19 v0.6.358 · 14-day watch closed clean 08-02 · 14-day post-v0.6.405 watch through 08-27
  • wdp wind-direction predicted-transition persistence (wd): 5 SHIP cells · FLIPPED 2026-07-27 v0.6.382 · 14-day watch CLOSED CLEAN 08-11 (Δ −2.7% n=16,416)
  • dpbp dew-point antecedent-error persistence (dp): pre_frontal + nw_flow + sw_flow, lead ≥ 6h · FLIPPED 2026-08-04 v0.6.391 · 14-day watch through 08-18
  • C1h trend-direction premium widening: 9 SHIP cells (cc/cl/cm/ch 12-23h & 24-47h + t/24-47h) · SHIPPED 2026-08-05 v0.6.393 · RE-CURATED 2026-08-08 v0.6.396 · 08-08 refresh: t/12-23h SKIPed by magnitude floor (−0.09%), t/24-47h newly SHIPs WIDEN (+11.0%); cc/cl/cm/ch 12-23h + 24-47h all still SHIP · confidence-layer only (no forecast modification)
Production vs raw
  • ✓ ch −46% · wg −16% · cc −15%
  • ✓ cm −7% · dp −6% · h −2%
  • ✓ ws long-lead — RESOLVED. BLEND_HOURS 24→4 (v0.6.384 07-28) fixed the root cause; ws then dropped from L3_FIELDS entirely 2026-08-08 v0.6.397 after walkforward cleared 7-day gate (pooled ws L3 impact +0.1%/+0.6% is noise). ws L3 asymmetric skip-table (v0.6.370) + SKIP_TABLE entries (v0.6.386) are now inert — kept for historical documentation.
Guards
  • 🟢 raw-baseline verifier: healthy (no drift)
  • 🟢 live-layer change gate active (7-day / 2-tool / per-cell / no-ENT)
🟡 What's improving
11 active Stage 1+3 candidates · all auto-run in daily digest · shipped items live under Post-ship watches below
C1h trend-direction ✓ Stage 3 wired 07-08 · per-cell ortho gate 07-10 · gated OFF
15 SHIP cells (all WIDEN) live-stamping on confidence.cells[field][band].c1h. Per-cell co-axis ortho gate (v0.6.321) suppresses cells non-orthogonal to the currently-firing co-axis — cl fires freely, ch 24-47h + t × 3 never fire.
→ Already live-stamping (confidence_layer reads the curated table each tick). The narrow-promote counter is auto-suppressed in the digest per KNOWN_LIVE_PIPELINES — the ship gate for user-visible band widening is C1 Stage 4 (last HOLD 36.36%).
dp depression regime frontal branch closed · nor_easter watch opened
Frontal decayed −2.19 → −1.98 → −1.51 → −0.87°F — below the 1.5°F action floor as of 07-09. Branch retired.
New: nor_easter +3.79°F ★ but n=279 — magnitude passes floor, sample thin. Gate: 3 consecutive reads with n growing.
C1d cloud disagreement ✓ Stage 3 wired 07-08 · gated OFF
13 SHIP + 1 MARGINAL cells stamping on confidence.cells[field][band].c1d. Fires when live cloud_inter_source_sigma ≥ Q3 (44.55). cc 24-47h MARGINAL NARROW −7.06% is a documented outlier — flagged for Stage 4 review.
→ Already live-stamping (curated table read per tick). Narrow-promote counter auto-suppressed in digest per KNOWN_LIVE_PIPELINES — ship gate for user-visible band widening is C1 Stage 4 (last HOLD 36.36%).
Pre-frontal cloud widening ⚙ Stage 2 curated 07-12 · counter wired · blocked on population (n=8 passages, THIN)
Latest ortho read: 7 SHIP cells (cell is SHIP iff ORTHOGONAL vs C1a AND vs C1e). Written to weather_collector/data/pre_frontal_curated.json. v0.6.372b matched-regime fix HOLDS PROMOTE with the corrected baseline; SHIP set expanded 5→7. Digest currently at streak 1/7 due to cell-set drift + population still THIN.
→ Real 7-day narrow-promote counter wired 2026-07-12 v0.6.328d. Blocked on population: only n=8 frontal passages in the current window, 17% pair-log join rate — both matched-regime KILL (h_hsf) and matched-regime PROMOTE (h_pre_front) share this THIN caveat. Do NOT act on either verdict until passage count ≥ 15 (probably 08-05 to 08-15). See [[project_c1e_hsf_kill_investigation]].
cl persistence gate (clp) ⚙ Stage 3 wired 07-24 v0.6.379 · ENABLED=False · streak walker seeded 08-10 v0.6.401b (day 3/7 · need 4 more, fire set n=4) · earliest flip 08-16 · verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09)
Successor to the retired cl_persistence_short_lead (all-9-regimes design gate not met). Broader regime × lead_band gate: 12 SHIP / 8 MARGIN / 16 SKIP / 1 THIN. Whitelist: persistence on calm/se_flow/unknown all-leads + nw_flow 24-47h + short-lead in sw_flow/ne_flow/nw_flow/pre_frontal + pre_frontal 6-11h. Wired after Lc. 7-day flip gate: SHIP cell-set Jaccard ≥ 0.8 across 7 daily reads. See [[project_cl_persistence_investigation]].
wg residual persistence gate ⚙ Stage 3 wired 07-14 v0.6.351 · 07-27 flip HELD · awaiting Stage 1 recovery
Long-lead-only: adds per-clock-hour L2-residual mean (prior 14d) on top of fc_l2, replaces L3-corrected wg on gate-fired cells. Short leads always SKIP (L2's Kalman handles close-in). 07-27 flip HELDh_wg_residual_persistence_stage1 flipped PROMOTE → MARGINAL 07-26 (+20.24%, was +17.74%). Persistence hypothesis's aggregate weakened; hold until Stage 1 recovers PROMOTE. (Divergence report drop-{wg,ws} was stale-tool artifact per [[feedback_walkforward_skip_table_blind]]; fix v0.6.381 landed 07-26.)
dp residual persistence gate ⚙ Stage 3 wired 07-25 v0.6.380 · ENABLED=False · gate closed 08-01 — verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09)
Cloned from wg template. Runs after wg_residual_persistence; replaces L3-corrected dp on gate-fired cells with (fc_l2 + fitted L2-residual). Stage 2 preview: 8 SHIP / 2 MARGIN / 26 SKIP / 1 THIN — SHIP cluster long-lead only (frontal 12-47h, nw_flow 24-47h, pre_frontal 12-47h, sw_flow 6-47h); zero SHIP at 0-5h. Best cells sw_flow 24-47h (−20.4%), sw_flow 12-23h (−18.6%). Clamp |Δ| > 10°F. Same Jaccard ≥ 0.8 streak gate as wg.
pp frontal × 6-11h Platt (fixed b=0.6) ⏸ PARKED 2026-08-02 · script retired to .skip.py 2026-08-10
07-27 Stage 1 SHIP (Brier lift −28.56%/−18.73%, n=292) re-flipped HOLD on 08-02 halves-strict re-cut — recalibration params don't transfer across time at this sample size. Root: Reliability ≈ 5% of Brier (0.005 of 0.086); Uncertainty + Resolution dominate. All Platt-family scripts (h_pp_platt_calibration, h_pp_platt_by_regime, h_pp_frontal_platt_stage1, h_pp_bin_calibration, h_pp_bias_persistence_stage0) retired to .skip.py on 08-10 with the daily-digest retirement pass (10 dead-verdict scripts total). Structural next step if pp ever reopens: physical-feature gating (CAPE, RH-500mb, front proximity), not calibration parameter fitting.
Lc recent-bias gate (item #3 on pipeline-to-good plan) ⚙ Stage 1 SHIP 08-09 v0.6.399 · per-field clearance added 08-12 v0.6.401i · ch streak 4/7 · earliest ch wire 2026-08-15
Applies existing lc_correction_table.json shift only where recent 3-day observed bias still agrees with the historical fit (sign match + magnitude ≥ 0.5×|hist|). Smaller surface than the EMA/Kalman rewrite alternative — no new lookup table, just a per-cell gate on the live shift. Today: ch = STAGE 1 PROMOTE, cm = STAGE 1 PROMOTE, cl = STAGE 1 PROMOTE (new as of 08-13), cc excluded (derived via Ccd). Per-field gate (08-12 fix): each field tracks its own 7-day streak — ch streak 5/7 (clears 08-15), cm streak 3/7 (clears 08-17), cl streak 1/7 (clears 08-19). Runtime table shipped 08-13 v0.6.407: weather_collector/data/lc_recent_bias_gate.json now emitted daily with per-cell gate_apply decisions + promoted_fields + fields_cleared. Stage 3 wire ready to consume on ch's clearance. Set-level `stable` check no longer blocks ch when cm joins the SHIP set. Gate history in .cache_lc_recent_bias_gate_history.json. Stage 3 wire when ch clears: modify cloud_saturation_correction to consult the gate before applying the shift on ch. See [[project_lc_regime_conditional]] + [[project_plan_pipeline_to_good]] item #3.
chp full-shape refinement (adopt L6-baseline 5-cell SHIP set) ⚙ 6-cell emergency demote LIVE 07-27 v0.6.382t · full-shape adoption 7-day gate closed 08-03 · verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09)
Preview curated JSON at weather_collector/data/ch_persistence_gate_curated_vs_l6.json. L6-baseline Stage 2 rebuild (h_ch_persistence_blend_stage2_vs_l6.py) says chp's honest SHIP set is 5 cells: calm/0-5 (−70%), pre_frontal/0-5 (−41%), sw_flow/0-5 (−30%), se_flow/0-5 (−26%), ne_flow/6-11 (−17%). Currently-live chp is at 14 SHIP + 8 MARGIN post-demote; 5 more halves-disagreement/small-magnitude cells still fire but should also fall away. 7-day live-layer change gate closed 08-03; verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09). Adopt condition: 7 daily reads of the L6-baseline Stage 2 continue to agree on the 5-cell SHIP set + no new frontal-only-recalibration interaction.
ws L3 hardcode-REPLACEMENT ⚫ INERT 2026-08-08 v0.6.397 — ws dropped from L3_FIELDS entirely
Now dead code. All ws L3 asymmetric skip-table + SKIP_TABLE logic is inert because ws is no longer in L3_FIELDS (walkforward 08-08 cleared 7-day gate; pooled ws L3 +0.1%/+0.6% is noise). Historical: L3 is a mean-bias subtraction that helps on over-forecast rows and hurts on under-forecast rows. Asymmetric fc-bin SKIP wired additively on top of the old hardcode (ne_flow all + sea_breeze 0-11) — additive can only add skips, never subtract. Reached 26 SKIP cells (was 44 at scaffold); 23 of 40 concentrated at 24-47h — matches [[project_ws_l3_long_lead_regression]]. REPLACEMENT flip stayed HELD across three re-runs (Jaccard 0.70 → 0.75 → 0.75, none clear 0.8). BLEND_HOURS 24→4 fix (v0.6.384 07-28) resolved the root long-lead regression. 07-28 v0.6.386 extended SKIP_TABLE with frontal 24-47 + pre_frontal 24-47 — those never got a chance to matter before ws left L3. (wg's asymmetric wiring shipped LIVE 07-20 v0.6.366 with 48 SKIP cells; still live — see in-flight table below.)
🔵 What's being evaluated next
Upcoming
Forward-looking only. For today/yesterday narrative see Recent activity below.
Fri 08-15
L6 Fix B rolling gate — 7-day gate from 08-08. Currently HOLD-GATE (1/7). Verdict promotes if 7 consecutive days SHIP with consistent ship_bins.
Sat 08-16
clp Stage 3 flip gate — earliest flip (streak walker seeded 08-10 v0.6.401b, fire set n=4, day 3/7). Flips if 7-day Jaccard ≥ 0.80 clears.
Mon 08-18
dpbp 14-day post-flip watch closes (v0.6.391, flipped 08-04). Triggers: any focus regime regressing worse than v0.6.387-baseline once n ≥ 200; overall dp MAE drifting up.
Tue 08-19
Lsb 14-day post-flip watch closes (v0.6.394, flipped 08-05). Triggers: sr Last-24h vs raw drags (hour 13 candidate cause); halves B stays < +5%; per-regime sr MAE at sea_breeze/<25cc worse than v0.6.383b baseline.
ongoing
C1d narrow-promote — auto-suppressed (KNOWN_LIVE_PIPELINES). Already live-stamping via confidence_layer; user-visible band-widening gate is C1 Stage 4.
held
pre-frontal C1e narrow-promote — HELD 08-11 fresh re-run. Down from 16 ortho cells (June) to 2 SHIP (ch 24-47h, cl 24-47h). Only 3 frontal passages in the 45d window; most cells THIN. Re-run when autumn front cadence returns.
held
frontal-t bias Stage 0 — Stage 1 gate day 1/7 (08-12 v0.6.401g). 3/4 bands HIT (0-5h +1.94, 6-11h +2.74, 12-23h +1.06); 24-47h SIGN_FLIPS. Half A concentrated in one 07-13/16 event. Rolling 7-day gate added to h_frontal_t_bias_stage0.py mirroring l6_fix_b_refit: requires ≥7 distinct days, no HOLD days, ship_bands STABLE, ≥1 band. Verdict escalates to STAGE 1 CLEAR — the trigger to scope frontal_t_bias.py specialist. Earliest clear 2026-08-19.
held
sr obs-recent override Stage 0 — 08-11 near-miss. Fired-subset lift +23.5% but pooled test lift 4.58% under 5% gate (77 fires, 91% non-Lsb). Two paths to Stage 1: loosen pooled gate OR tighten trigger to 300 W/m². Backlog #8.
blocked
wsbp — flip gate still HELD as of 08-08 (v0.6.388): calm regime shadow still below MIN_N_ANTECEDENT=20 in the 24h antecedent window. Predict: one more calm overnight pushes it over.
Post-ship watches — active
  • pr L2 regime-gated SHIP (v0.6.401, 08-10) — day X/14. First live pr L2 apply since 07-01 kill. K=1 additive with τ=8h, gated on (nw_flow, 0-5h) + (nw_flow, 6-11h) — both-halves winners from pr_l2_regime_lead_retro (Jaccard 0.50, A +21.8%/B +41.6% and A +10.3%/B +13.1%). Shadow-write remains unconditional so retro keeps evaluating skip cells for future promotion. Triggers: (1) pr Prod-vs-raw drifting worse than pre-08-10 baseline at either fire cell; (2) nw_flow winners flipping sign on next retro re-cut (would suggest 08-10 shadow window was atypical); (3) 7 consecutive daily reads with retro confirming → consider promoting pooled-only WINs (nw_flow/12-23h, pre_frontal/{0-5,6-11}, sw_flow/{0-5,6-11}). Scoreboard: pr dropped from MAE_UNCORRECTED_FIELDS; touched median n now 9. 08-12 re-verify: retro flipped STAGE 1 SHIP → MIXED (Jaccard 0.50 → 0.25) on the pooled-only candidate cells; shipped nw_flow cells both-halves-WIN with gain grown (0-5h A +25.1%/B +37.8%; 6-11h A +11.8%/B +10.7%). Retro script hardened with SHIPPED_CELLS constant + SHIPPED CELLS STATUS block + candidate-only Jaccard so future non-shipped flips don't masquerade as live-gate regressions.
  • chp emergency cell-skip (v0.6.405, 08-13) — day X/14 through 08-27. 9 mid/long-lead (regime, lead_band) cells force-skipped in ch_persistence_gate.py::_CELL_SKIP: calm/{12-23,24-47}, nw_flow/12-23, pre_frontal/12-23, se_flow/24-47, sea_breeze/{12-23,24-47}, sw_flow/{12-23,24-47}. h_ch_persistence_blend_stage2_vs_l6 flagged these as losing to L6 with Δ +9.7% to +34.9% (worst pre_frontal/24-47 Δ +78%). Prod-vs-L6 MAE gap on ch had widened monotonically +3.4 (08-06) → +12.7 (08-13). Prior [[project_ch_chp_regression_watch_08_07]] + [[project_ch_chp_midlead_band_watch_08_10]] closed clean 08-09/08-10 but the gap started widening after. Reversibility: remove any tuple from _CELL_SKIP. Triggers: (1) Prod-vs-L6 gap on ch not shrinking within 3 days (residual damage is Lc-side, not chp) → escalate per [[project_ch_chp_regression_watch_08_13]] playbook; (2) any of the 9 demoted cells appearing legitimate on next Stage 2 rebuild → widen or narrow the frozenset; (3) fitter emits new SHIP verdicts for these cells → note but the processor-level SKIP overrides.
  • chp diurnal gate (v0.6.401, 08-10) — day X/14. chp fire suppressed during daytime valid hours [10, 18) local in nw_flow + pre_frontal. 08-10 48h band-cut traced 12-23h + 24-47h scoreboard regression to chp_bias +20.51 at 12-23h nw_flow day (n=78) and +4.25 at 24-47h pre_frontal day (n=91); nighttime cells same regimes remain −6 to −26% vs raw wins. Physical: post-frontal / pre-frontal residual overnight clouds burn off by midday, chp persists them into daytime → over-forecast. Triggers: (1) diurnal_skips_by_band recording zero skips (means classification broken); (2) daytime nw_flow / pre_frontal chp cells appearing to help at n ≥ 30 (over-restriction — reconsider gate scope); (3) scoreboard ch 12-23h / 24-47h cells failing to recover within 48h.
  • walkforward L3/L4 (v0.6.397, 08-08) — day X/14. ws dropped from L3_FIELDS + 6 SKIP cells (wg L3 sea_breeze 12-23; cc L4 nw_flow 0-5/12-23 + pre_frontal 0-5/6-11/12-23). Triggers: ws Prod-vs-raw drifting worse than v0.6.396-baseline; cc L4 nw_flow/pre_frontal short-mid leads suddenly winning again at n ≥ 200; wg sea_breeze 12-23 regressing L3.
  • C1h re-curate (v0.6.396, 08-08) — day X/14. Refresh of 08-05 v0.6.393. 9 SHIP cells: cc/cl/cm/ch 12-23h + 24-47h unchanged; t/12-23h SKIPed (−0.09%), t/24-47h newly SHIP WIDEN (+11.0%). Data-only. Trigger: any of the 9 reversing sign on re-cut; t band flip 12-23h vs 24-47h recurring (week-to-week noise).
  • Lsb (v0.6.394, 08-05) — day X/14 through 08-19. sr sea_breeze cc<25 Lsr override LIVE, 4/4 lead-bands SHIP. Triggers: (1) hour 13 SKIP-on-test but bias still applied — add per-hour SHIP filter if sr Last-24h drags; (2) halves B +4% recency-weak, revisit if stays < +5%; (3) any sr Prod-vs-raw regression at sea_breeze/<25cc rows.
  • dpbp (v0.6.391, 08-04) — day X/14 through 08-18. First antecedent-error specialist LIVE. Gate: pre_frontal/nw_flow/sw_flow + lead ≥ 6h + prev_24h_dp_bias < -1.5°F → +2°F. Pooled dp MAE 3.108 → 2.808. Triggers: any focus regime regressing worse than v0.6.387-baseline at n ≥ 200; overall dp MAE drifting up. Gap: shadow key not stamped in pair log.
  • wg L3 SKIP_TABLE (v0.6.385 07-28 + v0.6.391 08-04) — SKIPs VALIDATED 08-12, but wg L2 concern surfaced. Reinvestigated 08-12: earlier "4/7 cells regressing" flag was a measurement mix. Today's h_wg_l3_regression_stage1 halves-verified numbers CONFIRM all 6 measurable SKIP cells (L3 vs L2 hurts by 5.9-28.4% on both halves): calm 12-23 +28.4%, calm 24-47 +9.4%, sea_breeze 6-11 +5.9%, sea_breeze 12-23 +7.9%, sea_breeze 24-47 +7.3%, frontal 12-23 +9.2%. The 2 remaining SKIP cells (calm 0-5, ne_flow 6-11) are n<200 — can't measure, no evidence against. Real concern (separate workstream): in the SKIP cells where the stack top-of-stack (L2, since L3 skipped) is still worse than raw L1 (calm 12-23 +23%, calm 24-47 +12%, sea_breeze 6-11 +14%, sea_breeze 24-47 +8%), the culprit is wg L2 (wind_blend), not the SKIP. Removing the SKIP would make it worse. Wg L2 in calm/sea_breeze cells is a new investigation — see project_wg_l2_windblend_cell_concern. 14-day post-ship watch on the SKIP_TABLE itself CLOSED CLEAN.
  • Lc emergency intervention (v0.6.389d-g + v0.6.390, 07-30) — active state, no close date. cl + cc BOTH off Lc via _FIELD_SKIP. cl held out (walk-forward: pool_vs_raw −22.15%, reg_vs_raw −30.37%). cc retired v0.6.390 architecturally → Ccd derives max(cl_l6, cm_l6, ch_l6). cm + ch untouched, +34-50% vs raw held-out. Next step for cl: EMA/Kalman shift tracker OR recent-bias gate. See [[project_lc_regime_conditional]] + [[project_lc_regime_stage1_pool_prereq]].
  • Lc regime-conditional Stage 1 gate — day X/7. 08-08 gate history reset to CHURN after pool-prereq fix demoted 11 false-positive cells (72 → 61 SHIP). See [[project_lc_regime_stage1_pool_prereq]].
  • wsbp — HELD (v0.6.388, 07-28). Sibling of dpbp, calm-only, sign-inverted. ENABLED=False. Preflight 08-04: calm regime n=0 in shadow window. Wait for calm regime accumulation. Cost of waiting: nothing (dpbp covers dp side).
  • l6_fix_b_refit — HOLD-GATE (rolling gate added 08-08). Was single-day SHIP +2.15% today, HOLD at +0.29% 26 days prior; no consistency check between. Rolling gate needs 7 distinct days + same ship_bins before verdict can promote. Lt stays on do-not-reopen. See [[project_l6_fix_b_rolling_gate]].
  • wg persistence-skill thin margin — pooled L4 skill +0.169 (< +0.20 hold margin). Drag at 24-47h. 07-27 ship candidates HELD; re-check post-Stage 1 recovery.
Closed / superseded / inert watches → Archive → Post-ship watches
Details for each cell live in the sections below · Engineering updates · Accuracy · Research & Diagnostics

Recent activity — rolling 3-day window

Recent activity — today + 2 prior days (older entries in docs/CHANGELOG.md, trimmed on next curation)
Badges: PIPELINE · DISCOVERY · INFRA · DASHBOARD · PREFLIGHT

Engineering updates — where we are

✅ Production stack
Core pipeline — general-purpose, can apply to any field
  • L1 raw HRRR / GFS — the starting point every other layer corrects.
  • L2 mesonet blend — additive bias on t / dp / h (Kalman gain) + KBOS+KBVY hourly[0] blend on cc / cl / cm / ch; direct selection on ws / wg. pr re-enabled 2026-08-10 v0.6.401 gated on nw_flow/{0-5h, 6-11h} (2026-07-01 v0.6.276 → 2026-08-10 disabled everywhere: pooled read masked regime-conditional wins).
  • L3 lead-decay — {wg, ch, cm, pp}. ws dropped from L3_FIELDS 2026-08-08 v0.6.397 (walkforward cleared 7-day gate; pooled ws L3 +0.1%/+0.6% is noise after BLEND_HOURS 24→4 fix resolved the long-lead regression). wg L3 uses asymmetric additive (v0.6.366) per (regime, band, fc-quartile) with SKIP_TABLE (v0.6.385 + .391). wg L3 08-11 investigation flag: 4/7 SKIP_TABLE SHIP cells regressing post-ship — see Post-ship watches → wg L3 SKIP_TABLE.
  • L4 diurnal — applies to {ch, cc}.
  • Cards read corrected hourly[0]. See Applicability map for live triggers + gates (e.g. L3 skip cells for ws in ne_flow / short-lead sea_breeze, Lsr skip regimes for sr in ne_flow / calm, sun-up threshold on Lsr).
Specialists — domain-scoped, parallel to the core stack
  • Lsr synoptic-regime solar (sr only). Shipped 06-28 v0.6.248. Per-(regime, hour) bias above sun-up. Skip: ne_flow, calm. Lsr layer section →
  • Lc cloud saturation-unbiasing (cm/ch only; cl + cc BOTH OFF). FLIPPED 2026-07-17 v0.6.355; emergency intervention 07-30 v0.6.389d-g + v0.6.390. cl held out via _FIELD_SKIP after walk-forward showed both shift-table shapes hurt cl; cc retired from Lc v0.6.390 as an architectural correction (Ccd derives cc from components instead). cm and ch untouched, +34-50% vs raw on held-out. See Post-ship watches above + [[project_lc_regime_conditional]] + Lc layer section →
  • Ccd cc from-derivation (specialist, cc only). LIVE 2026-07-30 v0.6.390 · ENABLED=True (flipped 6d ahead of gate). Runs LAST in the cloud pipeline; overwrites hourly.cloud_cover with max(cl_l6, cm_l6, ch_l6) except in SKIP_REGIMES = {"se_flow", "unknown"} (Pirate fallback). Held-out lift +8.48% pooled MAE vs pre-Ccd production (h_cc_derivation.py, 123K quads, halves +11.4%/+5.8%). 08-04 v0.6.391 — saturation guard added: if raw cc ≥ 90, keep raw cc (skip derivation). Traces to the last-24h loss where cm's Lc had bias +26.9 → −10.0 (over-corrected 10 pts), dragging max composite below METAR total sky cover at obs bin 95-100 (raw MAE ~2 → Ccd MAE ~30 across 4 regimes). Ccd still owns the mid + clear tail where corrections help. Retired cc's whole Lc surface — all three cc _CELL_SKIP entries removed; cc added to Lc's _FIELD_SKIP. See Post-ship watches above + [[project_lc_regime_conditional]] Ccd section.
  • chp ch persistence gate. FLIPPED 07-19 v0.6.358 · 07-27 v0.6.382t emergency demote of 6 mid-lead regression cells. Rollup now 14 SHIP + 8 MARGIN + 13 SKIP + 2 THIN. 14-day watch closed 08-02. Full shape refinement gate closed 08-03 · verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09). 08-13 v0.6.405 second emergency demote: 9 additional mid/long-lead cells force-skipped in processor _CELL_SKIP after Prod-vs-L6 gap widened +3.4→+12.7 over 8 days. 14-day watch through 08-27. See [[project_ch_chp_regression_watch_08_13]] + [[project_chp_midlead_regression_watch]].
  • clp cl persistence gate. Shipped 07-24 v0.6.379, ENABLED=False. Stage 2: 12 SHIP / 8 MARGIN / 16 SKIP. Streak walker seeded 08-10 v0.6.401b (fire set n=4, day 3/7); earliest flip 08-16 if 7-day Jaccard ≥ 0.80 clears. Verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09).
  • wg residual persistence gate wg only. Shipped 07-14 v0.6.351, ENABLED=False. 07-27 flip HELD (Stage 1 flipped PROMOTE → MARGINAL 07-26). See What's improving above.
  • wdp wd persistence gate. FLIPPED 2026-07-27 v0.6.382 · 5 SHIP cells. 14-day post-ship watch CLOSED CLEAN 08-11 (Δ −2.7% n=16,416). See Archive → Post-ship watches.
  • dp residual persistence gate dp only. Shipped 2026-07-25 v0.6.380, ENABLED=False. Stage 2: 8 SHIP / 2 MARGIN. Gate closed 08-01 — verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09).
Confidence — orthogonal; widens/narrows bands, doesn't shift values
  • C1 multi-axis confidence layer — gated off (ENABLED=False). Stamps weather_data["confidence"] every tick. Axes: C1a regime transition, C1b cluster spread, C1c pressure tendency, C1f precip-forecast presence, C1e hours-since-front (07-01 v0.6.272), C1d KBOS-vs-KBVY σ (07-08 Stage 3, 13 SHIP + 1 MARGINAL cells), C1h trend direction (v0.6.316 07-10, 15 SHIP cells). Flip after Stage 4 audit clears. Refined view primary since 07-16 v0.6.353e; last read HOLD 36.36% (legacy MIXED, multi-axis NOT READY per c1_stage4_mixture_check). Next re-audit gated on window roll past cl-poisoned days (08-04+). Note: C1h + C1d curated tables ARE read live at stamp time (confidence_layer.py), gated per-cell by co-axis ortho whitelist; downstream user-visible band widening is what "ENABLED=False" refers to. 08-11 Stage 3 wire-up: new by_xr_q sub-table in c1_confidence_premium_v2.json — per-field cross-run spread quintiles (same-vt max−min forecast across runs). 252 cells wired, 214 SHIP (103 WIDEN / 111 NARROW) per c1_curate_confidence_table_v2. 08-12 v0.6.401g — live stamper shipped: weather_collector/processors/cross_run_spread.py reads forecast_log.json (already accumulating per-run L1 via forecast_snapshot.py — no new writer needed) plus this tick's live L1, groups by (field, vt), max−min, buckets into xr_q using the curated edges. Stamps weather_data["cross_run_spread"] = {field: {vt: {spread, xr_q, n}}} for 7 promoted fields (dp, h, pr, t, wd, wg, ws). Verified 13:37 UTC. Consumer-less today by design; tomorrow's ship = wire the by_xr_q marginal multiplier in confidence_layer.py (mirrors C1h / C1d) against a full day of live data. See [[project_cross_run_spread_c1_axis]].
⏸ Built, not applied (gated off)
  • Marine-layer correction sandbox — stamps weather_data["marine_layer_correction"] every tick, ENABLED=False. In-bin cc bias collapsed — real break 06-30, not 07-07 (2026-07-16 diagnostic). marine_layer_collapse_diagnostic.py recomputes fresh per-obs-day (independent of the Fitter's cumulative aggregation): in-bin cc bias +42.7 (06-28) → +16.6 (06-30) → +3.9 (07-05) → never above +5 after. The 07-07 "cliff" reported by marine_layer_anomaly.py was the growing cumulative window finally reflecting the older shift, not the actual break. Split at 07-04 (cm HRRR-anomaly onset): in-bin Δ = −48.6, out-of-bin Δ = −5.0, ratio 9.8× — stratum-local, pre-HRRR, rules out HRRR-anomaly causation. Companion cl signal (marine_layer_cl_stage1 W29 −2.98 vs +12/+22/+17/+13 W25–W28) points seasonal. Hold OFF indefinitely. Will not re-arm when the cm HRRR anomaly clears — different event. Redesign candidate: time-of-year gating on the MLC bin.
🧪 Open architectural questions
  • L2-as-observation-only. Remove L2 from forecast pipeline, keep as training target for L3/L4/Lsr only. No date; architectural exploration.
  • Tight-τ cloud bias propagation across leads 1–3h (Stage 1 candidate). Pearson at lead 1h is +0.50/+0.57/+0.55/+0.66 for cc/cl/cm/ch — strong, clean; degrades by lead 3h. Architectural slot: short-lead τ-decayed bias propagation (~2h half-life) on top of the existing hourly[0] override. Needs Stage 0 MAE-impact estimate (probably small — first 3 leads only). Not blocking; logged for backlog. Spun out from resolved item KBOS+KBVY cloud blend (2026-06-30, ρ ≤ 0 at lead 6h — hourly[0]-only formalized as intentional).
Settled / historical questions (kept as pointers): SKIP_TABLE architecture (shipped v0.6.279, asymmetric extension v0.6.366/370 — see decay_apply.py and Applicability map), Lt Fix B refit (retired 07-13, [[project_lt_fix_b_answered]]), Per-field τ lesson (see [[feedback_tau_streak_gate_limits]]), Scorecard banner + C1e + per-row applied-layer stamping + per-lead Brier for PP (all shipped 07-01), dp L4 skip-table case (resolved 06-30). Full details in git log + CHANGELOG.
Convention: L = correction layer (applied); R = research hypothesis (logged, not applied); G/S/B/F = operational tools in Research section.

Forecast accuracy — how accurate is it?

Two lenses on the same question. First (Accuracy over time, below): the per-obs-day trajectory of Raw vs Production — is the stack drifting? Did a recent ship move the needle? Second (per-field tables, further down): the current-window breakdown by lead band per field — where in the 0-47h horizon is each correction layer helping or hurting? The over-time view is the drift detector; the per-band tables are the shipping-decision granularity. Production in both views IS a real per-row aggregate. Per-band tables: per_layer_mae_by_lead[field].production where the per-lead sample count clears n≥30, keyed on applied_layer stamps. Over-time chart: mae_over_time.json's "prod_real" series, bucketed per obs-day from the same applied_layer stamps (v0.6.371). Both surface the same "what users actually saw" quantity, just at different aggregation grains. Specialist lines (Lsr, Lc, chp, clp) remain visible on the over-time chart as intermediate series showing gate-fired-only MAE — smaller sample than Raw/L2/L3/Prod, useful for isolating each specialist's per-lead lift during 14-day watches.

How we measure whether the forecast is good — the metric framework

Core comparison: pipeline error vs. raw HRRR/GFS error on the same forecast-observation pairs. Every metric on this page averages those errors differently.

"Observed" per field (best local measurement, not absolute truth):

  • t, dp, h, ws, wg, pr: mesonet Kalman blend across 60 WU + 26 Tempest stations within ~2.5 mi, station-bias corrected.
  • cc, cl, cm, ch: mean of KBOS + KBVY METAR sky reports (octa → percent). Coarser than mesonet.
  • sr: median across valid Tempest solar_radiation_wm2.
  • pa: max WU precip_rate_in across stations (patchy-rain-aware).
  • pp: binary — any measurable rain this hour (0 or 100).
  • wd: station wind_direction, circular error.

Three metrics per NWS/ECMWF conventions:

  • MAE — mean |error|. "Typical-day miss." Historical page text usually means this.
  • RMSE — root mean squared error. Weights big misses harder; if RMSE improvement lags MAE, pipeline has hidden blow-ups.
  • Bias — signed mean error. Positive = over-forecast; negative = under. Catches systematic drift MAE masks.

pp uses Brier score, not MAE (binary obs vs. probability forecast). Brier = mean of (fc/100 − obs/100)². Same "% erased" framing.

Measurement gaps:

  • Persistence skill — shipped 2026-07-11 (h_persistence_skill.py). Baseline: 5 ADD VALUE (t/h/pr/ws/sr), 4 MIXED (dp/wg/cc/pp), 3 NO SKILL — cl/cm/ch lose to persistence at every band despite L3+L4. Scorecard integration shipped 07-12 v0.6.328 — the "vs Persistence" line in the scorecard banner (above) is live-updated from persistence_skill.json. Response for ch: ch persistence gate FLIPPED LIVE 07-19 v0.6.358 (14 SHIP cells post-emergency-demote 07-27 v0.6.382t (was 27) on refreshed windows; 14-day watch CLOSED CLEAN 08-02). Landmark answered 07-14 v0.6.351a: persist_only ties gate pooled but gate wins halves-hedge; keep the gate. cl persistence gate wired Stage 3 07-24 v0.6.379 (replaces retired cl_persistence_short_lead — narrow 0-5h hypothesis disproven by halves-verified Stage 2). 08-09 Stage 2 rerun — 6 SHIP / 22 SKIP whitelist shape ready for flip decision. SHIP cells: calm/0-5 (−35.7%), ne_flow/0-5 (−54.4%), ne_flow/12-23 (−16.6%), nw_flow/0-5 (−18.7%), pre_frontal/0-5 (−17.8%), sea_breeze/0-5 (−32.9%). Flip = cl_persistence_gate.ENABLED=True. 08-09 c1_stage4_mixture_check: cl/6-11h [transition] b3 no longer in DEGRADED list (07-29 escalation window CLOSED CLEAN, never reached 3× ≥+200%); cm/12-23h [transition] IMPROVED; t/24-47h [transition] newly DEGRADED (unrelated). Joint cl/cm correction hypothesis rejected 07-29 (cl has no correction stack; cm corrections are stable improvers; drift is raw-HRRR-side on top-fc transition rows and fitter absorbs via c1 confidence bands).
  • Climatology skill — long-lead reference. Needs a climatology dataset.
  • pp Brier reliability decomposition — does "30% chance" actually happen 30% of the time? Aggregate Brier only, no decomposition.

Historical wording caveat: section text written pre-v0.6.325 (2026-07-10) treats MAE as the only measure. Read alongside RMSE + bias — disagreements usually mean occasional big misses or systematic drift.

📈 Accuracy over time — Raw vs Production, per-day trajectory
Per-obs-day rollup. Chart shows the trajectory of Raw (bare HRRR/GFS) and every correction layer applicable to the selected field — legend adapts per field so noise-lines that don't do work for that field are hidden (e.g. sr shows Raw / Lsr; pa shows Raw / Prod only). Rolling 7-day mean overlays on Raw and the final applied layer smooth out day-to-day noise. Ship-date vertical annotations mark when live-layer changes landed. Source: mae_over_time.json — accumulating history, x-axis grows one day at a time (retention-independent; the pair log is capped at 30 days but this rollup persists prior days). Refreshed hourly by the myweather-publisher Cloud Function (v0.6.395f, hourly cron); prior to that it was tied to the daily digest. (v0.6.361: per-specialist attribution — Lsr, Lc, and post-Lc specialists (ch-persist, cl-persist) each plotted separately once ≥3 days of data have accumulated. New collector-side pair-log columns start accruing today; specialist lines will appear on the chart around 2026-07-22. v0.6.371: the 'Prod' line is now the real per-row aggregate keyed on applied_layer stamps — sample-comparable to Raw/L2/L3 on the same chart. Pre-v0.6.371 it was L4's daily MAE for non-specialist fields and a specialist line for cc/cl/cm/ch/sr (gate-fired-only, sample-mismatched).)
Detail view for selected field (Raw / L2 / L3 / Prod + rolling means + ship annotations):
Scan all fields (click any panel to focus the detail chart above):
📊 Per-field breakdown by lead band — MAE / RMSE / bias per correction layer

The bright white "Production" column is the forecast users actually see. Lower is better. For each field, one combined table shows how far off the forecast tends to be at each lead band — MAE (typical error), RMSE (occasional big misses that MAE averages away), bias (systematic drift, signed) — per correction layer. Compare a MAE row across columns to see where each layer helps or hurts.

How to read these tables
What you're looking at. One card per field. Each card shows a table with rows grouped by lead band (0-5h, 6-11h, 12-23h, 24-47h, ALL). Under each band, three metric rows: MAE (primary), RMSE (secondary), bias (secondary, signed). Columns are the correction layers actually applied to that field, plus the Production output on the right.
Only-applied layers. Layers that aren't applied to a field don't render as columns because their per-lead MAE array equals the previous applied layer's — noise. The badges above the table (L2 ✓, L3 ✓, L4 off, etc.) confirm the applied set. If Production isn't the lowest MAE, check the Applicability map for a gate that's wrong.
Per-metric usage:
  • MAE — average absolute error. The primary metric; drives L3/L4 whitelist decisions. Best cell in row highlighted green, worst red.
  • RMSE — squared-error root. Higher than MAE by a factor set by how tail-heavy the errors are. Watch for a band where RMSE jumps proportionally more than MAE — that's occasional big misses hiding behind an OK average.
  • bias — mean signed error (forecast − obs). Near-zero = calibrated on average. Persistent + or − by band signals systematic drift.
Layer labels:
  • Raw: bare HRRR/GFS forecast for this coordinate. Knows nothing about Wyman Cove.
  • Aggregate bias (L2): what 40+ nearby weather stations say the model is currently getting wrong, blended in by distance.
  • Lead decay (L3): historical per-lead bias correction from the pair log.
  • Diurnal (L4): hour-of-day bias correction. Final line for t / dp / h / ws / wg / pp.
  • Synoptic-regime (Lsr): per-regime W/m² delta on direct solar radiation. Only present on sr.
  • Cloud saturation (Lc): per-(field, value_bin) percentage-point shift, clamped to [0, 100]. Final line for cm / ch (unchanged); cl and cc both fully off as of 07-30 (v0.6.389f cl · v0.6.390 cc). cc now derived downstream via Ccd = max(cl_l6, cm_l6, ch_l6). FLIPPED 2026-07-17 v0.6.355; emergency intervention 07-30 — see Lc layer section for details.
L2 badge variants: L2 ✓ additive = bias added (t, dp, h, pr). L2 ✓ direct = station median replaces model wind (ws, wg). L2 n/a = no station network reports this field (clouds, solar, precip). Note for clouds: L2 is n/a as a forecast correction, but obs truth for cc/cl/cm/ch comes from a KBOS + KBVY METAR blend (v0.6.134) — feeds the joiner so L3/L4 can be evaluated, not applied at L2.
Bands match the walk-forward validator buckets — this is exactly the granularity that drives shipping decisions. Source: time_series_diagnostic.json::per_layer_{mae,rmse,bias}_by_lead, 7-day window. Design note (v0.6.350): previously carried a per-card chart + 3 separate tables. The chart mostly restated what the MAE band table already showed; killed for signal density.

Applicability map — what corrections trigger, why, and when they actually fire

Two lenses on the same object. Applicability (top) describes what's configured to fire and under what gates — built each tick from describe_applicability() in each correction module, so the page can never drift from the code. Runtime firing frequency (bottom, 7-day rolling) shows what's actually firing per operator × field × regime, so silent dormancy (config says "enabled" but the code path never mutates a value — the class of bug that hid ws L3 for 4 days after v0.6.279) surfaces as an ★ on cells with 0 fires despite ≥5 ticks in that regime.

How to read this
Three categories: General-purpose layers (L1–L4) can apply to any field; the per-field applicability set controls which actually trigger. Specialists (Lsr, Lt, Lc) are domain-scoped by construction — the physics of the correction binds them to a field type. Confidence (C1) doesn't modify forecast values; it widens or narrows uncertainty bands along orthogonal axes (transition, pressure tendency, mesonet spread, pre-frontal proximity) that cross multiple (field, band) cells.
applies when is the predicate in plain text. gated by names the module-level constant the gate reads (omitted if always-on). current state resolves the constant for the reader — what the gate is doing this very tick.
Terminology key: applies = correction applied to a (field, row) this tick · enabled = feature switch (the module-level constant is on) · active = runtime branch (which code path actually fired).
Source: weather_data.applicability_map.layers. Schema: weather_collector/data/applicability_map_schema.json. If this section says "not available," the collector hasn't shipped the block yet — it landed v0.6.260.
L2 — Aggregate bias (mesonet blend) general-purpose hand-curated · promote to describe_applicability()
field applies when gated by current state
t Always — additive bias from station network, scaled by Kalman gain K always on applies
dp Always — additive bias from station network always on applies
h Always — additive bias scaled by Kalman gain K always on (K-taper 1.0 → 0.4 by lead 24h) applies
pr RE-ENABLED 2026-08-10 v0.6.401 with regime gate — K=1 additive, τ=8h fitted, applied on (regime, lead_band) ∈ {(nw_flow, 0-5h), (nw_flow, 6-11h)}; every other cell stays raw. Cell selection from analysis/pr_l2_regime_lead_retro.py 08-10 (Jaccard 0.50 STAGE 1 SHIP, both-halves winners A +21.8%/B +41.6% and A +10.3%/B +13.1% on 8,596 shadow rows). History: DISABLED 2026-07-01 v0.6.276 after pooled Production +2.4% BAD — but that pooled read was regime-blind. The nw_flow short-lead win was real all along; needed regime-conditional cross-cut + halves verification to surface. Shadow-wire since 2026-07-29 v0.6.389 supplied the fresh data. regime-gated (current-tick regime × lead band) applies on nw_flow 0-11h; other cells raw
cc Always — Kalman-blended KBOS + KBVY METAR override on hourly[0] (current-hour value used by cards). Not propagated across forecast leads — promotion to a true per-lead L2 bias correction is queued (see Open architectural questions). always on (needs KBOS or KBVY at current hour) applies
cl / cm / ch Derived — same Kalman K applied to L/M/H splits on hourly[0] to keep them self-consistent with cc. Not propagated across leads, same as cc. always on (derives from cc) applies
ws / wg Always — direct selection from per-octant median (no per-station bias track); KBOS+KBVY authoritative-source floor when both agree at >1.4× the octant median always on (direct-selection, not additive) applies (direct)
sr / pp / pa N/A — no station network for this field; L2 row is structurally empty n/a — no obs network n/a
wd AlwaysSHIPPED 2026-07-20 v0.6.368a. Circular unit-vector blend in wind_blend.py: obs wd + fc wd → (sin, cos), weighted mean by linear decay over 24 leads, atan2 back to degrees. Solves the wrap-around that made "N/A — L2's linear math doesn't apply" true prior to this ship. Calm-floor guard WIND_DIR_MIN_SPEED = 3.0 mph skips cells where both obs and fc speed are below floor. raw_wind_direction preserved for baseline. Complementary wd persistence gate (Stage 2 07-20 v0.6.365) targets long-lead regime-transition cells where L2 decay has expired. always on (needs both obs+fc speed ≥3 mph) applies (circular)
Runtime firing frequency — 7-day rolling
Fires = the correction actually mutated a value at that lead (not "would have applied"). Skips = would-fire cells suppressed by the skip table (L3 ws in ne_flow / short-lead sea_breeze; Lsr sr in ne_flow / calm) or by a gate being OFF (MLC currently ENABLED=False → all matching cells count as skips). Rate/tick = fires ÷ ticks_in_regime; for L3/L4 with 48 leads, healthy is close to 48. marks (operator × field × regime) cells with 0 fires despite ≥5 ticks in that regime — silent-dormancy candidates.
Source: gate_firing_rollup.json (nightly digest via analysis/gate_firing_rollup.py) over per-tick gate_firing_log.jsonl (written by weather_collector/processors/gate_firing_log.py, v0.6.318).

L1 — Raw model (HRRR / GFS)

The bare government weather model. Knows nothing about Wyman Cove specifically. Everything below corrects what it gets wrong here.

About the raw model The starting point. Open-Meteo's HRRR (next 48h) and GFS (days 3–7) numerical weather models. Multi-kilometer grid resolution; the model has no specific knowledge of Wyman Cove. Every layer below corrects what the raw model gets wrong locally.
Curves below show the current raw model forecast per field.

L2 — Aggregate-bias correction (local station network)

What our 66 nearby weather stations say the model is getting wrong right now — and how much of that signal we trust.

How the network correction is built Two parallel aggregation paths under this layer, each suited to the noise behavior of its metric:
Temp / humidity / pressure (additive bias): per-station Kalman offset (rolling 48h) → per-octant 1/distance² × exp(-|elev_diff|/30)-weighted (station − model) bias → unweighted mean across non-empty octants. t and h scale by per-field network Kalman K (sec 2c); pr applies full strength.
Lead-decay: bias_applied(lead) = current_bias × exp(-lead/τ) per field. Fitter (03/15 local) refits τ on train/test split; loader adopts fitted τ only if it beat hardcoded default on held-out RMSE AND stays within 0.25×–4× default. Defaults: τ_t=4h, τ_h=240h, τ_pr=12h. Fields without τ (wind/clouds/solar) get flat L2. See sec 2d for live curves.
Wind / gust (direct selection): no per-station calibration, no additive bias. Per-octant MAX gust → MEDIAN across octants; linear-decay blend into next 24h (100% h0 → 0% h24). KBOS+KBVY floor: if both airports agree on speed >1.4× octant-median, defer to their median. Direction rejected if >60° off airport+buoy+Tempest consensus. Gust override allows single-source (METAR omits when steady). Physical floor: gust ≥ wind.
Cloud cover (Kalman METAR blend): KBOS+KBVY sky obs (BKN/SCT/FEW/OVC → percent + L/M/H) blended against HRRR hourly[0] via _kalman_gain_cloud. K=0.90 if airports agree within 20pp, 0.70 at 20-40pp, 0.50 at >40pp, 0.35 for single-source. Same K on total + L/M/H for self-consistency. Live values in cloud_l2_meta.

2a. Octant coverage — where this tick's stations came from

loading…
Per-station detail — mesonet map & Kalman-tracked offsets table
Per-station uptime — fetch success rates (rolling window)

2b. Network bias estimate (full, un-confidence-scaled)

loading…

2c. Network confidence (Kalman gain K)

loading…

2d. Lead-decay applied to L2 bias (v0.6.44)

loading…

2e. Post-aggregate-bias forecast — what gets passed to L3 · engineering view (pre-clamp)

⚠ These are internal pipeline values, not user forecasts. After the aggregate-bias correction is applied but before downstream layers and physical-bounds clamping (FIELD_BOUNDS in decay_apply.py). Values can legitimately fall outside physical ranges here — cloud_cover 121%, precip_probability −6%, precip_amount −0.025 in — because the bias offset is additive and the clamp comes downstream. If any of these look wrong for user display, check the L3 / L4 / cloud-clamp path, not this section.

L3 — Lead-decay correction

How the model's error tends to grow with each hour into the future — and the per-field nudge we apply to push back against that drift.

How decay correction works For each field and each lead hour, we learn from millions of historical pairs how far off the model usually is — then bend the forecast back toward truth by that amount. Some fields (wind, gusts, high & mid cloud, POP) genuinely benefit; others net-negative on held-out data and are paused (see Applicability below). Tracked over a 30-day window with an exponential recency weighting — default τ=14 days, with per-field overrides for fields where analysis/decay_tau_tuning.py shows ≥5% MAE improvement vs the default. Current overrides: pp (POP) at τ=28d (+11.1% held-out, 2026-06-21 v0.6.167; note: today's read shows only +3.2% vs τ=14, now below the 5% floor — flagged for next τ-audit day). pa (precip amount) reverted 2026-07-19 v0.6.358 after today's read landed +0.9% vs τ=14 (was +5.9% yesterday); pa's best-τ swung 28→42→7 across three IMPLEMENT reads. See [[feedback_tau_streak_gate_limits]].

Applicability

Held-out MAE audit picks the fields where L3 actually beats L2 — currently wind speed, gusts, high cloud, mid cloud. Other fields stay off because the correction at best ties (temperature, pressure) or actively hurts (humidity, dew point, solar, low cloud, precip). POP is the special case: it's evaluated by Brier score, not MAE, so the audit's MAE-based ⚠ rule is suppressed for it — the v0.6.20 calibration analysis showed flat-additive correction cuts Brier 5%. Currently applied: . Brier-evaluated: . See Applicability map for live triggers and per-field gates.

Live state — with vs without decay correction

Historical calibration — fitted correction curves per lead hour

Calibration history — decay curves over time

L4 — Diurnal correction

A separate correction for the part of model error that follows the sun — e.g., bias that's different at 3 AM than at 3 PM. Currently applied to high cloud and cloud cover.

How diurnal correction works Bins historical errors by hour-of-day (0–23) and fits the persistent pattern at each hour. Most fields don't have a clean diurnal signal once L2/L3 have run, so this layer's applicability set is narrow — currently {ch, cc} (cc added 2026-06-24 v0.6.214). Fits for fields outside the applicability set are still computed below for diagnostic purposes. See the Applicability map for the live trigger state.

Applicability

L4 corrects the portion of forecast error that repeats with time of day. It is the hardest layer to earn because the same hour-of-day bias must recur consistently across many days. Most fields fail that test because their dominant errors are driven by changing weather regimes (air mass, cloud regime, frontal timing, etc.) rather than the clock. Currently applied to {ch, cc} — the two fields showing a stable enough diurnal signal to pass the held-out audit. Fits for the remaining fields are retained in the calibration sections below as diagnostics. See Applicability map for live gates.

Live state — what L4 is doing this tick

No dedicated live-state panel yet. Current per-tick L4 deltas land in weather_data.hourly[*].corrected_<field>; the applied-vs-unapplied comparison is visible via the L4 lines on the per-field accuracy charts at the top of the page. Future build-out — placeholder for a dedicated live snapshot mirroring L3 §3.

Historical calibration — diurnal correction curves over time

Calibration history

L4 doesn't currently maintain a separate fit-history view — the curves-over-time panel above serves both "current fit" and "how it's changed" roles. Split out if/when a dedicated current-fit snapshot ships.

Lsr — Synoptic-regime correction (solar specialist)

A per-regime W/m² delta applied to direct solar radiation. The classifier reads current wind direction, speed, pressure trend, hour, and temp; the regime label keys into a calibrated per-regime delta. L1–L4 are trained on general bias correction; Lsr is the first layer trained on synoptic-state stratification — different "kinds of weather" get different corrections. Skip regimes (v0.6.280): ne_flow and calm return 0.0 — l5_solar_analysis showed Lsr hurt sr in those regimes (+32% and +10% worse) while helping everywhere else. Skip regimes are a targeted patch until Fix B / refit against L4 baseline lands.

How Lsr works — Classify → Lookup → Apply
1. Classify. regime_classifier.classify_synoptic_regime(wind_dir, wind_speed, pressure_in, pressure_trend_hpa_3h, hour_local, temp) returns one of nine labels — nw_flow, ne_flow, sw_flow, se_flow, sea_breeze, frontal, pre_frontal, nor_easter, calm (plus unknown when inputs are missing).
2. Lookup. Per-(regime, hour) Δ W/m² from _BIAS_BY_REGIME_HOUR[regime][hour_local] in solar_correction.py; falls back to _BIAS_FALLBACK_BY_REGIME[regime] when the hour cell has too few samples.
3. Apply. Added to every lead's direct_radiation where the lead's raw value is above SUN_UP_THRESHOLD (50 W/m²). Pre-sunrise / fully overcast leads get Δ = 0 regardless of regime.

Applicability

Lsr is a specialist — it only applies to direct solar radiation (sr). Other fields have no Lsr line on their accuracy charts by construction. See Applicability map for the live trigger predicate (ENABLED in solar_correction.py + the sun-up gate + skip regimes ne_flow and calm since v0.6.280).

Live state — what Lsr is doing this tick

Loading…
Read directly via curl -s https://data.wymancove.com/weather_data.json | jq .solar_correction. The per-lead Δ is also baked into the sr card on the Forecast Accuracy chart above.

Engineering status

Where we are (2026-07-28):
Lsr lives in production for sr. Skip regimes: ne_flow + calm. Structural raw-baseline verifier (v0.6.291) catches raw-column mutation drift automatically. Unit-mismatch resolution — Lsb Stage 3 wired 07-17 v0.6.354, ENABLED=False. Overrides Lsr's direct-beam output with (shortwave − bias(hod)) on sea_breeze rows where cc < 25 — narrowed 2026-07-28 v0.6.383b from the original two-sided (cc < 25) OR (cc >= 75) after 07-24 halves re-run showed the overcast half (75-100 cc-bin) actively regressed (Δ=−17.6%, halves +6%/−23%) while the clear-sky half held robustly (Δ=+35.8%, halves +33%/+40%). The "thick attenuation missed" hypothesis for overcast was wrong; data says model attenuation is fine or over-corrected. Middle-and-overcast cc bins (≥25) now all fall back to Lsr's original direct-beam correction. Narrowed-shape Stage 2 re-run (07-28) PROMOTES cleanly: pooled +31.50%, halves +43.78% / +22.69% (both above +10% ship gate), 4/4 lead-bands SHIP. Landed ENABLED=False with fresh 7-day live-layer gate 07-28 → 08-04. Broader unit-mismatch across other regimes (pre_frontal +12, unknown +367 W/m²) still open. Shortwave shadow-log infrastructure stays alive (underpins future sr L2 work once unit mismatch resolves per project_sr_unit_mismatch). Divergence-report LSR_ENABLED source bug fixed 07-20 v0.6.365 (was reading l5_solar_analysis, now routes through .cache_l5_gate_history.json) — status now correctly AGREE, live gate 100% SHIP for 30+ days.
Full Lsr shipped-history log (v0.6.248 → v0.6.291)

Developer notes — classifier, lookup tables, audit details

Regime classifier. regime_classifier.classify_synoptic_regime(...) returns one of nine labels; same source feeds C1a's transition axis. Lsr reads via state.regime_synoptic. Live label surfaces in Production Stack box, R2.
Per-regime delta table — what the lookup contains. Lsr's bias tables live in weather_collector/processors/solar_correction.py as two Python dicts:
Refit cadence: regenerate via python3 analysis/l5_recompute_biases_hourly.py after at least 7 days of accumulated DAYTIME pair rows (raw_solar ≥ 50 W/m²). The daily digest prints a drop-in replacement table; mid-trajectory refits invalidate the audit window — wait for a clean break.
Lsr vs L4 audit — held-out MAE. Primary view: the Forecast Accuracy chart above. The sr card carries five lines (Raw → L2 → L3 → L4 → Lsr); the green box on the Lsr line is the user-visible MAE on rows where Lsr fired. Note: on rows where Lsr's skip regimes fire (ne_flow, calm since v0.6.280) Lsr returns 0.0, so those rows contribute L4-value MAE to the Lsr aggregate — the Lsr line converges toward L4 as skip-regime coverage grows.
Fitter audit: verdict logged to conditional_audits.l5 each Fitter cycle, recency-weighted since v0.6.178 (exp(−age_days/14)). The trailing 7-day rolling gate lives in l5_gate_history.json and surfaces beneath the Lsr row in the S1 Shadow Tuner section below. First actually clean 7-day window closes ~2026-07-10 (7 days from the 07-03 per-lead delta deploy — the earlier ~07-05 date was against the pre-fix deploy timeline).
RESOLVED 07-20 v0.6.365: the divergence report's LSR_ENABLED claim was sourced from l5_solar_analysis — a candidate script testing a simpler regime-only bias lookup, NOT the live hourly Lsr. Its HOLD verdict meant "don't ship the candidate refinement," but the divergence table was rendering it as "retire live Lsr." Live Fitter cycle gate history (l5_gate_history.json) shows 100% SHIP across 30+ days, 13-21% MAE improvement per cycle. Divergence report now routes LSR_ENABLED through the live gate history and correctly shows AGREE.

Lc — Cloud saturation-unbiasing (cloud specialist)

A per-(field, value_bin) percentage-point shift applied to the L4-corrected forecast for cc / cl / cm / ch. The bias pattern Lc unwinds is saturation: the models systematically overshoot at 0-5% cloud and undershoot at high fractions. The same shape holds across all four cloud fields but with different magnitudes. Lc reads each L4-corrected value, keys into a curated (field, value_bin) → shift table where the fitter's SHIP verdict is set, adds the shift, and clamps to [0, 100]. Emergency intervention 2026-07-30 v0.6.389d-g — pooled shift table diverged from recent regime-conditional truth, causing 8× cl MAE blow-up on 07-30. Live SHIP surface after intervention: cl fully off (_FIELD_SKIP); cc ships at 0-5 (all regimes), 50-80 + 80-95 (all regimes except ne_flow), 95-100 KILLED universally; cm unchanged (20-50 / 50-80 / 80-95 / 95-100, all regimes); ch unchanged (20-50 / 50-80 / 80-95, all regimes). SKIP cells (bins that stayed within the ±5 pp noise floor across the 30-day window) pass through unchanged. See Engineering status below for the walk-forward evidence that motivated the interventions.

How Lc works — Preserve → Lookup → Shift → Clamp
1. Preserve. Before mutation, stash the L4-corrected array as hourly.<field>_post_l4. Pair log + debug page read this to attribute L4 vs (L4+Lc) cleanly.
2. Lookup. For each lead 0-47, determine the value bin from [0, 5, 20, 50, 80, 95, 100]. If the (field, bin) cell has SHIP verdict in lc_correction_table.json, read its Δ pp; else Δ = 0 (SKIP — the L4 value passes through unchanged).
3. Apply. Add Δ to the value and clamp to [0, 100]. Writes back to hourly.<field>.
Sign convention: the fitter's stored bias is (forecast − observed); the applied shift is −bias, pulling the forecast toward the observation.

Applicability

Lc is a specialist — applies only to cc / cl / cm / ch. Fires when the L4-corrected forecast falls in a SHIP-verdict value bin (16 of 24 cells; see Applicability map for the full per-cell shift table + verdict list).

Live state — what Lc is doing this tick

Loading…
Read directly via curl -s https://data.wymancove.com/weather_data.json | jq .cloud_saturation_correction. The per-lead Δ is also baked into cc / cl / cm / ch cards on the Forecast Accuracy chart above.

Engineering status

Where we are (2026-07-30 — emergency intervention day):
Overall Prod-vs-Raw compressed from ~−13% a week ago to −5.2% today, driven by cl and cc l6 MAE blowing up. On 2026-07-30: raw cl MAE 7.29 (model + obs both mostly-clear) but Lc-corrected cl MAE 56.96 (8× worse). Same on cc (raw 7.21 → l6 44.23). Diagnosed via 4 new analysis scripts (Stage 0 regime × bin sweep, Stage 1 halves-strict fit, walk-forward validator, rolling-window sweep) as an architectural failure of the shift-table itself: the historical fit table subtracts 46-88pp at overcast bins because model USED to over-forecast overcast heavily, but the model no longer does — bias shrunk 4-38× or sign-flipped in the last 3-10 days. No rolling window length recovers cl on held-out (best W=3d still −3.7% vs raw). Regime-conditional slicing doesn't fix it either (walk-forward: cl reg_vs_raw −30.4%). Emergency bandages shipped v0.6.389d-g: cl fully off Lc, cc/95-100 universal bin-skip, cc/ne_flow/50-80 + cc/ne_flow/80-95 regime-conditional demote. cm and ch untouched (walk-forward +34% and +50% vs raw on held-out — still helping). Original 14-day post-ship watch (07-17 → 07-31) cannot close routinely. New tracking axis is the regime-conditional Lc Stage 1 gate (7-day walk-forward stability starting today) — but even that's superseded by the walk-forward validator's finding that the naive regime-conditional shape isn't the fix. Architectural next step (unshipped): EMA/Kalman shift tracker OR recent-bias gate on the existing lc_fit table. Both multi-day. Contingency: per-bandage reversibility is a single-line edit to _FIELD_SKIP or _CELL_SKIP frozenset in cloud_saturation_correction.py. See [[project_lc_regime_conditional]] for the full pipeline state.
Prior state (2026-07-17):
FLIPPED 2026-07-17 v0.6.355 after 8/7-day gate clear, 16 SHIP cells stable 7 consecutive days (07-11 → 07-17), LC_ENABLED READY on divergence report, and no cc/cl/cm/ch ANOMALY in the pair-log anomaly detector. First live tick (15:27 EDT): 113 cells fired — cc 46/48, ch 39/48, cm 19/48, cl 9/48; mean |Δ| in the expected 28-42 pp range per field. Predicted biggest Prod MAE lifts: cl 80-95 −55%, cl 95-100 −47%, ch 50-80 −37%. Two-gates-per-layer cross-check via [[feedback_two_gates_per_layer]]. 07-31 clean close superseded by the 07-30 intervention above.
Full Lc shipped-history log (v0.6.298 → v0.6.355)

Developer notes — fit table, refit cadence

Fit table location. weather_collector/data/lc_correction_table.json. Structure: {field: {value_bin: {shift, verdict, n_samples, mae_pre, mae_post, delta_pct}}}. Loaded at import time in cloud_saturation_correction.py via _load_correction_table().
Refit cadence. analysis/lc_fit.py runs in the daily digest and rewrites the JSON in-place. SHIP-set stability (7 consecutive days with the same SHIP cells + no HOLD days) is the fitter's own gate; it lives in .cache_lc_gate_history.json.
Value-bin edges. [0, 5, 20, 50, 80, 95, 100] percentage-points — six bins per field × four fields = 24 cells. Bin 5-20 is currently SKIP for all four fields (mid-low cloud has near-zero systematic bias).

Research & Diagnostics — experimental signals + audit views (not applied to live forecast)

Diagnostics — audit live behavior

R0. Per-layer audit — is each layer + specialist earning its keep?
Held-out MAE per field per layer (leads 1–47), refit every Fitter cycle. Dim subtext = signed bias. Δ columns compare to layer below: green = beats + applied; amber = beats but NOT applied; red = loses; gray = tied. Banners fire when an enabled layer or specialist loses >3% (hidden regression) or a disabled specialist / layer wins >3% (missed opportunity / post-ship watch signal). Specialists column (v0.6.382n) chains Lsr → Lc → chp / clp / wdp for the fields each owns; each rung compares to the previous rung (or L4 if first). Production column shows the real per-row Production MAE (aggregated from applied_layer stamps). MAE-only caveat: RMSE tells a different story on L2-additive fields (wg Production improvement drops −33%→−26%, dp −17%→−13%, h −7%→−3%). Per-layer RMSE+bias column rewrite queued.
F1. Frontal passage log — detector instrumentation (last 14 days)
Live readout of detected frontal passages from frontal_events_log.json. Detector runs every tick in frontal_detection.py: passage fires when ≥2 of 3 signals hit (dp drop >8°F, wd shift >60°, pressure inflection with ≥0.02″ rise). Confidence 67% with 2 signals, 100% with 3. Sanity-check against real fronts before letting the briefing AI rely on the cause-attribution line. Live consumers: PWA front-passage card + C1e confidence axis + 10 analysis scripts joining on passage timestamps.
Loading detected passages...
R2. State-stratified accuracy — which regimes does the model fail in?
Per-field MAE + bias by regime. Big MAE spread across bins = regime-aware correction candidate. Addressed rows excluded from top-10 (shipped corrections absorb raw signal downstream); excluded rows in collapsible below. Temperature back in the top-10 since Lt retired 07-13. Refit twice daily; published to state_stratified_accuracy.json. MAE-only; RMSE would rerank L2-additive fields (dp/h/ws/wg).

Tools — evaluate candidates before promotion

S1. Shadow applicability tuner — what would auto-tuner have chosen?
Per Fitter cycle, log which fields a naive MAE auto-tuner would include in L3_FIELDS/L4_FIELDS, alongside production. Recommend ON if layer beats layer-below by ≥3% in any band AND bias no worse. Field-membership only — doesn't reason about per-field gates or skip tables (those in Applicability map). Precondition for automation is agreement after 90+ days; mismatches informative, not actionable. Also surfaces the live conditional_audits.r6 C1a-transition verdict per Fitter cycle.
Loading…
B1. Backtest sweep — alternative L3/L4 configs vs production
A/B comparison of candidate L3/L4 applicability configs vs production, computed by replaying the held-out pair log. Current production: L3 = {ws, wg, ch, cm, pp}, L4 = {ch, cc}. Run via python3 -m backtest.sweep --write-gcs (add --local-file ~/.cache/myweather/forecast_error_log.jsonl for fast iteration).
Loading sweep results...

Candidates — in-flight hypotheses (Stage 3 = wired ENABLED=False by definition)

C1 confidence stack — status
7 axes stamping every tick, applied=False until Stage 4 clears. Curated v3: 296 SHIP / 42 MARGINAL / 1048 SKIP across 39 axis-keys. C1a shipped as regime-transition axis (was R6; verdict logged under conditional_audits.r6 per Fitter cycle, surfaced via S1). Latest audit: HOLD 36.36% (8 CALIBRATED / 14 DRIFTED, 2 BRIER_EXEMPT). Prior cl "DEGRADED" reading was driven by the applied_layer poison fixed 08-02; next re-audit gated on window roll past cl-poisoned days (08-04+). Mixture-normalized view: only 2/112 cells are REAL DRIFT (per c1_stage4_difficulty_lens); legacy FAIL count is inflated by weather-mixture shift. Narrow-promote counters (today): pre-frontal 6/7 (2 SHIP, 1 to go). C1h + C1d auto-suppressed as KNOWN_LIVE_PIPELINES (already live-stamping). Ship gated on Stage 4. Per-cell co-axis ortho gate v0.6.321 in confidence_layer.py: cl fires freely, cc/cm/ch conditionally suppressed by the non-ortho co-axis, ch 24-47h + t × 3 never fire (REDUND both). Individual axes tracked in Backlog Group A below.
Backlog — Stage 0-2 hypotheses (pre-wire pipeline)
Stage 0-2 hypotheses (pre-wire). Promotion path: 0 exploration → 1 curated finding → 2 per-cell verification → 3 wired ENABLED=False → 4 shipped. Stage 3+ detail lives in Post-ship watches / layer sections, not here. Most surviving hypotheses measure forecast uncertainty (C1 axes), not forecast bias. Real bias candidates (marine layer, radiational cooling) overlap heavily with L2 and L4.

Group A — C1 multi-axis confidence extension

Individual axes. Five join the multi-axis Stage 4 calibration (C1a/b/c/f/e); C1h + C1d compose as marginal-premium tables (kept off the join to avoid cell-dilution). Stack-level status in C1 confidence stack card above.

Group B — Bias candidates (paced, individually)

Group C — Lower priority (dominated by existing layers)

In-flight candidates + current stage counters live in Current state → What's improving at the top of the page (single source of truth).

Group D — Methodological refinements (modify existing layers, not new ones)

Promotion rule: single-shot script in analysis/. Stage 2 verdict must hold across ≥2 reads spaced 3+ days apart. Group A → C1 axes (C1a/b/c/...); Group B → R-numbers if they clear the 7-window walk-forward gate.

Experiments — open design seeds + live telemetry probes

E1. Stage 0 explorations — open design seeds + data-limitation flags
Only genuinely-open items live here. Promoted (→ Stage 1+), killed by orthogonality, and settled-null items live in Archive — single source of truth.
E2. Lt — Cove microclimate telemetry probe (active experiment — correction returns 0.0; gradient log still appending)
Loading…
Status: active telemetry, correction dormant. The processor code runs each tick (compute_cove_correction() returns 0.0 in both branches; telemetry stamps for retro-check). cove_gradient_log.json continues to append per-tick waterfront-vs-inland gradient data on GCS (14-day rolling window) — cheap, kept alive in case seasonal shift ever changes the microclimate signal and a new refit becomes worth trying. The Lt row in the divergence report now reads from l6_fix_b_refit (07-14 v0.6.351b), which reports HOLD and matches production, so the row stays AGREE forever unless a future refit crosses the +1% gate.
Original design. A small Δ°F added to the temperature forecast based on how much the waterfront stations (Willow Rd, Neptune Rd) typically diverge from the inland-network median under different wind / sea-breeze / hour combinations. Lt was the first correction trained on a spatial differential between station subgroups (L1–Lsr are trained on forecast-vs-aggregated-obs errors).
Shipped 2026-06-26 v0.6.238. Cleared a 2-read confirmation gate on r5_cove_analysis.py (06-25 SHIP + 06-26 SHIP). Per-lead projection fix 2026-06-26 v0.6.237 (initial ship applied current-tick Δ to all 48 leads — wrong by 3–5°F at distant leads when the table crossed zero). Cooling branch disabled 2026-06-30 v0.6.259 (paired-MAE on 19,975 t rows: cooling branch made cooling rows −74.9% worse). Warming branch disabled 2026-07-01 v0.6.276 after per-row Production data (v0.6.269 applied-layer stamping + v0.6.275 backfill on 47,301 T pair rows) showed warming-branch rows ran ~40% worse than L2 on the same rows. Hypothesis at that time: double-counting — L2's Kalman blend for T is dominated by waterfront Tempests (Willow Rd, Neptune Rd at ~0.1–0.2 mi from the cove) — L2 already carries "waterfront bias" by station weighting; adding a (waterfront − inland) Δ on top re-adds the same signal.
Fix B — tried 2026-07-13 v0.6.329, failed the +1% gate. analysis/l6_fix_b_refit.py refit both lookup tables against (cove obs − L2 forecast at cove) instead of the raw (waterfront − inland) gradient — the design the section header used to promise. Ran on 202,321 pair rows (154k train / 47,757 held-out). Panel B (sb_off × hour) looked like a real overnight cove-cooling signal on training: 00-06h +0.9 to +2.1°F, 7 SHIP bins. But held-out MAE improvement was +0.29% vs the +1.0% ship gate. Panel A (sb_on × octant) came back with 0 SHIP bins after refit — the warming-branch signal was largely a fitting-against-raw-L1 artifact. Mechanism confirmed: L2's Kalman blend re-fits per-tick based on obs-vs-model bias and absorbs the same microclimate signal dynamically. A static hourly cove table adds a delta L2 already added → wash on held-out.
Reactivation criterion. Not queued for any specific date. Would require the held-out MAE on a re-run of l6_fix_b_refit.py against fresh pair-log data to clear +1.0% for a sustained window (3+ reads). Possible triggers: seasonal shift (fall/winter offshore vs summer sea-breeze patterns), sensor swap changing L2 station weighting, or a materially different mechanism than the one that failed. Until then, Lt telemetry runs dry.
Related files: weather_collector/processors/cove_correction.py, weather_collector/processors/cove_gradient_log.py, analysis/l6_fix_b_refit.py, analysis/l6_l2_double_counting.py, analysis/r5_cove_analysis.py, GCS cove_gradient_log.json (still updating, 14d rolling).

Archive — Retired ideas & historical investigations

Things we built, ran, and stopped running. Two kinds live here: hypotheses the data answered no, and settled tunings where the sweep concluded "current value is fine." Kept as institutional memory. Scripts in analysis/ can be re-run if conditions shift. Retired layers (code intact but permanently no-op) stay in R&D — see Lt above.

Post-ship watches — closed / superseded / inert
Watches that ran their 14-day window and closed clean, or were superseded by a follow-on ship, or went inert because the target was dropped. Active watches live under Current state → What's improving → Post-ship watches — active.
Recently ruled out — 2026-06-22 to 06-29 Stage 0/1 kills
One-shot smoke tests that landed at "no signal" or "captured by an existing axis." Re-run in 2-3 months if the seasonal regime shifts.
[HYPOTHESIS] Tide-phase corrections — does forecast error track the tide cycle? (RETIRED 2026-06-08)
Verdict: weak signal, mostly entangled with diurnal cycle. Per-field tide-phase curves were tracked across weeks; the signal that survived stratification was hard to distinguish from hour-of-day patterns we're already correcting in L4. Cost of keeping it running (NOAA fetch + 12-bin accumulator + GCS history per Fitter pass) wasn't justified. Analysis: analysis/tide_hypothesis.py.skip.py (renamed 07-18 — removed from digest run so it stops spurious-FAILing on NOAA fetch). Fitter module flag: RUN_TIDE_TRACKING. Frozen charts below show the final state at retirement; they will not update.
Companion view — error vs tide elevation over time (frozen)
Time-domain rendering of the same data as the bucketed chart above. Different angle on the same retired hypothesis. The frozen state below is the last fit before tide tracking was disabled.
Higher leads = forecast made further ahead. Switch to see if the tide pattern is lead-specific.
[HYPOTHESIS] Derived humidity — Magnus(T_corrected, T_d_corrected) vs network-blended humidity (RETIRED 2026-06-08)
Verdict: equivalent. Tested whether deriving humidity from corrected temperature + corrected dew point via Magnus outperforms the L2 network-blended humidity. 27k triples, identical MAE within noise. We kept the derived path anyway because it keeps the (T, T_d, RH, AH) quadruple internally consistent — but the hypothesis "derivation is more accurate" is closed. Analysis: analysis/derived_humidity.py.
[SETTLED TUNING] L3/L4 recency window τ — sweep over fit-window decay constant (RETIRED 2026-06-08)
Verdict: τ=14 days is fine within noise. Not a hypothesis — a parameter sweep over the Fitter's recency-weighting τ (how much old pairs count when fitting decay curves). Not the L2 lead-decay τ added in v0.6.44, which controls how a current bias is spread across forecast leads (see sec 2d). Tested τ ∈ {7, 14, 21} days across four reports. Held-out MAE differences under 2%, well below run-to-run variance. τ=14d stays. With L3/L4 mostly disabled in v0.6.45, this knob barely matters anymore. Analysis: analysis/decay_tau_tuning.py.
[HYPOTHESIS] R4 — HRRR vs GFS spread as confidence signal (RETIRED 2026-06-17, verdict: CLOSE)
Hypothesis: when HRRR and GFS disagree at a given forecast hour, actual error magnitude tends to be higher — i.e. |HRRR − GFS| per (field, lead) predicts |forecast − obs|. If true, the spread becomes a free uncertainty number that can widen displayed intervals and feed Gemini hedge language ("models disagree on tomorrow's high"). Data collection: HRRR L1 already in forecast_log.json; gfs_l1_log.json captures GFS L1 per tick for the same 0-48h window. Decision rule was: ship if median Spearman ρ > 0.25 for ≥3 fields, consistent across lead bands.
CLOSE verdict (2026-06-17, 112,877 joined pairs):
0 of 6 fields above the 0.25 ρ threshold. Maximum observed |ρ| = 0.012 (wind speed at 1-6h) — essentially zero correlation. HRRR vs GFS spread does NOT predict forecast error magnitude. Retired without auto-wiring. Manual script: analysis/r4_spread_analysis.py — re-run quarterly or after a model release.
[HYPOTHESIS] R5 — Cove warming — sea breeze across the peninsula heats Wyman Cove vs inland (RETIRED 2026-06-17, verdict: HOLD — L2 already captures it)
Reframed hypothesis (2026-06-13): Wyman Cove sits in the lee of the Marblehead peninsula on a S/SE/SW sea breeze. Marine air crosses ~2 miles of sun-heated land before reaching the cove, picking up surface heat in transit. Expected pattern: delta_wf_inland = waterfront_median − inland_median goes positive (cove warmer) when wind is from the S half AND sea breeze is active, with magnitude scaling to solar input (peaking ~12-14 EDT). Should flatten to zero when wind is from N/NE (cove is windward of peninsula) or after sunset (no surface heating). Original hypothesis ("waterfront cools during sea breeze") was geographically backwards and is closed.
Day-12 refit (1,732 ticks through 2026-06-24): matches the reframed model; magnitudes tightened further as the sample grew. NW flipped from neutral to weakly negative; E cooled further.
WindSea breezenmean Δ°F
Sactive186+1.5
SEactive88+2.0
SWactive79+1.1
Ninactive378−1.0
NEinactive103−1.0
Einactive86−1.3
NWinactive459−0.9
Diurnal curve under offshore/calm conditions shows clean morning-marine-cooling: trough around −3.7°F at 12:00 EDT (refit 06-24, n=1,732 entries over 12 days; cool air pool over Salem Sound persists; inland warms fast with sun while cove stays anchored to marine boundary). Both signals are physically coherent with the lee-warming model.
Data collection: cove_gradient_log.json captures waterfront-tagged Tempest median (Willow Rd, Neptune Rd — both at cliff-edge elevations on the harbor, confirmed by Joe), inland Tempest median (~18 stations), ambient T, wind dir/speed, salem_water_temp_f, buoy_water_temp_f, sb_active, sb_likelihood per tick (14-day retention).
Two-step plan:
Step 2 verdict: HOLD (run 2026-06-16, n=29,444 matched pairs)
The L2-overlap hypothesis was empirically confirmed. L2's 1/distance² × elevation station weighting for the cove is dominated by the two waterfront Tempests (Willow Rd, Neptune Rd at ~0.1–0.2 mi). L2's "cove bias" is effectively "waterfront bias" by construction. Layering R5's (waterfront − inland) delta on top double-counts the same signal — the cove obs is already waterfront-influenced via L2, so adding more waterfront delta pushes the forecast AWAY from the obs.
Decision (global R5, retired 2026-06-17): r5_audit.py's held-out test of R5 applied across the full pipeline showed it makes cove temp 20–22% worse — L2's waterfront-weighted station blend already captures the signal, so layering R5 on top double-counts. That global formulation stays retired.
Current status (Lt — microclimate correction, RETIRED 2026-07-13 v0.6.329): Both branches return 0.0. Cooling branch killed 2026-06-30 v0.6.259 (paired-MAE showed it materially increased cove temp MAE); warming branch killed 2026-07-01 v0.6.276 after real per-row Production exposed the same double-counting for the sea-breeze direction; Fix B refit tried 2026-07-13 (analysis/l6_fix_b_refit.py) and failed at held-out +0.29% vs +1.0% ship gate. Lookup tables + cove_gradient_log.json continue to log per-tick against a possible future reactivation (seasonal shift, sensor swap, or materially different mechanism). Full rationale in R&D → Lt.
One niche subtlety in the breakdown: long-lead (24-47h) sea-breeze forecasts get +7.85% MAE improvement with R5. At long leads, L2's τ=4h decay has long since faded, so R5 has something L2 doesn't. Not worth shipping a conditional correction for, but documented.