Forecast pipeline

loading… Scoreboard …
Right now — current conditions (click to expand)
Right now — current conditions
Field Base forecast (L1) Production Correction Field Base forecast (L1) Production Correction
Loading current tick…
Confidence: — · — stations reporting
Briefing source: — · —
Base forecast (L1) = the selector's chosen raw source fed into our correction stack (HRRR raw or NBM raw, per field per lead). Production = final Wyman Cove forecast (after L2/L3/L4/gates). Correction = Production − L1 for this tick.
Total Lift
Prod vs NBM raw (the field your default weather app already shows you). The referee's number.
7 Day
median
—
mean
—
24 Hour
median
—
mean
—
1 winning
ch
2 flat
cl · cm
8 losing
t · h · dp · ws · wg · wd · sr · cc
Pipeline Lift
Local correction stack value: Prod vs L1 selector's pick. Answers "are the pipelines pulling their weight on top of raw?"
7 Day
median
—
mean
—
24 Hour
median
—
mean
—
3 winning
cl · cm · ch
0 flat
—
8 losing
t · h · dp · ws · wg · wd · sr · cc
Selector Skill
Value Captured: of the routing gain a perfect picker could have captured, how much did we actually get? Magnitude-aware and n-weighted.
7 Day · Value Captured
median
—
mean
—
24 Hour · Value Captured
median
—
mean
—
9 winning
t · h · dp · ws · wg · wd · sr · cc · ch
0 flat
—
0 losing
—
Attribution — Routing · Cascade · Total (median, pp of user default)
Median across fields per column — the typical field's split. Three independent stats: medians don't add, so Total is its own median, not Routing + Cascade.
7 Day
routing
—
cascade
—
total
—
24 Hour
routing
—
cascade
—
total
—
Attribution — Total = Routing + Cascade (mean, pp of user default)
Additive split of Total Lift. Routing = pick vs user default (0 when the pick IS the default — most rows today). Cascade = local stack on top of that pick.
7 Day
routing
—
cascade
—
= total
—
24 Hour
routing
—
cascade
—
= total
—
Prod Trend
Change vs the prior equal window. Are we improving over time?
7 Day
median
—
mean
—
24 Hour
median
—
mean
—
4 improving
t · cm · ch · cc
3 flat
cl · pr · pp
4 regressing
h · dp · ws · sr
Notable Calls (7D)
Biggest wins and losses at the (field, lead-band) level. Referee highlights.
Best
loading…
Worst
loading…
2nd worst
—
3rd worst
—
Top-1 win + top-3 losses at (field, lead-band). Scored on current-config counterfactual Total Lift (v0.6.494) so killed layers don't drag the rankings. dp + cc excluded (derived); pa + pp excluded (non-MAE); n < 200 filtered.
National Source (7D · lower MAE wins)
Which national raw source is currently scoring better per field. Diagnostic only — Total Lift already scores against NBM raw (or HRRR raw for the 5 HRRR-only fields), so National Source doesn't set the baseline; it's context for interpreting where the pipeline's lift is coming from.
—
loading…
Health & Reliability (7D)
Trust indicators. Independent of who's winning.
Confidence loading…
High-conf cells loading…
Halves agree loading…
Halves agree below ~70% is a noisy week — read the score tiles above with caution. Halves agree = both halves of the 7d window show the same sign of lift.

Current state

Stack health trajectory — aggregate Prod vs 90d-ref raw, per day (12 fields, pa/pp excluded)
Per-obs-day cross-field Median + Mean + P25–P75 band of (1 − prod_MAE / raw_MAE_90d_ref) × 100. Denominator is each field's fixed 90-day-reference raw MAE — daily weather difficulty is absorbed by the fixed baseline, so the ratio only moves when Prod moves. Positive = Prod beating the 90d baseline; a ship that improves the stack pushes the 7d rolling median UP over the following week. Full per-field trajectory + drill-down: Accuracy over time ↓.
Trend windows
Per-field diagnostic — 7 day
Coach's table: each field's total, the pipeline lift on the selector's chosen stack (mirrors the tile), the two cascade skills (independent of routing), and the two selector-quality views. If Total Lift is red, this table shows which lever to pull. Live from per_field_scoring.json (refits hourly). Win Rate + Value Captured are diagnostic on the paired chosen/alt Prod pool — NBM-scope fields only. dp and cc omitted — both derived, no independent skill to score (see note below).
Field Total LiftProd vs Public Baseline Difficulty7d raw MAE ÷ 90d ref · >1 = harder than usual Pipeline Lift(L1 pick − Prod) / L1 pick HRRR Pipeline Skill(raw − Prod) / raw NBM Pipeline Skill(raw − Prod) / raw Win Ratewins / (wins + losses); ties excluded Value Captured% of oracle n
loading per_field_scoring.json…
dp and cc omitted — both derived at prod-time (dp = Magnus(prod_t, prod_h); cc = Ccd max(prod_cl, prod_cm, prod_ch)) with no independent skill chain. Future dp bias layer or cc composition tuner would restore their rows.
🪦 cfg X% annotation on a Total Lift cell = current-config counterfactual (v0.6.494). Computed by walking the cascade stamps and skipping layers currently ENABLED=False. Big delta between the as-shipped number and cfg = a recent kill is still poisoning the rolling window; expect it to fade over ~7-30 days as new rows accumulate. Snapshot 2026-08-26: skips error_l5_nbm (killed 08-25) + error_l6_nbm (scaffold). HRRR-side counterfactual deferred (needs per-field disable logic because HRRR layer slots are field-dependent).
Windows: all six lift/skill columns computed on the same paired 7d/24h pool (v0.6.492) — rows where every relevant residual (HRRR raw, NBM raw, selector output, Prod, both cascade Prods) exists. Total Lift, Pipeline Lift, and per-cascade Pipeline Skill columns share denominators (each vs its own reference: user default, L1 pick, own raw), so causal chaining is meaningful: if Total Lift moves, at least one of Pipeline Lift, a Pipeline Skill, or the selector's Win Rate moved on the same rows to cause it.
Per-field diagnostic — 24 hour
Same columns as the 7d table above, computed over the last 24 hours of paired residuals. Read this table for "did anything break today" — a field whose 24h row goes red while its 7d row stays green flagged a regression the daily view is still averaging away. Live from per_field_scoring.json's windows.24h.per_field block. Same aggregation rules as the 7d table (dp/cc derived-omitted; pa/pp non-MAE excluded from Median/Mean).
Field Total LiftProd vs Public Baseline Difficulty7d raw MAE ÷ 90d ref · >1 = harder than usual Pipeline Lift(L1 pick − Prod) / L1 pick HRRR Pipeline Skill(raw − Prod) / raw NBM Pipeline Skill(raw − Prod) / raw Win Ratewins / (wins + losses); ties excluded Value Captured% of oracle n
loading per_field_scoring.json…
Per-field diagnostic — 12 hour · fresh post-ship read (v0.6.615)
Same columns as the 24h table, computed over the last 12 hours of paired residuals. Read this AFTER a ship — the 24h and 7d windows lag any override/selector-fit change by up to 24h/7d respectively, so a fresh routing change or L2 fix reads as noise there. 12h reads first. Thin-window caveat: pair-log typically has an 8h backstamp lag, so the 12h window is really "recent 4h of ripe pair-log rows"; fields with n≤50 read as noise, not signal. Live from per_field_scoring.json's windows.12h.per_field block. Same aggregation rules as 7d/24h.
Field Total LiftProd vs Public Baseline Difficulty7d raw MAE ÷ 90d ref · >1 = harder than usual Pipeline Lift(L1 pick − Prod) / L1 pick HRRR Pipeline Skill(raw − Prod) / raw NBM Pipeline Skill(raw − Prod) / raw Win Ratewins / (wins + losses); ties excluded Value Captured% of oracle n
loading per_field_scoring.json…
Per-field pipeline architecture + status
Companion to the per-field scoring table above (numeric selection/correction/total lifts). This section shows the pipeline shape for each field — HRRR-side cascade, NBM-side cascade (if in scope), which side the selector picks — plus hand-curated narrative context (open work, recent ships, current state). Numeric cells live in the scoring table; this table is the architecture + journal.
Field Pipeline · HRRR Pipeline · NBM Selector Status
t temperature
L1 → L2 additive
raw_nbm → l2_nbm (l3_nbm dropped 08-25 v0.6.472 — walkforward −2.2% agg; l6_nbm v0.6.453 scaffold ENABLED=False)
loading… Stable. L2 additive only. HRRR wins; NBM l2_nbm minimal.
h humidity
L1 → L2 additive
raw_nbm → l2_nbm (l3_nbm dropped 09-05 v0.6.551 — sentry HOT + 14/15 cells help_fresh negative)
loading… Positive vs raw on 7d. Numeric values live in the per-field scoring table above (7d/24h/12h). Structural state (09-29): L3_NBM dropped 09-05 v0.6.551 (sentry HOT + 14/15 cells help_fresh negative), so NBM cascade is raw_nbm → l2_nbm only. HRRR side runs the full L1→L2 additive → L4 → chp/dpbp/lsb/C1 stack. Routing has been carrying the field vs cascade at various times; v0.6.588 escalation cell h/calm/24-47 live since 09-11. Current shadow: v0.7.6 L1 static blender covers h with universal ω=0.44 on 10 cells (SHADOW ENABLED=False); 09-29 retro shows h/nw_flow/24-47 +39.4% SHIP-READY (n=336, below gate 400 — narrow-flip decision ~10-02).
ws wind speed
L1 → L2 direct (L3 dropped 08-08 v0.6.397)
raw_nbm → l2_nbm (l3_nbm dropped 08-25 v0.6.472 — walkforward −2.5% agg)
loading… Stable — L2 blend + direct selection. wsbp calm-only sibling gate ENABLED=False (preflight n=0).
wg wind gust
L1 → L2 direct → L3 lead-decay → wg residual persistence (dormant)
raw_nbm → l2_nbm → l3_nbm
loading… Stable win. L3 asymmetric SKIP table live. wg residual persistence gate ENABLED=False (Stage 1 held MARGINAL).
dp dew point
L1 → L2 additive → L3 skip-table → dpbp → dp residual persistence (dormant) (derived from corrected t, h via Magnus)
raw_nbm → l2_nbm → l3_nbm (v0.6.440)
loading… Derived from corrected t + h. No independent skill chain. dpbp LIVE since 08-04.
cc cloud cover (derived)
L1 → L2 blend → Lc → Ccd = max(cl_l6, cm_l6, ch_l6) (SKIP_REGIMES: se_flow, unknown → Pirate fallback)
raw_nbm → l2_nbm (l3_nbm dropped 09-10 v0.6.577 — sentry HOT + walkforward DROP two-tool; l4_nbm dropped 09-08 v0.6.563)
loading… Derived via Ccd = max(cl_l6, cm_l6, ch_l6) except SKIP_REGIMES. Fully removed from NBM cascade layers 09-08 + 09-10. Combine walker tracking random-overlap replacement for the hardcoded max formula.
ch high cloud
L1 → L2 blend → L3 lead-decay → L4 diurnal → Lc (07-17) → chp (07-19)
raw_nbm → l2_nbm → l3_nbm → l4_nbm → chp_nbm (l4_nbm v0.6.451, chp_nbm v0.6.454)
loading… Best-performing field vs raw. Full stack L1→L2→L3→L4→Lc→chp. 10-cell _CELL_SKIP forces mid/long-lead cells back to L4 (v0.6.405/409 emergency demotes).
cl low cloud
L1 → L2 blend → Lc → clp (dormant; cl in _FIELD_SKIP; cl_l6 = raw)
not extracted (NBM CO does not publish cl)
HRRR (no NBM cascade) Off Lc (in _FIELD_SKIP) since 07-30. clp gate wired Stage 3, ENABLED=False. Open: EMA/Kalman shift tracker to un-skip Lc.
cm mid cloud
L1 → L2 blend → L3 lead-decay → Lc (L3 dropped 09-15 v0.6.625 — walkforward pooled net-negative; Lc 07-17)
not extracted (NBM CO does not publish cm)
HRRR (no NBM cascade) Stable win + Lc live. L3 DROP walkforward streak 2/7 — potential ship 09-15.
pp precip probability
L1 only (L3 dropped 07-04 — hurts Brier despite MAE gain)
not extracted (POP deferred pending unit audit)
HRRR (no NBM cascade) L1 only. All recalibration paths PARKED 08-02 (non-stationary). Reopen path: physical-feature gating (CAPE, RH-500mb, front proximity).
pr pressure
L1 → L2 additive (gated on nw_flow/0-5h + 6-11h since 08-10)
not extracted (NBM CO does not publish pressure)
HRRR (no NBM cascade) L2 SHIP gated on nw_flow/{0-5h, 6-11h} since 08-10 v0.6.401.
sr solar radiation
L1 → Lsr synoptic-regime (skip: ne_flow + calm) → Lsb (sea_breeze cc<25 override)
raw_nbm → l2_nbm (l3_nbm dropped 08-25 v0.6.472; l5_nbm KILLED 08-25 v0.6.471 — sentry +238% MAE + walkforward −126%/−145% agreed; sr has no HRRR L2 so l2_nbm ≈ raw_nbm)
loading… Lsr LIVE, skip ne_flow + calm. Lsb sea_breeze cc<25 override LIVE. NBM cascade minimal — l3_nbm/l5_nbm both killed for sr. Selector pick per band shown live in the Selector column above; direction can flip with recent-window MAE.
pa precip amount
L1 only (no L2/L3/L4 apply)
not extracted (APCP deferred pending unit audit)
HRRR (no NBM cascade) L1 only, no open work.
wd wind direction
L1 → L2 wind_blend → wdp (07-27)
raw_nbm → l2_nbm → l3_nbm (sin/cos) → wdp_nbm (08-19 v0.6.437)
loading… L2 wind_blend + wdp both LIVE. NBM cascade full (l2/l3/wdp_nbm).
Note: the narrative snippet below is the LEGACY per-layer MAE narrative — it reads mae_over_time.json's last_7d block (with an unweighted-across-leads fallback when that block is empty). It is NOT the top scoreboard's source. The scoreboard's Total Lift / Pipeline Lift / Notable Calls / High-Conf tiles all read per_field_scoring.json's paired-pool numbers (v0.6.492+), which don't touch this fallback. Kept here as a per-field cross-check; retire when the paired-pool numbers are trusted for the same purpose.
Loading 7-day per-layer narrative from mae_over_time.json…
Loading provenance…
What's running · improving · being evaluated next
What's running
Stack
  • L2 mesonet blend: t · h · dp · cc · cl · cm · ch · ws · wg (pr on nw_flow 0-11h since 08-10)
  • L3 lead-decay: wg · ch (pp dropped 2026-07-04 v0.6.304; ws dropped 2026-08-08 v0.6.397; cm dropped 2026-09-15 v0.6.625 — walkforward L3 fc −1.6% / obs −1.2% pooled, calm/24-47h LOSS −17.2%; wg L3 skip cells: calm 0-5 + 12-23 + 24-47, ne_flow 6-11, sea_breeze 6-11 + 12-23 + 24-47, frontal 12-23)
  • L4 diurnal: {ch} · cc dropped 2026-09-08 v0.6.563 (30d marginal lift +0.79% pooled, best skip-table only +1.38% — not worth the maintenance)
  • Lsr synoptic (sr): skip ne_flow + calm · Lt reclassified as telemetry experiment 07-27 (Fix B held-out +0.29%)
  • Lsb sr sea_breeze cc<25 Lsr override (sr): per-hour bias hours 12-19 · FLIPPED 2026-08-05 v0.6.394 · gate: regime == "sea_breeze" AND cc < 25 · 14-day watch CLOSED CLEAN 2026-08-19 (+11.4% pooled, n=15,985)
  • NBM parallel cascade — AT ARCHITECTURAL PARITY (v0.6.499, 2026-08-26): raw_nbm for 9 fields; native L2 for all 8 NBM-scope fields (t/h/dp additive Kalman+decay, ws/wg/wd wind_blend, cc/ch cloud_obs_blend hourly[0]; sr identity — no L2 either side); l3_nbm for wg/ch/sr (scope trimmed 08-25 v0.6.472; sr re-added 09-04 v0.6.549 then per-cell skipped 09-29 v0.7.11 on nor_easter/12-23h + /24-47h; h dropped 09-05 v0.6.551; cm dropped 09-15 v0.6.625; cc dropped 09-10 v0.6.577); l4_nbm for ch only (cc dropped 09-08 v0.6.563); l5_nbm for sr KILLED 08-25 v0.6.471; l6_nbm for t scaffold (ENABLED=False); chp_nbm for ch; wdp_nbm for wd. · Field-level pick distribution changes daily — see the per-field Selector column above and the National Source tile for real-time per-band picks. The static summary sentence that used to live here drifted from reality within days of every ship. · skip_table_nbm_curated.json 17 cells post-v0.7.11 (was 15 post-v0.6.622; +2 for sr × nor_easter on 09-29). Symmetric ADD + REMOVE audit loop runs daily in the digest.
  • Lc cloud saturation-unbiasing (cm · ch only) · FLIPPED 07-17 v0.6.355 · cl+cc both OFF via _FIELD_SKIP after 07-30 v0.6.389d-390 intervention
  • Ccd cc from-derivation · LIVE ENABLED=True 07-30 v0.6.390 (flipped 6d ahead of gate) · cc = max(cl_l6, cm_l6, ch_l6) except SKIP_REGIMES {se_flow, unknown}
  • chp high-cloud persistence-of-obs (ch): 14 SHIP cells post-07-27 demote, 10 additional cells force-skipped via _CELL_SKIP in processor 08-13 v0.6.405 · FLIPPED 07-19 v0.6.358 · 14-day watch closed clean 08-02 · 14-day post-v0.6.405 watch through 08-27
  • wdp wind-direction predicted-transition persistence (wd): 5 SHIP cells · FLIPPED 2026-07-27 v0.6.382 · 14-day watch CLOSED CLEAN 08-11 (Δ −2.7% n=16,416)
  • dpbp dew-point antecedent-error persistence (dp): pre_frontal + nw_flow + sw_flow, lead ≥ 6h · FLIPPED 2026-08-04 v0.6.391 · 14-day watch through 08-18
  • C1h trend-direction premium widening: 9 SHIP cells (cc/cl/cm/ch 12-23h & 24-47h + t/24-47h) · SHIPPED 2026-08-05 v0.6.393 · RE-CURATED 2026-08-08 v0.6.396 · 08-08 refresh: t/12-23h SKIPed by magnitude floor (−0.09%), t/24-47h newly SHIPs WIDEN (+11.0%); cc/cl/cm/ch 12-23h + 24-47h all still SHIP · confidence-layer only (no forecast modification)
Production vs raw
  • ✓ ch −46% · wg −16% · cc −15%
  • ✓ cm −7% · dp −6% · h −2%
  • ✓ ws long-lead — RESOLVED 08-08 (L2 blend fix, then dropped from L3_FIELDS). Details in Archive.
Guards
  • raw-baseline verifier: healthy (no drift)
  • live-layer change gate active (7-day / 2-tool / per-cell / no-ENT)
What's improving
15 active Stage 1+3 candidates · all auto-run in daily digest · shipped items live under Post-ship watches below
L1 blender (v0.7.3 · shadow-only) ⚠ v0.7.2 APPLY flip rolled back 2026-09-24 · BLENDER_APPLIED_FIELDS = frozenset() · curated table stale-window fit · re-curate needed · sibling: v0.7.6 L1 static blender (2026-09-26, separate mechanism — universal ω per field, L1 seat, cascade BYPASS, curated for h/dp only, ENABLED=False shadow). Both mechanisms are wired in parallel; both currently shadow. v0.7.6 09-29 shadow retro (post-v0.7.8): 2 SHIP-READY (h/nw_flow/24-47 +39.4%, dp/nw_flow/24-47 +8.6%, both halves-stable, n=336 each — below gate min_n=400, wait ~3d) / 13 HOLD / 7 KILL / 7 THIN on 20 curated cells. Universal apply-flip 10-03 still not on track — pre_frontal cells KILL, and nor_easter (629 dp + 471 h rows, 41% of stamps) is off-curated because its best-ω is HRRR-favoring (1.00 dp, 0.65 h) vs universal 0.27/0.44. Off-curated stamp rate dropped 84%→41% post-v0.7.8, residual is entirely nor_easter — a curation gap, not a plumbing bug. Narrow-flip candidate: nw_flow/24-47 h+dp only, decision around 10-02 when n crosses 400.
First architectural change at the L1 seat since the selector. Replaces binary picking with continuous per-obs blend weight ω ∈ [0,1] from an L2-regularized ridge on 12 features (ims, xr_spread, lead, hour circular, cc_inter_sigma, pressure_trend, wd circular, ws, cloud_low, solar). At runtime: forecast = ω · HRRR_terminal + (1−ω) · NBM_terminal, per (field, regime, band) cell. Fields in BLENDER_APPLIED_FIELDS (currently empty) apply the blend to the served forecast; every other field gets shadow telemetry only. Un-curated cells fall through to the selector unchanged.
→ Stale-window fit discovered 2026-09-24. The GCS forecast_error_log_backstamped.jsonl had been frozen since 2026-08-21 (uploaded manually once, then no auto-refresh). l1_blender_stage1.py reads that URL via cached_path(); every fit since Aug 21 saw data ending 2026-08-20T20:07. The 13 STABLE cells shipped v0.7.0 and today's v0.7.2 dp APPLY flip were both fit on that frozen window. Fresh-data refit results: 11 of 13 cells fail halves-stable. Survivors: h/pre_frontal/24-47 (A+25.7% / B+8.1%, ω̄ 0.44-0.55) and t/se_flow/24-47 (A+11.8% / B+14.9%, ω̄ 0.37-0.46). One new candidate not in the shipped set: wg/pre_frontal/12-23. v0.7.3 (this session) reverts BLENDER_APPLIED_FIELDS to empty and wires analysis/nbm_backstamp_append.py into the publisher CF (runs first each hour, appends new pair-log bytes to the GCS backstamped file via bucket.compose(), HWM tracked in gs://myweather-data/backstamp_hwm.json) so this staleness class can't recur silently. l1_blender_curated.json intentionally not re-curated in v0.7.3 — since applied-fields is empty, no cell fires and shadow telemetry now accumulates against fresh data. Re-curation is a follow-up ship. Related audit: per-obs classifier v2 (v0.6.644-646) fit on the same stale window; fresh-data refit shows STAGE 1 HOLD on both h and t with every cell marked degenerate (fNBM 67-99% — the fit rubber-stamps NBM rather than selecting per-obs).
NBM parallel cascade ✓ AT ARCHITECTURAL PARITY 2026-08-26 v0.6.499 · native L2 shipped for all 8 NBM-scope fields · all layer slots + specialists mirrored
Full mirror of the HRRR cascade on the NBM side, layer-for-layer: raw_nbm → l2_nbm → l3_nbm → l4_nbm (ch) → l5_nbm (sr) → l6_nbm (t, ENABLED=False) → chp_nbm (ch) / wdp_nbm (wd). Each layer has its own fitter (analysis/l{3-6}_nbm_*.py) writing a curated JSON that the collector loads at import; per-hour stamps land in the snapshot log and per-row error_l{n}_nbm columns land in the pair log. L1 selector picks argmin(MAE(HRRR-deepest-Prod), MAE(NBM-deepest-Prod)) per (field, lead-band). Bins reached L3 maturity via analysis/nbm_backstamp.py (walks pair log + backfilled NBM point extracts, reconstructs l2_nbm exactly from snapshotted HRRR values); deeper layers warm naturally over the next ~30 days as error_l3_nbm/error_l4_nbm/error_l5_nbm rows accumulate. See per-layer fit-status tiles below.
→ Live per-field / per-band picks in the National Source tile above. Router-scope aggregate lift last refit at v0.6.500 (2026-08-26) was +53.5% on n=73,023; refreshed daily by the pool table. v0.7.5 (09-26) added learned_gbm + ims_threshold precedence on top of the pool for ch and sr; v0.7.11 (09-29) added a per-cell L3 skip for sr × nor_easter that runs after the selector picks NBM. Static "which side wins" summaries drift within days of every ship — read the live tile.
C1h trend-direction ✓ Stage 3 wired 07-08 · per-cell ortho gate 07-10 · gated OFF
15 SHIP cells (all WIDEN) live-stamping on confidence.cells[field][band].c1h. Per-cell co-axis ortho gate (v0.6.321) suppresses cells non-orthogonal to the currently-firing co-axis — cl fires freely, ch 24-47h + t × 3 never fire.
→ Already live-stamping (confidence_layer reads the curated table each tick). The narrow-promote counter is auto-suppressed in the digest per KNOWN_LIVE_PIPELINES — the ship gate for user-visible band widening is C1 Stage 4 (last HOLD 65.00% — moved 57.14% → 65.00% on 08-15 re-cure).
dp depression regime frontal branch closed · nor_easter watch opened
Frontal decayed −2.19 → −1.98 → −1.51 → −0.87°F — below the 1.5°F action floor as of 07-09. Branch retired.
New: nor_easter +3.79°F ★ but n=279 — magnitude passes floor, sample thin. Gate: 3 consecutive reads with n growing.
C1d cloud disagreement ✓ Stage 3 wired 07-08 · gated OFF
13 SHIP + 1 MARGINAL cells stamping on confidence.cells[field][band].c1d. Fires when live cloud_inter_source_sigma ≥ Q3 (44.55). cc 24-47h MARGINAL NARROW −7.06% is a documented outlier — flagged for Stage 4 review.
→ Already live-stamping (curated table read per tick). Narrow-promote counter auto-suppressed in digest per KNOWN_LIVE_PIPELINES — ship gate for user-visible band widening is C1 Stage 4 (last HOLD 65.00% — moved 57.14% → 65.00% on 08-15 re-cure).
Pre-frontal cloud widening ⚙ Stage 2 curated 07-12 · counter wired · blocked on population (n=8 passages, THIN)
Latest ortho read: 7 SHIP cells (cell is SHIP iff ORTHOGONAL vs C1a AND vs C1e). Written to weather_collector/data/pre_frontal_curated.json. v0.6.372b matched-regime fix HOLDS PROMOTE with the corrected baseline; SHIP set expanded 5→7. Digest currently at streak 1/7 due to cell-set drift + population still THIN.
→ Real 7-day narrow-promote counter wired 2026-07-12 v0.6.328d. Blocked on population: only n=8 frontal passages in the current window, 17% pair-log join rate — both matched-regime KILL (h_hsf) and matched-regime PROMOTE (h_pre_front) share this THIN caveat. Do NOT act on either verdict until passage count ≥ 15 (probably 08-05 to 08-15). See [[project_c1e_hsf_kill_investigation]].
cl persistence gate (clp) ⚙ Stage 3 wired 07-24 v0.6.379 · ENABLED=False · streak walker tracked continuously via auto-rolling windows (v0.6.398, 08-09); flip when 7-day Jaccard clears — no fixed date target
Successor to the retired cl_persistence_short_lead (all-9-regimes design gate not met). Broader regime × lead_band gate: 12 SHIP / 8 MARGIN / 16 SKIP / 1 THIN. Whitelist: persistence on calm/se_flow/unknown all-leads + nw_flow 24-47h + short-lead in sw_flow/ne_flow/nw_flow/pre_frontal + pre_frontal 6-11h. Wired after Lc. 7-day flip gate: SHIP cell-set Jaccard ≥ 0.8 across 7 daily reads. See [[project_cl_persistence_investigation]].
wg residual persistence gate ⚙ Stage 3 wired 07-14 v0.6.351 · 07-27 flip HELD · awaiting Stage 1 recovery
Long-lead-only: adds per-clock-hour L2-residual mean (prior 14d) on top of fc_l2, replaces L3-corrected wg on gate-fired cells. Short leads always SKIP (L2's Kalman handles close-in). 07-27 flip HELD — h_wg_residual_persistence_stage1 flipped PROMOTE → MARGINAL 07-26 (+20.24%, was +17.74%). Persistence hypothesis's aggregate weakened; hold until Stage 1 recovers PROMOTE. (Divergence report drop-{wg,ws} was stale-tool artifact per [[feedback_walkforward_skip_table_blind]]; fix v0.6.381 landed 07-26.)
dp residual persistence gate ⚙ Stage 3 wired 07-25 v0.6.380 · ENABLED=False · gate closed 08-01 — verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09)
Cloned from wg template. Runs after wg_residual_persistence; replaces L3-corrected dp on gate-fired cells with (fc_l2 + fitted L2-residual). Stage 2 preview: 8 SHIP / 2 MARGIN / 26 SKIP / 1 THIN — SHIP cluster long-lead only (frontal 12-47h, nw_flow 24-47h, pre_frontal 12-47h, sw_flow 6-47h); zero SHIP at 0-5h. Best cells sw_flow 24-47h (−20.4%), sw_flow 12-23h (−18.6%). Clamp |Δ| > 10°F. Same Jaccard ≥ 0.8 streak gate as wg.
pp frontal × 6-11h Platt (fixed b=0.6) ⏸ PARKED 2026-08-02 · script retired to .skip.py 2026-08-10
07-27 Stage 1 SHIP (Brier lift −28.56%/−18.73%, n=292) re-flipped HOLD on 08-02 halves-strict re-cut — recalibration params don't transfer across time at this sample size. Root: Reliability ≈ 5% of Brier (0.005 of 0.086); Uncertainty + Resolution dominate. All Platt-family scripts (h_pp_platt_calibration, h_pp_platt_by_regime, h_pp_frontal_platt_stage1, h_pp_bin_calibration, h_pp_bias_persistence_stage0) retired to .skip.py on 08-10 with the daily-digest retirement pass (10 dead-verdict scripts total). Structural next step if pp ever reopens: physical-feature gating (CAPE, RH-500mb, front proximity), not calibration parameter fitting.
Lc recent-bias gate (item #3 on pipeline-to-good plan) — SHIPPED ON 08-15 v0.6.413 ⚙ Stage 1 SHIP 08-09 v0.6.399 · per-field clearance added 08-12 v0.6.401i · runtime table 08-13 v0.6.407 · Stage 3 wire SHIPPED OFF 08-14 v0.6.410 · toggle FLIPPED ON 08-15 v0.6.413
LC_RECENT_BIAS_GATE_ENABLED = True live in cloud_saturation_correction.py. Applies existing lc_correction_table.json shift only where recent 3-day observed bias still agrees with the historical fit (sign match + magnitude ≥ 0.5×|hist|). Smaller surface than the EMA/Kalman rewrite alternative — no new lookup table, just a per-cell gate on the live shift. 08-18 refresh (v0.6.430): ch per-field streak now 11 days; cl and cm remain CHURN (in and out of promoted set day-over-day). Deploy carries refreshed lc_correction_table.json + lc_recent_bias_gate.json into the collector image; runtime toggle unchanged. Live cells suppressed today = 0/3 — all three ch per-bin cells (20-50, 50-80, 80-95) have gate_apply=True because recent 3d bias tracks historical (or is thin) on every bin. cc excluded (derived via Ccd). Ship is a live no-op today but the self-healing mechanism is in place for any future per-bin drift — the anti-scar-tissue counterpart to hand-curated skip frozensets. Runtime telemetry verified 10:48 UTC tick: weather_data["cloud_saturation_correction"]["recent_bias_gate"] = {enabled: true, fields_cleared: ["ch"]}. See [[project_lc_regime_conditional]]. Follow-on: cl remains in _FIELD_SKIP (Lc architecture can't rescue cl regime) — needs EMA/Kalman fallback, separate workstream [[project_lc_cl_unskip_investigation]].
chp full-shape refinement (adopt L6-baseline 5-cell SHIP set) ⚙ 6-cell emergency demote LIVE 07-27 v0.6.382t · full-shape adoption 7-day gate closed 08-03 · verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09)
Preview curated JSON at weather_collector/data/ch_persistence_gate_curated_vs_l6.json. L6-baseline Stage 2 rebuild (h_ch_persistence_blend_stage2_vs_l6.py) says chp's honest SHIP set is 5 cells: calm/0-5 (−70%), pre_frontal/0-5 (−41%), sw_flow/0-5 (−30%), se_flow/0-5 (−26%), ne_flow/6-11 (−17%). Currently-live chp is at 14 SHIP + 8 MARGIN post-demote; 5 more halves-disagreement/small-magnitude cells still fire but should also fall away. 7-day live-layer change gate closed 08-03; verdict tracked continuously via auto-rolling windows (v0.6.398, 08-09). Adopt condition: 7 daily reads of the L6-baseline Stage 2 continue to agree on the 5-cell SHIP set + no new frontal-only-recalibration interaction.
What's being evaluated next
Upcoming
Forward-looking only. For today/yesterday narrative see Recent activity below.
Fri 10-03
KEY DATE — v0.7.5 + v0.7.6 fresh-data verdicts, scope narrowed. (a) v0.7.5 router-as-authority — ch-only now. 09-28 finding: sr's learned_gbm path never fires because l1_learned_selector_curated.json has no sr cells; v0.7.5 verdict is really about ch via ims_threshold. Attribute per-mechanism using v0.7.7's selector_mechanism stamp. Rollback = one flag flip. (b) v0.7.6 L1 static blender shadow retro — universal apply-flip NOT on track, but nw_flow/24-47 narrow-flip candidate emerging. 09-29 retro: 2 SHIP-READY (h/nw_flow/24-47 +39.4%, dp/nw_flow/24-47 +8.6%, both halves-stable, n=336) / 13 HOLD / 7 KILL / 7 THIN on 20 curated cells. Off-curated ratio dropped 84%→41% (residual all nor_easter, not in curated). Held on narrow flip: n=336 < gate min_n_rows=400 AND most of 7d predates v0.7.8's clean-plumbing. Wait ~3 days for n≥400 on strictly post-v0.7.8 rows; if halves hold, ship narrow (nw_flow/24-47 only, h+dp). Universal apply-flip still blocked — pre_frontal cells KILL in shadow, nor_easter needs per-regime ω (best-ω is HRRR-favoring, opposite of universal 0.44/0.27).
~10-02 (n gate)
v0.7.6 nw_flow/24-47 narrow-flip decision. Re-run analysis/l1_static_blend_shadow_verify.py when h/nw_flow/24-47 and dp/nw_flow/24-47 both cross n≥400 on strictly post-v0.7.8 rows. If both halves still positive ≥5%, ship narrow flip: modify l1_static_blend_curated.json to keep only nw_flow/24-47 cells per field, set ENABLED=True. Rollback = flag back to False.
~10-04 (n gate)
nor_easter static-blend re-evaluation. Today's nor_easter fit shows h/12-23 best-ω=0.65 with +28% lift halves-stable +23.5/+38.0, n=381 (below 400 gate). If n crosses 400 next 3-5 days: decide schema — extend l1_static_blend_curated.json to support per-regime ω override (nor_easter isn't compatible with the field's universal ω=0.44). Also revisit dp/nor_easter (best-ω=1.00, halves unstable). If no cells clear, hold and revisit when regime accumulates more sample.
Mon 10-05
v0.7.9 chp dynamic gate 7d verify. h_ch_persistence_blend_stage2_vs_l6 WATCH count should drop from 4 live losing cells to ≤1. If not, one of the 5 dynamic-only cells has a Simpson-paradox artifact and the flag should be reverted; investigate before shipping any additional cells.
next session
dp NWS coherence — SUPERSEDED by v0.7.6 blender. Original plan (09-13): route dp through NWS at the hourly-array level via decay_apply. Superseded 09-26 by the L1 static blender's coverage of dp with universal ω=0.27 on 10 (regime, band) cells; the analysis showed L1-blend-no-cascade beats every routing/blending option that keeps the cascade. If v0.7.6 apply-flip lands 10-03 clean, this workstream closes. If not, re-scope NWS coherence as fallback.
rolling
L1 by-regime walker — near-miss watch. First wire landed 09-11 v0.6.581 (ws/calm/12-23, ws/nw_flow/12-23). Under the fixed gate (sum(n_today) ≥ 60 across 3d), three PPP cells sit below threshold and will clear as sum_dn accumulates: wg/sw_flow/0-5 (sum_dn=32, lift +7.4%), ws/calm/0-5 (sum_dn=6, lift +21.4%), ws/calm/6-11 (sum_dn=2, lift +28.7%). Digest watch — no action.
rolling
Prune HRRR PBL morning-overshoot gate. v0.6.606 hardcoded gate is a stop-gap until walker's stagnant_high × t × 0-5 cell clears the wire naturally. Once the walker routes t at that regime/band, the named gate becomes redundant. Remove HRRR_PBL_MORNING_OVERSHOOT_* block from l1_selector.py; drop the hour_local kwarg call from forecast_snapshot.py. Was "4-7d" (09-13); regime rarely active recently.
Backlog
NBM sea-breeze specialist — dedicated workstream. Only if 09-11 walker read + skip curation loop don't close enough of the aggregate gap. Stack-health trajectory (top of page) is the arbiter — if 7d median plateaus after this week's ships, structural NBM specialist work is next; if it keeps trending up, keep tuning.
clp deferred
clp Stage 3 flip gate — 08-16 walker FAIL (min-J 0.250; fire set churned 7 → 3 cells over 7d window). No ENABLED flip. Walker continues; re-check when digest reports PASS.
ongoing
C1d narrow-promote — auto-suppressed (KNOWN_LIVE_PIPELINES). Already live-stamping via confidence_layer; user-visible band-widening gate is C1 Stage 4.
held
pre-frontal C1e narrow-promote — HELD 08-11 fresh re-run. Down from 16 ortho cells (June) to 2 SHIP (ch 24-47h, cl 24-47h). Only 3 frontal passages in the 45d window; most cells THIN. Re-run when autumn front cadence returns.
held
frontal-t bias Stage 0 — Stage 1 gate day 1/7 (08-12 v0.6.401g). 3/4 bands HIT (0-5h +1.94, 6-11h +2.74, 12-23h +1.06); 24-47h SIGN_FLIPS. Half A concentrated in one 07-13/16 event. Rolling 7-day gate added to h_frontal_t_bias_stage0.py mirroring l6_fix_b_refit: requires ≥7 distinct days, no HOLD days, ship_bands STABLE, ≥1 band. Verdict escalates to STAGE 1 CLEAR — the trigger to scope frontal_t_bias.py specialist. Earliest clear 2026-08-19.
held
sr obs-recent override Stage 0 — 08-11 near-miss. Fired-subset lift +23.5% but pooled test lift 4.58% under 5% gate (77 fires, 91% non-Lsb). Two paths to Stage 1: loosen pooled gate OR tighten trigger to 300 W/m². Backlog #8.
blocked
wsbp — flip gate still HELD as of 08-08 (v0.6.388): calm regime shadow still below MIN_N_ANTECEDENT=20 in the 24h antecedent window. Predict: one more calm overnight pushes it over.
Post-ship watches — active
  • v0.7.15 sr learned_gbm cells finally live (opened 09-29). 5 sr STABLE GBM cells (nw_flow/12-23, nw_flow/24-47, se_flow/12-23, se_flow/24-47, sw_flow/6-11) populated in l1_learned_selector_curated.json. Completes v0.7.5's sr side after 3 days as a silent no-op. Deploy 15:50 UTC clean. Watch: (a) overnight — first sr rows in the 5 covered cells stamp selector_mechanism=learned_gbm (was band_pool); (b) ~7d — sr Value Captured 7d trends positive as covered cells accumulate; (c) if any covered cell shows negative lift, that cell needs to be dropped and refit. See feedback_shipped_flag_verify_effect.
  • v0.7.16 v5 candidate writer schema fix (opened 09-29). One-line preventive fix in analysis/l1_selector_per_obs_classifier_stage1_v5.py: candidate JSON now emits band "12-23" instead of "12-23h" to match runtime _band_for_lead(). Prevents recurrence of the v0.7.5→v0.7.15 silent-no-op bug on future ships from this pipeline. Watch: next ship attempt from this script produces cells that fire without band-key mismatch.
  • v0.7.14 selector terminology rule (opened 09-29). "Selector" is the primary name in prose. Code entities keep their names (l1_selector.py, pick_source(), selector_source/_mechanism). "Router-as-authority" retained only as the v0.7.5 pivot's framing name. Motivated by v0.6.432 "L1 router" archive collision. See feedback_selector_is_primary_name. Watch: new prose in future sessions holds the line (no drift back to "router" as a synonym for selector).
  • v0.7.13 operator narrative structural cleanup (opened 09-29). 22 old post-ship watches (opened 08-30 through 09-15, ≥14d concluded) archived via display:none. Recent Activity today entry compressed 5k→1.7k chars. CLOSED CLEAN summary line below expanded to catalog every archived entry. Watch: next session's Recent Activity stays landmark-only (memory owns full detail).
  • v0.7.11 sr × nor_easter L3 bypass — TEMPORARY (opened 09-29). Added sr × nor_easter × 12-23h and 24-47h to l3_nbm in skip_table_nbm_curated.json. Circuit-breaker per feedback_fresh_fire_vs_circuit_breaker_frames: known layer/regime/direction, two-day worsening, n=220 above threshold. Watch: (a) tomorrow's pair-log 12-23h backstamps show applied_layer=l2_nbm (not l3_nbm); (b) nbm_regression_sentry sr.l3_nbm exits HOT; (c) 24h/12h sr Total Lift recovers toward raw NBM MAE. Re-review 2026-10-13: either remove (regime faded, walkforward proposes cleanly) or keep as formal skip. Reversal = delete the two entries in the JSON.
  • v0.7.12 debug narrative de-staled (opened 09-29). Humidity row narrative rewritten event-based (no drifting numbers); two NBM cascade "which side wins" summaries replaced with pointers to live tiles; cm-dropped note added to L3_NBM sentence. Watch: whether future ship notes stay landmark-only (no live numbers in prose) — recurring drift class.
  • v0.7.10 h_cc_derivation format guard (opened 09-29). Two pct()=None format crashes fixed in per-day and per-regime tables. Substantive verdict unchanged (long-standing PROMOTE: derived-random beats prod cc +47.9%). Watch: next digest shows h_cc_derivation OK (was FAIL(1)).
  • v0.7.9 chp dynamic gate flipped (opened 09-28). CHP_CELL_GATE_ENABLED = True in ch_persistence_gate.py. 9 cells cleared 7-day gate (5 dynamic-only new suppressions: ne_flow/6-11, ne_flow/12-23, ne_flow/24-47, nw_flow/24-47, se_flow/12-23). Watch: (a) h_ch_persistence_blend_stage2_vs_l6 WATCH count drops from 4 live losing cells to ≤1 within a week; (b) any of the 5 new dynamic-only cells flipping OUT of gate = churn signal, revisit threshold; (c) ch pooled Prod-vs-L6 delta should improve ~5-10% MAE per 09-21 66-day anchor. Reversal: flag back to False; static _CELL_SKIP retained.
  • v0.7.8 regime-source reconciliation (opened 09-28). forecast_snapshot.py field loop now classifies regime inline from entry values with cloud_cover (same signature as forecast_error_log). Watch: (a) 09-29 VERIFY PARTIAL PASS — off-curated ratio dropped 84%→41%; residual (1,126 rows) is entirely nor_easter, a regime not in the curated table at all. Fix confirmed working; the residual is a curation gap, not divergence. (b) Non-nor_easter off-curated: 26 rows. (c) New workstream: nor_easter needs per-regime ω (universal 0.44/0.27 hurt −13 to −152%; best-ω is HRRR-favoring). Schema extension pending n≥400 on h/nor_easter/12-23 (~10-04). (d) v0.7.6 apply-flip partially unblocked — 2 SHIP-READY cells emerging.
  • v0.7.7 mechanism attribution stamp (opened 09-27). {f}_selector_mechanism pair-log tag written by pick_source_with_mechanism(). Watch: (a) 10-03 v0.7.5 verdict now attributable per-mechanism — filter pair-log by selector_mechanism=='ims_threshold' for ch. Sr has zero learned_gbm rows because curated table has no sr cells — verdict scope narrowed to ch-only. (b) analysis/l1_static_blend_shadow_verify.py retro scorer available in digest.
  • sr regression — L3_nbm / nor_easter watch (opened 09-28, worsening 09-29). Post-v0.7.7 sr rows losing to raw_nbm on nor_easter cells. 09-29: nbm_regression_sentry sr.l3_nbm HOT +8.1%→−46.7% (Δ +54.8pp), worsened vs 09-28's +5.8%→−1.9%. NBM skip-add proposal l3_nbm sr nor_easter 12-23h now n=220 lift −47.2% (up from n=39 −74% on 09-28), above skip-threshold, but 14d+50d two-window verdict not yet cleared (nor_easter regime too new). Router correctly picks NBM (band_pool); downstream l3_nbm correction inflates error on low-solar. NOT a v0.7.5 issue — GBM never fires on sr. Still held on emergency skip per fresh-fire lucky-baseline discipline. Watch: 14d+50d gate should promote cleanly within ~1 week if regime persists.
  • L3_FIELDS drop cm (opened 09-15 v0.6.625). Walkforward gate cleared 7/7 on the drop proposal (L3 fc -1.6% / obs -1.2% pooled, no regime × band WIN, calm/24-47h LOSS -17.2%). Same fix pattern as ws v0.6.397, cc L4 v0.6.515. Watch: (a) next digest divergence report should show L3_FIELDS AGREE (was READY, drop cm); (b) pair-log rows for cm at calm/24-47h stop stamping applied_layer=l3; (c) cm pooled prod-vs-raw should not regress vs prior 30d — the L3 correction was net-negative, so removing it is a small positive expected.
  • Digest streak = same-proposal (opened 09-15 v0.6.625). build_executive_summary.py streak walkback now requires matching normalized verdict text. Watch: next digest walkforward_l3l4_validator shows 7/7 (or drops out of ship-eligible once drop-cm shipped verdict text changes). Any other ship-resolution script's streak count should stay reasonable — sudden resets on a stable proposal would indicate the normalizer is too aggressive (strip too much) or too lax (miss meta-clauses).
  • KILLED_LAYERS 09-05 pruned (opened 09-15 v0.6.626). Removed (ch, chp_nbm) and (h, l3_nbm) from nbm_regression_sentry.py registry. Both windows post-date the kill so sentry reads natural nominal/THIN. No runtime change. Watch: sentry verdict lines for ch.chp_nbm and h.l3_nbm should read CLEAN/THIN, never HOT/WATCH. (cc, l4_nbm) 09-08 and (cc, l3_nbm) 09-10 remain — prune 09-16 and 09-18 respectively when sustained window post-dates.
  • nbm_skip_add_audit filters shipped (opened 09-15 v0.6.627). Audit now drops proposals already in skip_table_nbm_curated.json before evaluating. Watch: next digest's CONFIRMED list should show only genuinely new proposals. If a shipped cell somehow appears, filter key format may not match — check the audit against a known-shipped cell that would previously have surfaced.
  • Publisher CF redeployed (opened 09-15). make deploy-publisher at 14:00:58 UTC picked up v0.6.615's 12h WINDOWS addition to analysis/per_field_scoring.py and analysis/scoreboard_v2.py. Publisher had been running 09-13 code for 2 days, blanking the 12h per-field diagnostic table. Watch: (a) hourly GCS publish continues to write a 3-window per_field_scoring.json; (b) any future analysis edit that touches a publisher-hosted script triggers a paired publisher redeploy — [[feedback_deploy_hygiene_publisher_pairs_analysis]].
  • pr L2 unwire nw_flow/6-11h (opened 09-12 v0.6.590). Three-tool ship: retro Δ-13.6% over 1,447 pairs since 08-13 (both halves negative -8.5/-17.7), layer-shape sentry pr/production@6-11h +10.6%, yesterday's Notable Calls pr 6-11h -8.7%. Watch: pair-log MAE at nw_flow/6-11h should return toward raw as new obs stamp applied_layer=l1; layer-shape sentry for pr 6-11h should clear within days. Stability re-read scheduled 09-19 — surviving nw_flow/0-5h should stay halves-positive.
  • inter_model_spread Stage 2 7-day gate (opened 09-12 v0.6.595). 33 SHIP cells cleared |premium| ≥ 30% halves-stable in 14d test window. Gate: SHIP set must stay stable across 7 daily reads before wire. Daily verification via h_inter_model_spread_c1_stage2.py output. Earliest wire flip to axis_6: 2026-09-19. First new C1-axis candidate to reach Stage 2 since cross_run_spread in June.
  • L1 3-way walker + runtime NWS routing (opened 09-13 v0.6.601, extends 09-12 v0.6.597). Walker adds NWS as 3rd direction; runtime pick_source() returns "nws" with `entry[f"{f}_nws"]` swap in forecast_snapshot.py. dp gated at wire via _NWS_FIELDS_WIRE_ELIGIBLE = {t, ws, wd, pp} pending option B design (Magnus consistency at hourly-array level cascades to h/AH/feels-like — see [[project_nws_dp_coherence_wire]]). Today: 3 dp cells cleared escalation but fall through to pool; runtime is a no-op today by design. Watch: first non-dp NWS cell (t/ws/wd/pp) to clear the 3-way fitter — the runtime wire fires immediately. Cell-stability across daily reads for the 5 dp cells (informational only until option B lands).
  • l3_nbm.wg.nw_flow/12-23h skip (opened 09-13 v0.6.602). Two-window CONFIRMED: 14d n=472 lift -3.8%, 50d n=1,545 lift -5.0% halves -8.87/-1.86. Watch: pair-log MAE at wg.nw_flow/12-23h should trend back toward raw as new obs land with the l3_nbm skip in place.
  • l3_nbm.wd.se_flow/12-23h skip (opened 09-14 v0.6.609). Two-window CONFIRMED: 14d n=953 lift -4.60%, 50d n=2,581 lift -9.79% halves -14.24/-5.26. Passes v0.6.574 two-window gate. Watch: pair-log MAE at wd.se_flow/12-23h should trend back toward raw as new obs stamp applied_layer=l2_nbm instead of l3_nbm.
  • l3_nbm.wd.se_flow/0-5h skip (opened 09-14 pm v0.6.622). Two-window CONFIRMED afternoon: 14d n=209 lift -3.6%, 50d n=808 lift -8.9% halves -13.4/-3.0. Completes wd.se_flow short-lead pattern — all three bands (0-5h, 6-11h, 12-23h) now skipped. Watch: pair-log MAE at wd.se_flow/0-5h should trend toward raw as new obs stamp applied_layer=l2_nbm. Same shape as v0.6.609.
  • h L2 soft_ramp retune (opened 09-14 v0.6.618). H_SOFT_RAMP_FLOOR 0.1→0.4, H_SOFT_RAMP_END 10→24 in corrected_hourly.py. Addresses digest top-alert τ-suspect (h/production helps 0-5h -43.7% but hurts 24-47h +6.6% under old shape). Source: h_l2_shape_sweep STAGE 1 PROMOTE 7/7 rolling days, halves-stable at 24-47h. Watch: (a) 09-21 the 7d walker window fully rotates post-ship — τ-suspect should clear; (b) per-band vs raw stays positive across all 4 bands (0-5h +48.22%, 6-11h +11.10%, 12-23h +2.72%, 24-47h +1.46% projected).
  • frontal_detector_health daily verdict (opened 09-14 v0.6.619). analysis/frontal_detector_health.py runs in daily digest, emits Verdict: HOLD/CLEAN line based on threshold-vs-observed distribution + type='cold' reachability + miss-rate. Currently HOLD — 4 events / 336h, type='cold' never classified, runtime missed 9/13 candidates (pre-v0.6.620 threshold data). Watch: verdict auto-flips HOLD→CLEAN as new obs accumulate under 4.0°F threshold; first CLEAN expected within days of first cold-front-shaped signal firing.
  • Frontal detector calibration + diagnostic logging (opened 09-14 v0.6.620/v0.6.621). DP_DROP_THRESHOLD 8.0→4.0°F (was above p99.9=6.3°F, unreachable). Diagnostic print(" frontal: score=…") when score≥1 for miss-rate traceability (v0.6.620 used logging.info which Cloud Run drops; v0.6.621 fixed to print(..., flush=True)). Watches: (a) first frontal: score=N sigs=... log line via gcloud functions logs read myweather-collector --region=us-east1 --gen2 | grep frontal whenever any signal fires; (b) first type='cold' event tagged in frontal_events_log.json; (c) if any logged score≥2 line does NOT produce a corresponding events-log write, the miss-rate root cause becomes traceable; (d) C1e "post-front" pool re-fits over ~4 weeks as c1_confidence_calibration_v2.py rolls its window — unblocks [[project_ch_24_47h_c1d_c1e_split]].
  • HRRR-wire cells first-gated live (opened 09-14 v0.6.609 deploy, walker generated 10:57 UTC, deploy 11:06 UTC). 3 cells cleared 3-day gate: cc/ne_flow/12-23, wd/se_flow/24-47, ws/sea_breeze/24-47. The last was already firing via 09-11 escalation clause — today formalizes gate-clearance. Watches: (a) pair-log MAE at each cell should trend toward the HRRR raw MAE as new obs stamp selector_source=hrrr; (b) any of the 3 flipping OUT of the walker within days = churn signal, revisit MIN_LIFT threshold; (c) 5 flipped_in_window cells (walker-noted) not eligible for wire without operator review.
  • stagnant_high regime label (opened 09-13 v0.6.605). New synoptic bin — ws<5 mph AND cc<0.40 AND |pt_3h|<0.5 hPa, checked before frontal/calm in classify_synoptic_regime(). Retroactive pair-log sweep at these thresholds: ws +16.6% NBM-vs-HRRR lift under stag (vs +3.7% non-stag); sr +35.4% (vs +13.8%); ch reverse anti-signal −46.1% vs +34.6%; t not clean. Stamp-only ship — walker consumes automatically. Watches: (a) first labeled rows on next joiner write (~16:07 UTC); (b) stag×ws + stag×sr cells escalation-wire earliest (magnitudes above 20% clause); (c) stag×ch cell should route to HRRR (reverse direction) when it clears the 3-way walker's hrrr-wire gate; (d) stag rate expected around 2% of pair-log rows in typical regimes, higher this week during current stagnant-high pattern.
  • HRRR PBL morning-overshoot routing gate (opened 09-13 v0.6.606). Named hardcoded gate: field=="t" AND regime=="stagnant_high" AND hour_local ∈ {4,5,6,7,8} (EDT) → NBM. Highest precedence in l1_selector.pick_source(). Diagnosis: pair-log dig showed HRRR MAE 2.74-3.40 with bias +2.48 to +3.13°F at UTC 09-11 (EDT 05-07); NBM MAE 0.59-0.67 same hours. Boundary-layer scheme mixes down aloft warm air too aggressively during morning heating under stagnant clear-air. Watches: (a) t 24h Total Lift trends toward zero as morning-hour rows use NBM; (b) selector_picks for t stagnant_high show NBM dominant in EDT 04-08 window; (c) walker stagnant_high × t × 0-5 cell clears wire (4-7d) → prune this named gate; (d) reversibility via HRRR_PBL_MORNING_OVERSHOOT_KILL.
  • h Stage 1 halves watch (opened 08-30 v0.6.520) — CLEARED same-session v0.6.522. Morning MARGINAL (halves −0.93 / +3.12) was methodology artifact. /code-review high exposed 3 real load-path bugs (fc_prod L2 contamination + test-window off-by-one + halves-crash-on-n=0); shared harness fix moved h to +24.04% STAGE 1 PROMOTE (halves +1.28 / +0.88 BOTH WIN). Sibling numbers moved: dp +21.25 → +28.83, wg → +36.53. Stage 2 harness extracted v0.6.524; h Stage 2 written same session. See [[project_h_residual_persistence_attribution_08_30]].
  • h Stage 2 walker watch (opened 08-30 v0.6.524; automated 08-31 v0.6.529). Day-1 rollup 9 SHIP / 8 MARGIN / 9 SKIP / 7 THIN. SHIP cells today: nw_flow all 4 bands, sea_breeze 6-11/12-23/24-47, sw_flow 12-23/24-47. Walker (h_h_residual_persistence_walker.py) accumulates per-cell verdicts to .cache_h_residual_persistence_walker_history.json and gates the Stage 3 flip on 7/7 consecutive days in {SHIP, MARGIN} per cell. Earliest clear ~2026-09-06.
  • Stage 1 harness halves-preference watch (opened 08-31 v0.6.528). Bug fixed: grid picked max-test-MAE (window=5d for h today, halves ✗) over halves-stable window=14d (halves ✓ +0.56/+4.10) and downgraded h verdict to MARGINAL. Now iterates every combo, computes halves, picks halves-passing max. h corrected to STAGE 1 PROMOTE at 14d (+24.51%); dp/wg unaffected. Wrote [[feedback_grid_select_halves_stable]]. Watch: does daily verdict now stay PROMOTE across the 7 walker days? Any override-line firings on wg/dp? Analysis-only fix — no runtime change.
  • h_residual_persistence Stage 3 shadow watch (opened 08-31 v0.6.530 collector deploy). Processor pre-staged (cloned from wg template with FIELD=h, HOURLY_KEY=corrected_humidity, [0,100] both-ends clamp for RH). Wired into collector.py ENABLED=False; all 4 v0.6.525 bugfixes baked in. Deploy-verified 06:33 UTC tick: nw_flow fires {0-5:5, 6-11:6, 12-23:12, 24-47:24} (47/47), clamped_out_by_band present with zeros, corrected_humidity_post_l3_pre_hrp absent (ENABLED=False semantics correct). 7-day shadow accumulation pairs with the walker cell watch. Ship-day (~09-06) is now a one-line ENABLED=False→True flip + redeploy.
  • Stage 3 processor bug-fix watch (opened 08-30 v0.6.525 collector deploy). 4 real bugs fixed in wg_residual_persistence.py + dp_residual_persistence.py: (1) telemetry pre/post-clamp mismatch (wg only — dp doesn't clamp), (2) _TABLE_CACHE never invalidated (mtime check + MYWEATHER_REFRESH now respected), (3) clamp-out silently pooled with normal skips (new clamped_out_by_band counter), (4) no_hourly_array early return skipped gate_firing_log.record_firing. Post-deploy verify 18:17 UTC tick on new revision myweather-collector-00544-fuj: clamped_out_by_band key live with expected zeros on nw_flow regime; table_generated_at reflects fresh v0.6.524 Stage 2 refit (mtime cache invalidation working). Watch 7 daily digest reads for gate_firing_rollup regressions vs pre-fix baseline.
  • Lc emergency intervention (v0.6.389d-g + v0.6.390, 07-30) — active state, no close date. cl + cc BOTH off Lc via _FIELD_SKIP. cl held out (walk-forward: pool_vs_raw −22.15%, reg_vs_raw −30.37%). cc retired v0.6.390 architecturally → Ccd derives max(cl_l6, cm_l6, ch_l6). cm + ch untouched, +34-50% vs raw held-out. Next step for cl: EMA/Kalman shift tracker OR recent-bias gate. See [[project_lc_regime_conditional]] + [[project_lc_regime_stage1_pool_prereq]].
  • wsbp — HELD (v0.6.388, 07-28). Sibling of dpbp, calm-only, sign-inverted. ENABLED=False. Preflight 08-04: calm regime n=0 in shadow window. Wait for calm regime accumulation. Cost of waiting: nothing (dpbp covers dp side).
  • l6_fix_b_refit — HOLD-GATE (rolling gate added 08-08). Was single-day SHIP +2.15% today, HOLD at +0.29% 26 days prior; no consistency check between. Rolling gate needs 7 distinct days + same ship_bins before verdict can promote. Lt stays on do-not-reopen. See [[project_l6_fix_b_rolling_gate]].
  • wg persistence-skill thin margin — pooled L4 skill +0.17 today (was +0.19 08-29; < +0.20 hold margin). BUT per-cell check 08-30: Prod-vs-persistence skill +0.215 pooled (above +0.20 margin, healthy at 6-11h/12-23h/24-47h; 0-5h at parity is structural — wind_blend uses recent obs). The "at-risk" line measures L4-vs-persistence; users-see production is fine. Not a real live-product regression. Re-flag only if Prod skill drops < +0.20 for 2+ days OR any of 6-11h/12-23h/24-47h halves specialist contribution.
CLOSED CLEAN 2026-08-26 → 2026-09-15 (14d+ post-ship, all watches ran to conclusion): pr L2 regime-gated, chp diurnal gate, walkforward L3/L4, C1h re-curate, Lsb, dpbp, wg L3 SKIP_TABLE, chp cell-skip variants, cc.l4_nbm HOT sentry, NBM skip-table 14d, L3_FIELDS drop cm (09-15 v0.6.625), digest streak same-proposal (09-15 v0.6.625), KILLED_LAYERS 09-05 prune (09-15 v0.6.626), nbm_skip_add_audit shipped-filter (09-15 v0.6.627), publisher CF redeploy (09-15), pr L2 unwire nw_flow/6-11h (09-12 v0.6.590), inter_model_spread Stage 2 gate (09-12 v0.6.595), L1 3-way walker + NWS routing (09-13 v0.6.601), l3_nbm skips wg.nw_flow/12-23h + wd.se_flow all-3-bands (09-13/14), h L2 soft_ramp retune (09-14 v0.6.618), frontal_detector_health + calibration (09-14 v0.6.619/620/621), HRRR-wire cells first-gated (09-14 v0.6.609), stagnant_high regime label (09-13 v0.6.605), PBL morning-overshoot gate (09-13 v0.6.606), h Stage 1/2/3 walker chain (08-30/31), Stage 3 processor bug-fixes (08-30 v0.6.525). Full history → Archive → Post-ship watches
🧪 Architectural backlog — longer-horizon items, no date target
  • MLC — marine-layer correction sandbox. Stamps weather_data["marine_layer_correction"] every tick, ENABLED=False. In-bin cc bias collapsed 06-30 (real break, not the 07-07 "cliff" the anomaly detector reported — that was cumulative-window artifact). Split at 07-04 (cm HRRR-anomaly onset): stratum-local, pre-HRRR, rules out HRRR-anomaly causation. Companion cl signal points seasonal. Hold OFF indefinitely. Will not re-arm when cm HRRR anomaly clears — different event. Redesign candidate: time-of-year gating on the MLC bin.
  • L2-as-observation-only. Remove L2 from forecast pipeline, keep as training target for L3/L4/Lsr only. Architectural exploration.
  • Tight-τ cloud bias propagation across leads 1–3h. Pearson at lead 1h is +0.50/+0.57/+0.55/+0.66 for cc/cl/cm/ch — strong, clean; degrades by lead 3h. Architectural slot: short-lead τ-decayed bias propagation (~2h half-life) on top of the existing hourly[0] override. Needs Stage 0 MAE-impact estimate (probably small — first 3 leads only). Not blocking; logged for backlog. Spun out from resolved KBOS+KBVY cloud blend (2026-06-30, ρ ≤ 0 at lead 6h — hourly[0]-only formalized as intentional).
Recent activity — today + 2 prior days (older entries in docs/CHANGELOG.md, trimmed on next curation)
Badges: PIPELINE · DISCOVERY · INFRA · DASHBOARD · PREFLIGHT

Applicability map — what corrections trigger, why, and when they actually fire

Two lenses on the same object. Applicability (top) describes what's configured to fire and under what gates — built each tick from describe_applicability() in each correction module, so the page can never drift from the code. Runtime firing frequency (bottom, 7-day rolling) shows what's actually firing per operator × field × regime, so silent dormancy (config says "enabled" but the code path never mutates a value — the class of bug that hid ws L3 for 4 days after v0.6.279) surfaces as an ★ on cells with 0 fires despite ≥5 ticks in that regime.

How to read this
Three categories: General-purpose layers (L1–L4) can apply to any field; the per-field applicability set controls which actually trigger. Specialists (Lsr, Lt, Lc) are domain-scoped by construction — the physics of the correction binds them to a field type. Confidence (C1) doesn't modify forecast values; it widens or narrows uncertainty bands along orthogonal axes (transition, pressure tendency, mesonet spread, pre-frontal proximity) that cross multiple (field, band) cells.
HRRR + NBM cascades render together below. HRRR-side layers (L2 mesonet blend, L3 decay, L4 diurnal, Lsr, Lc, chp/clp/wdp/dpbp/wsbp/lsb, C1) plus NBM-side layers (l3_nbm, l4_nbm, l5_nbm, l6_nbm, skip_table_nbm) are all unioned into the same block. Selector (l1_selector) picks between HRRR-Prod and NBM-Prod per (field, lead-band) after both cascades finish. L2_NBM as of v0.6.499 is a NATIVE fit (not a delta reconstruction) — described in the hand-curated block below.
applies when is the predicate in plain text. gated by names the module-level constant the gate reads (omitted if always-on). current state resolves the constant for the reader — what the gate is doing this very tick.
Terminology key: applies = correction applied to a (field, row) this tick · enabled = feature switch (the module-level constant is on) · active = runtime branch (which code path actually fired).
Source: weather_data.applicability_map.layers. Schema: weather_collector/data/applicability_map_schema.json. If this section says "not available," the collector hasn't shipped the block yet — it landed v0.6.260.
L2 — Aggregate bias (mesonet blend) general-purpose hand-curated · promote to describe_applicability()
field applies when gated by current state
t Always — additive bias from station network, scaled by Kalman gain K always on applies
dp Always — additive bias from station network always on applies
h Always — additive bias scaled by Kalman gain K always on (K-taper 1.0 → 0.4 by lead 24h) applies
pr RE-ENABLED 2026-08-10 v0.6.401 with regime gate — K=1 additive, τ=8h fitted, applied on (regime, lead_band) ∈ {(nw_flow, 0-5h), (nw_flow, 6-11h)}; every other cell stays raw. Cell selection from analysis/pr_l2_regime_lead_retro.py 08-10 (Jaccard 0.50 STAGE 1 SHIP, both-halves winners A +21.8%/B +41.6% and A +10.3%/B +13.1% on 8,596 shadow rows). History: DISABLED 2026-07-01 v0.6.276 after pooled Production +2.4% BAD — but that pooled read was regime-blind. The nw_flow short-lead win was real all along; needed regime-conditional cross-cut + halves verification to surface. Shadow-wire since 2026-07-29 v0.6.389 supplied the fresh data. regime-gated (current-tick regime × lead band) applies on nw_flow 0-11h; other cells raw
cc Always — Kalman-blended KBOS + KBVY METAR override on hourly[0] (current-hour value used by cards). Not propagated across forecast leads — promotion to a true per-lead L2 bias correction is queued (see Open architectural questions). always on (needs KBOS or KBVY at current hour) applies
cl / cm / ch Derived — same Kalman K applied to L/M/H splits on hourly[0] to keep them self-consistent with cc. Not propagated across leads, same as cc. always on (derives from cc) applies
ws / wg Always — direct selection from per-octant median (no per-station bias track); KBOS+KBVY authoritative-source floor when both agree at >1.4× the octant median always on (direct-selection, not additive) applies (direct)
sr / pp / pa N/A — no station network for this field; L2 row is structurally empty n/a — no obs network n/a
wd Always — SHIPPED 2026-07-20 v0.6.368a. Circular unit-vector blend in wind_blend.py: obs wd + fc wd → (sin, cos), weighted mean by linear decay over 24 leads, atan2 back to degrees. Solves the wrap-around that made "N/A — L2's linear math doesn't apply" true prior to this ship. Calm-floor guard WIND_DIR_MIN_SPEED = 3.0 mph skips cells where both obs and fc speed are below floor. raw_wind_direction preserved for baseline. Complementary wd persistence gate (Stage 2 07-20 v0.6.365) targets long-lead regime-transition cells where L2 decay has expired. always on (needs both obs+fc speed ≥3 mph) applies (circular)
L2_NBM — NATIVE fit (v0.6.499) general-purpose · NBM cascade hand-curated · not a standalone module
v0.6.499 (2026-08-26) — L2_NBM is now NATIVE for all 8 NBM-scope fields. Full delta-transfer retirement. Every family mirrors its HRRR counterpart, re-anchored against NBM raw:

All native math shares the same station observations / Kalman gain / decay curves the HRRR side uses. The only substitution is which raw forecast the correction anchors against. This is exactly the fix the 08-26 nbm_l2_delta_audit prescribed: the station-network bias IS model-independent, but the anchor (raw at h=0) is not, and the delta transfer conflated the two.

field applies when gated by current state
t / dp / h Always — receives HRRR L2's Kalman-scaled additive bias delta rides on HRRR L2 applies
ws / wg Always — direct-selection delta from HRRR L2 wind blend rides across rides on HRRR L2 applies (direct)
wd Always — circular unit-vector delta from HRRR L2 wind_blend rides across (calm-floor 3 mph inherited) rides on HRRR L2 applies (circular)
cc / ch hourly[0] only — inherits HRRR L2 KBOS+KBVY METAR delta on current hour; not propagated across leads (same limitation as HRRR L2 cloud fields) rides on HRRR L2 applies on hourly[0]; other leads pass through
sr Passes through raw — HRRR L2 has no sr correction (no station network), so delta is 0. n/a — HRRR L2 sr is n/a l2_nbm = raw_nbm
cl / cm / pp / pa / pr Not emitted by NBM CO grib — permanent HRRR-only. Selector never picks NBM for these; no L2_NBM row exists. n/a n/a
🔍 L2_NBM soundness — 30d field × lead measurement
Grades the HRRR-delta transfer per (field, lead-band). Positive lift = delta helps NBM (assumption vindicated); near-zero = neutral (delta washes at long lead by construction — K-taper); negative = delta HURTS raw NBM for this cell (Wyman Cove bias is model-dependent). ★ = both halves of window agree on sign. Source: analysis/nbm_l2_delta_audit.py → gs://myweather-data/nbm_l2_delta_audit.json.
loading nbm_l2_delta_audit.json…
L1 blender shadow status v0.7.3 · shadow only · fresh-data since 2026-09-24 16:00 UTC
Per-cell verdict on whether the blender's shadow forecast is beating what the selector actually served. SHIP-READY = 7d n≥50, lift ≥ +3% vs served, ω drift ≤ 0.20 · HOLD = insufficient fires or below threshold · KILL = 7d lift ≤ −3% (blend hurts) · THIN = zero fires (regime hasn't matched). Add cells to BLENDER_APPLIED_FIELDS in weather_collector/processors/l1_selector.py only after SHIP-READY on fresh data. Caveat: the 13 rows in the table reflect the stale-window l1_blender_curated.json shipped v0.7.0. Table rebuild after re-curation on fresh data will trim it to the 2-3 cells that survive halves-stable. Source: analysis/l1_blender_shadow_verify.py → gs://myweather-data/l1_blender_shadow_verify.json.
loading l1_blender_shadow_verify.json…
Runtime firing frequency — 7-day rolling
Fires = the correction actually mutated a value at that lead (not "would have applied"). Skips = would-fire cells suppressed by the skip table (L3 ws in ne_flow / short-lead sea_breeze; Lsr sr in ne_flow / calm) or by a gate being OFF (MLC currently ENABLED=False → all matching cells count as skips). Rate/tick = fires ÷ ticks_in_regime; for L3/L4 with 48 leads, healthy is close to 48. ★ marks (operator × field × regime) cells with 0 fires despite ≥5 ticks in that regime — silent-dormancy candidates.
Source: gate_firing_rollup.json (nightly digest via analysis/gate_firing_rollup.py) over per-tick gate_firing_log.jsonl (written by weather_collector/processors/gate_firing_log.py, v0.6.318). NBM operators added v0.6.458 (F4, 2026-08-21): L3_NBM, L4_NBM, L5_NBM, L6_NBM, CHP_NBM, WDP_NBM now emit per-tick fires/skips from forecast_snapshot.py; expect them to appear in the 7-day rollup once the rollup script's next digest cycle picks up the newly-accumulating log rows.

L1 — Raw model input (parallel HRRR + NBM cascades; selector picks per cell)

Two forecast sources run in parallel for every hour: Open-Meteo HRRR/GFS on one side, direct NBM (National Blend of Models) on the other. Each side runs its own full correction cascade. The 🎯 selector below picks per (field, lead-band) which cascade's output becomes user-visible. Multi-kilometer grid resolution on both — still knows nothing specifically about Wyman Cove; that's what the layers below add.

About the L1 sources (post-selector-arm v0.6.445)
HRRR-side (Open-Meteo). Open-Meteo HRRR (next 48h) + GFS (days 3-7), fetched every 10-min tick. Feeds the full HRRR cascade: L1 → L2 mesonet blend → L3 lead-decay → L4 diurnal → L5 Lsr (sr) / L6 Lc (cm/ch) / specialists (chp, wdp, clp, dpbp) / Ccd (cc derivation). Every field starts from a HRRR value at L1.
NBM-side (direct grib ingester). NBM CO 2.5km grib byte-range fetched hourly at :45 UTC by the nbm-ingester CF (v0.6.433). Extracts 9 fields at Wyman Cove: t / dp / ws / wd / wg / sr / cc / ch / h. Full parallel cascade shipped 2026-08-21: raw_nbm → l2_nbm (HRRR-delta reconstruction for all 9) → l3_nbm (per-lead bias, LIVE v0.6.445, scope trimmed 08-25 v0.6.472 to {wg, h, ch, cc}; sr re-added 09-04 v0.6.549) → l4_nbm for ch only (cc dropped 09-08 v0.6.563; diurnal hour-of-day) → l5_nbm for sr KILLED 08-25 v0.6.471 (sentry +238% MAE + walkforward −126% to −147% agreed) → l6_nbm for t (v0.6.453 scaffold, ENABLED=False mirroring HRRR L6) → chp_nbm for ch (v0.6.454) / wdp_nbm for wd (v0.6.437). clp_nbm N/A (NBM doesn't emit cl). Path B discipline (owner call 08-25): any NBM cascade layer that can't beat its own raw gets killed, not routed around — a broken cascade should be diagnosed loudly, not silently masked. Adding raw as a selector fallback was considered and rejected on those grounds.
🎯 L1 selector (Prod-vs-Prod, v0.6.440 fit / v0.6.445 armed). For each (field, lead-band), argmin of recent-window MAE(HRRR-Prod, deepest applied HRRR layer) vs MAE(NBM-Prod, deepest applied NBM layer — l3_nbm → l4_nbm → l5_nbm → l6_nbm → chp_nbm/wdp_nbm as each fires). Fires per l1_selector_table_curated.json. In-scope fields: t / ws / wg / wd / h / ch / sr / dp / cc (all 9 NBM emits). Out-of-scope (HRRR-only): cl / cm / pr / pa / pp — NBM doesn't publish them. Currently 10 cells pick NBM: wg 12-47h, wd 12-23h (added 08-26 v0.6.500 after skip-table expansion + native L2 refit), dp 6-47h, cc all leads.
v0.6.432 L1 router retired 2026-08-19 v0.6.437. Superseded by the selector on all counts: wider scope (9 fields vs 3), stronger evidence base (30-day scoreboard vs 14-day, plus 108-day backfill for L3 fit), post-cascade application (all corrections apply first, selector picks last).
Curves below show the current L1 forecast per field for the HRRR-side only. NBM raw is available on the per-field accuracy charts as a dashed line (raw_nbm) once the Fitter has a day of pair-log coverage.
🎯 L1 selector — per (field, lead-band) source picker (LIVE — v0.6.445, Prod-vs-Prod)
For each (field, lead-band), argmin of recent-window MAE(HRRR-Prod, deepest applied HRRR layer) vs MAE(NBM-Prod, deepest applied NBM layer — l3_nbm → l4_nbm → l5_nbm → l6_nbm → chp_nbm/wdp_nbm as each fires) picks the cascade output that becomes the user-facing forecast. Table refit by analysis/l1_selector_fit.py; stored at weather_collector/data/l1_selector_table_curated.json. Selector runs after all NBM stamping in forecast_snapshot; when picked "nbm", replaces top-level {field} with the deepest applied NBM layer's value (e.g. {field}_chp_nbm for ch when the specialist fires, else {field}_l4_nbm, else {field}_l3_nbm, ... falling through to {field}_raw_nbm). Fall-through to HRRR is safe.
loading l1_selector_table_curated.json…

Raw L1 forecast per field (post-router where the router fired)

L2 — Aggregate-bias correction (local station network)

What our 66 nearby weather stations say the model is getting wrong right now — and how much of that signal we trust.

How the network correction is built Two parallel aggregation paths under this layer, each suited to the noise behavior of its metric:
Temp / humidity / pressure (additive bias): per-station Kalman offset (rolling 48h) → per-octant 1/distance² × exp(-|elev_diff|/30)-weighted (station − model) bias → unweighted mean across non-empty octants. t and h scale by per-field network Kalman K (sec 2c); pr applies full strength.
Lead-decay: bias_applied(lead) = current_bias × exp(-lead/τ) per field. Fitter (03/15 local) refits τ on train/test split; loader adopts fitted τ only if it beat hardcoded default on held-out RMSE AND stays within 0.25×–4× default. Defaults: τ_t=4h, τ_h=240h, τ_pr=12h. Fields without τ (wind/clouds/solar) get flat L2. See sec 2d for live curves.
Wind / gust (direct selection): no per-station calibration, no additive bias. Per-octant MAX gust → MEDIAN across octants; linear-decay blend into next 24h (100% h0 → 0% h24). KBOS+KBVY floor: if both airports agree on speed >1.4× octant-median, defer to their median. Direction rejected if >60° off airport+buoy+Tempest consensus. Gust override allows single-source (METAR omits when steady). Physical floor: gust ≥ wind.
Cloud cover (Kalman METAR blend): KBOS+KBVY sky obs (BKN/SCT/FEW/OVC → percent + L/M/H) blended against HRRR hourly[0] via _kalman_gain_cloud. K=0.90 if airports agree within 20pp, 0.70 at 20-40pp, 0.50 at >40pp, 0.35 for single-source. Same K on total + L/M/H for self-consistency. Live values in cloud_l2_meta.

2a. Octant coverage — where this tick's stations came from

loading…
Per-station detail — mesonet map & Kalman-tracked offsets table
Per-station uptime — fetch success rates (rolling window)

2b. Network bias estimate (full, un-confidence-scaled)

loading…

2c. Network confidence (Kalman gain K)

loading…

2d. Lead-decay applied to L2 bias (v0.6.44)

loading…

2e. Post-aggregate-bias forecast — what gets passed to L3 · engineering view (pre-clamp)

⚠ These are internal pipeline values, not user forecasts. After the aggregate-bias correction is applied but before downstream layers and physical-bounds clamping (FIELD_BOUNDS in decay_apply.py). Values can legitimately fall outside physical ranges here — cloud_cover 121%, precip_probability −6%, precip_amount −0.025 in — because the bias offset is additive and the clamp comes downstream. If any of these look wrong for user display, check the L3 / L4 / cloud-clamp path, not this section.

L3 — Lead-decay correction

How the model's error tends to grow with each hour into the future — and the per-field nudge we apply to push back against that drift.

How decay correction works For each field and each lead hour, we learn from millions of historical pairs how far off the model usually is — then bend the forecast back toward truth by that amount. Some fields (wind, gusts, high & mid cloud, POP) genuinely benefit; others net-negative on held-out data and are paused (see Applicability below). Tracked over a 30-day window with an exponential recency weighting — default τ=14 days, with per-field overrides for fields where analysis/decay_tau_tuning.py shows ≥5% MAE improvement vs the default. Current overrides: pp (POP) at τ=28d (+11.1% held-out, 2026-06-21 v0.6.167; note: today's read shows only +3.2% vs τ=14, now below the 5% floor — flagged for next τ-audit day). pa (precip amount) reverted 2026-07-19 v0.6.358 after today's read landed +0.9% vs τ=14 (was +5.9% yesterday); pa's best-τ swung 28→42→7 across three IMPLEMENT reads. See [[feedback_tau_streak_gate_limits]].

Applicability

Held-out MAE audit picks the fields where L3 actually beats L2 — currently wind speed, gusts, high cloud, mid cloud. Other fields stay off because the correction at best ties (temperature, pressure) or actively hurts (humidity, dew point, solar, low cloud, precip). POP is the special case: it's evaluated by Brier score, not MAE, so the audit's MAE-based ⚠ rule is suppressed for it — the v0.6.20 calibration analysis showed flat-additive correction cuts Brier 5%. Currently applied: . Brier-evaluated: . See Applicability map for live triggers and per-field gates.

Live state — with vs without decay correction

Historical calibration — fitted correction curves per lead hour

Calibration history — decay curves over time

L4 — Diurnal correction

A separate correction for the part of model error that follows the sun — e.g., bias that's different at 3 AM than at 3 PM. Currently applied to high cloud and cloud cover.

How diurnal correction works Bins historical errors by hour-of-day (0–23) and fits the persistent pattern at each hour. Most fields don't have a clean diurnal signal once L2/L3 have run, so this layer's applicability set is narrow — currently {ch, cc} (cc added 2026-06-24 v0.6.214). Fits for fields outside the applicability set are still computed below for diagnostic purposes. See the Applicability map for the live trigger state.

Applicability

L4 corrects the portion of forecast error that repeats with time of day. It is the hardest layer to earn because the same hour-of-day bias must recur consistently across many days. Most fields fail that test because their dominant errors are driven by changing weather regimes (air mass, cloud regime, frontal timing, etc.) rather than the clock. Currently applied to {ch, cc} — the two fields showing a stable enough diurnal signal to pass the held-out audit. Fits for the remaining fields are retained in the calibration sections below as diagnostics. See Applicability map for live gates.

Live state — what L4 is doing this tick

No dedicated live-state panel yet. Current per-tick L4 deltas land in weather_data.hourly[*].corrected_<field>; the applied-vs-unapplied comparison is visible via the L4 lines on the per-field accuracy charts at the top of the page. Future build-out — placeholder for a dedicated live snapshot mirroring L3 §3.

Historical calibration — diurnal correction curves over time

Calibration history

L4 doesn't currently maintain a separate fit-history view — the curves-over-time panel above serves both "current fit" and "how it's changed" roles. Split out if/when a dedicated current-fit snapshot ships.

Lsr — Synoptic-regime correction (solar specialist)

A per-regime W/m² delta applied to direct solar radiation. The classifier reads current wind direction, speed, pressure trend, hour, and temp; the regime label keys into a calibrated per-regime delta. L1–L4 are trained on general bias correction; Lsr is the first layer trained on synoptic-state stratification — different "kinds of weather" get different corrections. Skip regimes (v0.6.280): ne_flow and calm return 0.0 — l5_solar_analysis showed Lsr hurt sr in those regimes (+32% and +10% worse) while helping everywhere else. Skip regimes are a targeted patch until Fix B / refit against L4 baseline lands.

How Lsr works — Classify → Lookup → Apply
1. Classify. regime_classifier.classify_synoptic_regime(wind_dir, wind_speed, pressure_in, pressure_trend_hpa_3h, hour_local, temp) returns one of nine labels — nw_flow, ne_flow, sw_flow, se_flow, sea_breeze, frontal, pre_frontal, nor_easter, calm (plus unknown when inputs are missing).
2. Lookup. Per-(regime, hour) Δ W/m² from _BIAS_BY_REGIME_HOUR[regime][hour_local] in solar_correction.py; falls back to _BIAS_FALLBACK_BY_REGIME[regime] when the hour cell has too few samples.
3. Apply. Added to every lead's direct_radiation where the lead's raw value is above SUN_UP_THRESHOLD (50 W/m²). Pre-sunrise / fully overcast leads get Δ = 0 regardless of regime.

Applicability

Lsr is a specialist — it only applies to direct solar radiation (sr). Other fields have no Lsr line on their accuracy charts by construction. See Applicability map for the live trigger predicate (ENABLED in solar_correction.py + the sun-up gate + skip regimes ne_flow and calm since v0.6.280).

Live state — what Lsr is doing this tick

Loading…
Read directly via curl -s https://data.wymancove.com/weather_data.json | jq .solar_correction. The per-lead Δ is also baked into the sr card on the Forecast Accuracy chart above.

Engineering status

Where we are (2026-07-28):
Lsr lives in production for sr. Skip regimes: ne_flow + calm. Structural raw-baseline verifier (v0.6.291) catches raw-column mutation drift automatically. Unit-mismatch resolution — Lsb Stage 3 wired 07-17 v0.6.354, ENABLED=False. Overrides Lsr's direct-beam output with (shortwave − bias(hod)) on sea_breeze rows where cc < 25 — narrowed 2026-07-28 v0.6.383b from the original two-sided (cc < 25) OR (cc >= 75) after 07-24 halves re-run showed the overcast half (75-100 cc-bin) actively regressed (Δ=−17.6%, halves +6%/−23%) while the clear-sky half held robustly (Δ=+35.8%, halves +33%/+40%). The "thick attenuation missed" hypothesis for overcast was wrong; data says model attenuation is fine or over-corrected. Middle-and-overcast cc bins (≥25) now all fall back to Lsr's original direct-beam correction. Narrowed-shape Stage 2 re-run (07-28) PROMOTES cleanly: pooled +31.50%, halves +43.78% / +22.69% (both above +10% ship gate), 4/4 lead-bands SHIP. Landed ENABLED=False with fresh 7-day live-layer gate 07-28 → 08-04. Broader unit-mismatch across other regimes (pre_frontal +12, unknown +367 W/m²) still open. Shortwave shadow-log infrastructure stays alive (underpins future sr L2 work once unit mismatch resolves per project_sr_unit_mismatch). Divergence-report LSR_ENABLED source bug fixed 07-20 v0.6.365 (was reading l5_solar_analysis, now routes through .cache_l5_gate_history.json) — status now correctly AGREE, live gate 100% SHIP for 30+ days.
Full Lsr shipped-history log (v0.6.248 → v0.6.291)

Developer notes — classifier, lookup tables, audit details

Regime classifier. regime_classifier.classify_synoptic_regime(...) returns one of nine labels; same source feeds C1a's transition axis. Lsr reads via state.regime_synoptic. Live label surfaces in Production Stack box, R2.
Per-regime delta table — what the lookup contains. Lsr's bias tables live in weather_collector/processors/solar_correction.py as two Python dicts:
Refit cadence: regenerate via python3 analysis/l5_recompute_biases_hourly.py after at least 7 days of accumulated DAYTIME pair rows (raw_solar ≥ 50 W/m²). The daily digest prints a drop-in replacement table; mid-trajectory refits invalidate the audit window — wait for a clean break.
Lsr vs L4 audit — held-out MAE. Primary view: the Forecast Accuracy chart above. The sr card carries five lines (Raw → L2 → L3 → L4 → Lsr); the green box on the Lsr line is the user-visible MAE on rows where Lsr fired. Note: on rows where Lsr's skip regimes fire (ne_flow, calm since v0.6.280) Lsr returns 0.0, so those rows contribute L4-value MAE to the Lsr aggregate — the Lsr line converges toward L4 as skip-regime coverage grows.
Fitter audit: verdict logged to conditional_audits.l5 each Fitter cycle, recency-weighted since v0.6.178 (exp(−age_days/14)). The trailing 7-day rolling gate lives in l5_gate_history.json and surfaces beneath the Lsr row in the S1 Shadow Tuner section below. First actually clean 7-day window closes ~2026-07-10 (7 days from the 07-03 per-lead delta deploy — the earlier ~07-05 date was against the pre-fix deploy timeline).
RESOLVED 07-20 v0.6.365: the divergence report's LSR_ENABLED claim was sourced from l5_solar_analysis — a candidate script testing a simpler regime-only bias lookup, NOT the live hourly Lsr. Its HOLD verdict meant "don't ship the candidate refinement," but the divergence table was rendering it as "retire live Lsr." Live Fitter cycle gate history (l5_gate_history.json) shows 100% SHIP across 30+ days, 13-21% MAE improvement per cycle. Divergence report now routes LSR_ENABLED through the live gate history and correctly shows AGREE.

Lc — Cloud saturation-unbiasing (cloud specialist)

A per-(field, value_bin) percentage-point shift applied to the L4-corrected forecast for cc / cl / cm / ch. The bias pattern Lc unwinds is saturation: the models systematically overshoot at 0-5% cloud and undershoot at high fractions. The same shape holds across all four cloud fields but with different magnitudes. Lc reads each L4-corrected value, keys into a curated (field, value_bin) → shift table where the fitter's SHIP verdict is set, adds the shift, and clamps to [0, 100]. Emergency intervention 2026-07-30 v0.6.389d-g — pooled shift table diverged from recent regime-conditional truth, causing 8× cl MAE blow-up on 07-30. Live SHIP surface after intervention: cl fully off (_FIELD_SKIP); cc ships at 0-5 (all regimes), 50-80 + 80-95 (all regimes except ne_flow), 95-100 KILLED universally; cm unchanged (20-50 / 50-80 / 80-95 / 95-100, all regimes); ch unchanged (20-50 / 50-80 / 80-95, all regimes). SKIP cells (bins that stayed within the ±5 pp noise floor across the 30-day window) pass through unchanged. See Engineering status below for the walk-forward evidence that motivated the interventions.

How Lc works — Preserve → Lookup → Shift → Clamp
1. Preserve. Before mutation, stash the L4-corrected array as hourly.<field>_post_l4. Pair log + debug page read this to attribute L4 vs (L4+Lc) cleanly.
2. Lookup. For each lead 0-47, determine the value bin from [0, 5, 20, 50, 80, 95, 100]. If the (field, bin) cell has SHIP verdict in lc_correction_table.json, read its Δ pp; else Δ = 0 (SKIP — the L4 value passes through unchanged).
3. Apply. Add Δ to the value and clamp to [0, 100]. Writes back to hourly.<field>.
Sign convention: the fitter's stored bias is (forecast − observed); the applied shift is −bias, pulling the forecast toward the observation.

Applicability

Lc is a specialist — applies only to cc / cl / cm / ch. Fires when the L4-corrected forecast falls in a SHIP-verdict value bin (16 of 24 cells; see Applicability map for the full per-cell shift table + verdict list).

Live state — what Lc is doing this tick

Loading…
Read directly via curl -s https://data.wymancove.com/weather_data.json | jq .cloud_saturation_correction. The per-lead Δ is also baked into cc / cl / cm / ch cards on the Forecast Accuracy chart above.

Engineering status

Where we are (2026-07-30 — emergency intervention day):
Overall Prod-vs-Raw compressed from ~−13% a week ago to −5.2% today, driven by cl and cc l6 MAE blowing up. On 2026-07-30: raw cl MAE 7.29 (model + obs both mostly-clear) but Lc-corrected cl MAE 56.96 (8× worse). Same on cc (raw 7.21 → l6 44.23). Diagnosed via 4 new analysis scripts (Stage 0 regime × bin sweep, Stage 1 halves-strict fit, walk-forward validator, rolling-window sweep) as an architectural failure of the shift-table itself: the historical fit table subtracts 46-88pp at overcast bins because model USED to over-forecast overcast heavily, but the model no longer does — bias shrunk 4-38× or sign-flipped in the last 3-10 days. No rolling window length recovers cl on held-out (best W=3d still −3.7% vs raw). Regime-conditional slicing doesn't fix it either (walk-forward: cl reg_vs_raw −30.4%). Emergency bandages shipped v0.6.389d-g: cl fully off Lc, cc/95-100 universal bin-skip, cc/ne_flow/50-80 + cc/ne_flow/80-95 regime-conditional demote. cm and ch untouched (walk-forward +34% and +50% vs raw on held-out — still helping). Original 14-day post-ship watch (07-17 → 07-31) cannot close routinely. New tracking axis is the regime-conditional Lc Stage 1 gate (7-day walk-forward stability starting today) — but even that's superseded by the walk-forward validator's finding that the naive regime-conditional shape isn't the fix. Architectural next step (unshipped): EMA/Kalman shift tracker OR recent-bias gate on the existing lc_fit table. Both multi-day. Contingency: per-bandage reversibility is a single-line edit to _FIELD_SKIP or _CELL_SKIP frozenset in cloud_saturation_correction.py. See [[project_lc_regime_conditional]] for the full pipeline state.
Prior state (2026-07-17):
FLIPPED 2026-07-17 v0.6.355 after 8/7-day gate clear, 16 SHIP cells stable 7 consecutive days (07-11 → 07-17), LC_ENABLED READY on divergence report, and no cc/cl/cm/ch ANOMALY in the pair-log anomaly detector. First live tick (15:27 EDT): 113 cells fired — cc 46/48, ch 39/48, cm 19/48, cl 9/48; mean |Δ| in the expected 28-42 pp range per field. Predicted biggest Prod MAE lifts: cl 80-95 −55%, cl 95-100 −47%, ch 50-80 −37%. Two-gates-per-layer cross-check via [[feedback_two_gates_per_layer]]. 07-31 clean close superseded by the 07-30 intervention above.
Full Lc shipped-history log (v0.6.298 → v0.6.355)

Developer notes — fit table, refit cadence

Fit table location. weather_collector/data/lc_correction_table.json. Structure: {field: {value_bin: {shift, verdict, n_samples, mae_pre, mae_post, delta_pct}}}. Loaded at import time in cloud_saturation_correction.py via _load_correction_table().
Refit cadence. analysis/lc_fit.py runs in the daily digest and rewrites the JSON in-place. SHIP-set stability (7 consecutive days with the same SHIP cells + no HOLD days) is the fitter's own gate; it lives in .cache_lc_gate_history.json.
Value-bin edges. [0, 5, 20, 50, 80, 95, 100] percentage-points — six bins per field × four fields = 24 cells. Bin 5-20 is currently SKIP for all four fields (mid-low cloud has near-zero systematic bias).

Research & Diagnostics — experimental signals + audit views (not applied to live forecast)

Diagnostics — audit live behavior

R0. Per-layer audit — is each layer + specialist earning its keep?
Held-out MAE per field per layer (leads 1–47), refit every Fitter cycle. Dim subtext = signed bias. Δ columns compare to layer below: green = beats + applied; amber = beats but NOT applied; red = loses; gray = tied. Banners fire when an enabled layer or specialist loses >3% (hidden regression) or a disabled specialist / layer wins >3% (missed opportunity / post-ship watch signal). Specialists column (v0.6.382n) chains Lsr → Lc → chp / clp / wdp for the fields each owns; each rung compares to the previous rung (or L4 if first). Production column shows the real per-row Production MAE (aggregated from applied_layer stamps). v0.6.456 (F3-A): applied_layer now stamps NBM layers (l3_nbm/l4_nbm/l5_nbm/l6_nbm/chp_nbm/wdp_nbm) on rows where the selector picked NBM — before v0.6.456 those rows carried the HRRR walker's stamp and Prod attribution silently misclassified. Historical rows carry the old stamp; only post-08-21 rows carry NBM-aware attribution. MAE-only caveat: RMSE tells a different story on L2-additive fields (wg Production improvement drops −33%→−26%, dp −17%→−13%, h −7%→−3%). Per-layer RMSE+bias column rewrite queued.
Forecast accuracy — Raw vs Production + per-field lead-band breakdown

Two lenses on the same question. First (Accuracy over time, below): the per-obs-day trajectory of Raw vs Production — is the stack drifting? Did a recent ship move the needle? Second (per-field tables, further down): the current-window breakdown by lead band per field — where in the 0-47h horizon is each correction layer helping or hurting? The over-time view is the drift detector; the per-band tables are the shipping-decision granularity. Production in both views IS a real per-row aggregate. Per-band tables: per_layer_mae_by_lead[field].production where the per-lead sample count clears n≥30, keyed on applied_layer stamps. Over-time chart: mae_over_time.json's "prod_real" series, bucketed per obs-day from the same applied_layer stamps (v0.6.371). Both surface the same "what users actually saw" quantity, just at different aggregation grains. Specialist lines (Lsr, Lc, chp, clp) remain visible on the over-time chart as intermediate series showing gate-fired-only MAE — smaller sample than Raw/L2/L3/Prod, useful for isolating each specialist's per-lead lift during 14-day watches.

How we measure whether the forecast is good — the metric framework

Core comparison: pipeline error vs. raw L1 error on the same forecast-observation pairs. L1 = HRRR-side (Open-Meteo HRRR/GFS) baseline. NBM cascade runs in parallel; the 🎯 L1 selector (Prod-vs-Prod as of v0.6.440, armed via backstamp v0.6.445) picks per (field, lead-band) which cascade's output the user sees. Every metric on this page averages those errors differently.

"Observed" per field (best local measurement, not absolute truth):

  • t, dp, h, ws, wg, pr: mesonet Kalman blend across 60 WU + 26 Tempest stations within ~2.5 mi, station-bias corrected.
  • cc, cl, cm, ch: mean of KBOS + KBVY METAR sky reports (octa → percent). Coarser than mesonet.
  • sr: median across valid Tempest solar_radiation_wm2.
  • pa: max WU precip_rate_in across stations (patchy-rain-aware).
  • pp: binary — any measurable rain this hour (0 or 100).
  • wd: station wind_direction, circular error.

Three metrics per NWS/ECMWF conventions:

  • MAE — mean |error|. "Typical-day miss." Historical page text usually means this.
  • RMSE — root mean squared error. Weights big misses harder; if RMSE improvement lags MAE, pipeline has hidden blow-ups.
  • Bias — signed mean error. Positive = over-forecast; negative = under. Catches systematic drift MAE masks.

pp uses Brier score, not MAE (binary obs vs. probability forecast). Brier = mean of (fc/100 − obs/100)². Same "% erased" framing.

Measurement gaps:

  • Persistence skill — shipped 2026-07-11 (h_persistence_skill.py). Baseline: 5 ADD VALUE (t/h/pr/ws/sr), 4 MIXED (dp/wg/cc/pp), 3 NO SKILL — cl/cm/ch lose to persistence at every band despite L3+L4. Scorecard integration shipped 07-12 v0.6.328 — the "vs Persistence" line in the scorecard banner (above) is live-updated from persistence_skill.json. Response for ch: ch persistence gate FLIPPED LIVE 07-19 v0.6.358 (14 SHIP cells post-emergency-demote 07-27 v0.6.382t (was 27) on refreshed windows; 14-day watch CLOSED CLEAN 08-02). Landmark answered 07-14 v0.6.351a: persist_only ties gate pooled but gate wins halves-hedge; keep the gate. cl persistence gate wired Stage 3 07-24 v0.6.379 (replaces retired cl_persistence_short_lead — narrow 0-5h hypothesis disproven by halves-verified Stage 2). 08-09 Stage 2 rerun — 6 SHIP / 22 SKIP whitelist shape ready for flip decision. SHIP cells: calm/0-5 (−35.7%), ne_flow/0-5 (−54.4%), ne_flow/12-23 (−16.6%), nw_flow/0-5 (−18.7%), pre_frontal/0-5 (−17.8%), sea_breeze/0-5 (−32.9%). Flip = cl_persistence_gate.ENABLED=True. 08-09 c1_stage4_mixture_check: cl/6-11h [transition] b3 no longer in DEGRADED list (07-29 escalation window CLOSED CLEAN, never reached 3× ≥+200%); cm/12-23h [transition] IMPROVED; t/24-47h [transition] newly DEGRADED (unrelated). Joint cl/cm correction hypothesis rejected 07-29 (cl has no correction stack; cm corrections are stable improvers; drift is raw-HRRR-side on top-fc transition rows and fitter absorbs via c1 confidence bands).
  • Climatology skill — long-lead reference. Needs a climatology dataset.
  • pp Brier reliability decomposition — does "30% chance" actually happen 30% of the time? Aggregate Brier only, no decomposition.

Historical wording caveat: section text written pre-v0.6.325 (2026-07-10) treats MAE as the only measure. Read alongside RMSE + bias — disagreements usually mean occasional big misses or systematic drift.

Accuracy over time — Raw vs Production, per-day trajectory
Per-obs-day rollup. Chart shows the trajectory of Raw (Open-Meteo L1 seed — HRRR/GFS) and every correction layer applicable to the selected field (for t/ws/wd, "L1r (router)" line shows the v0.6.432 NBM router output at leads ≥6h) — legend adapts per field so noise-lines that don't do work for that field are hidden (e.g. sr shows Raw / Lsr; pa shows Raw / Prod only). Rolling 7-day mean overlays on Raw and the final applied layer smooth out day-to-day noise. Ship-date vertical annotations mark when live-layer changes landed. Source: mae_over_time.json — accumulating history, x-axis grows one day at a time (retention-independent; the pair log is capped at 30 days but this rollup persists prior days). Refreshed hourly by the myweather-publisher Cloud Function (v0.6.395f, hourly cron); prior to that it was tied to the daily digest. (v0.6.361: per-specialist attribution — Lsr, Lc, and post-Lc specialists (ch-persist, cl-persist) each plotted separately once ≥3 days of data have accumulated. New collector-side pair-log columns start accruing today; specialist lines will appear on the chart around 2026-07-22. v0.6.371: the 'Prod' line is now the real per-row aggregate keyed on applied_layer stamps — sample-comparable to Raw/L2/L3 on the same chart. Pre-v0.6.371 it was L4's daily MAE for non-specialist fields and a specialist line for cc/cl/cm/ch/sr (gate-fired-only, sample-mismatched).)
Detail view for selected field (Raw / L2 / L3 / Prod + rolling means + ship annotations):
Scan all fields (click any panel to focus the detail chart above):
📊 Per-field breakdown by lead band — MAE / RMSE / bias per correction layer

The bright white "Production" column is the forecast users actually see. Lower is better. For each field, one combined table shows how far off the forecast tends to be at each lead band — MAE (typical error), RMSE (occasional big misses that MAE averages away), bias (systematic drift, signed) — per correction layer. Compare a MAE row across columns to see where each layer helps or hurts.

How to read these tables
What you're looking at. One card per field. Each card shows a table with rows grouped by lead band (0-5h, 6-11h, 12-23h, 24-47h, ALL). Under each band, three metric rows: MAE (primary), RMSE (secondary), bias (secondary, signed). Columns are the correction layers actually applied to that field, plus the Production output on the right.
Only-applied layers. Layers that aren't applied to a field don't render as columns because their per-lead MAE array equals the previous applied layer's — noise. The badges above the table (L2 ✓, L3 ✓, L4 off, etc.) confirm the applied set. If Production isn't the lowest MAE, check the Applicability map for a gate that's wrong.
Per-metric usage:
  • MAE — average absolute error. The primary metric; drives L3/L4 whitelist decisions. Best cell in row highlighted green, worst red.
  • RMSE — squared-error root. Higher than MAE by a factor set by how tail-heavy the errors are. Watch for a band where RMSE jumps proportionally more than MAE — that's occasional big misses hiding behind an OK average.
  • bias — mean signed error (forecast − obs). Near-zero = calibrated on average. Persistent + or − by band signals systematic drift.
Layer labels:
  • Raw: bare Open-Meteo HRRR/GFS forecast for this coordinate. Knows nothing about Wyman Cove. (For t/ws/wd at leads ≥6h, the user-facing L1 is the NBM router output — see the L1r line, added v0.6.432.)
  • Aggregate bias (L2): what 40+ nearby weather stations say the model is currently getting wrong, blended in by distance.
  • Lead decay (L3): historical per-lead bias correction from the pair log.
  • Diurnal (L4): hour-of-day bias correction. Final line for t / dp / h / ws / wg / pp.
  • Synoptic-regime (Lsr): per-regime W/m² delta on direct solar radiation. Only present on sr.
  • Cloud saturation (Lc): per-(field, value_bin) percentage-point shift, clamped to [0, 100]. Final line for cm / ch (unchanged); cl and cc both fully off as of 07-30 (v0.6.389f cl · v0.6.390 cc). cc now derived downstream via Ccd = max(cl_l6, cm_l6, ch_l6). FLIPPED 2026-07-17 v0.6.355; emergency intervention 07-30 — see Lc layer section for details.
L2 badge variants: L2 ✓ additive = bias added (t, dp, h, pr). L2 ✓ direct = station median replaces model wind (ws, wg). L2 n/a = no station network reports this field (clouds, solar, precip). Note for clouds: L2 is n/a as a forecast correction, but obs truth for cc/cl/cm/ch comes from a KBOS + KBVY METAR blend (v0.6.134) — feeds the joiner so L3/L4 can be evaluated, not applied at L2.
Bands match the walk-forward validator buckets — this is exactly the granularity that drives shipping decisions. Source: time_series_diagnostic.json::per_layer_{mae,rmse,bias}_by_lead, 7-day window. Design note (v0.6.350): previously carried a per-card chart + 3 separate tables. The chart mostly restated what the MAE band table already showed; killed for signal density.
F1. Frontal passage log — detector instrumentation (last 14 days)
Live readout of detected frontal passages from frontal_events_log.json. Detector runs every tick in frontal_detection.py: passage fires when ≥2 of 3 signals hit (dp drop ≥4°F over 60min — lowered from 8°F on 2026-09-14 v0.6.620 after frontal_detector_health.py found observed 60-min dp drop max=6.8°F, p99.9=6.3°F; original 8°F threshold was unreachable, wd shift >60°, pressure inflection with ≥0.02″ rise). Confidence 67% with 2 signals, 100% with 3. Diagnostic print(" frontal: score=…") log line fires whenever score≥1 (2026-09-14 v0.6.620/v0.6.621) — traceable via gcloud functions logs read for miss-rate diagnosis. Health-check verdict runs in daily digest as analysis/frontal_detector_health.py. Live consumers: PWA front-passage card + C1e confidence axis + 10 analysis scripts joining on passage timestamps.
Loading detected passages...
R2. State-stratified accuracy — which regimes does the model fail in?
Per-field MAE + bias by regime. Big MAE spread across bins = regime-aware correction candidate. Addressed rows excluded from top-10 (shipped corrections absorb raw signal downstream); excluded rows in collapsible below. Temperature back in the top-10 since Lt retired 07-13. Refit twice daily; published to state_stratified_accuracy.json. MAE-only; RMSE would rerank L2-additive fields (dp/h/ws/wg).

Tools — evaluate candidates before promotion

S1. Shadow applicability tuner — what would auto-tuner have chosen?
Per Fitter cycle, log which fields a naive MAE auto-tuner would include in L3_FIELDS/L4_FIELDS, alongside production. Recommend ON if layer beats layer-below by ≥3% in any band AND bias no worse. Field-membership only — doesn't reason about per-field gates or skip tables (those in Applicability map). Precondition for automation is agreement after 90+ days; mismatches informative, not actionable. Also surfaces the live conditional_audits.r6 C1a-transition verdict per Fitter cycle.
Loading…
B1. Backtest sweep — alternative L3/L4 configs vs production
A/B comparison of candidate L3/L4 applicability configs vs production, computed by replaying the held-out pair log. Current production: L3 = {ws, wg, ch, cm, pp}, L4 = {ch, cc}. Run via python3 -m backtest.sweep --write-gcs (add --local-file ~/.cache/myweather/forecast_error_log.jsonl for fast iteration).
Loading sweep results...

Candidates — in-flight hypotheses (Stage 3 = wired ENABLED=False by definition)

C1 confidence stack — status
7 axes stamping every tick, applied=False until Stage 4 clears. Curated v3: 296 SHIP / 42 MARGINAL / 1048 SKIP across 39 axis-keys. C1a shipped as regime-transition axis (was R6; verdict logged under conditional_audits.r6 per Fitter cycle, surfaced via S1). Latest audit: HOLD 65.00% (13 CALIBRATED / 7 DRIFTED, 2 BRIER_EXEMPT — re-cured 08-15 from 57.14%). Prior cl "DEGRADED" reading was driven by the applied_layer poison fixed 08-02; next re-audit gated on window roll past cl-poisoned days (08-04+). Mixture-normalized view: only 2/112 cells are REAL DRIFT (per c1_stage4_difficulty_lens); legacy FAIL count is inflated by weather-mixture shift. Narrow-promote counters (today): pre-frontal 6/7 (2 SHIP, 1 to go). C1h + C1d auto-suppressed as KNOWN_LIVE_PIPELINES (already live-stamping). Ship gated on Stage 4. Per-cell co-axis ortho gate v0.6.321 in confidence_layer.py: cl fires freely, cc/cm/ch conditionally suppressed by the non-ortho co-axis, ch 24-47h + t × 3 never fire (REDUND both). Individual axes tracked in Backlog Group A below.
Backlog — Stage 0-2 hypotheses (pre-wire pipeline)
Stage 0-2 hypotheses (pre-wire). Promotion path: 0 exploration → 1 curated finding → 2 per-cell verification → 3 wired ENABLED=False → 4 shipped. Stage 3+ detail lives in Post-ship watches / layer sections, not here. Most surviving hypotheses measure forecast uncertainty (C1 axes), not forecast bias. Real bias candidates (marine layer, radiational cooling) overlap heavily with L2 and L4.

Group A — C1 multi-axis confidence extension

Individual axes. Five join the multi-axis Stage 4 calibration (C1a/b/c/f/e); C1h + C1d compose as marginal-premium tables (kept off the join to avoid cell-dilution). Stack-level status in C1 confidence stack card above.

Group B — Bias candidates (paced, individually)

Group C — Lower priority (dominated by existing layers)

In-flight candidates + current stage counters live in Current state → What's improving at the top of the page (single source of truth).

Group D — Methodological refinements (modify existing layers, not new ones)

Promotion rule: single-shot script in analysis/. Stage 2 verdict must hold across ≥2 reads spaced 3+ days apart. Group A → C1 axes (C1a/b/c/...); Group B → R-numbers if they clear the 7-window walk-forward gate.

Experiments — open design seeds + live telemetry probes

E1. Stage 0 explorations — open design seeds + data-limitation flags
Only genuinely-open items live here. Promoted (→ Stage 1+), killed by orthogonality, and settled-null items live in Archive — single source of truth.
E2. Lt — Cove microclimate telemetry probe (active experiment — correction returns 0.0; gradient log still appending)
Loading…
Status: active telemetry, correction dormant. The processor code runs each tick (compute_cove_correction() returns 0.0 in both branches; telemetry stamps for retro-check). cove_gradient_log.json continues to append per-tick waterfront-vs-inland gradient data on GCS (14-day rolling window) — cheap, kept alive in case seasonal shift ever changes the microclimate signal and a new refit becomes worth trying. The Lt row in the divergence report now reads from l6_fix_b_refit (07-14 v0.6.351b), which reports HOLD and matches production, so the row stays AGREE forever unless a future refit crosses the +1% gate.
Original design. A small Δ°F added to the temperature forecast based on how much the waterfront stations (Willow Rd, Neptune Rd) typically diverge from the inland-network median under different wind / sea-breeze / hour combinations. Lt was the first correction trained on a spatial differential between station subgroups (L1–Lsr are trained on forecast-vs-aggregated-obs errors).
Shipped 2026-06-26 v0.6.238. Cleared a 2-read confirmation gate on r5_cove_analysis.py (06-25 SHIP + 06-26 SHIP). Per-lead projection fix 2026-06-26 v0.6.237 (initial ship applied current-tick Δ to all 48 leads — wrong by 3–5°F at distant leads when the table crossed zero). Cooling branch disabled 2026-06-30 v0.6.259 (paired-MAE on 19,975 t rows: cooling branch made cooling rows −74.9% worse). Warming branch disabled 2026-07-01 v0.6.276 after per-row Production data (v0.6.269 applied-layer stamping + v0.6.275 backfill on 47,301 T pair rows) showed warming-branch rows ran ~40% worse than L2 on the same rows. Hypothesis at that time: double-counting — L2's Kalman blend for T is dominated by waterfront Tempests (Willow Rd, Neptune Rd at ~0.1–0.2 mi from the cove) — L2 already carries "waterfront bias" by station weighting; adding a (waterfront − inland) Δ on top re-adds the same signal.
Fix B — tried 2026-07-13 v0.6.329, failed the +1% gate. analysis/l6_fix_b_refit.py refit both lookup tables against (cove obs − L2 forecast at cove) instead of the raw (waterfront − inland) gradient — the design the section header used to promise. Ran on 202,321 pair rows (154k train / 47,757 held-out). Panel B (sb_off × hour) looked like a real overnight cove-cooling signal on training: 00-06h +0.9 to +2.1°F, 7 SHIP bins. But held-out MAE improvement was +0.29% vs the +1.0% ship gate. Panel A (sb_on × octant) came back with 0 SHIP bins after refit — the warming-branch signal was largely a fitting-against-raw-L1 artifact. Mechanism confirmed: L2's Kalman blend re-fits per-tick based on obs-vs-model bias and absorbs the same microclimate signal dynamically. A static hourly cove table adds a delta L2 already added → wash on held-out.
Reactivation criterion. Not queued for any specific date. Would require the held-out MAE on a re-run of l6_fix_b_refit.py against fresh pair-log data to clear +1.0% for a sustained window (3+ reads). Possible triggers: seasonal shift (fall/winter offshore vs summer sea-breeze patterns), sensor swap changing L2 station weighting, or a materially different mechanism than the one that failed. Until then, Lt telemetry runs dry.
Related files: weather_collector/processors/cove_correction.py, weather_collector/processors/cove_gradient_log.py, analysis/l6_fix_b_refit.py, analysis/l6_l2_double_counting.py, analysis/r5_cove_analysis.py, GCS cove_gradient_log.json (still updating, 14d rolling).

Archive — Retired ideas & historical investigations

Things we built, ran, and stopped running. Two kinds live here: hypotheses the data answered no, and settled tunings where the sweep concluded "current value is fine." Kept as institutional memory. Scripts in analysis/ can be re-run if conditions shift. Retired layers (code intact but permanently no-op) stay in R&D — see Lt above.

Post-ship watches — closed / superseded / inert
Watches that ran their 14-day window and closed clean, or were superseded by a follow-on ship, or went inert because the target was dropped. Active watches live under Current state → What's improving → Post-ship watches — active.
Recently ruled out — 2026-06-22 to 06-29 Stage 0/1 kills
One-shot smoke tests that landed at "no signal" or "captured by an existing axis." Re-run in 2-3 months if the seasonal regime shifts.
[HYPOTHESIS] Tide-phase corrections — does forecast error track the tide cycle? (RETIRED 2026-06-08)
Verdict: weak signal, mostly entangled with diurnal cycle. Per-field tide-phase curves were tracked across weeks; the signal that survived stratification was hard to distinguish from hour-of-day patterns we're already correcting in L4. Cost of keeping it running (NOAA fetch + 12-bin accumulator + GCS history per Fitter pass) wasn't justified. Analysis: analysis/tide_hypothesis.py.skip.py (renamed 07-18 — removed from digest run so it stops spurious-FAILing on NOAA fetch). Fitter module flag: RUN_TIDE_TRACKING. Frozen charts below show the final state at retirement; they will not update.
Companion view — error vs tide elevation over time (frozen)
Time-domain rendering of the same data as the bucketed chart above. Different angle on the same retired hypothesis. The frozen state below is the last fit before tide tracking was disabled.
Higher leads = forecast made further ahead. Switch to see if the tide pattern is lead-specific.
[HYPOTHESIS] Derived humidity — Magnus(T_corrected, T_d_corrected) vs network-blended humidity (RETIRED 2026-06-08)
Verdict: equivalent. Tested whether deriving humidity from corrected temperature + corrected dew point via Magnus outperforms the L2 network-blended humidity. 27k triples, identical MAE within noise. We kept the derived path anyway because it keeps the (T, T_d, RH, AH) quadruple internally consistent — but the hypothesis "derivation is more accurate" is closed. Analysis: analysis/derived_humidity.py.
[SETTLED TUNING] L3/L4 recency window τ — sweep over fit-window decay constant (RETIRED 2026-06-08)
Verdict: τ=14 days is fine within noise. Not a hypothesis — a parameter sweep over the Fitter's recency-weighting τ (how much old pairs count when fitting decay curves). Not the L2 lead-decay τ added in v0.6.44, which controls how a current bias is spread across forecast leads (see sec 2d). Tested τ ∈ {7, 14, 21} days across four reports. Held-out MAE differences under 2%, well below run-to-run variance. τ=14d stays. With L3/L4 mostly disabled in v0.6.45, this knob barely matters anymore. Analysis: analysis/decay_tau_tuning.py.
[HYPOTHESIS] R4 — HRRR vs GFS spread as confidence signal (RETIRED 2026-06-17, verdict: CLOSE)
Hypothesis: when HRRR and GFS disagree at a given forecast hour, actual error magnitude tends to be higher — i.e. |HRRR − GFS| per (field, lead) predicts |forecast − obs|. If true, the spread becomes a free uncertainty number that can widen displayed intervals and feed Gemini hedge language ("models disagree on tomorrow's high"). Data collection: HRRR L1 already in forecast_log.json; gfs_l1_log.json captures GFS L1 per tick for the same 0-48h window. Decision rule was: ship if median Spearman ρ > 0.25 for ≥3 fields, consistent across lead bands.
CLOSE verdict (2026-06-17, 112,877 joined pairs):
0 of 6 fields above the 0.25 ρ threshold. Maximum observed |ρ| = 0.012 (wind speed at 1-6h) — essentially zero correlation. HRRR vs GFS spread does NOT predict forecast error magnitude. Retired without auto-wiring. Manual script: analysis/r4_spread_analysis.py — re-run quarterly or after a model release.
[HYPOTHESIS] R5 — Cove warming — sea breeze across the peninsula heats Wyman Cove vs inland (RETIRED 2026-06-17, verdict: HOLD — L2 already captures it)
Reframed hypothesis (2026-06-13): Wyman Cove sits in the lee of the Marblehead peninsula on a S/SE/SW sea breeze. Marine air crosses ~2 miles of sun-heated land before reaching the cove, picking up surface heat in transit. Expected pattern: delta_wf_inland = waterfront_median − inland_median goes positive (cove warmer) when wind is from the S half AND sea breeze is active, with magnitude scaling to solar input (peaking ~12-14 EDT). Should flatten to zero when wind is from N/NE (cove is windward of peninsula) or after sunset (no surface heating). Original hypothesis ("waterfront cools during sea breeze") was geographically backwards and is closed.
Day-12 refit (1,732 ticks through 2026-06-24): matches the reframed model; magnitudes tightened further as the sample grew. NW flipped from neutral to weakly negative; E cooled further.
WindSea breezenmean Δ°F
Sactive186+1.5
SEactive88+2.0
SWactive79+1.1
Ninactive378−1.0
NEinactive103−1.0
Einactive86−1.3
NWinactive459−0.9
Diurnal curve under offshore/calm conditions shows clean morning-marine-cooling: trough around −3.7°F at 12:00 EDT (refit 06-24, n=1,732 entries over 12 days; cool air pool over Salem Sound persists; inland warms fast with sun while cove stays anchored to marine boundary). Both signals are physically coherent with the lee-warming model.
Data collection: cove_gradient_log.json captures waterfront-tagged Tempest median (Willow Rd, Neptune Rd — both at cliff-edge elevations on the harbor, confirmed by Joe), inland Tempest median (~18 stations), ambient T, wind dir/speed, salem_water_temp_f, buoy_water_temp_f, sb_active, sb_likelihood per tick (14-day retention).
Two-step plan:
Step 2 verdict: HOLD (run 2026-06-16, n=29,444 matched pairs)
The L2-overlap hypothesis was empirically confirmed. L2's 1/distance² × elevation station weighting for the cove is dominated by the two waterfront Tempests (Willow Rd, Neptune Rd at ~0.1–0.2 mi). L2's "cove bias" is effectively "waterfront bias" by construction. Layering R5's (waterfront − inland) delta on top double-counts the same signal — the cove obs is already waterfront-influenced via L2, so adding more waterfront delta pushes the forecast AWAY from the obs.
Decision (global R5, retired 2026-06-17): r5_audit.py's held-out test of R5 applied across the full pipeline showed it makes cove temp 20–22% worse — L2's waterfront-weighted station blend already captures the signal, so layering R5 on top double-counts. That global formulation stays retired.
Current status (Lt — microclimate correction, RETIRED 2026-07-13 v0.6.329): Both branches return 0.0. Cooling branch killed 2026-06-30 v0.6.259 (paired-MAE showed it materially increased cove temp MAE); warming branch killed 2026-07-01 v0.6.276 after real per-row Production exposed the same double-counting for the sea-breeze direction; Fix B refit tried 2026-07-13 (analysis/l6_fix_b_refit.py) and failed at held-out +0.29% vs +1.0% ship gate. Lookup tables + cove_gradient_log.json continue to log per-tick against a possible future reactivation (seasonal shift, sensor swap, or materially different mechanism). Full rationale in R&D → Lt.
One niche subtlety in the breakdown: long-lead (24-47h) sea-breeze forecasts get +7.85% MAE improvement with R5. At long leads, L2's τ=4h decay has long since faded, so R5 has something L2 doesn't. Not worth shipping a conditional correction for, but documented.