Black Hat Hydrology Lab
1.21.0Real river data. A few educated guesses. Mother Nature keeps score. Follow our predictions for the San Juan River and its mountain snowpack. Right Now updates the active season-game question. The scoreboard keeps the original prediction, and Ask a Date explores expected river flow over the next 90 days.
Experimental forecasts for entertainment. For current conditions, visit the River Desk ↗.
Today’s experiment
The season-game prediction, updated with the latest measurements. Compare today’s estimate with the original call and see what could change the answer.
Checking the river’s next move…
Loading measurements for the active game question.
Calculating from the available measurements.
Building the active experiment…
Waiting for the newest observations.
The forecast feed is still loading.
The Lab is finding comparable seasons.
New observations, forecast weather or a meaningful change in the river’s direction can move the live estimate.
Run the Lab again after the next source update to see whether the evidence moved.
MEASUREMENTS AND METHODEvidence, model, unfinished predictions and what could change the answer
Waiting for enough evidence.
Waiting for enough evidence.
Right Now updates the active game question using newer measurements.
CFS measures water passing the gauge each second. The shaded band shows the model’s range; it is not a promise of conditions.
What will the river be doing on your date?
Pick a date or touch the curve. The answer and activity outlook follow your selection.
Choose a date to see the estimate.
Once the Lab finishes loading, your answer will use the newest river and snowpack inputs.
Possible flow over the next 90 days
TUBING-RANGE ODDS
RAFTING-FLOW ODDS
These are experimental flow-threshold odds. Trip and product guidance is shown only for dates inside Pagosa Outside’s May 1–Labor Day operating season.
Calculating the short-range experiment.
Calculating the 30-day experiment.
Scanning the 90-day river map for the next meaningful move.
Eight predictions. One river.
Each card records the original prediction. The Black Hat earns a point when the observed result falls within the scoring window; Mother Nature earns a point when it does not. Right Now updates the same question while we wait.
PO BLACK HAT LAB vs MOTHER NATURE
The 2027 game is on deck
- 1First snowUPCOMING
- 2Snowpack peakUPCOMING
- 3Rafting startsUPCOMING
- 4Season qualityUPCOMING
- 5Runoff peakUPCOMING
- 6Rafting lengthUPCOMING
- 7Tubing startsUPCOMING
- 8Tubing lengthUPCOMING
How the season game works
Loading issued predictions…
Each card has a defined measurement, issue date or river trigger, and scoring window.
The issued prediction remains fixed for scoring.
New measurements update the answer to the same game question.
A new water-year game begins each October 1. Open a card for what counts and how the result is scored.
Gold marks the original prediction, blue the historical reference, and green the observed result.
Game predictions and results
Each answer locks on its issue date. Open a card for its evidence, exact definition and scoring rule.
When will four inches of mountain snow stick?
- Original prediction
- Locks October 1, 2026
- Measurement
- Four inches of snow depth
Evidence & scoring +
Comparable four-inch snow seasons are weighted by the measured depth and its recent three-day change. The estimate updates as the season progresses. A seven-day observed event confirms the start date. Weather forecasts are supporting context.
Scoring rule: Average measured snow depth at Upper San Juan and Wolf Creek Summit reaches 4 inches, then stays above 0.5 inch for seven consecutive days including the first day. Both stations must report. The predicted start date must be within 14 days.
Loading issued prediction and result…
- Actual result
- Waiting for Mother Nature.
How much water will the snowpack hold at its peak?
- Original prediction
- Locks February 1, 2027
- Historical reference
- About 27.3 in · around April 6
Evidence & scoring +
The estimate uses the measured snow-water reserve, accumulation, precipitation and seasonal conditions to predict the largest remaining reserve. Snow already measured sets a lower bound.
Scoring rule: The highest two-station average snow-water equivalent. The amount must be within 20% or 2 inches of water, whichever tolerance is larger. The peak date is additional context.
Loading issued prediction and result…
- Actual result
- Waiting for Mother Nature.
When will the river average 400 CFS for a day?
- Original prediction
- Locks March 1, 2027
- Historical reference
- Around April 2
Evidence & scoring +
The estimate uses river flow, its recent direction, headwater snow and seasonal conditions to predict the first qualifying daily average.
Scoring rule: The first calendar day whose USGS average reaches at least 400 CFS. The predicted date must be within 10 days. A brief crossing does not count.
Loading issued prediction and result…
- Actual result
- Waiting for Mother Nature.
How strong will May–June river flow be?
- Original prediction
- Locks April 1, 2027
- Historical reference
- About 1,100 CFS average
Evidence & scoring +
The estimate uses peak snow-water reserve measured so far, the remaining reserve and the previous 30-day river average. It predicts May–June average flow; a short high-flow period can still occur in a mild season.
Scoring rule: Average river flow from May 1 through June 30. The prediction must be within 30% or 100 CFS, whichever tolerance is larger.
Loading issued prediction and result…
- Actual result
- Waiting for Mother Nature.
When and how high will spring flow peak?
- Original prediction
- Locks April 15, 2027
- Historical reference
- Historical daily means are context, not the scored instantaneous peak
Evidence & scoring +
The game’s timing and flow formulas estimate the crest date and magnitude. A historical time-of-day reference supplies the clock time. Right Now applies that same calculation to updated inputs; the daily-average chart is context.
Scoring rule: The highest individual USGS reading from April 1 through July 31, with the earliest timestamp used for a tie. Both the 168-hour time tolerance and the larger of 30% or 250 CFS flow tolerance must be met.
Loading issued prediction and result…
- Actual result
- Waiting for Mother Nature.
How many days will reach rafting-flow levels?
- Original prediction
- Locks May 1, 2027
- Historical reference
- About 63 at 300+ · 46 at 600+ · 39 at 800+
Evidence & scoring +
Separate estimates count days likely to reach each of the three flow levels using remaining snow, recent flow and the seasonal outlook.
Scoring rule: Count daily averages at or above 300, 600 and 800 CFS from May 1 through Labor Day. Two of the three predicted counts must finish within 10 days.
Loading issued prediction and result…
- Actual result
- Waiting for Mother Nature.
When will the river stay below 400 CFS for a full day?
- Original prediction
- Locks at the first post-peak daily average below 800 CFS with a falling three-day direction
- Automatic-rule replay
- 2012–2026 · typically mid-June
- Actual result
- Detected from a complete 24-hour run of instantaneous readings below 400 CFS.
Evidence & scoring +
The date estimate updates with river flow, its decline and remaining snow. Actual gauge readings confirm the 24-hour event. This flow benchmark does not announce trip availability.
Scoring rule: The completion of the first qualifying continuous 24 hours below 400 CFS after the spring crest. A reading at or above 400 or a gap longer than 30 minutes resets the clock. The predicted date must be within seven days.
Loading issued prediction and result…
How many tubing-range days will follow that opening?
- Original prediction
- Locks automatically at the first post-peak complete 24 hours below 400 CFS
- Daily-mean replay · separate test
- Not a validation of the continuous opening rule
- Actual result
- Count begins the next full calendar day.
Evidence & scoring +
The estimate starts with calendar days remaining and the season’s snow-water reserve. Right Now combines qualifying days already measured with the estimated remaining days to show the expected final total. Rain and short-term flow are supporting evidence.
Scoring rule: Count daily averages at least 30 and below 400 CFS, starting the next full calendar day after the confirmed opening and ending on Labor Day. The prediction must be within 10 days. These days need not be consecutive.
Loading issued prediction and result…
The Lab is checking the calendar.
2026 Black Hat Lab vs Mother Nature
All eight predictions are settled. The final 2026 score is Black Hat Lab 3, Mother Nature 5.
FINAL SCORE COMPLETE
All eight predictions are complete. Final score: Black Hat Lab 3, Mother Nature 5. Open the Film Room for every bet, the five misses and the model changes carried into 2027.
2026 BLACK HAT FILM ROOMFINAL SCORE · WHAT FOOLED US · WHAT CHANGED FOR 2027
An exceptionally dry, warm year made the first game a useful stress test
The water year opened with two rain-driven flood crests from October 10–14, 2025. The larger reached about 8,570 CFS and 12.82 feet at the Pagosa gauge—the third-largest flood in the local record. Winter then reversed the story: persistent snow arrived late, Wolf Creek Ski Area opened November 22, and the two-station headwater snowpack peaked at only 14.3 inches of stored water on March 6. Spring runoff was short and mild, and the San Juan reached its High-Water Mark at 903 CFS at 1:15 a.m. MDT on May 15 before dropping toward tubing flow unusually early. Low summer water removed many potential tubing days. The year’s flood-to-drought swing was exactly the kind of outlier that exposes formulas trained mostly on more ordinary seasons.
If today’s formulas had made the 2026 calls, the Lab would have gone 5–3
The official 2026 score remains Black Hat 3, Mother Nature 5. That record never moves. This separate retrospective runs the newest playbook from the information available at each issue date or event trigger. It is a learning exercise—not a replacement score and not a blind test, because the formulas were improved after the Lab had seen 2026.
Archived daily-mean replay (not the current 24-hour game rule): The newer threshold-day counter turns Prediction 6 into a hit, and the event-triggered tubing-opening model turns Prediction 7 into a hit. Prediction 7 now scores the first post-peak USGS daily mean below 400 CFS rather than demanding a perfect 24-hour gauge window or a manual operator date. For Prediction 8, the new native-flow baseline replays 2026 at 53 days versus 38 actual. That is still 15 days high and still a miss.
Every original guess stays in the record
The retrospective was reconstructed from information available on each posted issue date. Hits stay hits and misses stay misses. The two season-total cards reached their fixed Labor Day finish line and are now final.
When will the first snow stick on the mountains?
- Black Hat guess
- October 24, 2025
- Historical average
- October 24, 2025
- Actual result
- November 17, 2025
Result: Mother Nature waited 24 days longer than our guess.
Evidence & scoring +
What we’re predicting: When winter will stop teasing us and leave a lasting snow reserve in the San Juan headwaters.
Win condition: Land within 14 days of the first date the two-station average reaches one inch of stored water and remains above one-half inch for seven days.
How much water will the snowpack hold at peak?
- Black Hat guess
- 37.3 in · April 2, 2026
- Historical average
- 27.3 in · April 6, 2026
- Actual result
- 14.3 in · March 6, 2026
Result: We overshot the peak by 23 inches. That miss taught the Lab not to pull an exceptionally poor snow year back toward average; the current method includes a stricter warm-and-dry guardrail.
Evidence & scoring +
What we’re predicting: How much water winter will bank in mountain snow before melting begins—an early clue to runoff strength and season length.
Win condition: Predict the peak amount within 20% or two inches, whichever gives the wider scoring window.
When will the San Juan first reach rafting levels?
- Black Hat guess
- March 29, 2026
- Historical average
- April 2, 2026
- Actual result
- March 20, 2026 · 421 CFS
Result: The river arrived nine days earlier than our guess—close enough to put one on the Black Hat side.
Evidence & scoring +
What we’re predicting: When the river will first climb out of winter flow and reach Pagosa’s general rafting crossover.
Win condition: Land within 10 days of the first calendar day whose average flow reaches at least 400 CFS; a brief spike does not count.
Will it be a good rafting season?
- Black Hat guess
- 376 CFS average
MILD RAFTING YEAR - Historical average
- About 1,100 CFS
AVERAGE RAFTING YEAR - Actual result
- 341 CFS average
MILD RAFTING YEAR
Result: Only 35 CFS high. The Lab correctly warned that this would be a mild rafting year—one to catch early rather than postponing until June warmed up.
Evidence & scoring +
What we’re predicting: Whether Pagosa is headed for a high-water year locals should not miss, a dependable average rafting year or a mild season when guests should get on the river early instead of waiting for warm June weather.
Win condition: Predict the May 1–June 30 average within 30% or 100 CFS. Mild is below 75% of the historical average; high-water begins above 125%.
When and how high will spring runoff peak?
- Black Hat guess
- May 14, 2026 · 12:15 a.m.
984 CFS - Historical average
- May 22, 2026 · 12:15 a.m.
2,280 CFS - Actual result
- May 15, 2026 · 1:15 a.m. MDT
903 CFS
Result: The predicted crest arrived 25 hours early and the flow was 81 CFS high—a clean Black Hat win.
Evidence & scoring +
What we’re predicting: One complete timestamp and the flow of the spring crest—the classic runoff question and valuable high-water planning context.
Win condition: Get the flow within 30% or 250 CFS and place the complete predicted timestamp within seven days (168 hours).
How long will rafting season last?
- Black Hat guess
- 45 at 300+
31 at 600+ · 11 at 800+ - Historical average
- 63 at 300+
46 at 600+ · 39 at 800+ - Actual result
- 36 at 300+
8 at 600+ · 0 at 800+
Result: Only the 300-CFS total finished within 10 days of the locked guess. The card needed two of the three counts, so Mother Nature takes the point.
Evidence & scoring +
What we’re predicting: How many days will reach the flow benchmarks behind the season’s Splash & Dash, Mesa Canyon and West Fork planning.
Win condition: Finish within 10 days on at least two of the three counts after Labor Day.
When will tubing season start?
- Black Hat guess
- June 14, 2026
- Historical average
- June 24, 2026
- Actual result
- June 4, 2026
Result: Tubing conditions arrived 10 days earlier than our guess, putting this point on Mother Nature’s side.
Evidence & scoring +
What we’re predicting: When post-peak runoff will hand the river from spring rafting to warm-weather tubing.
Win condition: Land within seven days of the first complete post-peak 24-hour period with every available gauge reading from 30 through 399 CFS.
How long will tubing season last?
- Black Hat guess
- 77 days
- Historical average
- 66 days
- Actual result
- 45 days
25 too high · 60 too low
Result: The season finished with 45 tubing-range days, 32 fewer than the locked guess. Low water—not rafting water—removed most of the missing days, so Mother Nature takes the point.
Evidence & scoring +
What we’re predicting: How many recreation-season days will land in Pagosa’s broad tubing range—an estimate of the season’s usable tubing supply.
Win condition: Finish within 10 days of the total number of May 1–Labor Day daily averages from 30 through 399 CFS.
The misses were useful because they exposed different weaknesses
Early-season timing can be undone by a long warm, dry stretch. KCPW freeze persistence now runs beside the locked model in Right Now, but has not yet earned a fixed-card promotion.
The old tree model pulled an exceptionally poor winter back toward normal because it could not extrapolate beyond the low-snow years in its training history. This was the clearest model failure of the season.
The old May 10 issue date arrived too early to see how quickly the small runoff would collapse. The new event trigger waits for the post-peak handoff; it predicts June 3 against the automatic June 1 answer and flips this card to a hit.
Only one original threshold-day guess finished within 10 days. The new low-snow/CBRFC consistency guard replays the card at 35/17/4, close enough on all three counts.
Under the new prospective question, 98 calendar days remained after the automatic June 1 opening; 60 later averaged below 30 CFS, leaving 38 tubing-range days. The current snow-state playbook would have forecast 53. That is still 15 days high and still a miss. In the six-year CPC comparison, adding the broad precipitation tilt did not improve the hit count, so it stays an uncertainty signal rather than moving the center.
New ideas have to earn the hat
- Regularized low-snow peak-snow model that can extrapolate in extreme winters.
- Positive-log snow-state model for May–June river strength.
- CBRFC daily-mean formula for spring peak and its likely date window.
- Low-snow/CBRFC consistency guard for the three rafting-day counts.
- Event-triggered tubing-opening and remaining-days questions tied to the first post-peak complete 24 hours below 400 CFS.
- Provisional native-flow survival bands for remaining tubing days.
- Upper San Juan + Wolf Creek Summit as the primary snowpack pair.
- Specialized first-400 formula.
- The 2026 official score and original issue rules, which remain unchanged in the archive.
- CPC July–August precipitation outlook as an uncertainty signal for tubing length—not a center adjustment yet.
- Target-specific continuation experiments for the locked seasonal questions.
- KCPW refreeze and KPSO basin comparison.
- Freezing-level, inversion and melt-pulse routing experiments.
- Blanket CPC use across every formula.
- Outside-basin snow stations used everywhere.
- KCPW-adjusted annual peak time.
- Generic 90-day model replacing target-specific locked cards; it now belongs only to Ask + Explore.
The rule going forward: Right Now may learn continuously. A challenger only enters the starting lineup after it beats the incumbent in a no-peeking historical test.
WHAT IS THE BLACK HAT LAB?The original prediction, the latest estimate and the observed result
The Lab is Pagosa Outside’s experimental forecast page for the San Juan River and its mountain snowpack.
Right Now
The latest estimate for the active game question, using new measurements as they become available.
Ask a Date
A flow estimate and forecast range for a day within the next 90 days.
Season Game
Eight original predictions, their scoring rules and the observed results.
For current conditions and trip availability, visit the River Desk.
HOW THE LAB WORKS
The Lab combines river measurements, mountain snowpack, weather information and historical seasons. Each game question has a defined outcome: a date, peak amount, average flow or count of qualifying days.
The game saves the original prediction for scoring. Right Now updates the same question. Ask a Date provides a separate flow estimate for your chosen day.
DATA AND DEFINITIONS
Snow depth measures inches of snow on the ground. Snow-water equivalent (SWE) measures the water stored in that snow. First snow uses measured depth; peak snow-water reserve and runoff estimates use SWE.
CFS measures water passing the river gauge each second. An individual gauge reading and a full day’s average are different measurements. Each card states which one counts.
The water-year game runs October 1 through September 30. Displayed observation times use Mountain Time. A source’s measurement time can be earlier than the page’s refresh time.
HOW ESTIMATES ARE CALCULATED
Open a game card for its inputs, calculation and scoring rule. Right Now follows that question as measurements arrive. In summer, expected tubing-range days remain the main answer; short-term rain and flow are supporting evidence.
Ask a Date uses seasonal patterns and, when available, official short-term river forecasts. Their influence fades beyond the official forecast period. The shaded range shows uncertainty; dates adjusted with official forecasts do not carry a verified 80% coverage claim.
Six game predictions issue on scheduled dates. The tubing-start prediction issues at its post-peak below-800 CFS threshold with a falling three-day trend. The tubing-length prediction issues when the first qualifying continuous 24 hours below 400 CFS is confirmed. One falling reading does not establish the spring crest.
TESTING AND LIMITATIONS
Accuracy varies by question, season and forecast distance. Historical tests describe the target and method actually tested. A test of daily averages does not establish accuracy for an individual peak reading or a continuous 24-hour event.
The four-inch snow-depth method has a limited development comparison: about 11.1 days average date error versus 11.5 days for its baseline across nine seasons. The method was selected using those years; this is not independent validation. Its historical spread is not a calibrated confidence interval.
Average historical error is not a guaranteed forecast window. Exact peak clock time is experimental. See Past Results for completed games and clearly labeled historical model tests.
USGS Pagosa gauge: river readings and daily averages. The spring-peak game uses the highest individual reading; tubing opening uses a qualifying continuous 24-hour period below 400 CFS.
Upper San Juan and Wolf Creek Summit: measured snow depth and water stored in snow.
National Weather Service: weather observations and forecasts. Colorado Basin River Forecast Center: official river forecasts. NOAA Climate Prediction Center: broad climate outlooks, used only where the question’s method specifies them.
Each chart identifies its measurement and historical reference. For daily river comparisons, normal is the date-matched 1991–2020 median.
Enough Guessing?
Trade the sparks and speculation for actual San Juan River flow, water temperature, short-term forecasts and today’s river options.
FAQs
Stuff you want to know about Pagosa Outside's Black Hat Hydrology Lab Experiment
The Black Hat Hydrology Lab is Pagosa Outside’s public-facing river-forecasting experiment. It combines historical San Juan River flows, headwater snowpack, current conditions and weather information to make longer-range flow estimates and seasonal predictions. It is part river science, part local experience and part ongoing attempt to make better educated guesses than we used to make over coffee.
Read the whole backstory on our blog HERE.
The name is a wink, not a confession. In our version, “black hat” means taking public river, snowpack, weather and climate information traditionally scattered across separate sources, hacking it together in unconventional ways and testing whether it can produce a better-informed home-field advantage.
We are not breaking into government computers or hiding the methods: every source is public, every official guess is locked and every miss stays on the scoreboard. The only thing we are trying to outsmart is Mother Nature—and she still has the administrator password.
The River Desk shows current measurements, published government forecasts, snowpack, weather and today’s river options. The Black Hat Lab goes beyond those published forecasts and experiments with what might happen 30 to 90 days from now or later in the river season.
The Lab’s predictions are independent Pagosa Outside experiments—not official government forecasts. They should be treated as entertainment rather than instructions for river use, travel or business decisions. Use the River Desk for current conditions and published short-term forecasts.
The Lab estimates San Juan River flow 30, 60 and 90 days ahead and displays an experimental 90-day hydrograph. It calculates the probability of reaching Pagosa’s general tubing and rafting-flow ranges, operates a Live Seasonal Watch and publishes eight fixed predictions about important events during the water year.
Those seasonal predictions cover headwater snowpack, the beginning and strength of spring runoff, peak flow, rafting-flow days, Tubing Season Opening Day and tubing-range days.
The honest answer is: promising, but still experimental. In historical testing, the seasonal prediction methods scored 73 hits in 96 development tests. A later five-season test—using methods selected without seeing those results—scored 30 hits in 40 predictions. That works out to roughly three correct calls out of four.
The 30-day flow model reduced its typical error by approximately 28% compared with using historical averages alone, while the 90-day model improved by approximately 11%. The experimental ranges contained the actual result about four times out of five during historical testing.
Accuracy still changes considerably by season, forecast distance and type of prediction. One unusual storm, heat wave or stubborn snowpack can embarrass the entire contraption. That is why the scoreboard displays every public hit, miss and pending result instead of showing only the guesses that make us look clever.
The Lab uses river flow from the U.S. Geological Survey gauge at Pagosa Springs; snowpack from the Upper San Juan and Wolf Creek Summit Snow Telemetry stations; temperature and precipitation forecasts from the National Weather Service; selected climate outlooks from the National Oceanic and Atmospheric Administration; and historical guidance from the Colorado Basin River Forecast Center.
Middle Creek and Lily Pond may be used as regional storm-pattern checks on predictions where historical testing showed an improvement. Other stations—including Beartown and Grayback—were tested but did not improve the Pagosa predictions enough to earn a regular seat at the workbench.
The Lab does not ask a chatbot to stare into the river and invent a number. Its predictions come from repeatable formulas, historical comparisons and backtested models. The calculations run automatically, but the targets, inputs, scoring rules and operating thresholds were developed around the actual questions Pagosa river managers and guides try to answer.
The Lab compares current river flow, recent movement, seasonal timing, headwater snowpack, temperature and precipitation with similar conditions from previous water years. It produces a center estimate and an experimental range for each future date. The historical average is displayed as a baseline, not as a second prediction.
The center estimate is the Lab’s single best guess. The experimental range shows a broader group of reasonably plausible outcomes and is calibrated to contain the actual result about four times out of five during historical testing. It is not a best-case and worst-case limit, and the river is under no contractual obligation to remain inside it.
Each horizon is tested separately. A 60-day range can occasionally be wider than a 90-day range when the 60-day date lands during a volatile runoff or seasonal transition and the 90-day date lands during historically steadier flow.
The probability table estimates the chance that flow on each forecast date will fall within Pagosa’s general operating ranges:
30–399 cubic feet per second: tubing range
300 cubic feet per second or higher: possible Splash & Dash rafting flow
600 cubic feet per second or higher: possible Mesa Canyon rafting flow
800 cubic feet per second or higher: possible West Fork rafting flow
These are flow probabilities—not promises that a particular trip will operate. Water temperature, weather, debris, access and current river conditions still influence Pagosa Outside’s daily lineup.
Right Now in the Lab changes jobs as the water year progresses. During winter, Spring Runoff Watch follows headwater snowpack development and early runoff potential. As spring runoff begins rising, Predict the Peak Runoff Watch continuously updates the Lab's best guess for the peak date, time and flow. After a sustained decline indicates that the spring crest has likely passed, it becomes Tubing Opening Day Watch and looks for the first complete 24 hours with every available Pagosa gauge reading from 30 through 399 cubic feet per second. After the tubing transition, Monsoon Flow Watch follows possible rain-driven tubing boosts, low-water recoveries and returns to rafting flow. After Labor Day, it becomes Bonus Flow Watch.
Unlike the eight official predictions, Right Now in the Lab updates whenever the Lab runs and new information becomes available. Its changing estimates never rewrite the fixed guesses on the scoreboard.
The eight official predictions are:
First persistent headwater snowpack
Peak headwater snowpack
First 400-plus cubic-feet-per-second day
May–June river strength
Spring peak date, time and flow
Rafting-flow days from May 1 through Labor Day
Tubing Season Opening Day
Tubing-flow days from May 1 through Labor Day
Each question uses its own combination of snowpack, river flow, weather, climate outlooks, seasonal timing and historical comparisons. Predicting peak snowpack is not the same problem as predicting spring peak flow or tubing days, so one magic formula would not be very magical.
Each official call also has a fixed issue date selected to balance useful lead time with historical accuracy. Once published, the guess is locked and cannot be quietly changed when better information arrives.
Every prediction has a measurable finish line and scoring rule established before the outcome is known. A hit earns one point for the Black Hat Lab. A miss earns one point for Mother Nature. Unfinished predictions remain pending.
The scoreboard follows the October 1 through September 30 water year and divides the season into four quarters. A new game begins each October, while the previous season’s record remains visible. The Lab does not get to quietly drag its misses behind the shed.
A retrospective season recreates earlier predictions using only information that would have been available on each stated issue date. The model is trained on previous years and is not allowed to peek at what happened later. Retrospective results are clearly labeled because they were historical tests—not predictions publicly posted at the time.
Sometimes. Archived National Oceanic and Atmospheric Administration one-month and three-month outlooks improved selected predictions, including the first 400-plus cubic-feet-per-second day, May–June river strength and spring peak flow.
Those outlooks did not improve every part of the Lab. They are excluded from the 30-, 60- and 90-day flow estimates, hydrograph, operating-threshold probabilities and Monsoon Flow Watch because historical testing found that the other methods performed as well or better. The Live Seasonal Watch still uses the National Weather Service’s seven-day headwater forecast for short-range weather and storm information.
More data does not automatically produce a better prediction, so every new ingredient has to earn its way into the formula.
Absolutely. This is a continuing public experiment, and the build number shows which version is running. New water years create new tests, misses expose weak formulas and better data sources may improve individual predictions.
A method changes only when backtesting shows a real improvement without peeking at the answer. Past guesses and results remain visible so a new formula cannot rewrite history or polish an old miss into a hit.
Because once we put the whole river story onto one desktop, somebody asked the dangerous question: What happens after the official forecast ends—and can we guess it better than usual?
The Lab combines 25 years of river and snowpack history with current conditions, weather forecasts, tested formulas and the PO River Crew’s accumulated hunches to make longer-range and seasonal predictions. It is where actual science, river-guide instinct and old-timer theories are forced to share one questionable workbench while Mother Nature controls the scoreboard.
We built it to find out whether our educated guesses can beat ordinary historical averages—and because arguing about peak runoff at the coffee shop was not producing enough charts.
Read the whole backstory on our blog HERE.