Methodology
How the Forecast Works
One model produces every number on the forecast page. It is a probability machine, not a crystal ball. This page explains exactly what goes into it, in the order it runs, so you can judge the output for yourself and argue with it where it deserves an argument.
The short version
The model takes each competitive race, blends three independent reads on it (the ratings, the polling, and the state’s underlying partisan lean), adjusts for the national mood, and then simulates the entire map tens of thousands of times. A party’s odds are simply the share of those simulations it wins. Nothing here favors a party by design. The inputs are public, the weights follow from how reliable each input is, and the same recipe runs for a seat in a deep blue state and a seat in a deep red one.
How the forecast works
1. Ratings
For each race the model starts with the published competitiveness ratings (Cook Political Report and Sabato’s Crystal Ball, where available). A rating like “Lean R” or “Toss-Up” is translated into an expected vote margin using a lookup that was trained on real Senate outcomes from 2014 through 2024. In other words, “Lean R” means whatever “Lean R” races actually did over the last decade, not a number someone picked by feel.
2. Polling, weighted by age
Where a race has polling, the model folds it in. Fresh polls count for more than stale ones. The reliability of a poll degrades as it ages, so an old survey is widened toward uncertainty and pulled back toward the rating-based estimate. A race with no recent polling leans more heavily on its rating and its baseline (below). A race with fresh, plentiful polling lets the polls do more of the talking.
There is a hard floor under that. Poll age is counted in whole days, and a poll we can date is dropped once it reaches 46 days old. Up to that point the age penalty applies and the poll keeps a reduced say; past it the poll is dropped, the race runs on its rating and its baseline, and the race is marked as carrying no current polling rather than being handed a number that a survey from six weeks ago helped set. The line is hard, not a steep discount: a poll on its 45th day still counts, and on the next day it does not.
Some of the polling we hold arrives with no usable date on it. A missing date is not evidence that a poll is fresh, and we will not invent one, so those polls are dropped as well and the race runs on its rating and its baseline. The reason is plainer than it looks: a poll we cannot date can never reach the cutoff above, however old it really is, so exempting it would put a share of our polling permanently outside a rule we publish. Each run keeps its own record of the two kinds of exclusion separately, the stale and the undated, because the first is a poll that aged out and the second is a gap in what our sources gave us. That record is not yet something you can browse race by race on this site, and until it is, the count above is what we can show you.
Until this change an old poll was widened toward uncertainty but never removed, so a spring survey still helped set a race’s number in August. We measured the change before shipping it, by replaying the model against the facts as they stood on 24 August 2026: 119 of the 470 modeled races lost their polling to the cutoff, one of them a House district whose stored poll date read October 2016. In that replay the Senate topline did not move and the House probability of a Republican majority moved by less than a point. Those are the numbers of a test replay on one day’s data, not of the run you are reading.
The day after that cutoff shipped we stopped exempting the undated polls, and this one was not a small change. Replaying the model against the facts as they stood on 25 August 2026, 72 more races lost their polling, the count of races the model could use polling for fell from 122 to 50 out of 470, and the House moved from a Republican majority probability of 25 percent to 44 percent, which carried its label from Lean D to Toss-up. The Senate did not move at all, because only two of the 72 were Senate races. We are telling you this rather than letting the number move quietly: the shift is a change in what the model was willing to trust, not something a voter did. What those 72 polls had in common was that we could not date them; several claimed margins their districts make very hard to believe, and one stored date read 2026-05-00, which is not a day. We are not claiming the polls that survived are sound. Those are the numbers of a test replay on one day’s data, not of the run you are reading.
3. Partisan baseline
Every state has an underlying lean that does not move much year to year. The model captures it from the Cook Partisan Voting Index and the 2024 presidential margin, expressed relative to the country as a whole. This is the floor under a race. It keeps an unpolled seat in a lopsided state from drifting on thin evidence, and it gives the model a sane prior when the other two signals are quiet.
4. The national environment
Races do not move independently. When the national mood shifts toward one party, it tends to shift in many places at once. The model carries that shared mood as a single national number and applies it across the board, then lets each race vary around it. That is why a small national swing can move a lot of seats together, and why the forecast does not treat the map as a set of unrelated coin flips.
That national number starts from the generic congressional ballot (the “which party would you vote for” polling average). We do not feed it in raw, and here is the part we want a skeptical reader to see plainly: the generic ballot has typically overstated Democrats in recent cycles. So we shrink the raw number toward zero and apply a Republican-leaning correction sized to that pattern, a fixed slope-shrink plus a Republican-direction haircut (both disclosed, with their exact values, in the model parameters below) before the model uses it. If the raw average reads Democrats plus seven, the model works from closer to Democrats plus four. We would rather haircut a figure that flatters one side than run the raw number and let a known bias sit uncorrected.
We also widen the uncertainty on that national number the further we are from Election Day, on a fixed schedule set in advance. Six months out, the honest error bar on the national environment is wide (the generic ballot can be off by five points even on the eve of an election, and more than that this early). As November approaches, the band narrows on a pre-committed timetable, not because anyone decided the race got clearer. A forecast that quietly grew more confident as one party’s position improved would be telling you a story; the schedule keeps it honest.
5. Fusion, weighted by precision
For each race the model now holds three estimates (rating, poll, baseline), each with its own uncertainty. It combines them by trusting the more certain ones more. A confident, fresh poll outweighs a vague rating. A noisy poll defers to the baseline. The result is one combined estimate per race, with an honest error bar attached. No signal is allowed to dominate just because it is loud. It has to be precise to carry weight.
The three have to be measuring the same thing before they can be combined. Ratings and the partisan baseline describe how a seat leans compared with the country. A poll does not: it measures the race as it stands, national mood included. So the model subtracts the national number from the poll before combining, and adds it back once, in the simulation. Get that wrong and the national mood is counted twice for every polled race, which is exactly what this forecast did until 26 August 2026.
6. The simulation
With a combined estimate and error bar for every race, the model runs the whole map tens of thousands of times. Each run draws one shared national shock (so the races move together, as they do in real life) plus its own local noise for each seat. Every run produces a full seat count. Stack up all the runs and you get a distribution: not a single number, but the range of plausible outcomes and how often each one shows up.
7. Seats and probabilities
From that pile of simulations the model reports the median seat split (the middle outcome), an 80 percent range (the band that contains four out of five simulations), and the probability of a majority (the share of simulations in which a party reaches the threshold: 51 in the Senate, 218 in the House). Joint control of Congress is counted from the same simulations, not multiplied together after the fact, because the two chambers tend to move in the same direction in a given year.
Model parameters
These are the actual numbers the model runs on. The stable constants come from the model configuration; a current-run value (marked below) is read from the latest completed run, never re-typed by hand.
- Chamber correlation
- ρ = 0.80
- National shock (election eve)
- 2.5 / 3.0 pts
- Local error, by what we know
- 3.0 / 6.9 / 5.2 pts
- Poll error (combined)
- 3.54 pts
- Poll age penalty
- 0.05 pt / day
- Poll cutoff
- 45 days
- Baseline uncertainty
- 6.0 pts
- Generic-ballot shrinkage
- 0.87
- Thin-polling inflation
- up to 2.0 pts
- This run’s national mood
- + 4.1 pts
How tightly the House and Senate national moods move together in the simulation.
Senate / House uncertainty floor on the national mood, widening on a fixed schedule (floor + 2.5 × √(days ÷ 180)) the further out the model runs.
The uncertainty belonging to a race alone, on top of the national shock, for a seat with a current poll, a seat the raters call competitive, and a seat carried by its partisan lean. A current poll buys the most certainty. A contested seat carries the most uncertainty of its own, because that is where the money and the recruited challengers are; a landslide seat sits between them for the same reason in reverse. These are judgments anchored to published error statistics, not measurements of our own accuracy.
Sampling error (2.5 pts) and systematic/house-effect error (2.5 pts) combined in quadrature.
Charged against a poll for every day of its age, up to the cutoff below. The widest penalty a poll still in use can carry is 2.25 points.
A poll is excluded outright on its 46th day, and a poll that reaches us with no date we can read is excluded whatever its true age. Either way the race falls back to its rating and baseline and is counted as carrying no current polling. ; .
Fixed sigma on the partisan-baseline signal (Cook PVI and 2024 presidential margin).
Slope applied to the raw generic-ballot margin, plus a 1.75-point Republican-direction haircut for the ballot’s historical Democratic overstatement.
Extra uncertainty added when the national-environment read comes from few polls.
The regressed, Dem-positive national-shock mean actually used by the latest completed run (from forecast_runs.raw_json.modelParameters, not a re-typed constant).
The calibration review
Every simulated race carries two kinds of uncertainty: the part it shares with the whole map (the national shock) and the part that belongs to that race alone. For most of this forecast’s life the second part came out to zero for all 470 modeled races, because the model subtracted the national uncertainty out of each race instead of adding the two together. A freshly polled race and one nobody has polled in months carried the identical interval. That has been rebuilt. The two now add, and the size of a race’s own share is set by what we actually know about that race: whether it has a current poll, whether the raters consider it in play, or whether it rests on its partisan lean alone.
The House seat range still does not appear on this page. The reason has changed, and the reason we gave before was wrong, so it is worth being exact about both.
We are not publishing a range we cannot defend and hoping nobody checks. We are withholding it instead. The median seat count and the probability of a majority stay on the page, and the range and the shape of the distribution are gated. What would lift the gate is a reliability test over past election cycles the model has never seen: an eighty percent range is a claim about how often the answer lands inside it, and we have not yet measured how often ours does. Rebuilding the variance math was necessary for that test to be worth running. It was not the test.
This section used to say the directional verdicts were unaffected by the variance defect. That was too strong, and checking it while making a separate change is how we found out. Against the last reading published before the fix, the Senate label moved from Lean D back to Toss-up and its majority odds moved about seventeen points. The House label stayed at Toss-up. Every reading published before that date was taken with local uncertainty set to zero, which was the one thing this page told you we did not believe.
This section also used to say that the corrected math produced a House range tighter than a forecast this far out has a right to be, and that the range would return once the rebuild landed. The rebuild has landed, and that prediction did not survive contact with it: the corrected eighty percent range is about as wide as the one it replaced, not narrower. Withholding on the grounds that the number would look too confident was the wrong reason. The right reason, stated above, is that we have never measured whether our ranges cover what they claim to cover, and that is a separate piece of work.
Polling and rating sources
These are the model’s research and aggregation universe. A source appearing here means the pipeline draws on it as part of the model’s inputs, not that every source feeds every single run.
| Category | Sources |
|---|---|
| Ratings | Cook Political Report, Sabato’s Crystal Ball |
| Race polling | RealClearPolitics, FiveThirtyEight, Decision Desk HQ, 270toWin |
| Named pollsters | Marist, Quinnipiac, Emerson College |
| National environment | RealClearPolitics, FiveThirtyEight/ABC, Silver Bulletin, Race to the WH, fiftyplusone, Sabato’s Crystal Ball |
The model does not weight individual pollsters differently from one another; polling reliability is captured only through the age penalty above, not a per-pollster house- effect adjustment.
How to read the four views
The forecast page shows the same model from four angles. They are not four models. They are four ways of looking at one.
- The topline and chamber cards. The headline odds for Congress, plus a card for each chamber with its median seats and its majority odds. The Senate card also carries an 80 percent range; the House range is withheld, for the reason given under the calibration review below. This is the answer to “who is favored, and by how much.”
- How it moved. The change since the prior run. Toward Republicans, toward Democrats, or steady, with the size of the shift and the days between runs. This is the trend, not the level.
- What each signal says. Where the ratings alone, the polls alone, and the baseline alone each land, before fusion. When they agree, confidence is high. When they disagree, the gap between them is the uncertainty, shown plainly rather than hidden.
- What if the mood shifts. The seat counts under a few pinned national environments (say, a Democratic-leaning year, a neutral year, and a Republican-leaning year). This shows how sensitive the map is to the one thing nobody can poll yet: the mood on election day.
Where it can be wrong
A model is only as good as its inputs and its assumptions. Ratings can lag a fast-moving race. Polling can be sparse, herded, or systematically off, as it has been in recent cycles. The partisan baseline assumes the recent past still describes the present, which is exactly the assumption a realignment would break. The national-environment estimate is itself uncertain, and the correlation between chambers is a modeling choice, not a measured fact. The forecast reports ranges and probabilities precisely so that none of this is papered over. A 70 percent favorite loses about three times in ten. That is not the model failing. That is the model telling you the truth about how much it does not know.
A second model, kept off this page
There is now a second forecast running alongside this one. It is a structural model: it reads a district or state partisan lean, the national generic ballot, and incumbency, and it reads nothing else. No polls. No race ratings. It exists to disagree with the forecast above, because a model built from different inputs disagreeing is information, and two models built from the same inputs agreeing is not.
When it was built, it also treated uncertainty differently. The forecast above used to split each race's uncertainty into a national part and a local part by subtraction, which meant a race whose own signal was tighter than the national environment ended up with no independent uncertainty of its own. In the run compared below, that was true of all 470 races. The second model adds the two parts together instead. The forecast above now does the same, which is the second model having done the job it was built for.
The comparison has already turned up two defects in our own data. When the second model first ran, it disagreed with the forecast most sharply in a set of districts where the published forecast was leaning on polls carrying no date at all. Chasing that down found both problems: polls were being stored without recording which party led, and undated polls were never being aged out. The forecast above no longer uses either, and the two dated notes further up this page record what changed.
With those fixed, the two models moved a long way toward each other on the House and stayed far apart on the Senate. Against the forecast run of August 25, 2026, the published model put Republican odds of a House majority at 43.6 percent and the structural model at 44.6 percent, a gap of one point. On the Senate the published model said 41.9 percent and the structural model said 76.4 percent, a gap of 34.5 points. Two things drove the Senate gap and it is worth keeping them apart: the published forecast reads polls and race ratings that the structural model does not, and at the time of that comparison the two models divided up uncertainty differently. The gap was not a clean measurement of either one on its own.
Those figures describe one comparison against one run, and that run predates the variance rebuild described further up this page. They do not update themselves, the second model does not yet run on a schedule, and the numbers move whenever the generic ballot does. Treat them as a dated reading, not as a standing score. The published half in particular has moved since, and in both directions: repairing the polling supply took the Senate number down, and the two model fixes described above took it back up past where it started. The next run of the second model is what will replace these figures, and when it runs on its own cadence this section will take its numbers from the run instead of from a figure typed in by hand.
The second model is not calibrated. Its accuracy has not been tested against past cycles, no reliability curve has been run on it, and nothing about it changes the numbers on the forecast page. It does not appear anywhere a reader would mistake it for the publication's position, and the withheld House seat range stays withheld. It is a check on our own work, published because a check nobody can see is not a check.
One thing the check found is worth stating plainly, because it was a criticism of the forecast above rather than of the second model. In the run compared here, every one of the 470 races came out with no independent uncertainty of its own, which is the subtraction problem described earlier. That has now been fixed. Each race carries an independent error bar sized by what is actually known about it, so a seat with a current poll, a seat the raters call competitive, and a landslide seat no longer read alike. Every race’s interval widened and none narrowed. A second defect surfaced alongside it: the forecast was fusing a poll, which measures a race including the national mood, with two signals that already have the national mood removed, and then adding the national mood again on top. That pushed every polled race toward the party the national number favored, by about one and three-quarter points. Both are fixed. Together they moved the Senate from Lean D back to Toss-up, which is a change in what the model says produced by fixing how it counts rather than by anything a candidate did. The House stayed at Toss-up.
What it is not
It is not a bet, an endorsement, or a reason to tune out. The point of putting the reasoning on the page (signal by signal, with the error bars showing) is that a forecast you can inspect is a forecast you can hold accountable. Watch whether the races resolve inside the ranges it gave. That is the test, and we will keep score.