Skip to main content

METHODOLOGY

Full transparency on how scores are calculated, where data comes from, and why each analytical technique was chosen. No black boxes.

OUR APPROACH

Civitas is an open-data AI/ML platform that aggregates data from official U.S. government sources into unified transparency scorecards for senators, House representatives, presidents, and Supreme Court justices. Every score is computed from publicly available federal records. We do not editorialize, endorse, or oppose any candidate or party.

What the scores measure: for senators and House representatives, how well they carry out the will of the majority of their constituents— not the preferences of a few wealthy donors, and not party defection for its own sake; for presidents, how well they serve the country; for Supreme Court justices, how well they serve the law regardless of party. Every scoring dimension is justified against that yardstick: crossing party lines is credited only where it plausibly moves toward the state's median voter, and the funding dimensions exist because money concentrated in few hands is the main channel by which representation drifts away from the majority.

Scores reflect observable behavior — voting patterns, funding sources, legislative activity — not ideology. The formulas are symmetric across parties: the same voting record in the same seat produces the same score regardless of whether the member is a Democrat or Republican. The system is designed to be structurally non-partisan.

Every metric on the scorecard includes a [?] tooltip explaining what it measures and how to interpret it. Hover on desktop or tap on mobile. We believe no number should be presented without context — if you see a metric, you should be able to understand what it means and where it came from.

When data is missing or insufficient, scores default to a neutral 50 out of 100. No politician is penalized for something we cannot measure, and no politician receives a perfect score without evidence. This implements Bayesian shrinkage toward a neutral prior — a standard statistical technique for preventing extreme estimates from small samples. [19] Efron & Morris 1975

The Action Center extends this mission to daily civic engagement. It automatically surfaces trending issues from news analysis, provides objective summaries free of editorial opinion, and recommends non-partisan actions citizens can take to participate in their government — without assuming which side of any issue the reader supports.

CONGRESSIONAL SCORECARD METRICS

Every senator and House representative receives three sub-scores on a 0-100 scale, weighted into an overall Representation Score. Higher is better. All 100 senators and 435 House representatives are scored with the identical framework below — same formulas, same data sources (FEC, Congress.gov, GovInfo), same classification techniques — so scores are directly comparable across both chambers. House members are sourced from the same Congress.gov and FEC endpoints and processed in the same nightly pipeline run as senators; the House leaderboard supports pagination and party filtering to navigate the larger membership.

Campaign-promise tracking (kept/broken/partial) is still collected and shown on each member's profile, but is not folded into the weighted score below. For the audit history behind the current weights and dimensions — including why Promise Persistence was removed, why Funding Diversity was folded into Funding Independence, and why Donor Independence was removed from Constituent Alignment — see the scoring changelog.

Funding Independence (33%)

In short: rewards members whose campaigns are funded by lots of small individual donors rather than PACs or a handful of big donors and industries. The more spread-out and grassroots the money, the higher this score — regardless of party or chamber.

Measures five dimensions: (1) PAC dependency — a blend of the share of funding from PACs and how close contributing PACs are to their legal per-election cap, chamber-specific since Senate and House candidates rely on PAC money at structurally different rates; (2) the share of funding from small (<$200, unitemized) donors — the broadest possible funding base; (3) relative top-donor concentration — what fraction of the itemized external donor pool comes from the top 10 donors, with the member's own money and transfers from their own committees excluded; (4) source breadth — small-donor money counts fully, industry-classified money counts moderately, and opaque money counts least; and (5) industry concentration — the inverse Herfindahl-Hirschman Index (HHI) of industry donations, where funding concentrated in a single industry suggests potential regulatory capture. Components (4) and (5) were folded in from a separate Funding Diversity dimension in 2026-07 after finding the two dimensions correlated at r=0.72 across the Senate — the same underlying funding-profile signal under two labels, not two genuinely distinct ones. PAC dependency follows Stratmann (2005), [5] Stratmann 2005who found that PAC contributions are more strongly correlated with roll-call alignment than individual contributions. The concentration components apply the same intuition as HHI at the donor level, following Bonica (2014) [1] Bonica 2014 and the industrial-organization logic Rhoades (1993) [6] Rhoades 1993 built the HHI metric on.

"UNCLASSIFIED" money (committee transfers, joint-fundraising splits, donations lacking employer data — a real 32% median share of total funding across the Senate) is scored neutrally rather than penalized. It is a residual we cannot attribute at all, not evidence of concentration in one source — the same "missing data defaults to neutral" principle applied everywhere else on this page.

Constituent Alignment (33%)

In short: checks whether a member's voting actually matches what their state or district elected them to do. Voting with your party is notpenalized on its own — for a safe-seat member, that often IS representing your constituents. The score moves below neutral only for a clear, readable sign of a mismatch — a voting position toward the party's flank for a seat that isn't safe for that flank — and moves above neutral for the mirror case: a voting position genuinely in step with the seat.

Measures how a member's voting compares to what their state elected them to do — not raw defection from party. Each member's contested-vote break rate is scored against a seat-specific expectation derived from state partisan lean (Cook PVI [4] Carson et al. 2010): an aligned safe seat expects near-base-rate dissent (~3%), a swing seat ~8%, and a seat whose electorate leans toward the opposing party up to ~20%. Matching the expectation scores ~50 — a typical partisan for that seat. The score is deliberately asymmetric around that expectation (v6.6), under one governing principle: it moves off neutral only for readableevidence a member is representing their constituents, and treats behavior whose meaning can't be read as neutral. Below-expected loyalty is notpenalized: it floors at neutral, never below. A low defection rate is unreadable — it may be faithful representation of the coalition that elected the member (not the geographic median voter), it is the structural norm for both parties in the modern Senate, and being "out of step" is a matter of ideological position, not a loyalty rate — so we decline to score the rate itself rather than penalize it. This loyalty floor was the only behavioral change in v6.6, and it is deterministic: a below-expected loyalist with no other signal available scores exactly 50 on this component. v6.7 adds one legible exception: if a member's cosponsorship-derived ideology score places them in their own party's most extreme third, and their seat isn't safely aligned for that extremity, they are discounted below neutral (scaled by how unsafe the seat is — not penalized at all in a genuinely safe seat, where that extremity is the structural norm for both parties). This targets ideological POSITION, not the loyalty rate itself, so it doesn't reopen the rate-is-unreadable problem above — a member can be maximally loyal and still be flagged if their position is a clear outlier for their seat. This discount's maximum strength was reduced in v6.8 after a fairness audit found it overlapped with coalition breadth (below) more than intended — see the scoring changelog and the Known Limitations note below. Above-expected crossing is the readable side: it earns credit only where it plausibly moves toward the seat's political center, discounted by seat lean (full credit in opposed and swing seats, near-neutral in deep aligned seats, since there the center sits with the party). A further discount for members positioned on their party's ideological flank — whose crossings more likely point awayfrom the center (Kirkland & Slapin 2017) — was designed but is notshipped, and checking it against live data found a deeper problem than a missing calibration number: the members who actually cross party lines most often all read as ideologically centrist on this platform's own cosponsorship-based ideology measure, not flank-extreme — the opposite of what the discount assumes. Crossing behavior and this measure of ideology turn out to be linked rather than independent, so this specific fix is shelved, not just uncalibrated. This is a deliberately humble use of the delegate model, with partisan lean standing in for issue-level constituent opinion — a measurable, disclosed simplification (see Known Limitations below, including why this measures the rate and direction of a member's deviation rather than the distance between their position and their constituency's, so it cannot yet positively credit representation achieved through congruent loyalty). Note on composition: confirmation votes on nominations make up a large share of recent Senate roll calls and count at full weight — they are genuine, whipped party-line tests.

Before v4.2 this dimension was called Independent Voting and rewarded raw defection; it also exempted party-line votes on policy areas related to a member's top donor industries. That exemption is removed: donor industries are not a proxy for state interests, and it shielded exactly the votes most suspect for donor influence.

The score blends seat-relative vote alignment (70%, or 100% when roll-call ideal-point data is unavailable) with position congruence (30%, when available — v6.11): the member's DW-NOMINATE first-dimension position (Voteview; the standard roll-call-based position measure in political science) compared against a seat-conditional expectation — what a same-party member of a similarly-leaning seat typically holds, fit per chamber and per party from real data. Per-party fits deliberately avoid the swing-seat artifact a single pooled fit would create (Bafumi & Herron 2010). A position toward the party's flank relative to that norm scores below neutral (scaled by how unsafe the seat is — flank positions in genuinely safe seats are the structural norm and are not penalized); a position toward the seat's center scores above neutral (with the same seat-direction credit shape as surplus crossing). When active, this component supersedes the v6.7 cosponsorship-based discount above — same construct, better signal, and measuring it twice would repeat the exact double-count v6.8 fixed. Coalition breadth (20% here from v5 through v6.10) has moved to Legislative Effectiveness: bipartisan coalition-building is a legislative-effectiveness signal, not a constituent-alignment one — demand for bipartisanship varies with the seat's own makeup (Harbridge & Malhotra 2011), so a bipartisan member of a lopsided seat can be bipartisan and misaligned at once. Through v6.4 this dimension also included a Donor Independence component (25%, a heuristic based on the money associated with donor-vote topical overlaps) — removed in 2026-07 after finding it measured a close cousin of the Funding Independence signal (both keyed off total money raised and donor-industry concentration) while itself reducing to one of four fixed values for 85% of senators, since no data source discloses per-bill donor positions. Its freed weight now goes entirely to seat-relative vote alignment. We follow the methodological caution of Ansolabehere et al. (2003) [18] Ansolabehere et al. 2003 in interpreting donation-vote correlations generally: correlation does not prove causation. [5] Stratmann 2005

Legislative Effectiveness (34%)

In short: measures whether a member is actually getting legislative work done — introducing bills that matter, moving them through Congress, and building a network of cosponsors other members trust. Introducing a substantive bill earns real credit even before it passes, matching how political scientists actually measure legislative productivity.

Measures whether a member is producing tangible legislative outcomes, following Volden & Wiseman's (2014) [34] Volden & Wiseman 2014 real published methodology: each sponsored bill is weighted by significance (5x for substantive bills — S./H.R./joint resolutions; 1x for commemorative simple/concurrent resolutions) and credited cumulatively across every stage it reaches — introducing a bill earns real credit on its own, not just bills that go on to pass a chamber or become law. Three components: bill significance & advancement (60%) — this cumulative stage-credit per congress served, compared against an expected credit for a sponsor of this chamber/majority-minority status; legislative leadership (25%) — cosponsorship-network PageRank, see below; and bipartisan coalition attraction (15%, v6.11 — moved here from Constituent Alignment): the share of cosponsors a member attracts to their own bills from the other party, chamber-median-normalized. Harbridge-Yong, Volden & Wiseman (2023) show that attracting cross-party cosponsors robustly predicts lawmaking success for majority and minority members alike — and that it is specifically the attraction of bipartisan cosponsors, not the offering of cosponsorships across the aisle, that carries the effect, so this component uses a receive-only rate rather than the blended give-and-receive bipartisanship figure shown on profiles. Their evidence is a strong association rather than a clean causal proof, and the measure is an input to effectiveness rather than realized output — both disclosed reasons its weight stays modest. When cosponsorship data is missing, the split reverts to exactly the prior 70/30.

Because introduction itself earns credit, a member who sponsors many substantive bills can score well even before any of them advance further — this is Volden & Wiseman's real design, not a bug: their published methodology counts a bill's contribution at every stage it reaches, and most sponsored bills never advance at all (our own corpus measures Senate majority sponsors advancing bills at 3.6% vs. 2.4% for minority sponsors; House 6.4% vs. 2.4%). The expected-credit baseline accounts for that majority/minority gap, so scoring everyone against one absolute threshold doesn't silently penalize whichever party is out of power. The score explanation on each profile breaks the substantive-bill count into introduced-only / advanced-further / became-law so the volume-vs-advancement split is visible as real numbers.

Both the expected-credit baseline and the majority/minority-status adjustment above are calibrated separately for the House and Senate (v6.9) — the two chambers' real bill volumes and advancement rates genuinely differ, and comparing every member against one shared, pooled-across-chambers figure previously understated House members' effectiveness and overstated the Senate's. See the scoring changelog for the live-population numbers behind that fix.

KNOWN LIMITATIONS & DISCLOSURES

Every item below is an open engineering problem, not a settled tradeoff we've made peace with — several started as disclosures here and were later fixed outright (see the v6.8 entry below, and the scoring changelog for the full history). Where a limitation is fixable, we fix it and remove the disclosure. Where it isn't — no dataset exists, or fixing it would require an editorial judgment call the platform's no-hardcoded-conclusions rule resists — we name the specific reason why, so it can be revisited if that changes.

In short: Democratic and Republican senators finance their campaigns differently on average, so Funding Independence scores differ by party on average too — not because the formula treats parties differently, but because the underlying fundraising behavior really is different.

Scores correlate with funding style, and funding style correlates with party. In current data, Democratic senators take roughly half the PAC share of Republican senators (median ~10% vs ~17%) and raise about twice the small-donor share (~24% vs ~12%). Because Funding Independence measures those behaviors directly, average scores differ by party. The formulas are identical for everyone and contain no party term; the gap reflects measured funding behavior, not editorial judgment.

In short: a bigger campaign naturally looks more "independent" by percentage even with the same PAC dollars, simply because the total got bigger. We also check absolute PAC dollars, but no single number fully separates "independent" from "big."

Fundraising scale still matters. Larger campaigns naturally have smaller PAC sharesbecause PAC checks are legally capped while individual money is not. We mitigate this by scoring absolute PAC dollars alongside the share, but no single number fully separates "independent" from "big."

In short: election cycles run different lengths for different members (and for the House vs. the Senate), so funding scores are technically comparing different-sized snapshots of time across members.

Comparison windows differ by tenure and chamber.Funding metrics cover a member's two most recent election periods — roughly 8 years for a veteran senator, 2 for a freshman, 4 for House members — so cross-member comparisons weigh different spans of time.

In short: when we flag a donor whose industry overlaps with a vote, that shows where money and legislative activity intersect — it is not proof the donation influenced the vote.

Donor-vote connections are semantic overlaps, not lobbying records. They aggregate employee and PAC money associated with an organization and match it to vote topics by embedding similarity. They indicate where money and votes intersect; they do not establish influence.

In short: a senator whose donors cluster in one industry scores the same whether that industry is their state's home industry or an out-of-state special interest. That's deliberate, not an oversight — see below for why.

Concentrated industry funding is scored as capture risk even when it plausibly reflects a state's real economic base.A senator whose donations concentrate in, say, the auto industry in Michigan or agriculture in Kansas scores the same on Funding Independence's industry-concentration component as one captured by an unrelated out-of-state interest — this platform does not check whether a donor industry is also a major local employer. That is a deliberate choice, not an oversight: we considered and rejected a "this industry matters to the state" exemption for the same reason the v4.2 donor-industry voting exemption was removed (see the scoring changelog) — local economic dominance plausibly gives an industry moreleverage over a senator, not less, so exempting it would weaken the signal exactly where large-scale capture is most consequential. No public dataset can separate "this funding reflects genuine local interest" from "this funding is capture that happens to correlate with local economic weight" — concentration is scored as risk, full stop, following the same industrial-organization logic (Rhoades 1993) the HHI metric is built on.

In short: we estimate what a senator's state "expects" from how the state votes for president overall, not opinion on the specific bill in front of them — a broad stand-in for local opinion, not a precise one. We looked for a better public alternative and didn't find one that wasn't itself stale or a black box (see below).

Presidential-vote PVI doesn't capture issue-specific constituent opinion.A senator's expected break rate (see Constituent Alignment above) is calibrated to how their state votes for president, not to opinion on the specific issue a given vote concerns — a state's presidential lean says little about, say, local opinion on public land use in Utah or water rights in Arizona. We looked for a real, freely available substitute: the best candidate found (Tausanovitch & Warshaw's survey-based ideology estimates by district/state) still only produces a single composite left-right score, the same kind of proxy PVI already is — not per-issue opinion — and its public data is already several years stale. Actual issue-level constituent opinion at this scale would require building multilevel-regression-and-poststratification (MRP) modeling in-house over raw survey microdata: a statistics pipeline, not a lookup, and a genuine black box relative to every other formula on this page. We chose not to build one rather than trade this platform's auditability for a partial, hard-to-explain fix.

In short: the score can now give extra credit to a senator who is genuinely in step with their state — the long-disclosed gap where a well-matched loyalist could never score above neutral is closed as of v6.11 — but the fix measures "in step with the seat" against the seat's overall partisan lean, which is not the same thing as the voters who actually elected the member. That residual simplification is still disclosed below.

Constituent Alignment's positive-credit gap is closed (v6.11), with a disclosed residual.Through v6.10 the dimension could flag a loyalist whose position was a clear outlier for their seat, but had no way to reward the mirror case — a member whose positions genuinely match their seat scored the same neutral ~50 as an unreadable loyalist. The blocker named here in earlier versions was the yardstick: every member sits more extreme than their state's raw median (Bafumi & Herron 2010), so the median itself was the wrong target, and authoring one by hand would violate the no-hardcoded-conclusions rule. The v6.11 position-congruence component resolves that with a party-relative, data-derived target — what a same-party member of a similarly-leaning seat typically holds, fit from the live chamber — exactly the party- or coalition-relative benchmark this disclosure said was needed (Canes-Wrone, Brady & Cogan 2002). What remains, and stays disclosed: the target is derived from seat partisan lean, a one-dimensional electoral proxy — members systematically track their reelection constituency (primary voters and copartisans) rather than the geographic median (Fenno 1978; Clinton 2006), which is why congruence credit is seat-direction-scaled rather than taken at face value, and why issue-level opinion data (e.g. MRP estimates or CES roll-call-matched items) remains the named next step for this dimension.

In short: two of the checks that used to lower Constituent Alignment were computed from the same underlying data, so they were catching the same problem twice. v6.8 cut that overlap; v6.11 removes it structurally — the two signals no longer live in the same dimension, and the position check now uses an independent data source where available. A smaller cousin of the same caveat now applies inside Legislative Effectiveness.

Cosponsorship-derived signal overlap — mostly resolved, one residual.A 2026-07-21 audit found Constituent Alignment's position-mismatch discount and coalition breadth correlate at r=-0.76 (58% shared variance, n=99) — both were projections of the same cosponsorship network. v6.8 reduced the double-count; v6.11 removes its structural basis: coalition breadth has left Constituent Alignment entirely (it now scores legislative effectiveness, where the evidence supports it), and the position signal is measured from roll-call ideal points (DW-NOMINATE via Voteview, ingested automatically every pipeline run behind ingestion gates) — the genuinely independent second signal this disclosure previously said wasn't available — with the cosponsorship-SVD discount surviving only as a fallback for members the ideal-point data doesn't yet cover. The residual: within Legislative Effectiveness, the leadership component (cosponsorship PageRank) and the new bipartisan-coalition-attraction component are both computed from the cosponsorship network (network centrality vs. cross-party share — related data, different measures). Their combined weight is capped at 40% for that reason, and their live correlation is a standing post-run check. See the scoring changelog for the full account.

SPONSORSHIP ANALYSIS (LEADERSHIP & IDEOLOGY)

Every senator and representative also receives two metrics derived from cosponsorship networks — the pattern of which members sign onto each other's bills (within each chamber's own network; House and Senate cosponsorship are separate graphs). Ideology is purely informational context. Legislative Leadership is notpurely informational — it already feeds into Legislative Effectiveness above at 25% weight (30% in the fallback split used when cosponsorship data is missing); the number shown on a member's card is the same underlying score, displayed directly (with a tenure adjustment, see below) rather than hidden inside the composite.

Legislative Leadership (0-100)

Measures legislative influence using the PageRank algorithm [32] Brin & Page 1998applied to cosponsorship networks. When Senator A cosponsors Senator B's bill, that creates a directed link in the network. PageRank computes centrality: a senator whose bills attract many cosponsors — especially from other influential senators — receives a higher score. This mirrors GovTrack's leadership methodology. [33] Tauberer 2012

The algorithm uses power iteration with a damping factor of 0.85 and converges in ~50 iterations. Raw PageRank values are rescaled to [0, 1] using a logarithmic transformation to compress the heavy-tailed distribution, then displayed as 0-100.

Network centrality structurally takes years to build — a freshman senator's raw score is near-zero not because they lead poorly but because they haven't had time to accumulate cosponsorship connections yet. Both the score component and the displayed number shrink the raw value toward neutral 50 for senators with under 6 years in office, confidence-scaled to a full term, so a brand-new senator reads as "not enough track record yet" rather than "bad at leadership."

Raw cosponsorship-network centrality can't on its own tell a substantive bill from a message bill introduced with no real chance of passing — a senator who signs onto ten symbolic resolutions accrued the same network weight as one who cosponsors ten bills that actually became law. Since v6.2, each cosponsorship is weighted by what happened to the underlying bill: full weight if it became law, reduced weight if it passed a chamber or cleared committee, and further reduced (not zeroed — a stalled bill is still real evidence of a working relationship) if it never advanced.

Ideology Score (0-1)

Computes a behavioral ideological position using Singular Value Decomposition (SVD) on the cosponsorship matrix, following Tauberer (2012). [33] Tauberer 2012The second singular vector (first is trivially related to overall activity) captures the primary ideological dimension — the axis along which senators most differ in who they cosponsor. This is analogous to DW-NOMINATE [20] Poole & Rosenthal 1985but derived from cosponsorship patterns rather than roll-call votes.

The ideology score is oriented so that lower values correspond to progressive positions and higher values to conservative positions, calibrated by checking the mean score of each party. It serves as a Bayesian prior for the partisan depth calculation: when a senator has few recorded votes, the ideology score regularizes the estimate; as vote data accumulates, the prior weight drops to zero. [19] Efron & Morris 1975

Sponsorship Description

Combines the leadership and ideology scores into a human-readable label (e.g., "progressive Democratic leader" or "conservative Republican backbencher"). The label encodes three dimensions: ideological position (progressive/moderate/conservative), party affiliation, and influence tier (leader/rank-and-file/backbencher).

PRESIDENTIAL SCORECARD METRICS

Presidents are scored on four dimensions, 0-100 scale, computed entirely from live, historical, and expert-survey datasets — there is no hand-set or seeded score anywhere in this pipeline (2026-07 rewrite). A dimension a president has no real data source for is left blank (N/A) rather than filled with a fabricated or neutral placeholder, and the overall score renormalizes across whichever dimensions actually apply to that president. Identity data (name, party, term dates) is fetched live too, from the same UCSB roster used for the metrics below — nothing about a president's profile is typed into this codebase by hand.

Independence and Follow-Through were removed entirely (2026-07), not just disclosed as limitations. Both were always a one-time hand-set number with no live formula and, unlike every dimension below, no realistic path to one: Independence's obvious data source (OpenSecrets' cabinet/appointee revolving-door tracking) was itself discontinued in 2025, and Follow-Through would need the same platform-text-vs-action matching technique already tried four times and abandoned for senators' Promise Persistence (see the scoring changelog — v6.0). Rather than keep presenting a hand-set number as a computed score, they're gone. Their combined weight first redistributed proportionally across the remaining four (Public Mandate 15→23%, Effectiveness 20→31%, Competence 15→23%, Agency Alignment 15→23%), then a fifth dimension — Historical Legacy — was added shortly after (also 2026-07, following review that found presidents like Lincoln landing in the bottom half of the ranking despite every individual number being defensible on its own terms: nothing in the first four dimensions could credit “preserved the Union, ended slavery” at all).

Historical Legacy's weight went through two revisions before landing at 35%, both checked against the real 47-president dataset rather than picked by eye. Equal fifths (20%) let the other four dimensions — which individually barely correlate with historian judgment at all (Spearman 0.17 between the four mechanical dimensions alone and C-SPAN's own ranking) — outvote the one dimension that actually tracks it, putting Coolidge, McKinley, and Harding in the top 10 while Lincoln and Eisenhower fell out of it. Raising Historical Legacy to 50% fixed that, but introduced a different problem: at 50%, this platform's overall ranking correlated 0.96 with simply using C-SPAN's own ranking alone — the four mechanical dimensions were contributing almost nothing of their own. 35% is the point where the top of the ranking is already recognizable (FDR, Washington, Lincoln, Theodore Roosevelt, JFK, Eisenhower) while the mechanical dimensions still meaningfully move the rest of the list (correlation to a pure C-SPAN ranking: 0.89, not 0.96). Coolidge and McKinley still edge into the bottom of the top 10 at this weight — a disclosed, arguable disagreement with C-SPAN's own ranking, not something we kept tuning the weight to paper over. Each president's page also shows how many of the 4 dimensions actually have a score for them (as few as 2, for a short-tenure or currently- serving president) — a score built from partial data is not shown with the same implied confidence as one built from all 4.

A closer look at Coolidge's own numbers turned up a real hole in Competence (executive-order activity rate), the dimension covering administrative execution: Coolidge and Harding have nearly identical EO-rates (~216/year each), yet C-SPAN's own historians rate their actual administrative skill 596 vs. 334 (of 1000) — almost as far apart as two presidents get. Across all 44 rated presidents, EO-rate correlates just 0.097 (p=0.53) with C-SPAN's “Administrative Skill” category — statistically no different from noise. Using C-SPAN's Administrative Skill score directly instead wasn't a clean fix either: it's one of the ten categories C-SPAN itself sums into the same Final Score already driving Historical Legacy at 35%, so folding it into a second dimension would push this platform's true historian-derived weight toward ~51%, undoing the exact over-reliance-on-C-SPAN problem the 50%→35% revision above was built to avoid. Competence is removed entirely (2026-07) — same standard as Independence/Follow-Through: no defensible live signal, no fabricated one in its place. Its 16.25% is split evenly across the three remaining mechanical dimensions (21.67% each); Coolidge drops from the top 10 to #12, Harding to #26, McKinley to #17, while Lincoln and Eisenhower both stay in the top 10 — the same qualitative target that justified 35% still holds.

A closer look at how that 35% actually gets applied found it wasn't the real operative number for most presidents. The renormalization used to spread flatly across whichever dimensions had data — so a president missing Agency Alignment (everyone before Clinton, ~36 of 47) had Historical Legacy's EFFECTIVE weight rise to ~44.7%, and the four non-elected successors missing both Agency Alignment and Public Mandate (Tyler, Fillmore, Arthur, Andrew Johnson) had it rise to ~61.8%. 35% was only the true weight for 4 of 47 presidents. This is fixed (2026-07): Historical Legacy is now held at exactly 35% whenever at least two mechanical dimensions are present, with the mechanical dimensions renormalizing only among themselves for the rest. Below that floor — a single mechanical dimension alone — it falls back to the old flat renormalization instead, since one number isn't reliable enough to carry 65% of a score by itself: Fillmore's Effectiveness is 100/100 purely from a Gold-Rush-era GDP boom he had little to do with, which would have swapped his real (near-bottom, 19/100) historian rating for a top-10 placement under a flat 65% share. Re-checked against the real dataset under this corrected scheme: 35% still keeps Lincoln and Eisenhower in the top 10 and Coolidge/Harding/McKinley out of it, so the headline number didn't need to change — only how consistently it gets applied.

Public Mandate (21.67%)

Reflects approval trajectory and coalition retention. Gallup, this platform's original approval source, ended presidential approval tracking entirely in February 2026 after 88 years; approval data now comes from the American Presidency Project (presidency.ucsb.edu), which is still updated for the sitting president, aggregating AP-NORC/CNN-SSRS/Marist/Pew/Verasight. This covers every president from Truman onward — 70% average approval over the term, 30% the trend from term-start to term-end, both scored against real population statistics computed from every president's actual polling history. Presidents before Truman have no polling era at all, so their Public Mandate uses UCSB's historical election-margin data instead — the average margin of victory across their own election win(s), the pre-polling-era proxy. The five presidents who never won a presidential election in their own right have neither and show N/A for this dimension, not a fabricated number.

Effectiveness (21.67%)

Measures tangible economic outcomes: GDP growth (60%) and job creation (40%). GDP growth is computed for the full presidency — BEA/FRED for the modern era, MeasuringWorth's real-GDP series (1790-present) for earlier presidents, both producing the same “average annual growth over the term” figure, with the term's first calendar year excluded when per-year data allows it (that year mostly reflects the outgoing administration's policy). Job creation comes from BLS nonfarm payroll data, which only exists from 1939 onward — presidents before that are scored on GDP growth alone, renormalized to 100% of the formula's weight, not defaulted on the missing component.

Agency Alignment (21.67%)

Measures how well executive agency actions align with stated presidential priorities, via Federal Register rulemaking data — the count of final and proposed rules published during the term, and what fraction were finalized rather than left pending. This is a digitization wall, not a conceptual one: notice-and-comment rulemaking was a real, functioning practice well before the 1990s, but no machine-readable record of it exists that far back — checked directly (2026-07) rather than assumed: federalregister.gov's API returns zero results for any pre-1994 president, and govinfo.gov's own structured Federal Register data starts at year 2000. Earlier issues exist only as scanned page images with no structured document-type or agency tagging, and reconstructing rulemaking counts from those would mean OCR'ing and classifying decades of raw scanned text — the same kind of unreliable pipeline already rejected for Follow-Through and Competence's court-success-rate. Every president before Clinton shows N/A for this dimension, excluded from their overall score entirely rather than scored on a proxy.

Historical Legacy (35%)

Covers what none of the other three dimensions can: crisis leadership, moral authority, vision, and similar historical-consequence judgments that don't reduce to GDP growth, approval polling, or rulemaking volume. Sourced from C-SPAN's Presidential Historians Survey — ~142 professional historians in the 2021 cycle (the most recent; the 2025 cycle was explicitly postponed by C-SPAN, citing the risk of turning “historical analysis” into “punditry” with a former president returning to office), scored across ten categories and aggregated into one point total. This is categorically different from the hand-set Independence/Follow- Through values removed elsewhere: a real, external, periodically-run survey with a documented methodology, not a single number invented for this platform — the same “trust a well-documented external institution” category as citing BLS or Federal Register data, just survey-based rather than administrative-record-based. Only rates presidents whose terms were complete as of the 2021 cycle — every currently-serving or just-departed president shows N/A here, genuinely unrated by the survey's own cadence, not a fetch gap.

Being real, external, and methodologically documented does not make this survey unbiased, and we don't present it as neutral ground truth. Political scientists who study these historian-ranking surveys have documented real, specific patterns in them: professional historians as a field skew toward favoring presidents who expanded federal/executive power, which plausibly inflates FDR, Wilson, and LBJ relative to how a more ideologically mixed panel might rate them; and historians are reluctant to rank very recent presidents at all until enough distance has passed to assess their legacy, which is the direct reason Obama and George W. Bush's scores may still be unsettled and Biden and the current president have none. Weighting this survey at 35% means this platform's ranking inherits those biases at roughly that same strength, not zero.

SUPREME COURT JUSTICE SCORECARDS

Justices are scored on impartiality and ideological consistency using case-level voting data from the Oyez Project and official Supreme Court records. Case opinions link directly to the official supremecourt.gov slip opinion PDFs.

Justice scoring evaluates whether a justice applies consistent legal principles across cases or shifts positions based on the political valence of the parties involved. This is analogous to the independence metric used for senators but adapted to the judicial context where party loyalty is replaced by jurisprudential consistency.

ACTION CENTER

The Action Center surfaces the most important civic issues of the day using automated news analysis. It is designed to inform, not persuade — every summary is non-partisan and presents facts without editorial framing.

NEWS ANALYSIS PIPELINE

Eight RSS feeds from seven newsrooms — AP News, NPR (Politics and World desks, counted as one source), PBS NewsHour, BBC World, The Hill, Politico, and Roll Call — are parsed hourly; opinion and editorial sections are filtered out of every feed. Under common media-bias ratings this mix spans center to lean-left, with no right-of-center outlet currently included — a disclosed limitation of the source diet, not a neutral sample of all coverage. Each article is filtered for U.S. policy relevance using embedding cosine similarity against policy area prototypes — the same sentence-transformer model used throughout the platform. Articles that pass the relevance threshold are clustered by semantic similarity to group coverage of the same story across sources.

TRENDING TOPIC INTEGRATION

Clusters are ranked using a weighted combination of civic actionability (40%) — whether officials are named and how closely the story resembles the ingested corpus of civic documents — coverage breadth (35%), how many independent newsrooms cover the story, and trending relevance (25%), whether the topic aligns with what the public is actively discussing. Actionability leads because it is what makes an issue something a citizen can act on; trending is weighted least because it is the most volatile of the three signals. Trending signals are drawn from Google Trends and policy-relevant Reddit communities, cross-referenced with the news clusters via embedding similarity.

NON-PARTISAN SUMMARIZATION

The top-ranked issues are summarized by the LLM with explicit instructions to present objective facts, avoid opinion or editorial framing, and recommend actions that do not assume which side of an issue the reader supports. Recommended actions include contacting representatives, attending public hearings, and reviewing primary source documents — not advocating for or against any policy position.

CROSS-REFERENCING

When a ranked politician is involved in a trending issue, the Action Center links directly to their scorecard. Related government documents from the Explore database are matched using semantic search. Source articles include direct links to the original reporting.

GOVERNMENT ACTIVITY TABS

Dedicated tabs for all three branches of government — Legislative (Senate and House), Executive, and Judicial — display the most recent government documents: floor speeches, executive orders, proposed rules, court opinions, and notices, pulled directly from the Explore database.

NATIONAL MONITORS

When an issue persists in the news across multiple days, the system automatically creates a National Monitor — a dedicated tracking page for that ongoing concern. Monitors build a sourced timeline of developments, detect when separate news stories are facets of the same underlying event using embedding similarity, and merge duplicate monitors automatically. Monitors transition to "watching" status when coverage subsides and reactivate when new developments appear.

YEAR-IN-REVIEW TIMELINE

Each day's top issue is permanently recorded in a timeline that accumulates throughout the calendar year. The Timeline tab provides a month-by-month chronological view of what mattered most, with top policy themes calculated for each month and the year as a whole. At year's end, this becomes a complete "Year in Review" of the issues that shaped civic life.

ELECTIONS TAB

The Elections tab displays upcoming election dates, Senate races with incumbent scores linked to their scorecards, and an interactive U.S. map for selecting states. State-specific information helps users understand their local races in the context of national trends.

INTERACTIVE GLOBAL NEWS MAP

The World tab features a 3D interactive globe that visualizes U.S.-related international news coverage. Countries mentioned in current news feeds are highlighted with points scaled by article count. Clicking a country scrolls to recent headlines about U.S. relations with that nation, linking to the original source articles.

CONTENT-BASED PARTY ALIGNMENT

A bill's partisan alignment is determined by analyzing what the bill does, not how senators voted on it. This is a deliberate architectural decision grounded in political science methodology.

The standard approach in political science — roll-call-based ideology estimation (DW-NOMINATE) [20] Poole & Rosenthal 1985— assumes sincere voting. But as Clinton, Jackman & Rivers (2004) note, this assumption is routinely violated by logrolling (vote trading), whip pressure, omnibus packaging, and tactical compromises. [21] Clinton, Jackman & Rivers 2004A senator might vote for a bill they ideologically oppose to secure support for a different bill, or because party leadership made it a litmus test.

HOW IT WORKS

We implement a nearest-centroid classifier (Rocchio 1971) [22] Manning, Raghavan & Schütze 2008in sentence-embedding space. Each party's known platform positions on each policy area (taxes, healthcare, environment, etc.) are embedded as centroids using the same sentence-transformer model used throughout the pipeline. Bill text is then embedded and compared to both party centroids via cosine similarity.

The stance direction(pro/anti) disambiguates cases where both parties have positions on the same topic: a "pro" environment bill (strengthen EPA enforcement) aligns with the Democratic platform, while an "anti" environment bill (roll back regulations) aligns with the Republican platform. This encodes the saliency-plus-direction model from manifesto research. [23] Laver & Garry 2000

TWO-SIGNAL FUSION

Content analysis is the primary signal for party alignment. Vote tallies from roll-call data serve as a secondaryrefinement. When both agree, confidence is high. When they disagree, content wins unless the vote data shows a clear party-line split (which is itself informative — the bill was important enough to whip). This follows Snyder & Groseclose (2000) who demonstrated that vote outcomes reflect party discipline as much as ideology. [24] Snyder & Groseclose 2000

ADAPTIVE LEARNING

Platform position descriptions are seed prototypes bootstrapped from published party platforms. As the pipeline processes bills, sponsor party data from Congress.gov serves as supervised ground truth — bills sponsored by a single party are labeled examples that refine the classifier over time. This follows the self-training paradigm. [25] Yarowsky 1995

CLASSIFICATION AND NLP PIPELINE

The pipeline classifies thousands of entities (bills, donors, industries, votes) per run. We use a tiered strategy that reserves expensive techniques for cases where cheaper methods fail, following the computational parsimony principle. [12] Jurafsky & Martin 2023

BILL POLICY AREA CLASSIFICATION

Bills and votes are classified into 18 policy areas (healthcare, defense, energy, etc., plus a procedural catch-all for non-substantive motions) using a tiered adaptive strategy:

  • 1.Learning store exact match — bills classified in prior pipeline runs are recalled instantly by ID. This is the experience replay pattern. [10] Lin 1992
  • 2.kNN against reference corpus — the k=7 most similar previously-classified bills in the vector store are retrieved and the policy area is assigned by similarity-weighted majority vote. [9] Cover & Hart 1967This is retrieval-augmented classification: the reference corpus grows with each pipeline run, improving accuracy over time. [26] Lewis et al. 2020
  • 3.Embedding similarity against policy descriptions — cosine similarity between bill text embeddings and pre-computed policy area description embeddings. This is the cold-start fallback using nearest-centroid classification. [7] Reimers & Gurevych 2019

The policy taxonomy is based on the Congressional Research Service (CRS) policy area scheme used by Congress.gov. The approach follows the text-as-data paradigm reviewed in Grimmer & Stewart (2013). [27] Grimmer & Stewart 2013Stance derivation (pro/anti/neutral) uses embedding cosine similarity against direction prototypes — the bill text is compared to semantic signatures of supportive, restrictive, and reform-oriented legislative language, following the Comparative Agendas Project coding tradition. [28] Baumgartner & Jones 1993Zero LLM calls are used for bill classification.

DONOR AND INDUSTRY CLASSIFICATION

Donor classification uses a five-tier strategy:

  • 1.FEC metadata — structured fields from the Federal Election Commission API encode committee type and designation codes, providing ground-truth classification for PACs vs. individual donors.
  • 2.Semantic detection — embedding cosine similarity against category prototypes replaces ~200 lines of hardcoded string patterns. This generalizes to unseen entities because distributed representations capture semantic meaning. [29] Bengio et al. 2003
  • 3.Learning store lookup — previously classified entities are recalled instantly by name.
  • 4.Embedding cosine similarity — donor names are compared against pre-computed industry description embeddings. Industry descriptions include exemplar company names as anchoring tokens, following the zero-shot classification setup. [30] Yin, Hay & Roth 2019
  • 5.k-Nearest Neighbor (kNN) — remaining unclassified donors are classified by the k=7 most similar already-labeled entities using distance-weighted majority voting. [9] Cover & Hart 1967This mirrors prototypical networks for few-shot learning [14] Snell et al. 2017where classification is performed by comparing query embeddings to accumulated real examples.

The kNN approach was chosen over LLM-based classification after empirical testing showed the LLM hallucinated invalid categories (producing labels like "SPORTS" or "RESTAURANT" outside the valid taxonomy) and was orders of magnitude slower. The kNN classifier processes ~5,000 donors in under 5 seconds versus 40+ minutes for the LLM, with more consistent results.

LEARNING STORE AND ADAPTIVE CLASSIFICATION

All classifications are persisted in a learning store (SQLite table) that functions as an evolving knowledge base. On subsequent pipeline runs, previously classified entities are retrieved instantly without recomputation. This is analogous to experience replay in reinforcement learning [10] Lin 1992 — past decisions inform future ones, improving both speed and accuracy over time.

The learning store also feeds into the self-training loop [25] Yarowsky 1995 — high-confidence classifications from prior runs become labeled examples for kNN and reference corpus retrieval in future runs. The system literally gets better each time the pipeline runs.

To prevent stale data from persisting when analysis algorithms are updated, the pipeline implements version-aware artifact management. At the start of each run, a SHA-256 fingerprint of all analysis source files is compared to the stored hash from the previous run. If the code is unchanged, all learning data is preserved to promote self-training. If the code has changed, stale artifacts (LLM results, learned classifications, kNN reference corpus) are automatically cleared so updated algorithms start fresh. The API cache (raw data from government APIs) is never cleared.

SEMANTIC SEARCH (EXPLORE)

The Explore feature uses dense passage retrieval [11] Karpukhin et al. 2020 to enable free-text search over government documents — Senate and House floor speeches, presidential actions (executive orders, proclamations, memoranda), Supreme Court opinions, and Federal Register rulemaking documents. Bill text is notindexed here; it is used separately, title-only, for the kNN bill-classification step in the scoring pipeline. Each document gets a single embedding (no chunking) over its title, summary, and first 800 characters of body, encoded with Snowflake all-MiniLM-L6-v2 and stored in sqlite-vec for nearest-neighbour retrieval. This outperforms keyword search (BM25) for conceptual queries like "climate policy" where exact term overlap is low.

HOW AI IS USED

Civitas uses two types of AI models, each for the task it is best suited for. AI is never used to generate scores directly — all scores are computed by deterministic, auditable formulas.

EMBEDDING MODELS (CLASSIFICATION + SEARCH)

Two sentence-transformers, both 384-dimensional and both around 22M parameters. Snowflake Arctic-XS handles classification: bill policy areas, donor industries, party alignment, motion types, and the k-nearest-neighbour reference corpus. all-MiniLM-L6-v2 [8] Wang et al. 2020handles the semantic-search index and the Action Center's similarity gates. Sentence-transformers produce dense vector representations where cosine similarity correlates with semantic similarity [7] Reimers & Gurevych 2019 — making them ideal for classification-by-comparison tasks where category definitions exist.

The split is measured, not accidental. Arctic is retrieval-asymmetric: it packs same-register text into a narrow raw-cosine band (~0.55-0.87), which left several similarity thresholds unable to separate genuine matches from noise. Against this platform's own live failure cases, all-MiniLM-L6-v2 measured roughly 4x the separation margin on document anchoring and 3x on policy relevance. Classification stays on Arctic because its thresholds were calibrated against that model's geometry — moving a classification gate to a different embedding space without re-measuring the threshold is how thresholds quietly stop meaning anything.

LLM (NARRATIVE SYNTHESIS)

LFM2.5-1.2B-Instruct via llama.cpp [16] Gerganov 2023 handles tasks requiring natural language understanding and multi-step reasoning:

Campaign promise extractionParses platform text from senator websites to identify specific policy commitments and assess whether votes support or contradict them. House positions never touch the LLM: they are derived from sponsored bills and evaluated with deterministic embedding similarity only.
Voting pattern narrativeGenerates human-readable summaries of a senator's voting patterns across policy areas
Key vote reasoningExplains why specific votes were flagged as significant given a senator's donor profile and party dynamics
PAC identificationIdentifies the parent organization and industry behind opaque PAC names using world knowledge
Explore summariesOn-demand summaries of how a government document relates to a user's search query

WHAT AI DOES NOT DO

Score calculationAll sub-scores use deterministic formulas with no LLM input. The math is fully auditable.
Bill classificationPolicy areas, party alignment, and stance are all embedding-based — no LLM in the loop.
Donor classificationFEC metadata + embeddings + kNN handle all donor and industry classification.
Data fabricationThe LLM only analyzes data already fetched from official APIs. It does not generate or invent facts.
Partisan framingPrompts are explicitly structured to avoid editorial framing. The LLM analyzes behavior, not ideology.

WHY THESE TECHNIQUES WERE CHOSEN

We follow a strict hierarchy: structured metadata first, then embedding similarity, then kNN, then LLM — reserving each more expensive technique only for tasks the cheaper ones cannot handle. The pipeline contains zero hardcoded keyword lists, regex patterns, or string-matching heuristics for classification decisions. Every classification is made mathematically via embedding cosine similarity against natural-language prototypes. [12] Jurafsky & Martin 2023

Embeddings (not LLM) for classification: sentence embeddings excel at text classification tasks when labeled examples or category descriptions exist. They are deterministic, fast, and avoid the hallucination risks inherent in generative models. [13] Minaee et al. 2021The kNN classifier further leverages accumulated labeled data as a growing reference set — a well-established approach in few-shot and semi-supervised learning settings. [14] Snell et al. 2017

Content analysis (not votes) for party alignment: roll-call votes confound ideology with legislative strategy. Analyzing what a bill does relative to published party platforms recovers ideological alignment more accurately, following the manifesto analysis tradition. [31] Laver, Benoit & Garry 2003

LLM for narrative synthesis: tasks like promise-vote cross-referencing and PAC identification require world knowledge and multi-step reasoning that embeddings alone cannot provide. These are inherently generative tasks suited to language models. [15] Wei et al. 2022

MODEL AND ARCHITECTURE

The inference model is LFM2.5-1.2B-Instruct, a compact open-weight language model running natively via llama.cpp [16] Gerganov 2023 compiled with ARM-specific optimizations (cortex-a76, dot-product, fp16). This provides faster inference compared to containerized runtimes, generating ~14 tokens/second on the Raspberry Pi 5 CPU. Results are cached in a local database so each unique analysis is computed at most once.

The embedding models are Snowflake Arctic-XS for classification (bills, donors, industries, party alignment) and all-MiniLM-L6-v2 [8] Wang et al. 2020 for the search index and similarity gates — both 384-dimensional, both in the 22M-parameter class, so carrying two costs little. Vectors are stored in sqlite-vec, a single-file SQLite extension that runs on the Pi with no separate vector server. Every model runs entirely on-device with no external API calls.

DATA SOURCES AND APIs

All data is sourced from official US government APIs and public records. No data is purchased, scraped from paywalled sources, or fabricated.

CONGRESSIONAL DATA (SENATE & HOUSE)

Congress.gov APIBill text, voting records, member data, sponsored legislation, and bill sponsor party affiliation — both chambers
FEC API (fec.gov)Campaign finance data: individual contributions, PAC donations, committee filings, disbursements, and committee type codes — both chambers
GovInfo APIFull bill text for policy area classification, Congressional Record floor proceedings for advocacy analysis — both chambers
Senate.govOfficial senator websites scraped for platform text and campaign promises, roll-call vote records with per-member votes — Senate only; House campaign promises are instead derived from sponsored legislation (see AI Usage above), since House platform text isn't available the same way

PRESIDENTIAL DATA

Federal Register APIExecutive order counts and metadata from federalregister.gov (Clinton onward, no API key required)
BLS APIBureau of Labor Statistics public API — total nonfarm employment payrolls for jobs-created calculations
C-SPAN Historians SurveyPresidential Historians Survey (2021 cycle) — the basis for the Historical Legacy dimension
American Presidency Project (UCSB)Presidential roster and identity data, approval polling for modern presidents (Truman onward), and pre-polling-era election margins. Replaced Gallup, which ended presidential approval tracking in February 2026
BEA NIPA Tables / FREDBureau of Economic Analysis GDP growth data for the modern era
MeasuringWorthReal-GDP series (1790-present) for presidents predating BEA coverage

EXPLORE FEATURE

Congressional Record (GovInfo)Senate and House floor proceedings — speaker-attributed transcripts from daily CREC packages
Federal RegisterExecutive orders, presidential memoranda, and proclamations with full text and metadata, plus proposed and final rules — including ones still open for public comment, surfaced with their comment link and deadline
Oyez / supremecourt.govSupreme Court opinions, indexed alongside the legislative and executive documents
Semantic SearchDocuments embedded with all-MiniLM-L6-v2 into a sqlite-vec table for dense passage retrieval — one embedding per document, no chunking

SUPREME COURT DATA

Oyez Project APICase metadata, justice votes, oral argument transcripts, and decision breakdowns
supremecourt.govOfficial slip opinion PDFs linked directly from case records

RATE LIMITING

The pipeline respects all API rate limits: Congress.gov at 1.2 requests/second, FEC at 0.25 req/s, GovInfo at 1.0 req/s, and BLS at 25 queries/day. Data is cached for 72 hours to minimize redundant API calls.

ENVIRONMENTAL AND ETHICAL CONSIDERATIONS

LOCAL-FIRST ARCHITECTURE

The entire Civitas stack runs on a single Raspberry Pi 5 (16GB RAM) with an NVMe SSD. There are no cloud GPU instances, no third-party AI API calls, and no data sent to external services for processing. The LLM, embedding model, vector database, SQLite database, backend API, and frontend all run on the same device.

ENERGY FOOTPRINT

A Raspberry Pi 5 draws approximately 5-12 watts under load. Running the full data pipeline (100 senators, ~100 LLM calls) takes several hours but consumes roughly the energy of a single LED light bulb. By comparison, a typical cloud GPU instance (NVIDIA A100) draws 250-400 watts. [17] Patterson et al. 2021This project demonstrates that meaningful AI analysis does not require industrial-scale compute. The trade-off is speed: what a cloud GPU processes in minutes takes hours on a Pi. We consider that an acceptable trade for a nightly batch pipeline.

DATA PRIVACY

No accounts, no cookies, no third-party analytics trackers, and no advertising networks. The site does record anonymized visit counts on its own server, to understand usage — a salted hash of (IP, browser, date) that rotates daily so the same visitor is unrecoverable across days, plus per-page view counts. Raw IP addresses and user agents are never stored. This data is never shared, sold, or transmitted anywhere. All data displayed is derived exclusively from public government, academic, and economic-history records. The only outbound network requests are to official government APIs (congress.gov, fec.gov, api.bls.gov, federalregister.gov) and, for presidential data with no government API equivalent, UCSB's American Presidency Project (presidency.ucsb.edu) and MeasuringWorth's historical GDP dataset (measuringworth.com).

OPEN-WEIGHT MODEL

We deliberately chose LFM2.5, an open-weight model, over proprietary alternatives like GPT-4 or Claude. This means: no per-token API costs that could make the project financially unsustainable, no dependency on a third-party company's continued service, full auditability of the model's behavior, and no user queries or government data leaving the device.

LIMITATIONS AND HONESTY

This project has real limitations and we believe in stating them clearly:

  • -A 1.2B parameter model is less capable than larger models. It occasionally produces imprecise promise analysis. We mitigate this with caching, post-processing heuristics, and deterministic overrides where the model output can be verified against structured data.
  • -Presidential scoring used to include two dimensions, Independence and Follow-Through, that were a one-time hand-set number for every president with no live formula behind them at all. We removed both entirely (2026-07) rather than keep presenting a hand-set number as a computed score, and rebuilt every remaining dimension — plus each president's identity data — on real live and historical datasets, with no seeded or hand-typed fallback left anywhere in the pipeline. A third dimension, Competence, was later removed too: its only live component (executive-order activity rate) measured no relationship (Spearman 0.097) with real administrative-skill judgment. See the Presidents methodology below for the full account. Agency Alignment has no machine-readable rulemaking data before Clinton (a real digitization wall in the underlying government sources, checked directly rather than assumed) and shows N/A for earlier presidents rather than a proxy score.
  • -Correlation between donations and votes does not prove causation. A senator who receives PAC money and votes favorably may be doing so for policy reasons unrelated to the donation. We follow the methodological caution urged by Ansolabehere et al. (2003). [18] Ansolabehere et al. 2003
  • -Content-based party alignment depends on the quality of platform position descriptions. While these are seeded from published party platforms and refined by sponsor data, edge cases involving bipartisan or cross-cutting legislation may be misclassified.
  • -FEC data has inherent reporting delays. Campaign finance filings may lag real-time donations by weeks or months.
  • -Embedding-based classification, while fast and consistent, lacks the world knowledge that a large model or human expert would bring. Edge cases involving shell companies or deliberately obscure entity names may be misclassified.

TECHNICAL STACK

HardwareRaspberry Pi 5 (16GB), NVMe SSD
BackendPython 3.13, FastAPI, SQLAlchemy, SQLite
FrontendNext.js 16, React 19, TypeScript, Tailwind CSS
Embedding ModelsTwo, both 384-dim / ~22M params: Snowflake Arctic-XS for classification, all-MiniLM-L6-v2 for the search index and similarity gates
LLM Runtimellama.cpp (native ARM build), LFM2.5-1.2B-Instruct
Vector Databasesqlite-vec (vec0 virtual tables in a local SQLite file, cosine distance)
ContainersDocker Swarm (single node) — zero-downtime start-first rolling updates behind an in-stack nginx reverse proxy, with automatic rollback on a failed health check
Pipeline ScheduleNightly at 3:00 AM via APScheduler
Data Caching72-hour TTL with persistent SQLite cache
Learning StoreSQLite table for persistent classification memory, version-aware invalidation on code change
Pipeline OptimizationProducer-consumer threading: embedding prefetch overlaps LLM inference, context compression for prompts
API PaginationServer-side paginated voting records with filter support
Sponsorship AnalysisPageRank (leadership) + SVD (ideology) on cosponsorship matrix
ClassificationZero hardcoded rules — all classifications via embedding similarity or kNN
Metric TooltipsEvery scorecard metric has a [?] tooltip explaining what it measures
Branches CoveredSenate (100), House (435), Presidents (historical + modern), Supreme Court (9 justices)
Action CenterHourly news analysis with national monitors for ongoing concerns and year-in-review timeline tracking
News SourcesAP News, NPR (Politics + World), PBS NewsHour, BBC World, The Hill, Politico, Roll Call — opinion sections filtered; mix spans center to lean-left (no right-of-center outlet currently included)
Trending IntegrationGoogle Trends RSS + Reddit policy subreddits, cross-referenced via embedding similarity
Globe Visualizationreact-globe.gl — interactive 3D globe for international news mapping

HOW IT'S BUILT

Civitas is built and maintained by a single developer as a hobby project, running on a home server. It exists to prove that meaningful civic accountability tools do not require venture capital, a team of engineers, or enterprise cloud infrastructure. Anyone with the knowledge, time, and a modest machine can build something like this.

SERVERRaspberry Pi 5 — a $80 single-board computer
LOCAL LLMLFM2.5-1.2B-Instruct via llama.cpp (Ollama as fallback) · runs entirely on-device, zero API cost
DATABASESQLite · no cloud database, no managed service
DEPLOYMENTDocker Swarm rolling updates on a single machine
EXTERNAL APIsCongress.gov, FEC.gov, Federal Register — all free and open
MONTHLY COST~$5–10 (electricity)
CLOUD SERVICESNone
VENTURE CAPITALNone

The pipeline runs overnight, the site serves from a home IP address, and the entire codebase is documented above. If you want to build something similar, everything you need to know about the methodology is on this page.

REFERENCES

  1. [1]Bonica, A. (2014). Mapping the Ideological Marketplace. American Journal of Political Science, 58(2), 367-386. doi:10.1111/ajps.12062
  2. [2]Naurin, E. (2011). Election Promises, Party Behaviour and Voter Perceptions. Palgrave Macmillan. doi:10.1057/9780230304598
  3. [3]Martin, S. (2011). Using Parliamentary Questions to Measure Constituency Focus. Political Studies, 59(2), 472-488. doi:10.1111/j.1467-9248.2011.00885.x
  4. [4]Carson, J. L., Koger, G., Lebo, M. J., & Young, E. (2010). The Electoral Costs of Party Loyalty in Congress. American Journal of Political Science, 54(3), 598-616. doi:10.1111/j.1540-5907.2010.00449.x
  5. [5]Stratmann, T. (2005). Some Talk: Money in Politics. A (Partial) Review of the Literature. Public Choice, 124(1-2), 135-156. doi:10.1007/s11127-005-4750-3
  6. [6]Rhoades, S. A. (1993). The Herfindahl-Hirschman Index. Federal Reserve Bulletin, 79, 188-189.
  7. [7]Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of EMNLP-IJCNLP 2019, 3982-3992. doi:10.18653/v1/D19-1410
  8. [8]Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., & Zhou, M. (2020). MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. Proceedings of NeurIPS 2020. arXiv:2002.10957
  9. [9]Cover, T. & Hart, P. (1967). Nearest Neighbor Pattern Classification. IEEE Transactions on Information Theory, 13(1), 21-27. doi:10.1109/TIT.1967.1053964
  10. [10]Lin, L.-J. (1992). Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine Learning, 8(3-4), 293-321. doi:10.1007/BF00992699
  11. [11]Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of EMNLP 2020, 6769-6781. doi:10.18653/v1/2020.emnlp-main.550
  12. [12]Jurafsky, D. & Martin, J. H. (2023). Speech and Language Processing (3rd ed. draft). Stanford University.
  13. [13]Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., & Gao, J. (2021). Deep Learning-Based Text Classification: A Comprehensive Review. ACM Computing Surveys, 54(3), 1-40. doi:10.1145/3439726
  14. [14]Snell, J., Swersky, K., & Zemel, R. (2017). Prototypical Networks for Few-Shot Learning. Proceedings of NeurIPS 2017, 4077-4087. arXiv:1703.05175
  15. [15]Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Proceedings of NeurIPS 2022. arXiv:2201.11903
  16. [16]Gerganov, G. (2023). llama.cpp: Inference of LLaMA model in pure C/C++. GitHub. github.com/ggerganov/llama.cpp
  17. [17]Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., & Dean, J. (2021). Carbon Emissions and Large Neural Network Training. arXiv:2104.10350.
  18. [18]Ansolabehere, S., de Figueiredo, J. M., & Snyder, J. M. (2003). Why Is There So Little Money in U.S. Politics? Journal of Economic Perspectives, 17(1), 105-130. doi:10.1257/089533003321164976
  19. [19]Efron, B. & Morris, C. (1975). Data Analysis Using Stein's Estimator and Its Generalizations. Journal of the American Statistical Association, 70(350), 311-319. doi:10.2307/2285814
  20. [20]Poole, K. T. & Rosenthal, H. (1985). A Spatial Model for Legislative Roll Call Analysis. American Journal of Political Science, 29(2), 357-384. doi:10.2307/2111172
  21. [21]Clinton, J., Jackman, S., & Rivers, D. (2004). The Statistical Analysis of Roll Call Data. American Political Science Review, 98(2), 355-370. doi:10.1017/S0003055404001194
  22. [22]Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press. Ch. 14: Vector Space Classification.
  23. [23]Laver, M. & Garry, J. (2000). Estimating Policy Positions from Political Texts. American Journal of Political Science, 44(3), 619-634. doi:10.2307/2669268
  24. [24]Snyder, J. M. & Groseclose, T. (2000). Estimating Party Influence in Congressional Roll-Call Voting. American Journal of Political Science, 44(2), 193-211. doi:10.2307/2669305
  25. [25]Yarowsky, D. (1995). Unsupervised Word Sense Disambiguation Rivaling Supervised Methods. Proceedings of ACL 1995, 189-196. doi:10.3115/981658.981684
  26. [26]Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Proceedings of NeurIPS 2020. arXiv:2005.11401
  27. [27]Grimmer, J. & Stewart, B. M. (2013). Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts. Political Analysis, 21(3), 267-297. doi:10.1093/pan/mps028
  28. [28]Baumgartner, F. R. & Jones, B. D. (1993). Agendas and Instability in American Politics. University of Chicago Press.
  29. [29]Bengio, Y., Ducharme, R., Vincent, P., & Jauvin, C. (2003). A Neural Probabilistic Language Model. Journal of Machine Learning Research, 3, 1137-1155.
  30. [30]Yin, W., Hay, J., & Roth, D. (2019). Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment Approach. Proceedings of EMNLP 2019, 3914-3923. doi:10.18653/v1/D19-1404
  31. [31]Laver, M., Benoit, K., & Garry, J. (2003). Extracting Policy Positions from Political Texts Using Words as Data. American Political Science Review, 97(2), 311-331. doi:10.1017/S0003055403000698
  32. [32]Brin, S. & Page, L. (1998). The Anatomy of a Large-Scale Hypertextual Web Search Engine. Proceedings of the 7th International World Wide Web Conference, 107-117.
  33. [33]Tauberer, J. (2012). Open Government Data: The Book. GovTrack.us methodology for ideology and leadership scoring via cosponsorship analysis. govtrack.us/about/analysis
  34. [34]Volden, C. & Wiseman, A. E. (2014). Legislative Effectiveness in the United States Congress: The Lawmakers. Cambridge University Press.

Questions about our methodology? Disagree with a score? We welcome scrutiny. This project is built on the belief that transparency is non-negotiable.