Five public sources, pulled nightly: STOCK Act trades, bills, contract awards, lobbying filings, market prices.
Each trade is joined against the bills, contracts, and lobbying inside the member’s committee jurisdiction.
Correlations roll up to one record per member, ticker, and trade, weighted into a 0 to 100 composite.
Two ML heads add context: one distills the rule score into SHAP factors, one asks whether the trade beat the market.
Nightly site refresh, a weekly CC0 data drop, and an editorial brief when a correlation is worth explaining.
Methodology
Capitol_Signals is a correlation engine for congressional financial activity. Each week it ingests:
- Stock trade disclosures filed under the STOCK Act (via QuiverQuant)
- Federal contract records (USAspending.gov)
- Congressional bills and committee assignments (Congress.gov + unitedstates/congress-legislators)
- Lobbying disclosures filed under the Lobbying Disclosure Act (lda.gov), used as sector context only
- Market price data (yfinance)
The engine joins these across indices to produce correlations (trade to contract and trade to bill) and aggregates them into rollups (one rollup per member-ticker-trade) with a 0-100 composite correlation score. A rollup’s composite is the score of its single strongest pair.
The composite score
The score weights six factors defined in correlation/config.py. The weights are published here so a reader can reproduce any score from the factor values in the data sidecar.
| Factor | Weight | What it measures |
|---|---|---|
Sector alignment (jurisdiction) | 0.2632 | Whether the trade’s sector or industry matches the event’s, and whether a committee the member sits on has jurisdiction over that sector. 1.0 = industry match with committee jurisdiction; 0.7 = sector match with committee jurisdiction; 0.5 = industry match without committee jurisdiction, or the member sits on the committee the event went through with no sector match; 0.35 = sector match with no committee jurisdiction; 0 = no link. Sector names are compared across the yfinance and GICS vocabularies. |
Event timing (timing) | 0.2632 | Exponential decay on the number of days between the trade and the event (half-life 10 days, floor 0.05). |
Pre-event positioning (trade_before_event) | 0.1577 | 0 when the trade follows the event. When the trade is on or before the event date: a purchase scores 1.0; a sale ahead of a contract award to the company scores 0 (it is the opposite of positioning for the award); a sale ahead of a bill scores 0.5, because the bill’s direction for that ticker is not known; an exchange or a trade of unknown direction scores 0.5. Before 2026-08-28 every pre-event trade scored 1.0 regardless of direction, and half of them were sales. |
Trade size (amount) | 0.1053 | STOCK Act disclosure band, from 0.1 for $1,001 to $15,000 up to 1.0 for the largest band. |
Disclosure lag (disclosure_lag) | 0.1053 | Days between transaction and disclosure relative to the 45-day statutory window; unknown lag scores 0.5. |
Multiple signals (multiple_signals) | 0.1053 | Number of distinct events linked to the trade by a sector or committee match (events with no link do not count), on a logarithmic scale that reaches 1.0 at 20 events: 1 event = 0.23, 3 = 0.46, 6 = 0.64, 10 = 0.79. The rollup records the count (distinct_signal_count) and whether the retrieval cap of 100 bills or 200 contracts was reached (signal_query_capped), in which case the count is a floor. Before 2026-08-28 the factor counted raw pairs, junk included, and saturated at three, so it was 1.0 on 99.4% of pairs. |
Composite = sum of (factor value x weight) x 100, clamped to 0-100, computed per trade-event pair. A trade’s headline score is the mean of its three strongest pairs among those with a sector or committee link (fewer if fewer exist); a trade whose matches have no such link scores 0. The best single pair (max_pair_score), the average over all pairs (avg_pair_score), and the number of linked pairs (linked_pair_count) are stored beside it. Before 2026-08-28 the headline was the single best pair, so one anomalous match set the public number. Tiers: 70 and above is reported as high (Critical on the site), 60 to 69 as medium (Elevated), below 60 as low. These thresholds were re-baselined on 2026-08-28 against the engine 1.3.0 distribution of 1,000 scored trades (median 42, 90th percentile 60, 98th percentile 68, maximum 81): high is the top 1.5% of scored trades and medium the top 11%. Under the previous 80/60 thresholds the medium tier covered 56% of scored trades. A brief is generated only for a rollup at 60 or above. Lobbying intensity is not a weighted factor, and abnormal price movement was removed from the composite in v1.2 to eliminate data leakage; realized abnormal returns are reported separately.
Candidate events are collected in a window that lets the trade precede the event by up to 90 days and follow it by at most 30 days (event dates from 30 days before to 90 days after the trade date). Before 2026-08-28 the window was applied the other way round, so 74% of scored pairs were trades that occurred after the event. The timing factor decays with the number of days between the two in either direction.
Point values quoted in a brief (for example “26.3 points” for timing) are the factor value times its weight. Because several factors saturate, identical point values recur across unrelated briefs; they are weights, not measurements of strength.
Machine-generated text
The “What the engine saw” paragraph in each brief is written by a local language model (qwen3:8b via Ollama) from the factor values above and is reviewed before publication. It is a summary of the engine’s own inputs, not independent research. Briefs published before 2026-08-27 carried model text that described the sector-alignment factor as “committee overlap” and cited unrelated lobbying filings as evidence; that text was withdrawn and replaced with a dated correction note on each affected page.
ML annotations
A second layer runs two XGBoost heads in annotation mode alongside the composite. Both see only the six factor values above; neither sees the ticker, the date, prices, or market state.
- Composite-distillation head: trained to reproduce the composite score from the same six inputs. Its rank correlation with the composite is about 0.94 on published data. It is a re-derivation of the rule, not an independent opinion, and its SHAP attributions describe which rule factors drove the composite. It cannot corroborate the composite.
- Forward-CAR head: trained to predict 30-day cumulative abnormal return versus SPY from the same six factors. On the project’s own published snapshot its rank correlation with realized 30-day abnormal return is +0.32 on trades it was trained on and +0.04 (not statistically different from zero) on trades it was not. Treat
ml_score_caras an unvalidated experiment, not as evidence that the market rewarded a trade. The realized return is theabnormal_return_30dfield.
Both heads use purged k-fold cross-validation at training time. The production model artifact is dated 2026-05-20 and has not been retrained since; scores are normalized to the 5th to 95th percentile of that artifact’s in-sample predictions and clip at 0 and 100.
ml_agreement enum. Each rollup carries a label summarizing how the two heads relate. diverged is the most common label (about half of scored rollups) and carries no measured relationship to realized returns.
| Value | Meaning |
|---|---|
high-confidence | Both model heads score the correlation high. This label describes model agreement, not factual confidence in misconduct. |
agreed | Heads concur within tolerance. |
diverged | The distillation and forward-CAR model estimates disagree. |
partial | Only one head produced a score (legacy artifact without the second head). |
unscored | Rollup predates dual-head deployment or lacked sufficient features. |
Validation to date. A backtest runs every Saturday against realized returns and stores its result in the engine. As of 2026-08-27 every weekly run since 2026-06-01 (13 runs) has graded the composite “No signal”: the rank correlation between composite score and 30-day abnormal return is between -0.01 and +0.05 with no p-value below 0.2. The composite is a descriptive index of timing and sector alignment in public records. It has not been shown to predict returns and is not presented as a trading signal.
Member signal score. Member pages show a separate 0-100 “Signal score” built from a member’s whole trading record, with weights in analytics/config.py: average correlation score 0.35, share of trades in a committee’s jurisdiction 0.20, average disclosure lag 0.10, count of high-severity correlations 0.10, maximum correlation score 0.10, ML score 0.15. A seventh factor, sector-shift frequency, is held at weight 0 until its detector is calibrated. This score is published in the weekly snapshot as signal_score.
What correlation does and does not mean. A high composite score means several public signals align in time and topic. It does not prove the member used non-public information, breached fiduciary duty, or violated the STOCK Act. Trades are attributed to the filing member; STOCK Act filings can cover spouse and dependent accounts, and the source feed does not distinguish them.
Source code, data sources, and licensing
- Pipeline source: the repository is private while the 2026-08 audit remediation is completed; the methodology, weights, and every input to a published score are on this page and in the data files. Requests for source access go to the editorial address below.
- STOCK Act trade data: Quiver Quantitative Hobbyist API. Used under QuiverQuant’s terms; Capitol Signals does not re-license the underlying trade records.
- Bills + member data: Congress.gov API (public domain)
- Contracts: USAspending.gov (public domain)
- Lobbying: LDA API (public domain)
- Market data: yfinance (Yahoo Finance, terms apply; derived return figures are published, price series are not)
License. Capitol Signals’ own outputs (composite and ML scores, correlation records, member signal scores, editorial text, and this methodology) are dedicated to the public domain under CC0 1.0. That dedication covers only what Capitol Signals created. Records reproduced from upstream sources keep their original terms: Congress.gov and USAspending records are public domain; QuiverQuant and Yahoo Finance data are used under their respective terms and are not re-licensed by this dedication.
Capitol_Signals is built and maintained by Brian Chaplow. No subscription, no copy-trading, no affiliate links.
Corrections and retractions
Send corrections to [email protected]. A factual error in a published brief is corrected in place with a dated note in the brief body; a brief whose headline claim does not survive correction is retracted with the note left in place of the body. Withdrawn machine-generated text is replaced by a dated note rather than silently removed. This page records material changes to the scoring method with their dates.
- 2026-08-27: sector alignment now requires a member committee with jurisdiction for values above 0.5; sector names compared across vocabularies; bill referral-committee sector lists no longer count as a sector match. Machine-generated narratives withdrawn from the five briefs published before this date. The new rule was applied to every historical pair the same day, and pairs the nightly engine had stopped producing but never deleted (188) were removed.
- 2026-08-28 (engine 1.3.0): the candidate-event window now runs from 30 days before to 90 days after the trade (it was the reverse); the pre-event bonus depends on trade direction and event type; multiple signals counts distinct linked events on a log scale instead of raw pairs saturating at three; the headline score is the mean of the three strongest linked pairs instead of the single best pair. Every historical pair was re-matched and rescored under these rules (1,047 eligible trades, 19,185 pairs, 1,000 scored trades). Tier thresholds re-baselined: high 70 (was 80), medium 60 (unchanged), and the site now uses the same two numbers for Critical and Elevated (it used 80 and 70).
Per-brief data sidecar schema
Every editorial brief at /briefs/{slug}/ ships a co-located data/ directory with the trades, correlations, and scores it cites, generated when the brief was drafted. Files:
| File | Columns | Notes |
|---|---|---|
| trades.csv / trades.jsonl | trade_id, bioguide_id, member_name, ticker, transaction_date, amount_range, amount_min, amount_max, sector, industry | The headline trade plus any sibling cited |
| correlations.csv / correlations.jsonl | correlation_id, trade_id, bioguide_id, ticker, transaction_date, composite_score, severity_tier, scored_at | Supporting rollups for same member + ticker |
| scores.csv / scores.jsonl | correlation_id, trade_id, composite_score, ml_score_distill, ml_score_car, ml_agreement, ml_shap_top3_distill, ml_shap_top3_car, ml_shap_top3 (alias of the distill column), abnormal_return_30d, abnormal_return_60d, abnormal_return_90d | Composite + ML head outputs; abnormal returns are versus SPY |
Reproducibility: every file is keyed by the rollup’s correlation_id. Scores are rescored nightly, so a sidecar captured at draft time can differ from the live value by the time the brief is read; the sidecar is the record of what the brief was written against.
Weekly snapshot schema
Every Monday at 02:30 ET (06:30 UTC) the engine writes a full snapshot to /data/{YYYY}-W{WW}/. Row counts are in each snapshot’s README. The snapshot contains:
| Table | Source index | Why included |
|---|---|---|
| trades | cs-trades | Primary signal |
| correlations | cs-correlations (rollups only) | Rollup correlations |
| member-profiles | cs-member-profiles | Aggregated member-level metrics, including signal_score |
File formats per table: {table}.csv, {table}.jsonl, {table}.xlsx, plus trades.sqlite.
Snapshots 2026-W28 through 2026-W35 shipped a trades table truncated to 1,000 rows by a publisher fault; the fault was fixed on 2026-08-27 and later snapshots carry the full table.
What we exclude
The weekly snapshot does NOT include:
- cs-bills: reference data, not signal data; committing it weekly would balloon the site repository.
- cs-lobbying: same scope-trim reasoning. Lobbying filings are sector context only and are not linked to individual members or trades.
Reproducibility
Each weekly snapshot directory is committed to the site repository under public/data/{YYYY}-W{WW}/. The directory contents are deterministic for a given engine state: columns are sorted, rows are sorted by primary key, and floats are rounded to 4 decimals before serialization.
To reproduce a specific brief’s claims:
- Find the brief at
/briefs/{slug}/. - Read the sidecar at
/briefs/{slug}/data/. - Cross-reference with the weekly snapshot at
/data/{iso_year_week}/and any later snapshot, since scores are rescored nightly. - Recompute the composite from the factor values with the weight table above.
Member opt-out posture
Capitol Signals correlates public data. Member trades are disclosed under the STOCK Act; contracts are public via USAspending; bills are public via Congress.gov; lobbying is public via lda.gov. The project does not invite member opt-out requests because the underlying disclosures are mandated by law. A current member of Congress who identifies a factual error in a brief should use the corrections address above; corrections are noted in the brief body.
Slug and iso_year_week convention
Brief slugs follow the pattern {YYYY-WW}-{lastname_lower}-{ticker_lower}, where iso_year_week refers to the ISO week the underlying trade occurred (transaction_date), NOT the brief publication week. For example, the Moskowitz GD brief at /briefs/2026-W13-moskowitz-gd/ derives W13 from the underlying trade on 2026-03-23 (ISO week 13), even though the brief was published in 2026-W21. This keeps the slug pinned to the editorial subject so the same correlation never gets a different slug across republications, and so auto-generated slugs from generate_weekly_brief.py:emit_astro_brief are deterministic. The derivation order is transaction_date first, scored_at fallback.