How RiskDiff works
Every year a public company lists the risks it faces in the Risk Factors section of its annual report (10-K). RiskDiff lines up this year's list against last year's and shows exactly which sentences changed. The filing text is the only evidence: every quote is verbatim, and every AI statement points to the sentences it is based on.
How it works
Step 1
Pull both filings from SEC EDGAR
The company's two most recent annual reports (10-K), straight from the SEC. Amendments are excluded.
✦ Step 2
Match each risk factor to last year's version
Each risk is paired with its counterpart from last year, even when it moved or was renamed. Unpaired risks are new or dropped.
Step 3
Find the exact sentences that changed
Sentence by sentence and word by word: what was added, removed, or rewritten.
✦ Step 4
Summarize, with every statement cited
A short plain-English summary of each change. Every point links to the filing sentences behind it.
What changed is always computed by fixed rules in code, so the same two filings always give the same result. AI helps in three narrow places only: recognizing that a reworded risk is the same risk as last year's (step 2), wording short summaries (step 4), and answering four fixed yes/no questions about each change. A “yes” adds a fixed, published number of points only when it cites the filing sentences behind it, and never when a rule in code already scored the same fact. Any summary statement that does not cite valid filing sentences is removed before it is shown. Each change also gets a wording change level (High 60+, Medium 25–59, Low below 25, on a 0 to 100 scale): how much the language changed, not how risky the company is.
What RiskDiff does not do
- No predictions. It reports what the filings say. It never forecasts events or outcomes.
- No risk ratings. A change in disclosure language is not a change in actual risk. The wording change level measures wording only.
- No investment advice. It never suggests acting on a disclosure change.
Technical detailThe full pipeline, the published scoring weights, thresholds and caps, blocked wording, and the evaluation status. RiskDiff reads Item 1A (Risk Factors) of each 10-K.
The pipeline (nine stages)
Nine stages. Seven are deterministic code; two use models, and neither decides what changed.
- Stage 0Fetch filings from EDGARDeterministic
- Stage 1Extract Risk FactorsDeterministic
- Stage 2Segment risk factorsDeterministic
- ✦ Stage 3Embed risk blocksAI-assisted
- Stage 4Match risks across yearsDeterministic
- Stage 5Sentence-level diffDeterministic
- Stage 6Deterministic signalsDeterministic
- ✦ Stage 7AI rubric + interpretationAI-assisted
- Stage 8Score + publishDeterministic
0 Fetch filings from EDGAR
Looks up the company on SEC EDGAR and downloads its two most recent original 10-K filings (amendments excluded). Every request identifies itself and stays within SEC fair-access limits.
1 Extract Risk Factors
Finds the Risk Factors section in each filing, skipping the table of contents and stopping where the next section begins.
2 Segment risk factors
Splits the Risk Factors section into risk blocks (a heading plus its body), then into sentences with stable ids and character offsets back into the filing text.
3 Embed risk blocks · ✦ AI-assisted
A local embedding model turns each risk block into a vector so that equivalent risks can be recognized even when the wording moved.
4 Match risks across years
Pairs old and new risk blocks with an optimal assignment over heading and body similarity. Pairs below the published threshold become new or removed risks.
5 Sentence-level diff
Aligns sentences inside each matched pair, computes word-level insertions and deletions, and classifies the change (new, removed, expanded, reduced, reworded, largely unchanged).
6 Deterministic signals
Rule-based detectors with citations: hedged wording that became actual wording, newly named entities, new quantities, and the share of added or removed text.
7 AI rubric + interpretation · ✦ AI-assisted
A language model answers narrow yes/no rubric questions and writes short interpretations. A citation validator drops every claim that does not cite valid filing sentences.
8 Score + publish
Adds up published rubric weights into a 0 to 100 change score with an itemized breakdown, then writes the comparison file with a reproducible analysis id.
Scoring weights
Each change gets a 0 to 100 score: a sum of published weights for signals that fired, each with cited sentences. It is shown as a change level (High at 60 or more, Medium from 25, Low below 25). It ranks how much the disclosure language changed and is not a measure of risk.
Scoring weights not yet generated.
The pipeline exports its weights, rubric questions and thresholds to contract/rubric.json. This table will appear once that file exists. Every published comparison still carries its own itemized score breakdown.
Evaluation
Matcher precision and recall on hand-labeled filing pairs, threshold calibration, and determinism checks.
Evaluation runs on the team's machine with real filings; this box shows the latest saved results.
Evaluation statusEVAL.md: matcher precision and recall, thresholds, determinism
Evaluation not yet run.
Results are written to EVAL.md by uv run python -m eval.eval. No numbers are shown until a real evaluation exists.
What RiskDiff never does
- Never decides what changed with a language model. Diffs, change types and scores are deterministic code.
- Never shows an AI claim without a valid citation to filing sentences. Uncited or invalid claims are dropped before publishing.
- Never characterizes a company's overall risk level, and never turns a disclosure change into a risk percentage.
- Never gives trading or portfolio advice, and never suggests acting on a disclosure change.
- Never fabricates filings, sentences or numbers. The development fixture is labeled on every screen and is never published.
- Never renders filing text as HTML. Filing text is displayed as plain text only.
- Never re-downloads a cached filing, and stays within SEC fair-access limits.