Skip to content
RiskDiff

How RiskDiff works

Every year a public company lists the risks it faces in the Risk Factors section of its annual report (10-K). RiskDiff lines up this year's list against last year's and shows exactly which sentences changed. The filing text is the only evidence: every quote is verbatim, and every AI statement points to the sentences it is based on.

How it works

  1. Step 1

    Pull both filings from SEC EDGAR

    The company's two most recent annual reports (10-K), straight from the SEC. Amendments are excluded.

  2. ✦ Step 2

    Match each risk factor to last year's version

    Each risk is paired with its counterpart from last year, even when it moved or was renamed. Unpaired risks are new or dropped.

  3. Step 3

    Find the exact sentences that changed

    Sentence by sentence and word by word: what was added, removed, or rewritten.

  4. ✦ Step 4

    Summarize, with every statement cited

    A short plain-English summary of each change. Every point links to the filing sentences behind it.

What changed is always computed by fixed rules in code, so the same two filings always give the same result. AI helps in three narrow places only: recognizing that a reworded risk is the same risk as last year's (step 2), wording short summaries (step 4), and answering four fixed yes/no questions about each change. A “yes” adds a fixed, published number of points only when it cites the filing sentences behind it, and never when a rule in code already scored the same fact. Any summary statement that does not cite valid filing sentences is removed before it is shown. Each change also gets a wording change level (High 60+, Medium 25–59, Low below 25, on a 0 to 100 scale): how much the language changed, not how risky the company is.

What RiskDiff does not do

  • No predictions. It reports what the filings say. It never forecasts events or outcomes.
  • No risk ratings. A change in disclosure language is not a change in actual risk. The wording change level measures wording only.
  • No investment advice. It never suggests acting on a disclosure change.
Not investment advice. RiskDiff shows changes in disclosed risk language, not changes in actual risk.
Technical detailThe full pipeline, the published scoring weights, thresholds and caps, blocked wording, and the evaluation status. RiskDiff reads Item 1A (Risk Factors) of each 10-K.

The pipeline (nine stages)

Nine stages. Seven are deterministic code; two use models, and neither decides what changed.

Deterministic code AI-assisted (never decides what changed)
In ticker, e.g. NVDApublished comparison JSON, every claim cited Out
  1. Stage 0Fetch filings from EDGARDeterministic
  2. Stage 1Extract Risk FactorsDeterministic
  3. Stage 2Segment risk factorsDeterministic
  4. ✦ Stage 3Embed risk blocksAI-assisted
  5. Stage 4Match risks across yearsDeterministic
  6. Stage 5Sentence-level diffDeterministic
  7. Stage 6Deterministic signalsDeterministic
  8. ✦ Stage 7AI rubric + interpretationAI-assisted
  9. Stage 8Score + publishDeterministic
Each stage is a pure function of its inputs and a stage version, cached on disk. Same filings and same versions produce byte-identical output and the same analysis id.
  1. 0 Fetch filings from EDGAR

    Looks up the company on SEC EDGAR and downloads its two most recent original 10-K filings (amendments excluded). Every request identifies itself and stays within SEC fair-access limits.

  2. 1 Extract Risk Factors

    Finds the Risk Factors section in each filing, skipping the table of contents and stopping where the next section begins.

  3. 2 Segment risk factors

    Splits the Risk Factors section into risk blocks (a heading plus its body), then into sentences with stable ids and character offsets back into the filing text.

  4. 3 Embed risk blocks · ✦ AI-assisted

    A local embedding model turns each risk block into a vector so that equivalent risks can be recognized even when the wording moved.

  5. 4 Match risks across years

    Pairs old and new risk blocks with an optimal assignment over heading and body similarity. Pairs below the published threshold become new or removed risks.

  6. 5 Sentence-level diff

    Aligns sentences inside each matched pair, computes word-level insertions and deletions, and classifies the change (new, removed, expanded, reduced, reworded, largely unchanged).

  7. 6 Deterministic signals

    Rule-based detectors with citations: hedged wording that became actual wording, newly named entities, new quantities, and the share of added or removed text.

  8. 7 AI rubric + interpretation · ✦ AI-assisted

    A language model answers narrow yes/no rubric questions and writes short interpretations. A citation validator drops every claim that does not cite valid filing sentences.

  9. 8 Score + publish

    Adds up published rubric weights into a 0 to 100 change score with an itemized breakdown, then writes the comparison file with a reproducible analysis id.

Scoring weights

Each change gets a 0 to 100 score: a sum of published weights for signals that fired, each with cited sentences. It is shown as a change level (High at 60 or more, Medium from 25, Low below 25). It ranks how much the disclosure language changed and is not a measure of risk.

Scoring weights not yet generated.

The pipeline exports its weights, rubric questions and thresholds to contract/rubric.json. This table will appear once that file exists. Every published comparison still carries its own itemized score breakdown.

Evaluation

Matcher precision and recall on hand-labeled filing pairs, threshold calibration, and determinism checks.

Evaluation runs on the team's machine with real filings; this box shows the latest saved results.

Evaluation statusEVAL.md: matcher precision and recall, thresholds, determinism

Evaluation not yet run.

Results are written to EVAL.md by uv run python -m eval.eval. No numbers are shown until a real evaluation exists.

What RiskDiff never does

  • Never decides what changed with a language model. Diffs, change types and scores are deterministic code.
  • Never shows an AI claim without a valid citation to filing sentences. Uncited or invalid claims are dropped before publishing.
  • Never characterizes a company's overall risk level, and never turns a disclosure change into a risk percentage.
  • Never gives trading or portfolio advice, and never suggests acting on a disclosure change.
  • Never fabricates filings, sentences or numbers. The development fixture is labeled on every screen and is never published.
  • Never renders filing text as HTML. Filing text is displayed as plain text only.
  • Never re-downloads a cached filing, and stays within SEC fair-access limits.