Prevalence + subgroup measurement in a clinical-text pipeline
Verdict
- Your pipeline's judge is itself measurable. A naive keyword judge labeled all 292 notes; 120 stratified gold labels were enough to measure its error profile per phenomenon — and the profiles differ: substance-use errors are false-positive-heavy (negation blindness: "denies alcohol or drug use" still flags), suicidality errors are false-negative-heavy. You cannot correct what you have not measured.
- Calibration matters more than the raw rate. Rogan-Gladen correction with Monte-Carlo error propagation moved estimates toward the (mock-mode knowable) truth; the 95% interval covered the true subgroup prevalence in 24/24 subgroup estimates — a single-realization, in-sample mock check (4 of the 24 are near-vacuous wide intervals), evidence of conservatism, not a calibration study.
- Judge error is subgroup-dependent — and that is measurable too. On the note-type axis, ONE pooled calibration misfits: for substance use in Emergency Room reports the pooled-calibrated estimate reads 52.6% against a true prevalence of 23.0% (ER notes are dense with negated screens). Calibrating each note type against its own gold slice covers the truth in all four strata — at the price of honestly wider intervals.
- Small subgroups get honest widths, not fake precision. The smallest populated bucket (0-17, n=27) reports a ±24-point interval — visible, not hidden.
- Drift check: no alarm across the batch split for either phenomenon (proxy drift alarm, not a validity certificate).
The corpus REAL DATA
| Section | n notes |
|---|---|
| Discharge Summary | 102 |
| Emergency Room Reports | 74 |
| Pain Management | 63 |
| Psychiatry / Psychology | 53 |
| Total | 292 |
195/292
age stated in text (high conf.); 97 → "unstated" bucket
218/292
sex stated (184 explicit, 34 pronoun-inferred = low conf.)
120
stratified gold sample (proportional by section, fixed seed)
Documented substance use
91%
measured sensitivity
(21/23 gold positives)
(21/23 gold positives)
92%
measured specificity
(89/97 gold negatives)
(89/97 gold negatives)
89% / 87%
true se/sp on full corpus
knowable in mock mode only
knowable in mock mode only
calibrated prevalence (Rogan-Gladen + MC, 95% CI) raw judge flag-rate MOCK JUDGE true prevalence (knowable in mock mode only)
Sex (as stated in note text)
Table view
| Bucket | n | Judge flags | Raw rate | Calibrated (95% CI) | True (mock-only) |
|---|---|---|---|---|---|
| female | 115 | 42 | 36.5% | 34.4% [22.2%, 48.6%] | 23.5% |
| male | 103 | 35 | 34.0% | 31.3% [18.8%, 45.7%] | 28.2% |
| unstated | 74 | 9 | 12.2% | 4.7% [0.0%, 16.9%] | 10.8% |
Age band (as stated in note text)
Table view
| Bucket | n | Judge flags | Raw rate | Calibrated (95% CI) | True (mock-only) |
|---|---|---|---|---|---|
| 0-17 | 27 | 11 | 40.7% | 39.7% [17.9%, 65.4%] | 25.9% |
| 18-39 | 47 | 19 | 40.4% | 39.2% [21.7%, 59.4%] | 27.7% |
| 40-64 | 63 | 27 | 42.9% | 42.2% [26.3%, 60.8%] | 36.5% |
| 65+ | 58 | 20 | 34.5% | 31.9% [16.5%, 50.0%] | 22.4% |
| unstated | 97 | 9 | 9.3% | 1.2% [0.0%, 11.2%] | 8.2% |
MTSamples section (note type)
Table view
| Bucket | n | Judge flags | Raw rate | Calibrated (95% CI) | Stratum gold (pos/neg) | Pooled-cal (for contrast) | True (mock-only) |
|---|---|---|---|---|---|---|---|
| Discharge Summary | 102 | 5 | 4.9% | 4.9% [0.0%, 16.6%] | 2/40 | 0.0% [0.0%, 4.4%] | 4.9% |
| Emergency Room Reports | 74 | 38 | 51.4% | 35.7% [2.1%, 74.0%] | 5/25 | 52.6% [37.1%, 71.0%] | 23.0% |
| Pain Management | 63 | 5 | 7.9% | 8.0% [0.0%, 26.6%] | 2/24 | 0.0% [0.0%, 11.2%] | 9.5% |
| Psychiatry / Psychology | 53 | 38 | 71.7% | 81.2% [53.1%, 100.0%] | 14/8 | 77.3% [59.8%, 100.0%] | 67.9% |
Drift check across batch split
| Batch | n | Judge flags | Flag rate (Wilson 95%) |
|---|---|---|---|
| batch-1 | 147 | 45 | 30.6% [23.7%, 38.5%] |
| batch-2 | 145 | 41 | 28.3% [21.6%, 36.1%] |
Suicidality mention (affirmed)
83%
measured sensitivity
(5/6 gold positives)
(5/6 gold positives)
92%
measured specificity
(105/114 gold negatives)
(105/114 gold negatives)
85% / 93%
true se/sp on full corpus
knowable in mock mode only
knowable in mock mode only
calibrated prevalence (Rogan-Gladen + MC, 95% CI) raw judge flag-rate MOCK JUDGE true prevalence (knowable in mock mode only)
Sex (as stated in note text)
Table view
| Bucket | n | Judge flags | Raw rate | Calibrated (95% CI) | True (mock-only) |
|---|---|---|---|---|---|
| female | 115 | 17 | 14.8% | 9.5% [0.0%, 25.8%] | 7.8% |
| male | 103 | 18 | 17.5% | 13.3% [1.1%, 33.3%] | 9.7% |
| unstated | 74 | 2 | 2.7% | 0.0% [0.0%, 2.4%] | 1.4% |
Age band (as stated in note text)
Table view
| Bucket | n | Judge flags | Raw rate | Calibrated (95% CI) | True (mock-only) |
|---|---|---|---|---|---|
| 0-17 | 27 | 6 | 22.2% | 20.5% [0.6%, 56.7%] | 14.8% |
| 18-39 | 47 | 9 | 19.1% | 15.9% [0.3%, 43.3%] | 14.9% |
| 40-64 | 63 | 15 | 23.8% | 22.2% [7.0%, 51.9%] | 11.1% |
| 65+ | 58 | 4 | 6.9% | 0.0% [0.0%, 12.6%] | 1.7% |
| unstated | 97 | 3 | 3.1% | 0.0% [0.0%, 2.0%] | 1.0% |
MTSamples section (note type)
Table view
| Bucket | n | Judge flags | Raw rate | Calibrated (95% CI) | Stratum gold (pos/neg) | Pooled-cal (for contrast) | True (mock-only) |
|---|---|---|---|---|---|---|---|
| Discharge Summary | 102 | 0 | 0.0% | 0.0% [0.0%, 100.0%] | 0/42 | 0.0% (clipped at 0: raw rate < judge FP floor — artifact, not certainty) | 0.0% |
| Emergency Room Reports | 74 | 1 | 1.4% | 1.4% [0.0%, 100.0%] | 0/30 | 0.0% (clipped at 0: raw rate < judge FP floor — artifact, not certainty) | 2.7% |
| Pain Management | 63 | 0 | 0.0% | 0.0% [0.0%, 100.0%] | 0/26 | 0.0% (clipped at 0: raw rate < judge FP floor — artifact, not certainty) | 0.0% |
| Psychiatry / Psychology | 53 | 36 | 67.9% | 44.3% [0.0%, 100.0%] | 6/16 | 82.1% [58.0%, 100.0%] | 34.0% |
Drift check across batch split
| Batch | n | Judge flags | Flag rate (Wilson 95%) |
|---|---|---|---|
| batch-1 | 147 | 17 | 11.6% [7.3%, 17.7%] |
| batch-2 | 145 | 20 | 13.8% [9.1%, 20.3%] |
Why this shape matters for a fairness dashboard
- Same phenomenon, per-subgroup estimates, honest interval widths. Subgroup comparisons on raw judge output confound judge error with true difference. Measuring the judge and correcting per subgroup separates the two — and puts a defensible CI on every cell of the dashboard, with n shown everywhere.
- Labeling cost scales with the gold sample, not the corpus. Here 120 gold labels served 292 notes. The same design serves a 17,000-note evaluation with a few hundred adjudicated labels per refresh, re-golding only when the drift alarm fires.
- The judge's error profile is subgroup-relevant. This demo uses one shared se/sp (a stated limitation); a real engagement measures se/sp per subgroup on adjudicated labels — differential judge error across subgroups is precisely the fairness failure mode worth watching.
Honest limitations
- The judge is a mock. A deterministic lexicon judge with a hidden
error profile stands in for an LLM judge so the pipeline runs keyless
end-to-end. An LLM-judge interface is wired (
camh_subgroup/judge.py) but deliberately not exercised: no API keys on disk, no real-LLM numbers are claimed anywhere in this report. - Mock-mode gold is gold by construction. The reference labels are a deterministic negation-aware rule set; "truth" here validates the measurement machinery, not any clinical claim. In a real engagement the gold is adjudicated clinician labels under a stated codebook, with inter-annotator agreement (κ) reported — and annotator error is a measurement-validity risk the CIs do not capture.
- MTSamples is not psychiatric EHR data. It is a public, US, de-identified transcription teaching corpus, mixed-specialty, with no timestamps and no race/ethnicity or housing-status fields — so the subgroup axes here (sex, age band, note type) are the ones the text itself supports. The drift split is alphabetical, a labeled stand-in.
- Shared se/sp across subgroups (small gold sample): subgroup intervals inherit a common calibration; per-subgroup calibration needs more gold per stratum.
- Demographic extraction is itself a model. Every extracted field carries a confidence label (34 sex values are pronoun-inferred, low confidence); extraction errors move notes between buckets and are not propagated into the CIs.
- What a real clinical deployment needs: the institution's own notes in its own environment, its clinicians' adjudicated gold under its own codebook, per-subgroup calibration sized to its strata, a real time-ordered drift split, and PHIPA/HIPAA-appropriate de-identification review before any data leaves the institution — this prototype makes no compliance claim of any kind.