Prevalence + subgroup measurement in a clinical-text pipeline

Working prototype · generated 2026-09-15 · REAL DATA note text & extracted demographics (MTSamples, public de-identified transcription samples) · MOCK judge (deterministic; no LLM calls, no keys) · research-phase measurement demo — measurement, not a guarantee, and not a clinical tool.

Verdict

The corpus REAL DATA

292 real, public, de-identified medical transcription sample notes fetched politely from mtsamples.com on 2026-07-18, each with its source URL preserved (data/notes.jsonl). De-identified at source; ages/sexes below are exactly what the note text states. See DATA-README.md for provenance and terms understanding.

Sectionn notes
Discharge Summary102
Emergency Room Reports74
Pain Management63
Psychiatry / Psychology53
Total292
195/292
age stated in text (high conf.); 97 → "unstated" bucket
218/292
sex stated (184 explicit, 34 pronoun-inferred = low conf.)
120
stratified gold sample (proportional by section, fixed seed)

Notes without an extractable value form their own "unstated" subgroup — they are measured, not silently dropped. In this corpus the unstated bucket has visibly lower phenomenon rates (shorter, administrative note types), which is itself a finding a dashboard should show.

Documented substance use

Judge labeled all 292 notes MOCK JUDGE · calibrated against a stratified gold sample of 120 notes (41% of corpus) — the other 172 notes cost zero labels.

90%
measured sensitivity
(19/21 gold positives)
90%
measured specificity
(89/99 gold negatives)
90% / 85%
true se/sp on full corpus
knowable in mock mode only
calibrated prevalence (Rogan-Gladen + MC, 95% CI) raw judge flag-rate MOCK JUDGE true prevalence (knowable in mock mode only)

Sex (as stated in note text)

0%10%20%30%40%50%femalen = 115female (n=115): calibrated 33.3% [20.2%, 48.6%] · raw judge 36.5% · true (mock-only) 21.7%33.3%malen = 103male (n=103): calibrated 30.1% [16.9%, 45.4%] · raw judge 34.0% · true (mock-only) 25.2%30.1%unstatedn = 74unstated (n=74): calibrated 2.6% [0.0%, 15.3%] · raw judge 12.2% · true (mock-only) 9.5%2.6%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)True (mock-only)
female1154236.5%33.3% [20.2%, 48.6%]21.7%
male1033534.0%30.1% [16.9%, 45.4%]25.2%
unstated74912.2%2.6% [0.0%, 15.3%]9.5%

Age band (as stated in note text)

0%10%20%30%40%50%60%70%0-17n = 270-17 (n=27): calibrated 38.8% [16.0%, 65.8%] · raw judge 40.7% · true (mock-only) 22.2%38.8%18-39n = 4718-39 (n=47): calibrated 38.2% [20.0%, 59.6%] · raw judge 40.4% · true (mock-only) 27.7%38.2%40-64n = 6340-64 (n=63): calibrated 41.3% [24.8%, 61.1%] · raw judge 42.9% · true (mock-only) 34.9%41.3%65+n = 5865+ (n=58): calibrated 30.7% [14.4%, 49.8%] · raw judge 34.5% · true (mock-only) 17.2%30.7%unstatedn = 97unstated (n=97): calibrated 0.0% [0.0%, 9.6%] · raw judge 9.3% · true (mock-only) 7.2%0.0%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)True (mock-only)
0-17271140.7%38.8% [16.0%, 65.8%]22.2%
18-39471940.4%38.2% [20.0%, 59.6%]27.7%
40-64632742.9%41.3% [24.8%, 61.1%]34.9%
65+582034.5%30.7% [14.4%, 49.8%]17.2%
unstated9799.3%0.0% [0.0%, 9.6%]7.2%

MTSamples section (note type)

This axis is calibrated per stratum: each note type against its own slice of the gold sample (counts in the table). Judge error is note-type-dependent, so the pooled se/sp misfits individual strata — compare the "pooled-cal" column. Strata with few or zero gold positives report near-[0%, 100%] intervals: the honest statement that the judge's sensitivity is unmeasured there, not fake precision. Where a calibration cannot reproduce a stratum's flag rate at all — the corrected estimate falls outside [0%, 100%] — the cell prints undefined rather than a zero-width interval.

0%10%20%30%40%50%60%70%80%90%100%Discharge Summaryn = 102Discharge Summary (n=102): calibrated 4.9% [0.0%, 16.6%] · raw judge 4.9% · true (mock-only) 3.9%4.9%Emergency Room Reportsn = 74Emergency Room Reports (n=74): calibrated 33.7% [0.0%, 85.6%] · raw judge 51.4% · true (mock-only) 18.9%33.7%Pain Managementn = 63Pain Management (n=63): calibrated 8.0% [0.0%, 26.6%] · raw judge 7.9% · true (mock-only) 9.5%8.0%Psychiatry / Psychologyn = 53Psychiatry / Psychology (n=53): calibrated 80.0% [43.2%, 100.0%] · raw judge 71.7% · true (mock-only) 64.2%80.0%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)Stratum gold (pos/neg)Pooled-cal (for contrast)True (mock-only)
Discharge Summary10254.9%4.9% [0.0%, 16.6%]2/400.0% [0.0%, 2.5%]3.9%
Emergency Room Reports743851.4%33.7% [0.0%, 85.6%]4/2652.1% [35.9%, 72.1%]18.9%
Pain Management6357.9%8.0% [0.0%, 26.6%]2/240.0% [0.0%, 9.6%]9.5%
Psychiatry / Psychology533871.7%80.0% [43.2%, 100.0%]13/977.6% [59.1%, 100.0%]64.2%

Drift check across batch split

alphabetical-by-title within section; stand-in for a time/batch split (MTSamples carries no timestamps). A flag-rate shift is a proxy drift alarm — it says "re-gold and re-estimate", not "the estimate is wrong".

BatchnJudge flagsFlag rate (Wilson 95%)
batch-114745 30.6% [23.7%, 38.5%]
batch-214541 28.3% [21.6%, 36.1%]

Two-proportion z = 0.44 → no alarm at this split.

Suicidality mention (affirmed)

Judge labeled all 292 notes MOCK JUDGE · calibrated against a stratified gold sample of 120 notes (41% of corpus) — the other 172 notes cost zero labels.

83%
measured sensitivity
(5/6 gold positives)
92%
measured specificity
(105/114 gold negatives)
85% / 93%
true se/sp on full corpus
knowable in mock mode only
calibrated prevalence (Rogan-Gladen + MC, 95% CI) raw judge flag-rate MOCK JUDGE true prevalence (knowable in mock mode only)

Sex (as stated in note text)

0%10%20%30%40%50%femalen = 115female (n=115): calibrated 9.5% [0.0%, 25.8%] · raw judge 14.8% · true (mock-only) 7.8%9.5%malen = 103male (n=103): calibrated 13.3% [1.1%, 33.3%] · raw judge 17.5% · true (mock-only) 9.7%13.3%unstatedn = 74unstated (n=74): calibrated 0.0% [0.0%, 2.4%] · raw judge 2.7% · true (mock-only) 1.4%0.0%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)True (mock-only)
female1151714.8%9.5% [0.0%, 25.8%]7.8%
male1031817.5%13.3% [1.1%, 33.3%]9.7%
unstated7422.7%0.0% [0.0%, 2.4%]1.4%

Age band (as stated in note text)

0%10%20%30%40%50%60%70%0-17n = 270-17 (n=27): calibrated 20.5% [0.6%, 56.7%] · raw judge 22.2% · true (mock-only) 14.8%20.5%18-39n = 4718-39 (n=47): calibrated 15.9% [0.3%, 43.3%] · raw judge 19.1% · true (mock-only) 14.9%15.9%40-64n = 6340-64 (n=63): calibrated 22.2% [7.0%, 51.9%] · raw judge 23.8% · true (mock-only) 11.1%22.2%65+n = 5865+ (n=58): calibrated 0.0% [0.0%, 12.6%] · raw judge 6.9% · true (mock-only) 1.7%0.0%unstatedn = 97unstated (n=97): calibrated 0.0% [0.0%, 2.0%] · raw judge 3.1% · true (mock-only) 1.0%0.0%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)True (mock-only)
0-1727622.2%20.5% [0.6%, 56.7%]14.8%
18-3947919.1%15.9% [0.3%, 43.3%]14.9%
40-64631523.8%22.2% [7.0%, 51.9%]11.1%
65+5846.9%0.0% [0.0%, 12.6%]1.7%
unstated9733.1%0.0% [0.0%, 2.0%]1.0%

MTSamples section (note type)

This axis is calibrated per stratum: each note type against its own slice of the gold sample (counts in the table). Judge error is note-type-dependent, so the pooled se/sp misfits individual strata — compare the "pooled-cal" column. Strata with few or zero gold positives report near-[0%, 100%] intervals: the honest statement that the judge's sensitivity is unmeasured there, not fake precision. Where a calibration cannot reproduce a stratum's flag rate at all — the corrected estimate falls outside [0%, 100%] — the cell prints undefined rather than a zero-width interval.

0%10%20%30%40%50%60%70%80%90%100%Discharge Summaryn = 102Discharge Summary (n=102): calibrated 0.0% [0.0%, 100.0%] · raw judge 0.0% · true (mock-only) 0.0%0.0%Emergency Room Reportsn = 74Emergency Room Reports (n=74): calibrated 1.4% [0.0%, 100.0%] · raw judge 1.4% · true (mock-only) 2.7%1.4%Pain Managementn = 63Pain Management (n=63): calibrated 0.0% [0.0%, 100.0%] · raw judge 0.0% · true (mock-only) 0.0%0.0%Psychiatry / Psychologyn = 53Psychiatry / Psychology (n=53): calibrated 44.3% [0.0%, 100.0%] · raw judge 67.9% · true (mock-only) 34.0%44.3%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)Stratum gold (pos/neg)Pooled-cal (for contrast)True (mock-only)
Discharge Summary10200.0%0.0% [0.0%, 100.0%]0/42undefined — calibration refuted here (corrected estimate -10.7% [-27.5%, -4.1%], entirely below 0)0.0%
Emergency Room Reports7411.4%1.4% [0.0%, 100.0%]0/30undefined — calibration refuted here (corrected estimate -8.8% [-24.5%, -0.6%], entirely below 0)2.7%
Pain Management6300.0%0.0% [0.0%, 100.0%]0/26undefined — calibration refuted here (corrected estimate -10.3% [-26.9%, -3.2%], entirely below 0)0.0%
Psychiatry / Psychology533667.9%44.3% [0.0%, 100.0%]6/1682.1% [58.0%, 100.0%]34.0%

Drift check across batch split

alphabetical-by-title within section; stand-in for a time/batch split (MTSamples carries no timestamps). A flag-rate shift is a proxy drift alarm — it says "re-gold and re-estimate", not "the estimate is wrong".

BatchnJudge flagsFlag rate (Wilson 95%)
batch-114717 11.6% [7.3%, 17.7%]
batch-214520 13.8% [9.1%, 20.3%]

Two-proportion z = -0.57 → no alarm at this split.

Why this shape matters for a fairness dashboard

Honest limitations