Prevalence + subgroup measurement in a clinical-text pipeline

Working prototype · generated 2026-07-18 · REAL DATA note text & extracted demographics (MTSamples, public de-identified transcription samples) · MOCK judge (deterministic; no LLM calls, no keys) · research-phase measurement demo — measurement, not a guarantee, and not a clinical tool.

Verdict

The corpus REAL DATA

292 real, public, de-identified medical transcription sample notes fetched politely from mtsamples.com on 2026-07-18, each with its source URL preserved (data/notes.jsonl). De-identified at source; ages/sexes below are exactly what the note text states. See DATA-README.md for provenance and terms understanding.

Sectionn notes
Discharge Summary102
Emergency Room Reports74
Pain Management63
Psychiatry / Psychology53
Total292
195/292
age stated in text (high conf.); 97 → "unstated" bucket
218/292
sex stated (184 explicit, 34 pronoun-inferred = low conf.)
120
stratified gold sample (proportional by section, fixed seed)

Notes without an extractable value form their own "unstated" subgroup — they are measured, not silently dropped. In this corpus the unstated bucket has visibly lower phenomenon rates (shorter, administrative note types), which is itself a finding a dashboard should show.

Documented substance use

Judge labeled all 292 notes MOCK JUDGE · calibrated against a stratified gold sample of 120 notes (41% of corpus) — the other 172 notes cost zero labels.

91%
measured sensitivity
(21/23 gold positives)
92%
measured specificity
(89/97 gold negatives)
89% / 87%
true se/sp on full corpus
knowable in mock mode only
calibrated prevalence (Rogan-Gladen + MC, 95% CI) raw judge flag-rate MOCK JUDGE true prevalence (knowable in mock mode only)

Sex (as stated in note text)

0%10%20%30%40%50%femalen = 115female (n=115): calibrated 34.4% [22.2%, 48.6%] · raw judge 36.5% · true (mock-only) 23.5%34.4%malen = 103male (n=103): calibrated 31.3% [18.8%, 45.7%] · raw judge 34.0% · true (mock-only) 28.2%31.3%unstatedn = 74unstated (n=74): calibrated 4.7% [0.0%, 16.9%] · raw judge 12.2% · true (mock-only) 10.8%4.7%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)True (mock-only)
female1154236.5%34.4% [22.2%, 48.6%]23.5%
male1033534.0%31.3% [18.8%, 45.7%]28.2%
unstated74912.2%4.7% [0.0%, 16.9%]10.8%

Age band (as stated in note text)

0%10%20%30%40%50%60%70%0-17n = 270-17 (n=27): calibrated 39.7% [17.9%, 65.4%] · raw judge 40.7% · true (mock-only) 25.9%39.7%18-39n = 4718-39 (n=47): calibrated 39.2% [21.7%, 59.4%] · raw judge 40.4% · true (mock-only) 27.7%39.2%40-64n = 6340-64 (n=63): calibrated 42.2% [26.3%, 60.8%] · raw judge 42.9% · true (mock-only) 36.5%42.2%65+n = 5865+ (n=58): calibrated 31.9% [16.5%, 50.0%] · raw judge 34.5% · true (mock-only) 22.4%31.9%unstatedn = 97unstated (n=97): calibrated 1.2% [0.0%, 11.2%] · raw judge 9.3% · true (mock-only) 8.2%1.2%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)True (mock-only)
0-17271140.7%39.7% [17.9%, 65.4%]25.9%
18-39471940.4%39.2% [21.7%, 59.4%]27.7%
40-64632742.9%42.2% [26.3%, 60.8%]36.5%
65+582034.5%31.9% [16.5%, 50.0%]22.4%
unstated9799.3%1.2% [0.0%, 11.2%]8.2%

MTSamples section (note type)

This axis is calibrated per stratum: each note type against its own slice of the gold sample (counts in the table). Judge error is note-type-dependent, so the pooled se/sp misfits individual strata — compare the "pooled-cal" column. Strata with few or zero gold positives report near-[0%, 100%] intervals: the honest statement that the judge's sensitivity is unmeasured there, not fake precision.

0%10%20%30%40%50%60%70%80%90%100%Discharge Summaryn = 102Discharge Summary (n=102): calibrated 4.9% [0.0%, 16.6%] · raw judge 4.9% · true (mock-only) 4.9%4.9%Emergency Room Reportsn = 74Emergency Room Reports (n=74): calibrated 35.7% [2.1%, 74.0%] · raw judge 51.4% · true (mock-only) 23.0%35.7%Pain Managementn = 63Pain Management (n=63): calibrated 8.0% [0.0%, 26.6%] · raw judge 7.9% · true (mock-only) 9.5%8.0%Psychiatry / Psychologyn = 53Psychiatry / Psychology (n=53): calibrated 81.2% [53.1%, 100.0%] · raw judge 71.7% · true (mock-only) 67.9%81.2%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)Stratum gold (pos/neg)Pooled-cal (for contrast)True (mock-only)
Discharge Summary10254.9%4.9% [0.0%, 16.6%]2/400.0% [0.0%, 4.4%]4.9%
Emergency Room Reports743851.4%35.7% [2.1%, 74.0%]5/2552.6% [37.1%, 71.0%]23.0%
Pain Management6357.9%8.0% [0.0%, 26.6%]2/240.0% [0.0%, 11.2%]9.5%
Psychiatry / Psychology533871.7%81.2% [53.1%, 100.0%]14/877.3% [59.8%, 100.0%]67.9%

Drift check across batch split

alphabetical-by-title within section; stand-in for a time/batch split (MTSamples carries no timestamps). A flag-rate shift is a proxy drift alarm — it says "re-gold and re-estimate", not "the estimate is wrong".

BatchnJudge flagsFlag rate (Wilson 95%)
batch-114745 30.6% [23.7%, 38.5%]
batch-214541 28.3% [21.6%, 36.1%]

Two-proportion z = 0.44 → no alarm at this split.

Suicidality mention (affirmed)

Judge labeled all 292 notes MOCK JUDGE · calibrated against a stratified gold sample of 120 notes (41% of corpus) — the other 172 notes cost zero labels.

83%
measured sensitivity
(5/6 gold positives)
92%
measured specificity
(105/114 gold negatives)
85% / 93%
true se/sp on full corpus
knowable in mock mode only
calibrated prevalence (Rogan-Gladen + MC, 95% CI) raw judge flag-rate MOCK JUDGE true prevalence (knowable in mock mode only)

Sex (as stated in note text)

0%10%20%30%40%50%femalen = 115female (n=115): calibrated 9.5% [0.0%, 25.8%] · raw judge 14.8% · true (mock-only) 7.8%9.5%malen = 103male (n=103): calibrated 13.3% [1.1%, 33.3%] · raw judge 17.5% · true (mock-only) 9.7%13.3%unstatedn = 74unstated (n=74): calibrated 0.0% [0.0%, 2.4%] · raw judge 2.7% · true (mock-only) 1.4%0.0%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)True (mock-only)
female1151714.8%9.5% [0.0%, 25.8%]7.8%
male1031817.5%13.3% [1.1%, 33.3%]9.7%
unstated7422.7%0.0% [0.0%, 2.4%]1.4%

Age band (as stated in note text)

0%10%20%30%40%50%60%70%0-17n = 270-17 (n=27): calibrated 20.5% [0.6%, 56.7%] · raw judge 22.2% · true (mock-only) 14.8%20.5%18-39n = 4718-39 (n=47): calibrated 15.9% [0.3%, 43.3%] · raw judge 19.1% · true (mock-only) 14.9%15.9%40-64n = 6340-64 (n=63): calibrated 22.2% [7.0%, 51.9%] · raw judge 23.8% · true (mock-only) 11.1%22.2%65+n = 5865+ (n=58): calibrated 0.0% [0.0%, 12.6%] · raw judge 6.9% · true (mock-only) 1.7%0.0%unstatedn = 97unstated (n=97): calibrated 0.0% [0.0%, 2.0%] · raw judge 3.1% · true (mock-only) 1.0%0.0%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)True (mock-only)
0-1727622.2%20.5% [0.6%, 56.7%]14.8%
18-3947919.1%15.9% [0.3%, 43.3%]14.9%
40-64631523.8%22.2% [7.0%, 51.9%]11.1%
65+5846.9%0.0% [0.0%, 12.6%]1.7%
unstated9733.1%0.0% [0.0%, 2.0%]1.0%

MTSamples section (note type)

This axis is calibrated per stratum: each note type against its own slice of the gold sample (counts in the table). Judge error is note-type-dependent, so the pooled se/sp misfits individual strata — compare the "pooled-cal" column. Strata with few or zero gold positives report near-[0%, 100%] intervals: the honest statement that the judge's sensitivity is unmeasured there, not fake precision.

0%10%20%30%40%50%60%70%80%90%100%Discharge Summaryn = 102Discharge Summary (n=102): calibrated 0.0% [0.0%, 100.0%] · raw judge 0.0% · true (mock-only) 0.0%0.0%Emergency Room Reportsn = 74Emergency Room Reports (n=74): calibrated 1.4% [0.0%, 100.0%] · raw judge 1.4% · true (mock-only) 2.7%1.4%Pain Managementn = 63Pain Management (n=63): calibrated 0.0% [0.0%, 100.0%] · raw judge 0.0% · true (mock-only) 0.0%0.0%Psychiatry / Psychologyn = 53Psychiatry / Psychology (n=53): calibrated 44.3% [0.0%, 100.0%] · raw judge 67.9% · true (mock-only) 34.0%44.3%
Table view
BucketnJudge flagsRaw rateCalibrated (95% CI)Stratum gold (pos/neg)Pooled-cal (for contrast)True (mock-only)
Discharge Summary10200.0%0.0% [0.0%, 100.0%]0/420.0% (clipped at 0: raw rate < judge FP floor — artifact, not certainty)0.0%
Emergency Room Reports7411.4%1.4% [0.0%, 100.0%]0/300.0% (clipped at 0: raw rate < judge FP floor — artifact, not certainty)2.7%
Pain Management6300.0%0.0% [0.0%, 100.0%]0/260.0% (clipped at 0: raw rate < judge FP floor — artifact, not certainty)0.0%
Psychiatry / Psychology533667.9%44.3% [0.0%, 100.0%]6/1682.1% [58.0%, 100.0%]34.0%

Drift check across batch split

alphabetical-by-title within section; stand-in for a time/batch split (MTSamples carries no timestamps). A flag-rate shift is a proxy drift alarm — it says "re-gold and re-estimate", not "the estimate is wrong".

BatchnJudge flagsFlag rate (Wilson 95%)
batch-114717 11.6% [7.3%, 17.7%]
batch-214520 13.8% [9.1%, 20.3%]

Two-proportion z = -0.57 → no alarm at this split.

Why this shape matters for a fairness dashboard

Honest limitations