Resilience
DC Medicaid reports hid data on 399,086 people
Two public DC Medicaid reports carried hidden row-level data on 399,086 people for up to three years. How summary reports leak, and how to audit yours.
The District of Columbia Department of Health Care Finance (DHCF), which runs DC Medicaid and the DC Healthcare Alliance, is mailing breach letters to people whose records sat inside two reports on its public website. The reports showed only summary figures such as enrollment counts, but the personal data behind those figures could be reached by anyone who knew where to look. DHCF told the US Department of Health and Human Services (HHS) that 399,086 people are affected: beneficiaries enrolled between 2023 and July 2026.
Nobody broke in. DHCF says it found the problem itself on July 21, 2026, removed both reports, and has no evidence that anyone viewed or misused the data. The exposure window still ran from 2023 until July 2026, and the same mistake is easy to make with any spreadsheet, dashboard or PDF that an organisation publishes from row-level data. If your team posts reports built from customer, patient or staff records, this one applies to you.
What was exposed
The notice letter lists Medicaid ID number, date of birth, provider name, race, gender, ward (one of DC's eight wards) and ethnicity. It says names, Social Security numbers and financial account information were not included. Letters go to the beneficiary, a parent for a child, or a relative of a beneficiary who has died.
Without names the data is harder to abuse, which is DHCF's own argument for why misuse is "less likely". It is not anonymous, though. A Medicaid ID is the key used to bill for care, and date of birth plus ward plus race, gender and ethnicity narrows a person down quickly, especially in a small ward or an uncommon demographic group. Provider name adds a health relationship that the person did not choose to make public.
DHCF reported the incident to the HHS Office for Civil Rights as a HIPAA (Health Insurance Portability and Accountability Act) breach, and it appeared on the OCR breach portal in the last days of September. The letters offer no credit monitoring. They point people to free annual credit reports and free fraud alerts from Equifax, Experian and TransUnion.
How a summary report hides row-level data
DHCF's own description is short: the reports "were intended to display only summary information" and "did not show anyone's personal details on the screen", but the "underlying personal information that supported these reports may have been reachable". HIPAA Journal describes the data as held in "hidden fields". DHCF has not said what software produced the reports or what a visitor had to do to reach the data, so what follows is a general explanation of how this class of exposure happens, not a reconstruction of what happened in DC.
Most reporting tools work the same way. You load the detailed records, one row per person, and the tool aggregates them into the counts and charts the reader sees. The problem is that many tools keep the detailed rows inside the published file or behind the published page, so the aggregation can be refreshed, filtered or drilled into. Hiding the rows only changes what is on screen; the file still holds them.
- Excel PivotTables. A PivotTable stores a pivot cache, a full copy of its source rows inside the workbook. Deleting or hiding the source sheet does not remove it. With "Enable show details" on, double-clicking any total opens a new sheet with every record behind it. Even with that off, an .xlsx file is a zip archive, and the cached records sit in
xl/pivotCache/pivotCacheRecords1.xmlfor anyone who unzips it. - Hidden sheets, rows and columns. Excel's Hide and Very Hidden settings are display options. The cells are still in the file and come back with Unhide, the VBA editor or a look inside the zip.
- BI dashboards. Tableau has separate permissions for downloading summary data and full data. Full data means every row and column in the source. Power BI reports have a matching "summarized data" or "underlying data" export option. A dashboard built on a person-level table, with the wrong permission set, will hand those rows to a viewer.
- Web charts fed by an API. Some interactive pages download the detailed rows as JSON and count them in the browser. The chart shows totals; the browser's network tab shows everyone.
- PDF exports. A black box drawn over a table hides nothing from a copy and paste, and attachments or layers can carry the source data along with the page.
What to do
Start with an inventory. Search your public site and file stores for spreadsheets, BI embeds and PDFs, for example with a site search such as site:yourdomain.example filetype:xlsx and the same for xls, csv and pdf, plus a list of every public Tableau or Power BI view. Flag anything whose source is person-level data.
Then open the files the way an outsider would, not the way the author does:
- For each workbook, rename the .xlsx to .zip and list its contents. Any
pivotCacheRecordsfile, or aworksheetsentry you do not see as a tab, needs a look. Excel's Document Inspector (File, Info, Check for Issues, Inspect Document) reports hidden rows, columns and worksheets. - For each PivotTable that must stay, open PivotTable Options, Data tab, and clear both "Save source data with file" and "Enable show details", then save and check the zip again.
- For each dashboard, check the download permissions. In Tableau, deny Download Full Data on public views. In Power BI Desktop, under File, Options and settings, Options, Current file, Report settings, limit export to summarized data or turn it off, and confirm the tenant export settings in the admin portal match.
- For each interactive web chart, load the page with the browser developer tools open and read the network responses. If you see one record per person, the aggregation must move to the server.
- For PDFs, select all text and paste it into a plain editor, and check for attachments and layers.
If you find an exposure, take the file down, then check web server or CDN access logs for downloads of that exact path over the period it was live. Those logs tell you whether anyone fetched it, which decides how you notify. Also search caches and archives such as the Wayback Machine for copies, and ask for removal if you find any.
The real fix is to publish only aggregated data. Build the public file from a summary table, not from the record-level extract with the detail hidden. Suppress small counts so a single person cannot be picked out of a ward or age group, a step that public health agencies already use for statistical releases.
Why this keeps happening
In every mechanism above, the software behaves as documented. The gap is in the review step: the report is signed off by someone who looks at the rendered output, while the risk sits in the file or the data feed. A pre-publication check that opens the file the way an outsider would, run by someone other than the author, catches most of these. DHCF's reports were reachable for up to about three years, from 2023 until July 2026, before the agency spotted the problem itself.
Our attack surface management work maps what your organisation exposes publicly, including documents and dashboards, and our compliance and audit readiness team can build a publication review into your data handling. Open the chat and Yaali, our AI agent, will pass your question to the engineer who would do the work.
Sources: DHCF data incident notice, DHCF sample notification letter, SecurityWeek, SC Media, Security Affairs, HIPAA Journal, Microsoft: show details in a PivotTable, Tableau permission capabilities.