# Healthcare Exclusion Screening Source Audit 2026

Edition: 2026-10-02. Author and funder: Exclia. License: CC BY 4.0 for this
methodology and the exported observations. Cite Exclia, edition, sample size and
limitations when sharing. Government source materials retain their own terms.

## Question and population

How does Exclia configure collection for the exclusion sources declared in its
repository? This is original descriptive analysis of configuration, not a survey
of healthcare employers or a measurement of exclusion prevalence or accuracy.
The unit is one source declaration, regardless of the number of URLs or formats.
Include every entry in SOURCE_REGISTRY at commit
071d45d15681ffe3028dd62670eee4e52509dfdf; exclude no entries. N = 3:
LEIE, SAM, TX_MEDICAID. These are a census of this registry, a convenience sample
of the wider source landscape. Observation date records extraction, not a live
agency check. No network requests or production database queries were made.

## Collection and reproduction

Check out the recorded commit in a clean separate checkout. Copy
scripts/research/build-source-audit.ts from this edition into that checkout,
install its workspace dependencies with Bun, and run from the repository root:

    bun scripts/research/build-source-audit.ts 071d45d15681ffe3028dd62670eee4e52509dfdf 2026-10-02

Compare the resulting CSV and JSON with this edition's downloads. The JSON
records the revision, input path, date, sample size and all observations. The
script refuses tracked changes to research input packages. It projects declared
fields without fetching linked agency files. Review each row against the pinned
source code; keep historical editions immutable and publish corrections as a
new edition with a change log. No independent external reviewer has verified
this edition.

## Data dictionary

- id: registry key; one row per key.
- name: Exclia's declared display name.
- jurisdiction: federal or two-letter state code.
- accessMethod: declared download, api, scrape or portal category.
- formats: declared file formats, pipe-separated in CSV; multiple formats do not create extra rows.
- intervalDays: configured expected collection/update interval, integer days.
- graceDays: configured freshness allowance beyond that interval, integer days.
- declaredBlocker: declaration's disabledReason text, or null in JSON / empty in CSV when absent. Absence is not proof of operational readiness.
- urls: declared source URLs; array in JSON, pipe-separated in CSV. These are unverified at extraction time.

## Analysis and findings

Count each accessMethod once per row; divide by N = 3. Download, API and portal
each account for 1/3 (33.3%, rounded to one decimal). Two of three declarations
configure a 30-day interval (66.7%); one configures 1 day (33.3%). One of three
contains an explicit declaration blocker (33.3%). All denominators include all
three rows. There are no missing access methods or interval values. This edition
uses no weighting, inferential statistics, confidence intervals or significance
tests. Configured intervals are not measured agency refresh frequencies.

## Privacy and limitations

Only public source configuration is exported. No exclusion records, names of
screened individuals, identifiers, customer rosters, credentials, job outcomes
or tenant data are collected. The sample is small and chosen by Exclia's own
implementation; Exclia has a commercial interest in screening software. Results
cannot establish national coverage, live source availability, legal compliance,
match quality, false-positive rates, time savings or customer performance.
The Texas blocker is historical declaration text, not a fresh reachability test.
Operational enablement and terms review are stored elsewhere and were not read.

## Protocol for a future workflow benchmark (not collected)

Before recruitment, version and publish the protocol, recruitment window and
eligibility criteria: US healthcare organizations with a documented recurring
screening process. Recruit across organization sizes and disclose channels,
invitations, responses, exclusions and attrition. Participation is opt-in with
explicit consent for aggregate publication; do not ingest rosters or individual
match evidence. One organization is one unit; repeated monthly observations
are reported separately, never counted as independent organizations.

Collect organization size band, source set, scheduled and completed screening
counts, observation window, staff minutes, potential-match count, reviewed-match
count and resolution time using a fixed dictionary. Report missingness per
field. Define completion rate as completed/scheduled screenings; review rate as
reviewed/potential matches. A zero denominator is unavailable, never zero percent.
Do not call potential matches confirmed exclusions or infer accuracy without an
independently adjudicated reference set.

Use coded organization IDs and a separately secured consent register; restrict
raw access to named researchers, honor withdrawal before aggregate release, and
delete raw submissions and the consent register within 90 days of publication.
Suppress groups smaller than 10 organizations and complementary cells that
would allow reconstruction. Publish only aggregates with per-metric sample
sizes, date windows, medians and interquartile ranges where appropriate, and
selection, self-reporting and missing-data limitations. Review disclosure risk
and calculations before release. Publish CSV aggregates, accessible HTML tables,
exportable charts, methodology and a correction log together. No customer
benchmark findings are claimed until that collection and review are complete.
