# The information in diagnostic tests: retrospective protocol record

**Record version:** 1.0.2 | **Prepared:** 2026-09-08 | **Source manuscript:** 1.10.1

**Status:** Prepared draft; not registered. This records completed work and awaits owner confirmation before registry submission.

## Research question

Across reproducibly accessible diagnostic accuracy reviews, how much uncertainty about a binary target condition is a test or threshold expected to resolve at a stated starting probability, and how does this information relate to conventional test-performance measures?

## Rationale

Diagnostic testing changes what is known about a target condition. Sensitivity and specificity describe test performance; likelihood ratios and post-test probabilities describe an observed result. Mutual information expresses the expected reduction in uncertainty before either result is known. Earlier studies established diagnostic applications of this measure. This study assembles a broad empirical reference scale from accessible diagnostic-review data so readers can interpret a diagnostic profile at a stated starting probability and compare its information with conventional accuracy measures. Information describes learning about a binary target. Clinical value also depends on consequences, available actions, and patient preferences.

## Objectives

- Build an empirical reference distribution of expected diagnostic information from source-defined tests and thresholds.
- Compare uncertainty resolved at 5%, 20%, and 50% starting disease probability and across a 1%-50% grid.
- Examine agreement and differences between diagnostic information and Youden's J using source-comparable examples.
- Distinguish expected information before testing from the probability update and uncertainty change after each result.

## Scope and known results

This is a diagnostic evidence map and secondary methodological synthesis. The completed analysis contains 273 pooled estimates from 210 reviews and 4,104 study-result appearances. At 20% starting probability, the median profile resolved 29.2% of starting uncertainty. These findings were known when this retrospective record was prepared; the accessible sample does not represent all diagnostic medicine.

## Eligibility

### Include

- Diagnostic-review data with explicit study identifiers, source links, and non-negative integer TP, FP, FN, TN counts.
- Groups with at least five unique study identifiers, one result per identifier, and at least three studies with both disease states represented.
- Aggregate sensitivity plus specificity at least one in the source-defined orientation, with one canonical representation after exact row-signature deduplication.
- Historical Cochrane groups linked to review-author-reported pooled findings; modern Cochrane main analyses; at most one canonical main table per OSF, Zenodo, or PMC review under the recorded priority rules.

### Exclude

- Too few studies, repeated identifiers within a group, too few studies representing both disease states, or below-chance performance in the original orientation.
- Counts reconstructed from rounded accuracy measures, pooled-only or consequence tables, and unsupported narrative transcription.
- Duplicate representations and alternatives not selected by source-specific primary designation; preserve their status in the audit record.

## Outcomes

- Distribution of 100 x I(D;T)/H(D) at 5%, 20%, and 50% starting probabilities, with corresponding information in bits; each pooled estimate has equal weight.
- Information profiles from 1% to 50% in one-percentage-point steps.
- Spearman correlations with Youden's J and descriptive near-tie comparisons.
- CEA information crossings and Ottawa ankle rule result branches.
- Sensitivity to zero-cell correction and the regularized direct-binomial estimator.

## Sources and exact retrieval strategies

### Fixed historical and modern Cochrane sources

Cochrane reference dataset and published review packages

Retrieve CL145_open_set_20181101.zip from https://zenodo.org/records/1303259 (SHA-256 f921cf6f1a9236c86f89e7c646497639f45e8ff8fcf56008c1b27422829fec4f), and the already-held published packages for CD014911 and CD015089. Parse available study-level 2x2 tables with scripts/cochrane_dta_atlas.py at the pinned commit. This is fixed-corpus acquisition, not a new exhaustive Cochrane search.

63 historical reviews and two selected modern reviews. The count and checksum identify the normalized corpus manifest. The original acquisition date needs owner confirmation.

Saved export: https://github.com/ShuhanCS/diagnostic-information/blob/0c8cff2d06233e0fb7dad4e3067e69c15ca6e5c7/analysis/cochrane_atlas/corpus_manifest.csv

Records: 65; recorded run date: not confirmed. SHA-256: `1a17f47b78c0e2ba7c2f326111e18d2a387d04780b2a36e7b7fe54a874f51321`.

### OSF public deposits

Frozen OSF metadata inventory and public-file inspection

Run each recorded query separately:
"diagnostic test accuracy"
"diagnostic accuracy"
"sensitivity and specificity"
sensitivity specificity meta-analysis
"true positive" "false positive"
"2x2" diagnostic

Count and checksum identify the saved normalized discovery manifest, not a new search or included-review count. Only publicly retrievable structured diagnostic review data were eligible. OSF entries here are other researchers' evidence, not registrations of this study.

Saved export: https://github.com/ShuhanCS/diagnostic-information/blob/0c8cff2d06233e0fb7dad4e3067e69c15ca6e5c7/analysis/open_repository_atlas/osf/source_manifest.csv

Records: 1152; recorded run date: 2026-08-17. SHA-256: `1f46eedb7af4d6d0724d55156ba98ebb92572f3b0c4378e521400d2cd634ec44`.

### Zenodo public deposits

Zenodo public metadata API

Run each recorded query separately:
"diagnostic test accuracy"
"diagnostic accuracy" AND "systematic review"
"diagnostic accuracy" AND "meta-analysis"
"test accuracy" AND "meta-analysis"
"sensitivity and specificity" AND "meta-analysis"
"true positive" AND "false positive" AND "meta-analysis"
"2x2" AND diagnostic AND "meta-analysis"

Count and checksum identify the saved normalized discovery manifest, not a new search or included-review count. Only publicly retrievable structured diagnostic review data were eligible. OSF entries here are other researchers' evidence, not registrations of this study.

Saved export: https://github.com/ShuhanCS/diagnostic-information/blob/0c8cff2d06233e0fb7dad4e3067e69c15ca6e5c7/analysis/open_repository_atlas/zenodo/source_manifest.csv

Records: 799; recorded run date: 2026-08-17. SHA-256: `0fa8e301ef599db6657089e96ef341e816f790e444a2c5fe28da202d4e61a44c`.

### Figshare public deposits

Figshare public metadata API

Run each recorded query separately:
"diagnostic test accuracy"
"diagnostic accuracy" AND "systematic review"
"diagnostic accuracy" AND "meta-analysis"
"test accuracy" AND "meta-analysis"
"sensitivity and specificity" AND "meta-analysis"
"true positive" AND "false positive" AND "meta-analysis"
"2x2" AND diagnostic AND "meta-analysis"

Count and checksum identify the saved normalized discovery manifest, not a new search or included-review count. Only publicly retrievable structured diagnostic review data were eligible. OSF entries here are other researchers' evidence, not registrations of this study.

Saved export: https://github.com/ShuhanCS/diagnostic-information/blob/0c8cff2d06233e0fb7dad4e3067e69c15ca6e5c7/analysis/open_repository_atlas/figshare/source_manifest.csv

Records: 262; recorded run date: 2026-08-17. SHA-256: `b38dfd72097e51a97cc52f0d679097e3897e29249ad7a7cf11c50a14bb0efb4f`.

### Dryad public deposits

Dryad public metadata API

Run each recorded query separately:
"diagnostic test accuracy"
"diagnostic accuracy" AND "systematic review"
"diagnostic accuracy" AND "meta-analysis"
"test accuracy" AND "meta-analysis"
"sensitivity and specificity" AND "meta-analysis"
"true positive" AND "false positive" AND "meta-analysis"
"2x2" AND diagnostic AND "meta-analysis"

Count and checksum identify the saved normalized discovery manifest, not a new search or included-review count. Only publicly retrievable structured diagnostic review data were eligible. OSF entries here are other researchers' evidence, not registrations of this study.

Saved export: https://github.com/ShuhanCS/diagnostic-information/blob/0c8cff2d06233e0fb7dad4e3067e69c15ca6e5c7/analysis/open_repository_atlas/dryad/source_manifest.csv

Records: 11; recorded run date: 2026-08-17. SHA-256: `e2e5e7913d05a578b36fbaaca5f107bf5d2bb1d0b8905ad2953c43e4e1a5d52f`.

### SRDR+ public deposits

DataCite SRDR metadata and fixed public SRDR inventory

publisher:"Systematic Review Data Repository"; retain the fixed inventory of 172 projects and apply the source-specific discovery and diagnostic-candidate rules in scripts/open_repository_dta.py.

Count and checksum identify the saved normalized discovery manifest, not a new search or included-review count. Only publicly retrievable structured diagnostic review data were eligible. OSF entries here are other researchers' evidence, not registrations of this study.

Saved export: https://github.com/ShuhanCS/diagnostic-information/blob/0c8cff2d06233e0fb7dad4e3067e69c15ca6e5c7/analysis/open_repository_atlas/srdr/source_manifest.csv

Records: 172; recorded run date: 2026-08-17. SHA-256: `0576f936db6c27b7cde5147f346289b08c399d7ef90e5f96902a066c669ff88a`.

### PubMed Central open-access review tables

NCBI E-Utilities and PMC JATS XML

(("diagnostic test accuracy"[Title/Abstract]) OR ("diagnostic accuracy"[Title/Abstract])) AND ("2x2"[All Fields] OR "2 x 2"[All Fields] OR ("true positive"[All Fields] AND "false positive"[All Fields] AND "false negative"[All Fields] AND "true negative"[All Fields])) AND (systematic review[Title/Abstract] OR meta-analysis[Title/Abstract]) AND open access[filter] AND 2020:2026[Publication Date] NOT protocol[Title]

The fixed query returned 1,518 articles on 2026-08-17; all were retrieved. Only explicit identifiers and integer TP, FP, FN, TN columns were normalized. Count and checksum identify articles.csv.

Saved export: https://github.com/ShuhanCS/diagnostic-information/blob/0c8cff2d06233e0fb7dad4e3067e69c15ca6e5c7/analysis/pmc_open_atlas/articles.csv

Records: 1518; recorded run date: 2026-08-17. SHA-256: `c0dbf4a9a700ef837e76b1f8551a1cb09c237d93a1ce1c8b9465ff731b81f8f3`.

## Coverage and limits

Public OSF, Zenodo, Figshare, Dryad, and SRDR deposits were examined for reusable diagnostic-review data. Dryad and SRDR contributed no primary pooled estimates. This accessible evidence sample is not an exhaustive search of unpublished studies or all diagnostic medicine.

## Deduplication

Retain source identifiers and compare exact normalized count-row signatures. Keep one result per identifier within a group; repeated identifiers make that group ineligible. The same study may contribute to different diagnostic profiles, so study-result appearances are not unique studies.

## Full-text and table selection

Inspect structured review deposits and machine-readable PMC tables for explicit study-level counts. Preserve article, table, file and review identifiers and exclusion reasons. No complete two-reviewer full-text screening process is claimed.

## Data extraction

Parse explicit study identifiers and integer TP, FP, FN, TN cells while retaining source location, orientation and checksums. Do not reconstruct counts from rounded accuracy metrics. Apply canonical workbook/table rules. Compare fitted estimates with usable source summaries and retain the recorded author audit; that audit is not independent adjudication.

## Risk of bias

No new uniform study-level QUADAS-2 assessment across the corpus is documented. Incomplete source risk-of-bias and clinical-context reporting limit interpretation. This record does not invent appraisal judgments or equate numerical model checks with risk-of-bias assessment.

## Statistical synthesis

Pool each eligible source-defined test or threshold separately with bivariate random-effects REML on sensitivity/specificity logits, adding 0.5 to all four cells in every study. Back-transform the fitted mean logits. Compute entropy, mutual information and percentage uncertainty resolved at the specified starting probabilities. Give each pooled estimate equal reference-distribution weight. Report pooled-mean intervals and plug-in predictive ranges conditional on fitted heterogeneity. Bootstrap reference medians over pooled estimates with 2,000 percentile replicates and a fixed seed; this does not propagate within-estimate uncertainty. Repeat analyses with correction only in zero-containing studies and with the regularized direct-binomial estimator. Treat within-review contrasts and CEA crossings as descriptive; no paired superiority inference or clinical utility ranking is supported. Distinguish signed entropy change, posterior-to-prior KL divergence, log2 likelihood ratio, and expected mutual information.

## Certainty

No new GRADE ratings are claimed for this methodological reference scale. Report source fidelity, numerical diagnostics, estimator sensitivity and sample limitations separately from certainty about clinical effects or recommendations.

## Timing and methods history

Retrospective draft prepared on 2026-09-08 after analysis and manuscript preparation. Commit 857cd0f on 2026-08-16 documents an executed Cochrane atlas synthesis. The synthesis field records this evidenced run as a provisional timing anchor, not a confirmed first-ever analysis date. Formal search, screening, extraction and original planned dates remain unconfirmed and unfilled. Repository/PMC expansion runs are dated 2026-08-17; manuscript 1.10.1 was committed on 2026-08-30. This record is not prospective. Primary-designation rules developed during staged audits were fixed before the final reference analysis; this does not establish that they preceded all data inspection. The owner must confirm actual milestones before submission. PRISMA-P organizes this retrospective record and does not imply prior adherence or registration.

## Screening

One author-directed computational selection workflow and a recorded author audit are documented. Two independent human screeners are not established. Fixed source priorities resolved representations; ambiguous groups were excluded or retained outside the primary sample. The owner must confirm this retrospective description.

Codex assisted retrieval, parsing and code drafting. Deterministic rules selected profiles; agent outputs do not constitute independent human screening.

## People and funding

The signed-in Publishing account identifies the protocol owner. Contact details are kept in the private workspace.

The author list and contribution details are being prepared. Approval of this exact protocol remains pending.

National Academy of Medicine, Agreement No. 2026A008797. The manuscript states that the funder had no role in design, acquisition, analysis, interpretation, writing, or submission.

The research team will confirm the contributor list and complete its declarations of interests before formal submission.

## Automation and source history

This new record has not received owner or coauthor approval. The manuscript reports SH review of earlier AI-assisted work; that does not approve this protocol. Source repositories are private unless access is granted. The proposed CC-BY-4.0 license applies only to protocol text after owner acceptance, not third-party data. Cochrane data retain their original restrictions. Unfilled dates require confirmation. No registration receipt, DOI, independent peer review, or prospective registration is claimed.

The pinned source commit is `0c8cff2d06233e0fb7dad4e3067e69c15ca6e5c7`. Methods and source links are in `protocol.json`; the contributor list remains in preparation; export counts and checksums are in `source-evidence.json`.

## Complete the registration

1. Confirm the original planned dates and actual search, screening, extraction, and synthesis start dates. The synthesis anchor currently records an evidenced run, not an attested onset.
2. Confirm the Cochrane acquisition date and source-specific retrieval descriptions.
3. Confirm contributors, disclosures, protocol-text license and human-reviewed fields.
4. Import `protocol.json` in the existing Systematic Reviews form, resolve highlighted fields, inspect the exact hash and approve through the human owner account.
5. Complete the OSF provider step and verify its public URL and DOI before adding a registration statement to the manuscript.

This branch uses OSF as the authoritative registry. A local draft, Git commit, PDF or example is not itself registration. OSF distinguishes registration from preregistration and permits registration at later stages: https://help.osf.io/article/330-welcome-to-registrations

## Inspect or import

From the repository root:

```powershell
node reviews-cli/src/conduct.mjs reviews validate frontend/public/review-pilots/diagnostic-information/protocol.json
node reviews-cli/src/conduct.mjs reviews verify frontend/public/review-pilots/diagnostic-information/protocol.json
```

Validation will identify the unconfirmed dates; no values were invented to force a passing result. Open the branch preview at `/systematic-reviews` and choose **Import JSON**. This record uses the existing general schema and `mapping_review` type.

## Draft correction in record 1.0.2

The author list is in preparation following the owner's correction. Public contact details were removed. Methods, recorded searches, and unconfirmed milestone dates are unchanged. This correction is not an author attestation or a registered amendment.
