Overview
At a glance
- Two markers: entropy production · irreversibility
- Fixed 12-study cohort, 24-item matrix
- Source-traceable accepted records
- Blinded three-stream IRR follow-up
- 75% agreement; κ 0.50–0.52
A conservative measurement-layer audit protocol for two explicitly defined markers — obs_entropy_production_rate and obs_irreversibility — applied to a fixed 12-study cohort.
The completed output is a 24-item study × marker matrix built from accepted records produced through machine extraction, source-PDF review, and deterministic export. Within the audited cohort, explicit irreversibility is broader than explicit entropy-production measurement, and all entropy-production-positive cases fall within the irreversibility-positive subset: entropy production is positive in 4 of 12 studies, irreversibility in 9 of 12, with no entropy-positive / irreversibility-negative cases.
A separate post-collection inter-rater reliability follow-up was conducted on the same 24 items. Three retained streams — a blinded GPT baseline run, a refined-instrument GPT rerun, and a Claude Opus sensitivity run — each produced 18 matches out of 24, or 75 percent agreement, with kappa values from 0.50 to 0.52. The main contribution is methodological: the protocol separates marker measurement from thematic inference and makes each accepted record source-traceable.
What it claims / What it does not claim
What it claims
A nested pattern, within this cohort. In the audited 12 studies, explicit irreversibility is the broader marker and every entropy-production-positive case falls within the irreversibility-positive subset — a descriptive result about the audited cohort.
A reproducible protocol. Fixed marker rules can be applied conservatively and traceably across heterogeneous studies, with marker measurement kept separate from thematic inference and every accepted record source-traceable.
What it does not claim
No population-level estimate. The cohort was kept deliberately small for auditability and rule stabilization; the matrix should not be generalized beyond it, and larger-cohort testing is a distinct next-stage robustness exercise.
Not human IRR. The reliability follow-up is independent AI-assisted secondary rating under firewall controls, and the kappa values are read cautiously given the small 24-item matrix and uneven marker prevalences.
Figures
