Public falsification benchmark

A controlled test of evidentiary-state measurement.

A synthetic medication-administration workflow with known evidence, known constraints, and a predetermined answer key. The same 80 operational events are evaluated across six evidentiary stages.

The records do not change. Only the admitted evidence used to interpret those records changes.

80operational events
6evidentiary states
480predetermined classifications
480 / 480matched
The scenario

A medication workflow records that administration was verified. A separate infusion-pump system records when infusion began. Both are ordinary operational records produced by different systems.

The material relationship
Medication administration verification must be established before the infusion pump starts.
The question we measure

Given the admitted evidence, has the required ordering been established to a single operationally defensible state, or do materially different states remain supported?

The experiment

Same 80 events. Six evidentiary stages.

Five additional pieces of evidence characterize the behavior of one medication-administration workstation clock, MED-WS-11. They are admitted sequentially. After each evidence addition, Hartford Logic evaluates the same 80 events again.

S080 / 0No additional clock evidence
S170 / 10Pre-correction characterization
S265 / 15Breakpoint + post-correction evidence
S370 / 10Better post-correction evidence
S475 / 5Additional local pre-correction evidence
S580 / 0Strong near-window synchronization evidence

More evidence did not monotonically produce more resolution.

The number of unresolved events first increased as additional evidence exposed uncertainty that the baseline record did not contain, then decreased as later evidence constrained the admissible states. The system is evaluating what the currently admitted evidence permits.

Worked example

One event. Same records. Different admitted evidence.

The example below uses one MED-WS-11 event to show how the classification changes when the admitted evidence changes.

Recorded times
Medication verification17:10:32recorded time
Infusion pump start17:10:35recorded time
Stage 2

Materially different temporal states remain supported.

Offset ≤ ±3.20 s
Rate ≤ ±60 ppm

R(t) = O + P|t − tₐ|
R = 3.20 + (60 × 10⁻⁶ × 32)
R = 3.20192 s

Admissible verification interval:

17:10:28.79808 → 17:10:35.20192

The admissible interval crosses the 17:10:35 pump-start boundary.

Stage 3

One material temporal state is supported.

Offset ≤ ±2.40 s
Rate ≤ ±60 ppm

R(t) = O + P|t − tₐ|
R = 2.40 + (60 × 10⁻⁶ × 32)
R = 2.40192 s

Admissible verification interval:

17:10:29.59808 → 17:10:34.40192

Every admissible verification time is before the 17:10:35 pump-start boundary.

The records did not change. The event did not change. The evidence changed what could be established about the event.
The result
Predetermined480classifications
Observed480classifications
480 predetermined classifications. 480 observed classifications. 480 matches.

What the result establishes

Under the published evidence requirements and constraints, Hartford Logic matched all 480 classifications predetermined for the controlled benchmark before first execution. Repeated execution with unchanged evidence and configuration produced identical substantive results.

Scope: This is a controlled test of deterministic evidence-state measurement. It is not a validation of every evidentiary domain or a real-world clinical determination.

Why temporal evidence?

Temporal uncertainty provides a precise, bounded, and independently calculable way to test whether changing evidence changes what a fixed record can support.

The clock is the experimental mechanism. Evidence-state measurement is what is being tested.

When multiple published temporal constraints apply to an event, their admissible intervals are intersected. For this controlled benchmark, pump-start timestamps are treated as the reference side with no additional temporal uncertainty.

Benchmark demonstration

Watch the published evidence move through the measurement.

Hartford Logic v1.0 Demo

This 4:37 demonstration uses the same operational records and temporal evidence contained in Benchmark v1.0. It shows the benchmark evidence being admitted across stages and the resulting RESOLVED and MULTIPLE STATES classifications. The demonstration interface exposes what is needed to understand the measurement without exposing Hartford Logic's internal implementation.

Public challenge

Don't take our word for it.

The public benchmark is designed to be challenged without access to Hartford Logic's implementation. The records, admitted evidence, mathematical specification, predetermined answer key, worked examples, validation material, and cryptographic integrity information are available in the public package.

You do not need Hartford Logic's implementation to challenge the published classifications.
Challenge a one-state classificationProduce a materially different operational state satisfying the same admitted evidence and constraints.
Challenge a multiple-state classificationDemonstrate that the same evidence and constraints force a unique material state.
From controlled evidence to external evaluation

The next phase is external.

This benchmark was designed to make the measurement inspectable and falsifiable under controlled conditions. The next phase is evaluating where the same measurement adds value in real evidence-dependent workflows.

Technical evaluation, integration, and licensing inquiries: tim@hartfordlogic.com.

HARTFORD LOGICWHERE EVIDENCE ENDS AND JUDGMENT BEGINS.