Public evaluation

Validation & Benchmarks

The demo suite is designed for blind diagnostic evaluation: the assistant sees evidence, not the planted answer, and the evaluator judges whether the reasoning is actually supported.

Five evidence domains

Different failure shapes, same evidence discipline.

The public suite spans more than 20,000 miner/evidence records across CSV, XLSX, JSONL, and NDJSON.

01

Fleet restart / recovery

Startup and recovery telemetry with Unknowns and explicit source units.

02

Network reachability

Shared segment evidence designed to separate unreachable devices from reachable zero-hash states.

03

Power & thermal

Topology, feed loading, and thermal evidence with competing infrastructure hypotheses.

04

Repair history

Longitudinal events for recurring-fault and repair-outcome reasoning.

05

Miner-log triage

Individual synthetic cases across healthy, startup, thermal, pool, PSU, hashboard, and other symptoms.

What a good run does

Accuracy is necessary, but not sufficient.

Extract faithfully

Preserve units, Unknown values, source meaning, and factual distinctions such as unreachable versus reachable zero-hash.

Reason transparently

Use useful derived metrics with visible denominators, recognize shared versus miner-specific patterns, and resist premature component diagnosis.

Reduce uncertainty

State confidence appropriately and finish with one discriminating read-only next check.

Current validation status

Claims should match evidence.

Synthetic demo suiteSIMULATED
Evidence methodologyTESTED FROM REAL EXPORT
Codex / Agent Skill packageSTRUCTURALLY VALIDATED
Claude pluginCLAUDE CLI VALIDATED
Physical miner integrationNOT PART OF v1.1.0

The repository intentionally does not publish hidden scoring answers or planted-case answer keys.

Blind use

Do not tell the assistant the planted pattern.

Use the public demo prompts, provide the evidence file, and judge whether the reasoning follows the supplied evidence rather than whether the assistant can restate an answer key.

Questions about validation, private pilots, or benchmark collaboration can be sent to austin@wnclogiclab.com.

View technical benchmark document →