MANDORKAWORKPROOF
THE MANDORKA INDEX
● CALIBRATION PHASE

Measure the quality of AI-assisted work.

The Mandorka Index is intended to become a public benchmark of observed work-control signals by role and track. We will not publish a dramatic number just because a dashboard can calculate one.

No index score published yetVersion-aware cohortsTrack-level reportingUncertainty disclosed
WHAT WILL BE MEASURED

Not “who uses AI.” Who controls the output.

ERROR CONTROL

Accuracy

How reliably people detect factual, numerical and operational errors before downstream use.

DECISION QUALITY

Judgment

How consistently they prioritize material risk, impact and constraints under pressure.

CONSTRAINT CONTROL

Instruction following

Whether explicit requirements survive AI-assisted drafting and execution.

EVIDENCE CONTROL

Verification

Whether unsupported claims are challenged instead of converted into confident output.

OUTPUT CONTROL

Communication

Whether the final message is clear, bounded and appropriate to the situation.

INTEGRITY

Assessment conditions

Published cohorts will distinguish legacy records from current session controls and assessment versions.

PUBLICATION GATE

No benchmark until the benchmark deserves trust.

GATE A

Enough usable observations

Cohorts must be large enough to make the summary meaningful rather than anecdotal.

GATE B

Comparable versions

Materially different assessment versions will not be silently pooled into one headline number.

GATE C

Quality checks

Abuse, incomplete sessions and data-quality problems must be excluded by documented rules.

GATE D

Validation context

Where outcome studies exist, descriptive benchmark data will be separated from predictive-validity findings.

DATA FLYWHEEL

Every valid proof can improve the instrument.

More completed assessments create better item diagnostics and calibration opportunities. Better calibration can improve employer utility. Employer use can create stronger validation data. That loop is the long-term Mandorka moat — not selling raw participant data.