Assessment network
Current work tracks
- AI Operations Control
- Executive Assistant Precision
- Research Verification
- Support Triage & Judgment
- Data Quality Control
- Client Communication
- Project Operations
- E-commerce Operations
- Sales Operations Control
- Finance Operations Control
- Recruiting Operations Control
- Security & Access Operations
- Marketing Operations Control
- Editorial & Content Quality
- Customer Success Operations
- Procurement & Vendor Control
- People Operations Control
- Compliance & Policy Operations
- Logistics & Fulfillment Control
- Automation & No-Code Operations
Session selection
Each session receives five tasks from the selected eight-task bank. Task order, option order and public identifiers are randomized for the session.
The browser receives only the text and opaque option IDs required to answer. Authoritative answer mappings remain server-side.
Score dimensions
Accuracy — 30
Detects material errors, inconsistencies and unsafe outputs.
Judgment — 25
Measures prioritization and decision quality under competing signals.
Instruction following — 20
Measures whether explicit constraints survive execution.
Verification — 15
Measures rejection of claims not justified by the supplied evidence.
Communication — 10
Measures clear, bounded communication without invented commitments.
Authoritative scoring
Scoring is deterministic and server-side. Single-choice tasks require the correct selection. Multi-select tasks award bounded partial credit for correct selections and penalize false positives.
A stored session variant and the same answer set produce the same score. Material changes to tasks or scoring receive a new assessment version rather than rewriting historical WorkProofs.
Integrity controls
- Cloudflare Turnstile before protected session allocation and final submission
- single-use, short-lived server sessions
- client/session fingerprint binding
- server-side rate limiting and daily retake limits
- server-only answer keys and opaque per-session option IDs
- unique session-to-attempt relationship to reject replay
Score bands
- 90–100 — Exceptional control
- 80–89 — Strong operator
- 70–79 — Capable
- 60–69 — Developing
- Below 60 — Needs stronger judgment
No invented percentile. No invented predictive validity.
Mandorka does not publish percentile rankings before a meaningful comparison cohort exists and does not claim that WorkProof predicts future job performance until validation evidence justifies that claim.
Public standard and validation
The WorkProof Standard publishes proof requirements and responsible-use boundaries. The Validation Program tests the signal against real job-relevant outcomes. WorkProof Standard · Validation Program
Limitations
A WorkProof records observed performance in a defined simulation. It does not establish identity, background, long-term reliability, every domain skill or guaranteed future work performance. It is evidence for a decision, not the decision itself.