Science you can defend.
Structured, multi-method assessment validated to r ≈ 0.65 predictive validity — with explainable, rubric-anchored AI scoring behind every decision.
BPSSIOPISO/IEC 42001ISO/IEC 27001SOC 2GDPRISO 27701The science behind the number.
Selected findings from the peer-reviewed occupational psychology literature and the leading evidence bodies — referenced inline.
How Nancy builds a defensible decision.
Six steps, every project — from the first competency conversation to the re-assessment that closes the loop.
- Competency designSHRM- and BPS-aligned framework specific to the role family.
- Battery compositionModules selected from the nine-service suite per role.
- Adaptive testingIRT-calibrated items, dynamic difficulty, proctoring.
- AI scoring, human-in-loopAuto-scoring with human review and bias checks before release.
- IDP / LDP outputPlans generated per candidate and per cohort.
- Re-assess & learnQuarterly cadence, programme effect measured against baseline.
No black box. No facial analysis.
Nancy scores what a candidate says and does against a rubric — never face, tone, or background. Every decision traces back to the evidence that produced it.
No facial analysis. Nancy scores language and structured responses — never expression, tone, or appearance.
No black-box score. Every score traces to a rubric anchor and a specific evidence excerpt from the transcript.
Human-in-the-loop. A trained reviewer checks every AI score against the rubric before it is finalised.
Full evidence trail. Transcript, rubric, and reviewer sign-off are preserved for every decision, permanently.
DIF testing on every item. Each question is checked for differential item functioning across protected groups, not just at the battery level.
Threshold < 0.05. Items that exceed the threshold are revised or retired before they reach a candidate.
Audit trail preserved. Every test, result, and revision decision is logged and reviewable — nothing is discarded.
Every item is tested before it ships.
Every item is tested for differential item functioning across protected groups against a threshold of < 0.05 — items that fail are revised or retired before deployment.
See how this bias discipline scales for national programmesThe objections we hear from scientists and lawyers.
What does r ≈ 0.65 mean in practice?
r ≈ 0.65 is the correlation between assessment score and later job performance — near the practical ceiling Schmidt & Hunter's meta-analysis found for personnel selection. In plain terms, roughly 2.4× more predictive than an unstructured panel interview.
Is AI scoring a black box?
No. Every score traces to a specific rubric anchor and evidence excerpt from the transcript, and a human reviewer signs off before a decision is finalised.
How is bias monitored?
Every item is tested for differential item functioning (DIF) across protected groups before it ships, against a threshold of < 0.05. Items that fail are revised or retired.
Can decisions survive an audit or legal challenge?
Yes — every transcript, rubric, score, and reviewer decision is preserved in an auditable record, reviewable internally, externally, or under regulatory and legal scrutiny.
Are norms regional?
Yes. Aptitude and cognitive norms are calibrated regionally, including GCC-specific norming, rather than applied from a single global sample — consistent with BPS and SIOP guidance on local validation.
Bring us your hardest hiring decision.
We'll walk through the validity evidence, the bias controls, and the audit trail behind a live Nancy engagement.
