The science

Science you can defend.

Structured, multi-method assessment validated to r ≈ 0.65 predictive validity — with explainable, rubric-anchored AI scoring behind every decision.

BPSSIOPISO/IEC 42001ISO/IEC 27001SOC 2GDPRISO 27701
Evidence base

The science behind the number.

Selected findings from the peer-reviewed occupational psychology literature and the leading evidence bodies — referenced inline.

r ≈ .65
Nancy's structured, multi-method battery reaches r ≈ 0.65 predictive validity — roughly 2.4× an unstructured panel interview.
Source · Schmidt & Hunter meta-analysis; Sackett, Zhang & Berry — Journal of Applied Psychology
0.27 0.65 Unstructured Nancy battery
36%
Average reduction in adverse impact across protected groups when validated psychometrics replace narrative interviews — sustained over 24 months.
Source · BPS Psychological Testing Centre · SIOP Principles · Harvard Business School working paper
−36%
72%
Cost-of-hire reduction when AI-led structured interviews replace conventional panels at high volume — without compromising predictive validity.
Source · McKinsey People & Org · Deloitte Global Human Capital Trends · internal Nancy AI benchmark
72% cost reduction
< .05
Every scored item is tested for differential item functioning across protected groups before it ships — items above threshold are revised or retired.
Source · BPS Psychological Testing Centre · SIOP Principles
0.05 threshold Every item tested — all below threshold
Methodology

How Nancy builds a defensible decision.

Six steps, every project — from the first competency conversation to the re-assessment that closes the loop.

  1. Competency design
    SHRM- and BPS-aligned framework specific to the role family.
  2. Battery composition
    Modules selected from the nine-service suite per role.
  3. Adaptive testing
    IRT-calibrated items, dynamic difficulty, proctoring.
  4. AI scoring, human-in-loop
    Auto-scoring with human review and bias checks before release.
  5. IDP / LDP output
    Plans generated per candidate and per cohort.
  6. Re-assess & learn
    Quarterly cadence, programme effect measured against baseline.
Explainable AI

No black box. No facial analysis.

Nancy scores what a candidate says and does against a rubric — never face, tone, or background. Every decision traces back to the evidence that produced it.

  • No facial analysis. Nancy scores language and structured responses — never expression, tone, or appearance.

  • No black-box score. Every score traces to a rubric anchor and a specific evidence excerpt from the transcript.

  • Human-in-the-loop. A trained reviewer checks every AI score against the rubric before it is finalised.

  • Full evidence trail. Transcript, rubric, and reviewer sign-off are preserved for every decision, permanently.

  • DIF testing on every item. Each question is checked for differential item functioning across protected groups, not just at the battery level.

  • Threshold < 0.05. Items that exceed the threshold are revised or retired before they reach a candidate.

  • Audit trail preserved. Every test, result, and revision decision is logged and reviewable — nothing is discarded.

Bias & fairness

Every item is tested before it ships.

Every item is tested for differential item functioning across protected groups against a threshold of < 0.05 — items that fail are revised or retired before deployment.

See how this bias discipline scales for national programmes
Engineered to the standards the world's most consequential decisions rely on
  • BPS
    Governs the psychometric validity of every instrument.
  • SIOP
    Professional principles for personnel selection decisions.
  • ISO/IEC 42001
    AI management system governance for the scoring engine.
  • ISO/IEC 27001
    Information security management across the platform.
  • SOC 2
    Independently audited trust services controls.
  • GDPR
    Lawful basis, consent, and data-subject rights.
  • ISO 27701
    Privacy information management, extending the security baseline.
Questions

The objections we hear from scientists and lawyers.

What does r ≈ 0.65 mean in practice?

r ≈ 0.65 is the correlation between assessment score and later job performance — near the practical ceiling Schmidt & Hunter's meta-analysis found for personnel selection. In plain terms, roughly 2.4× more predictive than an unstructured panel interview.

Is AI scoring a black box?

No. Every score traces to a specific rubric anchor and evidence excerpt from the transcript, and a human reviewer signs off before a decision is finalised.

How is bias monitored?

Every item is tested for differential item functioning (DIF) across protected groups before it ships, against a threshold of < 0.05. Items that fail are revised or retired.

Can decisions survive an audit or legal challenge?

Yes — every transcript, rubric, score, and reviewer decision is preserved in an auditable record, reviewable internally, externally, or under regulatory and legal scrutiny.

Are norms regional?

Yes. Aptitude and cognitive norms are calibrated regionally, including GCC-specific norming, rather than applied from a single global sample — consistent with BPS and SIOP guidance on local validation.

Ready to see the evidence

Bring us your hardest hiring decision.

We'll walk through the validity evidence, the bias controls, and the audit trail behind a live Nancy engagement.