V5.0: Evidence-Based Multi-Model Jury
Three independent models with evidence-weighted scoring and jury consensus logic. Current production system adopted 2026-06-01.
V4.0's anchor overcalibration and single-model bias revealed that rigor requires: (1) evidence-based scoring (not exemplar anchoring), (2) multi-model consensus (not single-point estimates), (3) formal inter-rater validation, and (4) human governance gates.
Three-Model Jury
V5.0 uses three independent LLMs (Claude Sonnet 4.6, GPT-4o, Gemini Pro) without shared context. Each model scores independently on the same evidence package.
Evidence Framework
Evidence classified into 8-tier hierarchy: Peer-reviewed journals → Books → Government reports → Interviews → News → Documentary → Grey literature → AI-derived.
Consensus Logic
- jury_mean: Average of three scores
- jury_spread: Max - Min (indicates agreement strength)
- consensus_strong: Boolean (spread ≤ 5 points)
Human Review Gate
Jury proposals require human review before acceptance based on spread thresholds.
Scoring Outputs (Generated in Parallel)
V5.0 generates two scoring outputs simultaneously for each organization:
- Young's Original Score: 0-10 binary checklist across the 10 criteria. How many of Young's indicators are present?
- Composite Score: 0-100% formula-based: (Breadth ÷ 10) × (Mean Intensity ÷ 10) × 100. Measures evidence-weighted prevalence and intensity.
Both scores are generated from the same jury consensus and evidence package. Users receive both perspectives on each organization.
| Adoption Date | 2026-06-01 |
| Organizations Scored | 565+ (all with jury provenance) |
| Status | ✓ Current Production |