System Overview
A comprehensive, auditable, multi-model evaluation framework designed for rigorous research at scale
Multi-Model Jury
Independent Consensus
Four independent AI models (Claude, GPT-4o, Gemini, Llama) score the same evidence independently. Jury consensus calculated via mean, median, spread. Models never see each other's outputs, eliminating groupthink.
Evidence Framework
8-Tier Hierarchy
Not all sources are equal. Evidence classified into 8 tiers from peer-reviewed journals to AI-derived synthesis. Standardized tier assignment ensures methodological consistency across 1000+ organizations.
Dual-Track Scoring
Complementary Frameworks
Two independent frameworks run in parallel: Young's behavioral indicators (C1-C10) + Lifton's system-level totalism (C11). Divergence between tracks reveals new insights; both provided to users.
Human Review Gate
Non-Negotiable Governance
All jury proposals require human review before acceptance. Spread thresholds determine routing: 0-5pt strong consensus → accept; 6-20pt → review criteria; >20pt → revise evidence. Preserves human judgment.
Complete Auditability
Provenance Chain
Every score includes: run_id linking to jury votes, methodology_version, per-model scores, evidence citations, human review decision, timestamp. Score history immutable. Full traceability from decision to evidence.
Formal Validation
Research-Grade Metrics
Inter-rater reliability calculated via Krippendorff's alpha (≥0.70 threshold), ICC(2,k) correlation, pairwise agreement analysis. All results include ECRA reliability statements for publication readiness.
Rigor Without Single Points of Failure
- Multi-model consensus: Three independent AI models eliminate single-model bias. Jury spread measures agreement strength; high spread triggers deeper review.
- Evidence-based, not heuristic: Organizations scored against framework, not calibration anchors. No ceiling effects or systematic inflation.
- Human governance: Every score passes human review gates. System amplifies human judgment, not replaces it.
Transparent & Auditable
- Complete provenance chain: Trace any score back to evidence sources, jury votes, and human decision.
- Immutable history: Score history never overwrites; tracks how assessments evolve as new evidence emerges.
- Publication-ready: Validation metrics (α, ICC) and ECRA statements included; ready for peer review and academic citation.
Generalizable Across Domains
- Framework-agnostic: The system architecture is reusable. Swap "Young's 10 criteria" for any other framework; the jury, evidence, governance, and validation machinery stays the same.
- Extensible: Dual-track and configural scoring allow multiple frameworks to run in parallel without conflict.
- Documented evolution: See how the system improved over six versions; understand the reasoning behind each methodology choice.
Current System (V4.0+)
Deep dive into modern methodology versions with full technical details, validation results, and trade-offs.
Historical Evolution (V0-V3)
Understand the system's origins and how it evolved from Young & Reed's framework through manual assessment to automated jury consensus.