Research System Architecture

System Overview

A comprehensive, auditable, multi-model evaluation framework designed for rigorous research at scale

Input Layer: Organizations + Evidence PackagesCSV/API Feeds • Manual Intake • External Data SourcesEvidence Framework: 8-Tier Source HierarchyPrimary Sources → Secondary → Grey Literature → AI-DerivedMulti-Model Jury Consensus (4 Independent Models)ClaudeGPT-4oGeminiLlamaParallel Scoring TracksYoung TrackC1-C10 BehaviorScore: 0-10 & 0-100%Lifton TrackC11 TotalismSystem PermanenceConfigural TrackStructural AnalysisExternal MappingHuman Review Gate & GovernanceConsensus Validation • Spread Thresholds • Acceptance/Rejection • Audit TrailSupabase (Source of Truth)Three Scoring Outputs • Provenance Chain • Validation Metrics • Immutable Score History

Multi-Model Jury

Independent Consensus

Four independent AI models (Claude, GPT-4o, Gemini, Llama) score the same evidence independently. Jury consensus calculated via mean, median, spread. Models never see each other's outputs, eliminating groupthink.

Evidence Framework

8-Tier Hierarchy

Not all sources are equal. Evidence classified into 8 tiers from peer-reviewed journals to AI-derived synthesis. Standardized tier assignment ensures methodological consistency across 1000+ organizations.

Dual-Track Scoring

Complementary Frameworks

Two independent frameworks run in parallel: Young's behavioral indicators (C1-C10) + Lifton's system-level totalism (C11). Divergence between tracks reveals new insights; both provided to users.

Human Review Gate

Non-Negotiable Governance

All jury proposals require human review before acceptance. Spread thresholds determine routing: 0-5pt strong consensus → accept; 6-20pt → review criteria; >20pt → revise evidence. Preserves human judgment.

Complete Auditability

Provenance Chain

Every score includes: run_id linking to jury votes, methodology_version, per-model scores, evidence citations, human review decision, timestamp. Score history immutable. Full traceability from decision to evidence.

Formal Validation

Research-Grade Metrics

Inter-rater reliability calculated via Krippendorff's alpha (≥0.70 threshold), ICC(2,k) correlation, pairwise agreement analysis. All results include ECRA reliability statements for publication readiness.


Rigor Without Single Points of Failure

  • Multi-model consensus: Three independent AI models eliminate single-model bias. Jury spread measures agreement strength; high spread triggers deeper review.
  • Evidence-based, not heuristic: Organizations scored against framework, not calibration anchors. No ceiling effects or systematic inflation.
  • Human governance: Every score passes human review gates. System amplifies human judgment, not replaces it.

Transparent & Auditable

  • Complete provenance chain: Trace any score back to evidence sources, jury votes, and human decision.
  • Immutable history: Score history never overwrites; tracks how assessments evolve as new evidence emerges.
  • Publication-ready: Validation metrics (α, ICC) and ECRA statements included; ready for peer review and academic citation.

Generalizable Across Domains

  • Framework-agnostic: The system architecture is reusable. Swap "Young's 10 criteria" for any other framework; the jury, evidence, governance, and validation machinery stays the same.
  • Extensible: Dual-track and configural scoring allow multiple frameworks to run in parallel without conflict.
  • Documented evolution: See how the system improved over six versions; understand the reasoning behind each methodology choice.

Current System (V4.0+)

Deep dive into modern methodology versions with full technical details, validation results, and trade-offs.

V4.0: Anchor HeuristicV5.0: Evidence JuryV5.1: ValidationV5.2: Case StudyV6.0: LiftonV6.1: Permanence

Historical Evolution (V0-V3)

Understand the system's origins and how it evolved from Young & Reed's framework through manual assessment to automated jury consensus.

V0: Young & ReedV1: Dual-MetricV3: Streamlit EraFull Timeline