Research System Architecture

An AI-Driven Evaluation Framework

How a rigorous, multi-model system evolved through systematic iteration and principled decision-making

This documentation explains a reusable framework for AI-driven research evaluation scoring methodologies. Rather than domain-specific, this system is designed to work across any research domain requiring: consistent evidence assessment, multi-model consensus, human review governance, and rigorous validation.

Current production system (V5.0 + V6.0): Generates three scoring outputs simultaneously for each organization:

  • Young's Original Score: 0-10 binary checklist (Young & Reed's 10 criteria)
  • Composite Score: 0-100% formula-based (evidence-weighted across all criteria)
  • Lifton's Totalism Score: System-level analysis of ideological totalism (C11 criterion)

All three are generated in parallel by the jury consensus process. Users receive all three perspectives on each organization. These scoring methodologies describe HOW organizations are evaluated, not changes to previously published scores or organizational records.

The system evolved through six versions (V4.0 → V6.1) plus one deliberate rejection (V5.2). Each iteration solved specific problems and introduced new capabilities. This documentation is for researchers, academics, and practitioners interested in building or understanding similar evaluation systems.

The pages below cover two additional things beyond methodology history: how evidence about a candidate organization actually gets gathered, and how the people and organizations named in this dataset are protected and governed.

V4.0Anchor HeuristicSingle ModelV5.0Evidence JuryCurrent ProdV5.1ValidationPilotV5.2Deepseek❌ RejectedV6.0Lifton FrameworkProd ReadyV6.1PermanenceProposed

1. Jury Consensus Mechanism

Every score is generated by independent evaluation from four AI models (Claude, GPT-4o, Gemini, Llama). Models score the same evidence independently without seeing each other's scores, eliminating groupthink and single-model bias. The jury mean becomes the proposed score; jury spread (max - min) measures agreement strength.

  • 0–2 point spread: High confidence; immediate acceptance
  • 3–5 point spread: Moderate confidence; accepted with validation flags
  • 6–20 point spread: Low confidence; triggered evidence re-review by human
  • >20 point spread: Model disagreement signals; evidence re-briefing required

2. Evidence Framework (8-Tier Hierarchy)

Not all sources are weighted equally. Evidence is classified into eight tiers from peer-reviewed journals to AI-derived synthesis. Each tier has calibrated weight in jury scoring. This prevents unsourced claims from inflating scores.

  • Tier 1–2: Peer-reviewed scholarship, government records, court filings
  • Tier 3–4: Investigative journalism, books, institutional documentation
  • Tier 5–6: News reports, interviews, documentary evidence
  • Tier 7–8: Grey literature, social media, AI-derived synthesis

3. Human Review Gate (Non-Negotiable Governance)

No score enters the dataset without human review. AI proposes; humans decide. The review gate checks:

  • Does the score align with the body text evidence?
  • Are N/A designations structurally justified?
  • Do cited sources actually support the claims?
  • Is the assessment consistent across the ideological spectrum?
  • Has jury spread triggered additional evidence review?

4. Immutable Audit Trail & Provenance Chain

Every score includes complete provenance tracking:

  • run_id: Links to the specific jury evaluation run
  • per_model_scores: Individual scores from each of the four models
  • jury_spread: Consensus confidence metric
  • methodology_version: Which version (V5.0, V6.0, etc.) generated the score
  • evidence_citations: Complete source list with tier classification
  • human_review_decision: Accept, modify, or reject
  • timestamp: When the score was finalized
  • score_history: All previous versions preserved (immutable); users can see how assessment evolved

5. Logging & Change History

The system logs every change at the database level:

  • Score modifications: Old score → new score, with timestamp and reason
  • Evidence updates: Which sources were added/removed and why
  • N/A rule changes: When criteria are marked as inapplicable and justification
  • Jury re-runs: When jury was asked to re-evaluate due to spread/evidence questions
  • Human review actions: Accept/modify decisions recorded with reviewer identifier
  • Methodology version bumps: When an org was re-scored under new methodology version

6. Formal Validation Metrics

Every dataset release includes inter-rater reliability statistics:

  • Krippendorff's alpha (≥0.70 threshold): Measures agreement strength across the four models
  • ICC(2,k) correlation: Intraclass correlation for intensity scores
  • Pairwise agreement: How often any two models agreed within ±2 points
  • ECRA statements: Explicit reliability claims suitable for academic publication

7. Why This Architecture Matters

Auditability: Every score is traceable to evidence and decision-makers. This enables peer review, independent verification, and publication in peer-reviewed venues.

Reproducibility: Jury consensus is documented. Others can examine whether the models agreed, whether evidence tier was appropriate, and whether human review was applied fairly.

Immutability: Score history preserves old versions. If methodology changes, past scores are not overwritten—they're preserved for comparison and back-analysis.

Generalizability: This architecture is domain-agnostic. The same system can evaluate organizations, policies, research claims, or any domain requiring evidence-based consensus scoring.