Research System Architecture

Evolution Timeline

How a research evaluation system evolved through systematic iteration, testing, and principled rejection.

V4.0: Single-Model Anchor Heuristic

Problem solved: Establish first systematic scoring approach.

Innovation: Calibration exemplars injected into prompts to anchor scoring range.

Limitation: Anchor overcalibration (~15-20pt ceiling), single-model bias, no evidence provenance.

V5.0: Evidence-Based Multi-Model Jury

Problem solved: V4.0's anchor bias and single-model limitations.

Innovation: Three independent models (Claude, GPT-4o, Gemini) scoring evidence-weighted packages. Jury consensus with formal spread thresholds.

Key decision: Human review gate is non-negotiable. ≥2/3 jury required for acceptance.

Scoring outputs (generated in parallel): Young's Original Score (0-10 binary) + Composite Score (0-100% formula). Both provided for each organization.

Adoption: 2026-06-01. 565+ organizations scored with V5.0 provenance.

V5.1: Formal Validation Metrics

Problem solved: Lack of formal inter-rater reliability (IRR) metrics for academic credibility.

Innovation: Added Llama (fourth model, open-weights). Formal Krippendorff's alpha and ICC(2,k) calculations. ECRA statements on all results.

Status: Pilot phase. Awaiting calibration audit before publication.

V5.2: The Deepseek Experiment (Rejected)

Problem attempted to solve: Could we improve consensus with a fifth diverse model?

Hypothesis: Deepseek (Chinese, open-weights) would add methodological diversity and improve jury agreement.

Results: Jury spread increased 40%, consensus dropped 42%, Krippendorff's α fell to 0.61 (below 0.70 threshold).

Key lesson: Not all diversity improves consensus. Systematic rejection is sometimes the right decision. Deepseek was a consistent outlier, not a productive addition.

V6.0: Framework Extension (Lifton Totalism)

Problem solved: V5 only covered Young & Reed's 10 criteria. Need complementary framework for totalism.

Innovation: Dual-track system. Young track (C1-C10) + Lifton track (C11 totalism). Configural (non-compensatory) scoring combining both signals.

Scoring outputs (generated in parallel): Young's Original Score (0-10 binary) + Composite Score (0-100%) + Lifton's Totalism Score (0-10). Three perspectives provided for each organization.

Validation: Krippendorff's α = 0.81 (≥0.80 threshold met). Young's YCDI concordance r = 0.78 (p < 0.001).

Adoption: 2026-06-12. Production-ready. 536+ organizations fully scored.

V6.1: Permanence-Aware Refinement

Problem solved: V6.0 conflates totalizing intensity with system permanence. US Marines penalized (temporary structure) despite high intensity (9.7).

Innovation: Three-component scoring: Intensity (1-10) × Permanence Multiplier (0.75-1.0) + Behavioral Durability Adjustment (-0.30 to +0.50).

Status: Proposed. Ready for PR review. Awaiting veteran affairs domain expert input.


Core Principles

  • Rigor > Speed: Each version prioritizes methodological validity over rapid deployment.
  • Transparency > Elegance: Document decisions, rejections, and trade-offs openly. Governance is visible.
  • Auditability > Automation: Human review gate is preserved even as jury complexity increases.
  • Generalizability > Domain-Specificity: System architecture is reusable across research domains.

Key Decisions

  • Multi-model > Single-model: Jury consensus reduces single-model bias. But diversity must improve metrics, not worsen them (V5.2 rejection).
  • Evidence tiers are essential: Not all sources are equal. 8-tier hierarchy ensures methodological rigor.
  • Human review is non-negotiable: No fully automated scoring. Governance gates preserve accountability.
  • Formal metrics enable publication: Krippendorff's alpha, ICC, ECRA statements transform jury scoring into publishable research.
  • Framework complementarity works: Young + Lifton dual-track outperforms single framework. Separate components reveal nuance single metrics miss.

Before Testing

  • Define success criteria in advance (e.g., α ≥ 0.70, jury spread must not increase).
  • Pilot on representative sample (≥50 organizations) before committing to production.
  • Measure against multiple dimensions, not just one metric.

If Results Are Negative

  • Reject decisively and document why (like V5.2).
  • Don't force adoption for cosmetic reasons.
  • Systematic outliers are warning signs, not acceptable diversity.

Incremental vs. Revolutionary

  • Small improvements (V5.0 → V5.1) are valid.
  • Complementary frameworks (V5 → V6 adding Lifton) are valid.
  • Refining separable components (V6.0 → V6.1) is valid.
  • But diversification for its own sake (V5.2) is not.

Each methodology is documented separately for deep-dive reference:

V4.0V5.0V5.1V5.2 Case StudyV6.0V6.1