Evolution Timeline
How a research evaluation system evolved through systematic iteration, testing, and principled rejection.
V4.0: Single-Model Anchor Heuristic
Problem solved: Establish first systematic scoring approach.
Innovation: Calibration exemplars injected into prompts to anchor scoring range.
Limitation: Anchor overcalibration (~15-20pt ceiling), single-model bias, no evidence provenance.
V5.0: Evidence-Based Multi-Model Jury
Problem solved: V4.0's anchor bias and single-model limitations.
Innovation: Three independent models (Claude, GPT-4o, Gemini) scoring evidence-weighted packages. Jury consensus with formal spread thresholds.
Key decision: Human review gate is non-negotiable. ≥2/3 jury required for acceptance.
Scoring outputs (generated in parallel): Young's Original Score (0-10 binary) + Composite Score (0-100% formula). Both provided for each organization.
Adoption: 2026-06-01. 565+ organizations scored with V5.0 provenance.
V5.1: Formal Validation Metrics
Problem solved: Lack of formal inter-rater reliability (IRR) metrics for academic credibility.
Innovation: Added Llama (fourth model, open-weights). Formal Krippendorff's alpha and ICC(2,k) calculations. ECRA statements on all results.
Status: Pilot phase. Awaiting calibration audit before publication.
V5.2: The Deepseek Experiment (Rejected)
Problem attempted to solve: Could we improve consensus with a fifth diverse model?
Hypothesis: Deepseek (Chinese, open-weights) would add methodological diversity and improve jury agreement.
Results: Jury spread increased 40%, consensus dropped 42%, Krippendorff's α fell to 0.61 (below 0.70 threshold).
Key lesson: Not all diversity improves consensus. Systematic rejection is sometimes the right decision. Deepseek was a consistent outlier, not a productive addition.
V6.0: Framework Extension (Lifton Totalism)
Problem solved: V5 only covered Young & Reed's 10 criteria. Need complementary framework for totalism.
Innovation: Dual-track system. Young track (C1-C10) + Lifton track (C11 totalism). Configural (non-compensatory) scoring combining both signals.
Scoring outputs (generated in parallel): Young's Original Score (0-10 binary) + Composite Score (0-100%) + Lifton's Totalism Score (0-10). Three perspectives provided for each organization.
Validation: Krippendorff's α = 0.81 (≥0.80 threshold met). Young's YCDI concordance r = 0.78 (p < 0.001).
Adoption: 2026-06-12. Production-ready. 536+ organizations fully scored.
V6.1: Permanence-Aware Refinement
Problem solved: V6.0 conflates totalizing intensity with system permanence. US Marines penalized (temporary structure) despite high intensity (9.7).
Innovation: Three-component scoring: Intensity (1-10) × Permanence Multiplier (0.75-1.0) + Behavioral Durability Adjustment (-0.30 to +0.50).
Status: Proposed. Ready for PR review. Awaiting veteran affairs domain expert input.
Core Principles
- Rigor > Speed: Each version prioritizes methodological validity over rapid deployment.
- Transparency > Elegance: Document decisions, rejections, and trade-offs openly. Governance is visible.
- Auditability > Automation: Human review gate is preserved even as jury complexity increases.
- Generalizability > Domain-Specificity: System architecture is reusable across research domains.
Key Decisions
- Multi-model > Single-model: Jury consensus reduces single-model bias. But diversity must improve metrics, not worsen them (V5.2 rejection).
- Evidence tiers are essential: Not all sources are equal. 8-tier hierarchy ensures methodological rigor.
- Human review is non-negotiable: No fully automated scoring. Governance gates preserve accountability.
- Formal metrics enable publication: Krippendorff's alpha, ICC, ECRA statements transform jury scoring into publishable research.
- Framework complementarity works: Young + Lifton dual-track outperforms single framework. Separate components reveal nuance single metrics miss.
Before Testing
- Define success criteria in advance (e.g., α ≥ 0.70, jury spread must not increase).
- Pilot on representative sample (≥50 organizations) before committing to production.
- Measure against multiple dimensions, not just one metric.
If Results Are Negative
- Reject decisively and document why (like V5.2).
- Don't force adoption for cosmetic reasons.
- Systematic outliers are warning signs, not acceptable diversity.
Incremental vs. Revolutionary
- Small improvements (V5.0 → V5.1) are valid.
- Complementary frameworks (V5 → V6 adding Lifton) are valid.
- Refining separable components (V6.0 → V6.1) is valid.
- But diversification for its own sake (V5.2) is not.