ISMS Copilot

Engineering & Research

How we evaluate AI models for compliance: benchmarks, methods, and findings.

The silent zero: when a missing judge score becomes a measurement
Engineering

The silent zero: when a missing judge score becomes a measurement

In our 2026-07-17 ablation of GRC document-generation strategies (GLM-5.2 and Claude Opus 4.8 as writers and as 1-10 rubric judges), pass 2 had 79 of 446 recorded judgment rows without an overall score. The aggregator counted each one as 0. The same zero-filling had manufactured a 3.25-point effect in pass 1, and when it vanished we wrote a judge-bias theory to explain why. The paired offset between the two judges on the same documents was 0.12.

·14 min read