An independent research studio working on how generative video and AI systems are measured.
Benchmarks report a number. We work on what that number rests on: who rated what, how many times, on which scale, and whether anyone can recompute it from the data that was released.
We recover the individual votes behind a published reliability figure and recompute it, then report what the release does and does not support.
Rating protocols built so the coefficient can be estimated at all: panel size set from the target, shared items across raters, identifiers that persist.
Open software that turns a vote matrix into the declarations a reader needs, and flags the cases where no coefficient should be reported.