Why We Published the Scoring Method Before the Scores
In one sentence
A benchmark is only citable if a stranger can reproduce it, which requires publishing the method before the results and including yourself in the cohort.
The problem with vendor benchmarks
Most vendor-published industry indices share a structure. The vendor defines the metric. The vendor runs the measurement. The vendor's clients do well. The methodology is described in a paragraph that cannot be reproduced.
Nobody is necessarily lying. But there is no way for a reader to distinguish an honest one from a dishonest one, so a rational reader discounts all of them. The format has spent its own credibility.
What we did instead
Three commitments, in the order they matter.
The method is public and versioned before any scores are
The scoring specification states the four dimensions, their exact weights, what each measures, and how the weighted total is computed. It carries a version number. When the formula changes in a way that breaks comparability, the version increments and old runs keep their original version stamp.
This means you can disagree with us specifically. "Your evidence dimension is overweighted at 20%" is a real critique we can engage with. That is a better conversation than "trust our score."
Every run ships its raw inputs
A published score without the prompts, engines, and timestamps that produced it is an assertion. Each run records which engines responded, which were configured but unreachable and why, how many prompts ran, and when it started and finished.
If an engine was down, the run says so. It does not quietly score three engines and present it as four.
We score ourselves in the same cohort
Ramola appears in the Index under the same probes as everyone else, and the result is published whatever it says.
This is not modesty. It is the only structural reason to believe us. A benchmark whose author is exempt from it is a sales asset wearing a lab coat.
Runs are immutable
A run id is permanent. Re-scoring produces a new run rather than editing an old one, because the point of a citable artifact is that a number someone quoted last quarter still resolves to that number.
If we get something wrong, we publish a correction as a new run with a note — we do not silently rewrite history.
What this costs us
Occasionally we will rank below a competitor in our own Index, and that will be public.
We think that is a fair price. The alternative is another vendor benchmark nobody has a reason to cite, which would make the whole exercise worthless — including to us.