Skip to main content
Methodology

Measurement confidence

Measurement confidence combines coverage, source fidelity, repeatability, freshness, and validation. These are technical indicators, not guarantees that an answer is true.

Coverage of stored data Answer repeatability Input data quality
Visualization for the page: Measurement confidence
Definition and protocol

Coverage of stored data

Coverage uses the value returned by scoring or, as a fallback, whether answer results exist. An empty set receives zero.

Read the value together with run status, provider configuration, and the number of successfully collected answers.

Methodology rule

Answer repeatability

The same question asked of an AI model several times can produce different answers. That is why we repeat queries and measure result stability — the more the answers agree, the higher the measurement confidence. High answer variance lowers confidence and is reported openly as a property of the environment, not a measurement error.

Methodology rule

Input data quality

The pipeline stores run statuses, performs automated report checks, and can route a report to manual QA.

There is not yet a complete gate that rejects every incomplete answer read before scoring. Quality requires reviewing provider statuses and raw answers.

Methodology rule

Run and history freshness

The current run has a timestamp, and audit history contributes to the change component. Missing history lowers a supporting freshness dimension.

Older individual answers do not currently receive a separate age-decay score. Measurement date must be considered during interpretation.

Methodology rule

Component directness

Components receive a real, proxy, or declared source status. Source fidelity also contributes to overall run confidence.

The status is not yet a complete metric for every individual answer. In particular, a model name does not prove its consumer application was measured.

Methodology rule

Confidence bands

Instead of a deceptively precise number, we communicate confidence in readable bands — from a complete measurement to a significantly limited one. The band immediately tells you how much weight the score and its changes can bear. The criteria for band assignment are published in the methodology.

Methodology rule

Partial measurement

A partial run can result from a missing provider, limit, refusal, API failure, or too few repetitions. Raw answers and provider logs help establish the cause.

Until failures are fully separated from brand absence, not every “absent” value should be interpreted as a real observation.

Methodology rule

Effect on interpretation

Confidence changes how a score should be read: with high confidence a shift of a few points can be a signal, with low confidence it sits within measurement noise. Reports pair every score delta with its confidence band to prevent over-interpretation. It is a safeguard against conclusions the data does not support.

Next step

Assess result reliability

Measurement confidence combines coverage, source fidelity, repeatability, freshness, and validation. These are technical indicators, not guarantees that an answer is true.