Answer repeatability
The same question asked of an AI model several times can produce different answers. That is why we repeat queries and measure result stability — the more the answers agree, the higher the measurement confidence. High answer variance lowers confidence and is reported openly as a property of the environment, not a measurement error.
Input data quality
The pipeline separates valid answers, failures, and unavailable AI platforms before scoring. It stores statuses, provider errors, platform context, and performs automated report QA.
A partial but non-zero measurement may be shown only with explicit coverage and interval width.