Answer repeatability
The same question asked of an AI model several times can produce different answers. That is why we repeat queries and measure result stability — the more the answers agree, the higher the measurement confidence. High answer variance lowers confidence and is reported openly as a property of the environment, not a measurement error.
Input data quality
The pipeline stores run statuses, performs automated report checks, and can route a report to manual QA.
There is not yet a complete gate that rejects every incomplete answer read before scoring. Quality requires reviewing provider statuses and raw answers.