Check your understanding: Reliability, Resilience & Standards
10 questions from Part 12 of the guide. Score instantly; every explanation links to the chapter that derives the answer.
The guide's core critique of legacy tier standards for AI clusters is that they…
A service objective requires live maintenance and survival of the defined unplanned-fault case. What should the owner require before selecting a resilience topology?
Two resilience proposals quote different topology premiums. What normalization must happen before the owner can compare them?
In the cited provider comparison, what training-goodput figures define the sensitivity case?
On a 16,384-H100 training run, how often did unplanned interruptions occur?
In the cited high-interruption training scenario, which recovery intervention has the largest modeled goodput leverage?
At a training failure, what establishes the RPO boundary?
Active-active geographic failover for an AI service costs roughly…
A provider cites a 99% node / 95% rack SLA example. What can the buyer validly infer?
What share of impactful data-center outages does power cause?