Spotting validator skew before your skip rate climbs
Skip rate is a lagging indicator. By the time dashboards turn red, vote credits may already be depressed for the epoch. Skew metrics — how far your node trails the cluster — often move first.
Root distance
Compare your node's root slot to a quorum of public RPC endpoints. A consistent gap of more than a few dozen slots during steady state suggests replay or disk contention. Sudden widening across all your peers points to a cluster event; widening only on your host points local.
Gossip round-trip time
Log gossip ping samples to a simple time-series store. Rising RTT with a shrinking peer count often precedes turbine delays. If peers cluster on one provider, RTT can look healthy while path diversity is poor.
Banking-stage histograms
Enable banking-stage timing metrics if your build supports them. A long tail on transaction execution during non-leader slots still matters — it competes with replay threads for CPU. Leader-slot spikes combined with replay lag are a reliable skip predictor.
A simple weekly drill
Every Monday, record root distance, peer count, median gossip RTT, and current skip rate. Plot four weeks. Operators who do this often spot gradual disk degradation before automatic alerts trigger.
When to commission a review
If root distance variance exceeds your historical band for two consecutive weeks and you cannot attribute it to a known upgrade, external eyes help. Our health review starts with these four metrics plus your vote credit trend.
See also: Glossary — Root bank and Reading slot-leader schedules.