Which infrastructure signals earn a permanent place on the wall
More metrics rarely mean clearer diagnosis. Keep the few that explain response time and uptime decisions.
Infrastructure signal mapping engagements often begin with boards full of unused panels. Teams collect because collection is cheap; interpretation is not.
We keep a signal when it has answered a past incident question, correlates with a user-facing symptom, or guards a known capacity cliff. Everything else is a candidate for archive, not deletion — history still matters for rare events.
A useful wall for response-time work typically includes request latency by journey, dependency error rates, host and container saturation, and path health for critical network hops. Fancy derived scores without owners rarely survive the next quarter.
After pruning, document who watches each remaining signal and what action a threshold implies. Ownership turns a chart into operational practice.