Dashboards & Metrics / Field Kit

Build the dashboard before the incident. Per domain: what to measure, what "trouble" looks like, how to read the panel at a glance, and how to wire it up in Dynatrace (the tool we use). A good dashboard answers one question — "is this healthy?" — in three seconds.

Start here

The rules that make every domain dashboard below readable — including by someone seeing it for the first time.

Measure the signals, not everything

  • Golden signals (services): Latency, Traffic, Errors, Saturation.
  • RED (request-driven): Rate, Errors, Duration.
  • USE (resources): Utilization, Saturation, Errors.
  • Pick 4–8 tiles per screen. If everything is on the dashboard, nothing is.

Make it readable in 3 seconds

  • Most important tile top-left; green/amber/red so status reads before numbers do.
  • Show rate of change, not just absolutes — "climbing" matters more than "high".
  • Annotate deploys on the timeline. Half of all incidents line up with one.
  • One dashboard = one question. "Is the AG healthy?" not "everything about SQL."

Dynatrace — how to reach it

  • OneAgent auto-discovers hosts, processes, and services — start from the entity, not a blank chart.
  • Davis AI does baselining + anomaly detection; let it flag the deviation instead of hand-setting every threshold.
  • Build custom dashboards from metric tiles; query with DQL (Grail) for logs/events.
  • Use management zones to scope a team/brand, and SLOs for the "are we within budget" tile.
  • Databases & middleware come via Hub extensions (deployed on an ActiveGate) — add the SQL Server / PostgreSQL / Oracle extension to get their metrics.

Dashboards & Metrics Field Kit — a ReliabilityOps tool. Tips are tool-agnostic where possible and Dynatrace-specific where noted; confirm entity/metric names against your Dynatrace version and extensions. Offline, single-file; nothing on this page calls the network.