Level 4 · Eyes That Never Sleep
Session 34: Buy the Eyes or Build Your Own (Dynatrace vs the DIY Path)
What if two companies can see exactly the same incident symptoms, but one sees them in 10 seconds and the other in 40 minutes — and the fast one pays ten times as much?
1. Imagine If: Two Ways to Have Eyes
Imagine two factories with an identical problem: machine number seven is vibrating abnormally.
The first factory has a team that installed sensors on every machine themselves, wired them to a homemade dashboard, and wrote the alarm rules. It took six months to set up. But once it was running, everyone in the factory understood how to read it, and there's no vendor bill.
The second factory called in a company. That company's team came, installed sensors inside every machine without taking anything apart, and connected them to their own system. Done in two days. The moment a machine vibrates strangely, the system doesn't just beep — it points straight at it: "bearing number three, wearing out since 14:22, about 40 hours of life left".
Now the honest question: which one is more right? The answer depends on what's scarcer at that company — engineer time or money.
2. What Actually Happens
2.1 Two Paths, Not Two Products
This chapter doesn't compare two brands. It compares two observability strategy paths, represented by two examples:
- The buy path: Dynatrace (representing the class of integrated commercial observability platforms).
- The build path: OpenTelemetry + Prometheus + Loki + Tempo + Grafana (representing the class you build and run yourself).
Either one can win. What's wrong is choosing without realizing what you're actually buying.
2.2 The Buy Path: Dynatrace
OneAgent is the core of the model: a single agent installed on the host/container injects instrumentation into the application automatically. Engineers don't add a single line of code. Zero application changes, zero redeploys. That's its biggest economic argument — not the features, but the absence of manual instrumentation work.
Some of its signature capabilities:
| Capability | What it means in practice |
|---|---|
| OneAgent auto-instrumentation | Instrumentation without changing code; broad coverage from day one |
| PurePath | A whole request recorded across services — like a trace, but with much more complete capture |
| Davis AI | Causal analysis that points to the root cause, rather than just firing alerts |
| Grail + DQL | Storage and a query language designed for large observability volumes without separating metrics, logs, and traces |
| Automatic topology | Dependencies between components are discovered by the agent itself, with no manual documentation |
Pricing is based per host and per data volume — so every new node and every log spike shows up directly on the bill. That's the core trade-off: the bigger the scale, the bigger the cost, growing linearly with the number of hosts.
The biggest risk isn't the price, but vendor lock-in in two directions: 1. The team's expertise shifts to the tool, not to the concepts. 2. Data and configuration are hard to move. Leaving the platform becomes a major project.
2.3 The Build Path: OTel + Prometheus + Loki + Tempo + Grafana
On the other side is the open path:
| Component | Role |
|---|---|
| OpenTelemetry (OTel) | A standard and SDK for producing metrics, logs, and traces; vendor-neutral |
| Prometheus | Metrics storage and querying (pull model, PromQL language) |
| Loki | Lightweight log storage that indexes labels, not text content |
| Tempo | Distributed trace storage |
| Grafana | A single interface that brings all three together |
The strengths are concrete: infrastructure cost is generally much lower at large scale, the data stays yours, and the expertise your team gains is expertise in observability concepts — which applies anywhere.
The costs are concrete too, and often underestimated:
- The operational burden piles up on your team. Upgrades, scaling storage, tuning retention, fixing noisy alerts — all of it is internal work.
- Manual instrumentation. You have to touch code or libraries to name spans, choose attributes, and propagate
traceparent. Coverage becomes as wide as the work you put in, not automatic. - Sampling is almost mandatory. Storing everything is too expensive, so you have to decide: record 1%, 10%, or everything at a high cost. Every sampling decision is a compromise that can hide rare bugs.
- You build the correlation yourself. Linking metrics → traces → logs so you can click from "p99 went up" to "this query" is integration work, not something built in.
2.4 Three Real Trade-off Axes
An honest comparison isn't "fast vs cheap", but three axes:
Axis 1 — Auto-instrumentation vs manual instrumentation. Automatic: broad coverage from day one, without touching code. Manual: full control over what gets measured, with no vendor dependency. The first wins when engineer time is scarce; the second wins when you need precise control and want to stay portable.
Axis 2 — Full capture vs sampling. Full capture: rare bugs still get recorded, but storage costs can explode when traffic rises. Sampling: costs stay under control, but you're betting that the bug that matters isn't in the part of the sample that got thrown away. This is a policy choice, not a technical one.
Axis 3 — Vendor support vs team independence. Vendor: there's someone to call at three in the morning, there's an SLA, there are feature releases. Independence: no surprise bills and no lock-in, but if your observability pipeline breaks, there's no ticket to escalate — you fix it yourself.
2.5 Why Level 4 Exists as a Separate Level
Level 3 (Sessions 22–31) talks about failure: retry storms, circuit breakers, cascading failure, and observability as a concept. Level 4 talks about how we see: who ships the data, what it costs, who displays it.
That distinction matters for performance testers, because the observability pipeline itself can become a source of incidents:
- An agent running out of memory on a production node.
- A log buffer that fills up and then silently drops data (Session 32).
- A log index that goes read-only and takes the application down with it (Session 33).
- A vendor that sends a bill that tracks incident volume.
- A dashboard query that overloads the cluster for everyone.
The sign of a mature team: their observability pipeline has metrics, alarms, and load tests just like any other production service.
3. What If We Try... (The Naive Solution)
"Just use the free one, it's bound to be cheaper."
Free in license, not free in time. Every hour the team spends maintaining the pipeline is an hour not spent testing the system. At small scale, the build path clearly wins. At large scale with a small team, the math can flip.
"Just use a commercial platform, everything's sorted."
Everything's sorted until the billing meter starts ticking with volume. A commercial platform also doesn't remove the need to understand the concepts: if you don't know what p99 or a trace is, a beautiful dashboard won't help anyone.
"Use both."
This is actually a sensible strategy at many companies: a commercial platform for teams that need fast answers, and OTel as a portable instrumentation layer so the data isn't completely tied to one vendor. What you want to avoid is running both fully by accident — paying twice for the same coverage.
4. The Official Name
| Term | Practical meaning |
|---|---|
| Dynatrace | A commercial observability platform with auto-instrumentation |
| OneAgent | A single installable agent that instruments applications without changing code |
| Auto-instrumentation | Instrumentation done by the platform, not the developer |
| PurePath | A recording of one request across services; like a trace with broader coverage |
| Davis AI | A causal analysis engine for pinpointing the root cause |
| Grail / DQL | Storage and query language for large-scale observability data |
| Vendor lock-in | A dependency that makes switching platforms expensive |
| OpenTelemetry (OTel) | An open standard for instrumenting metrics, logs, and traces |
| Prometheus / PromQL | Metrics storage and query language |
| Loki | Log storage with a label index, not a content index |
| Tempo | Distributed trace storage |
| Grafana | An interface that unifies metrics, logs, and traces |
| Sampling | Recording only part of the data because of cost constraints |
| MTTD / MTTR | Mean time to detect / mean time to recover from an incident |
5. In Our World
The questions to answer before picking a path:
- Who's on call at three in the morning? If there's someone you can call, that has value and the price is reasonable. If your team is already strong operationally, the build path is legitimate.
- How much engineer time does it take to add manual instrumentation to 40 services? If the answer is six months and the project needs observability next month, automation is the most expensive feature that's still worth buying.
- How much will log volume grow over the next three years? This decides whether a per-host-per-GB pricing model still makes sense.
- Is there regulation that forbids data from leaving your environment? This often forces the build path, no negotiation.
- Can the data be exported easily? Get a written answer before you sign up, not when you want to leave.
6. The Performance Tester's Lens
This is the session that touches the performance tester's work most directly, because the observability path is an extra load on the system under test.
- Agent overhead under load. Measure the agent's CPU and memory consumption and latency impact while the system runs at peak load, not while idle. This is the number asked for most often and available least often.
- Trace propagation completeness. Verify that
traceparentsurvives across 100% of services, with no broken spans. One broken hop wipes out visibility in the most important part. - Observability overhead under load. Make sure the collector/agent stays below the agreed threshold (example target: <2% CPU, bounded memory) even at 20,000 QPS of traffic.
- Drop rate and buffer drain time. How long does the pipeline hold up when the observability backend is completely down before it starts dropping data?
- MTTD & MTTR before vs after. This is the metric that justifies the investment. Without comparison numbers, the claim "observability speeds up recovery" is just an opinion.
- Mandatory negative tests: shut down the observability backend while a load test is running; inject a full disk into the log cluster; bombard it with expensive dashboard queries. All three test whether the observability pipeline will take down the production system — or just lose data.
7. Question for the Next Round
"Now we can see everything that happens. The last question isn't about tools, it's about people: we have 31 sessions of knowledge, and teams at very different levels. How do we make all of this actually stick — and what standard questions should a performance tester ask every time a new design doc lands on their desk?"
(Session 30 already prepared the list of standard questions — this is where all of Levels 0–4 come back together.)