UNDER PRESSURE

Level 4 · Eyes That Never Sleep

Session 33: The Price of an Index (Elasticsearch, ILM, & Kibana)

What if storing a single log costs more than running a single business transaction — and the team only realizes it when the third month's bill arrives?

Session 33 / 349 min read

1. Imagine If: A Library with One Magic Door

Imagine a library with one million books, and a librarian who can find any book in one second, from just a single word inside it. Impressive. Until you ask: how long does it take them to PUT one new book on the shelf?

It turns out one new book takes five minutes — because for every word in it, they have to write a small index card and slip it into the already-tidy card rack. More words, more cards.

And the last question, the most uncomfortable one: what's the rent on the building? Because that card rack isn't free, and books that are never read still take up space.

That's Elasticsearch. Its strength (instant search) and its burden (expensive writes, expensive storage) come from the same source: the index.


2. What Actually Happens

2.1 The Logs Are Collected — Now They Have to Be Searchable

By Session 32, the logs have been successfully shipped out of the pods. But they're still scattered across dozens of nodes. The incident question is always the same: "which service did the request with this ID fail in?"

If logs are just text files on servers, the answer takes hours: log into each server, grep, wait, repeat. Elasticsearch turns that into a single query.

2.2 Inverted Index: Why Search Is Instant, Why Writes Are Expensive

The index at the back of a book is the closest analogy. Instead of reading the whole book to find the word "timeout", you open the index page: timeout — p. 12, 45, 78.

Elasticsearch builds that structure for every field in every document. The consequences:

Operation Cost Why
Search Very cheap The index points straight to the matching documents
Write Expensive Every new document has to be broken into tokens and build/update index segments
Add a field Very expensive Every new field = a new index structure across all documents

A practical number: writing to Elasticsearch can be 10–30× more expensive than reading. That's why Elasticsearch isn't a transactional database, and why log volume has a direct cost consequence.

2.3 Shards and Replicas

An Elasticsearch index isn't stored in one place. It's cut into shards — pieces that can be searched in parallel across several nodes. Each shard has a replica (copy) for reliability and for serving reads.

Consequences that often get forgotten:

  • The number of shards is set when the index is created and can't be casually changed later. Pick wrong at the start = reindex all the data (an operation that can run for days).
  • Too many small shards is worse than a few large ones: every shard carries overhead (memory, file handles, search cost at the coordinator). This is the over-sharding disease, very common in unmanaged installations.
  • Too few shards means searches can't run in parallel, and a single shard gets too big to move.

2.4 Mapping Explosion — The Silent Killer

This is a classic log failure, and it happens because of a good habit: applications sending structured JSON.

The problem is, without limits, every new field name in the JSON automatically becomes a new column in the index. Some real examples:

  • Error logs include error.context.request.payload.user_id_12345 as a field name. Every user = one new column.
  • Audit logs include dynamic resource names.
  • One tenant adds custom fields, and the index template applies them to everyone.

The result: an index that should have 40 columns suddenly has 40,000 columns. Nodes run out of memory just holding the metadata. The cluster becomes unstable.

The defenses: - Dynamic mapping is turned off for log indices, or restricted to strict. - Index templates with a list of allowed fields. Unknown fields go into a single text-typed additional column. - Limits on the application side: logs are a contract, not a JSON dumping ground. Fields that aren't needed don't get sent. - Mapping explosion has a hard limit (default 1000 fields per index) that, once reached, causes documents to be rejected — meaning logs are lost, and you only feel it during an incident.

2.5 ILM: Hot, Warm, Cold, Delete

A log that's one hour old has different needs from a log that's one year old. Index Lifecycle Management moves indices between storage classes as they age:

Phase Typical age Storage Why
Hot 0–3 days Fast NVMe/SSD Actively being read during incidents
Warm 3–30 days Standard SSD / fast HDD Still searched, but less often
Cold 30–90 days Large HDD / object storage Slow searches, rarely accessed
Delete >90 days Deleted Retention is over

This isn't just a technical optimization — it's a business decision. The question "how long do we keep logs" has different answers for the security team (audit), the engineering team (debugging), and the finance team (the bill).

2.6 Disk Watermark → Read-Only Index

This is the least intuitive and most painful failure.

Elasticsearch has disk watermarks: 85% (low), 90% (high), 95% (flood). Once the disk passes 90% (high watermark), the cluster stops allocating new shards to that node. Once it passes 95%, Elasticsearch switches the index to read-only.

The effect isn't "logs can't be read". The effect is:

  1. The application tries to write logs → rejected.
  2. The application fails too, because some applications wait for logs to be delivered and mark the request as failed.
  3. The log pipeline (Session 32) piles up in the buffer.
  4. Buffer full → logs are silently dropped, or the node runs out of memory.

So one full disk in the log cluster can chain into a total production incident. What needs to be in place: cluster disk monitoring as a first-tier alarm, and a retention policy aggressive enough that the watermark is never approached.

2.7 Kibana: A Window, Not a Warehouse

None of this is useful without a place to ask questions. Kibana is where Discover (browsing raw logs), dashboards, visualizations, and saved searches live.

An important note for performance testers: searches in Kibana have a cost. A query with a leading wildcard (*timeout) or a very wide time range without field filters will scan every shard — and in a large cluster, one person leaving an auto-refreshing dashboard open can overload the log cluster for everyone. This is a real incident pattern: "Kibana is slow" almost always means "someone is firing an expensive query".


3. What If We Try... (The Naive Solution)

"Store all logs forever — storage is cheap, right?"

Storage is cheap, but search isn't. An index that keeps growing increases the number of shards, adds nodes, and raises the cluster's RAM cost. Unlimited retention is the most expensive way to store logs that will never be read.

"Let's use one big index for everything."

One big index means one ILM operation for everything, a single point of failure, and no way to move part of the data to a different storage class. Elasticsearch recommends per-day or per-size indices (rollover) precisely so that retention and migration can be granular.

"Let's index every field so we can search anything."

This is the road to mapping explosion. Indexing every field = the highest cost for the smallest chance of being searched. Fields that are never used for filtering are better left unindexed.


4. The Official Name

Term Practical meaning
Elasticsearch A Lucene-based search engine and index store
Inverted index A "word → list of documents" structure; the source of search speed and write cost
Shard A piece of an index that can be searched in parallel; the count is fixed when the index is created
Replica A copy of a shard; for reliability and for serving reads
Over-sharding Too many small shards; memory and search overhead
Mapping The field schema (name, type, whether it's indexed)
Mapping explosion The index column count blows up because of unlimited dynamic fields
Index template An index blueprint that sets mappings and settings ahead of time
ILM Index Lifecycle Management: hot → warm → cold → delete
Rollover Switching writes to a new index when a size/age threshold is reached
Disk watermark Disk thresholds (85/90/95%); crossing them makes the index read-only
Read-only index An index that rejects writes — the trigger for a chain of application failures
Kibana The search, dashboard, and visualization interface on top of Elasticsearch
ELK / EFK Elasticsearch+Logstash/Beats+Kibana / Elasticsearch+Fluent*+Kibana

5. In Our World

Signs of a mature log installation:

  • Indices use templates with explicit mappings; dynamic mapping is turned off.
  • Rollover + ILM manages the data's journey from hot to delete.
  • Retention is set together with the business owners, not left at the default.
  • Cluster disk alarms are high-priority alarms, not a side note.
  • There's control over expensive queries (limited, scheduled, or banned by policy).
  • The number of shards per index is calculated, not left to grow wild.

Signs of a fragile installation: a single index named logs-* that grows forever, mappings that change every week, and nobody knows how long the retention is.


6. The Performance Tester's Lens

The numbers to ask for when Elasticsearch shows up in a design doc:

  1. Ingest rate — how many documents/second have to be swallowed at peak load, and what's the real capacity?
  2. Write latency vs search latency — two different numbers with different characteristics; ask for both.
  3. Shard count and its growth — how many shards per index now, and how many three months from now?
  4. Watermark thresholds and retention time — when is the disk predicted to cross 90%? This can be calculated from the growth rate.
  5. Test scenarios: - Send 2× the normal log volume during a load test and see whether ingest falls behind (lag). - Fill the disk close to the watermark in a test environment and prove what happens to the application — does it actually fail, or degrade. - Run expensive Kibana queries at the same time as high load and measure the impact on the cluster.
  6. Field count per index — if the number is already in the thousands and climbing, a mapping explosion is underway.

7. Question for the Next Round

"We now know what it costs to assemble everything ourselves. But some companies pay ten times as much and get a root-cause answer in ten seconds. What exactly are they buying — and what are they giving up?"

(Continued in Session 34: Buy the Eyes or Build Your Own — Dynatrace vs the DIY Path)