UNDER PRESSURE

Level 4 · Eyes That Never Sleep

Session 32: The Invisible Pipes (Sidecar, Fluent Bit, & Fluentd)

What if the application already writes its logs correctly, but when an incident hits, the logs from a dead pod vanish along with it — and nobody even knows that file ever existed?

Session 32 / 349 min read

1. Imagine If: A Factory with Private Notebooks

Imagine a factory with twelve production machines. Each machine has a private notebook sitting next to it. Every time a machine does something odd, the operator writes it down in that notebook: what time it was, what looked wrong, and what they did about it.

Now imagine the factory has this rule: as soon as a machine breaks and gets replaced with a new one, the old machine's notebook is thrown out together with the machine. Nobody copies it. Nobody archives it.

One night, machine number 7 breaks at 03:12. In the morning the team arrives, the replacement machine is neatly installed, and production is running again. Everyone is relieved — until someone asks: "so why did machine 7 actually break?" And the answer is: nobody knows, because the only record became trash at 03:13.

That is exactly what happens to container logs when we let them live inside the pod.


2. What Actually Happens

2.1 A Container Is a Process, Not a Durable Box

Pods and containers are designed to die without regret. Pods get killed, replaced, moved to another node, scaled down — all of this happens dozens of times a day without anyone noticing. That's a feature, not a bug.

The consequence for logs: anything stored inside the container's filesystem has the same lifespan as that container. Once the container is killed, whatever it wrote is gone.

2.2 Why Logs Are Collected from stdout, Not from Files

The 12-Factor standard (factor 11) says it plainly: treat logs as event streams, not files. The application writes to stdout and stderr, and then the container runtime (Docker, containerd) captures that stream and stores it outside the container.

Why this matters:

Aspect Write to a file inside the pod Write to stdout
Data lifespan Same as the container's lifespan Managed by the runtime, survives when the container dies
Rotation The application has to handle it itself The runtime handles it
Collection You have to get inside the pod Just read the stream from outside
Pod restart File is lost, needs a special volume Picked up by the agent right away
Consistency Every application has a different format One point of control

The trap: a container only captures one stream per main process. A multi-process application that writes to internal files (for example nginx logs, or a Java app server with its own log file) makes that stream invisible to the runtime — and that's the most common cause of "why don't my application logs show up in Kibana".

2.3 Three Placement Patterns for the Log Collector

The question isn't "which tool should we use", but where that tool is placed.

Pattern How It Works Strengths Cost
Sidecar One companion container per pod, reading the main container's stream Clean isolation, can have per-application config, doesn't share resources Wasteful: 1 pod × 3 containers = 3× the number of pods to run. For 50 pods, that's 50 extra processes
DaemonSet / Node Agent One agent per node, reading every container on that node Efficient: 3 nodes = 3 agents, not 50. Centralized resources Shares node CPU/RAM with the applications. Config has to be generic. If the agent dies, all logs on that node are stuck
Library in the Application The application sends logs straight to the backend (SDK) Fewest hops, lowest latency Mixes business concerns with log plumbing. Changing the backend = changing code + redeploying 30 services

For most systems, DaemonSet is the right default, and sidecar is used as the exception for applications that need special treatment (for example ones that write to internal files and need a sidecar to tail them).

2.4 Fluent Bit vs Fluentd: Forwarder and Aggregator

These two names often get used interchangeably, but their roles are different.

  • Fluent Bit — written in C, a footprint of about 1 MB, low RAM usage, designed to be installed on every node as a forwarder. Its job is narrow: pick up from the source, tag it, pass it on.
  • Fluentd — written in Ruby (and C for parts of the core), hundreds of plugins, rich in transformations (parsing, enrichment, complex routing, filtering), designed as the aggregator in the middle.

A healthy architecture has two tiers:

[Node 1] Fluent Bit ─┐
[Node 2] Fluent Bit ─┼──► [Fluentd Aggregator] ──► Elasticsearch
[Node 3] Fluent Bit ─┘         (1-2 instances)

Putting Fluentd on every node is the wrong choice economically: you pay for Ruby's memory consumption on every node just for work that a single centralized instance could do.

2.5 Silent Failures — The Most Dangerous Part

Log pipelines fail in a way that doesn't scream. Unlike an application that errors out and returns a 500, a broken log pipeline just gets quieter.

  1. Buffer full → logs get dropped. The forwarder has a limited buffer. If the backend is slow or down, the buffer fills up, and the default policy usually drops the oldest data so it doesn't hold the application back. The result: during the biggest incident, it's precisely the logs of that incident that go missing.
  2. Unlimited retries. If it's configured to "never drop", the buffer keeps growing until the node runs out of memory — and what falls over isn't just the log pipeline, but the applications on the same node too.
  3. Node disk full. A file-based buffer (filesystem storage) eats the node's disk. Full disk = containers can't write = the applications on that node die with it.
  4. Split multi-line logs. A Java/YAML/Python stack trace is one log event but gets recorded as 40 separate lines. Without a multiline parser that understands the pattern, a search in Kibana returns 40 meaningless documents, and the root cause is never read in one piece.
  5. The agent becomes a single point of failure. Every pod on a node depends on one agent. The agent crashes, has a bad config, or gets OOM-killed — and all logs on that node disappear without an alarm.

The principle to hold on to: a log pipeline must have metrics for itself — how many lines come in, how many go out, how many are dropped, how full the buffer is. Without that, you'll never know that half your logs are already gone; you'll only find out when you need them most.


3. What If We Try... (The Naive Solution)

"Let's just keep logs in a file and mount a volume so they persist."

This works for one application living on one server. But once you move into container orchestration: pods move between nodes, pod names change on every restart, and volumes that aren't managed well pile up garbage — who cleans up the log files from 200 pods that died last month? This solution moves the problem around instead of solving it.

"We'll read the log file straight from the pod over SSH."

This is how people worked before log pipelines existed. The consequences: you need access, you need the pod to still be alive, and in the most important incidents (a pod that's already been killed), this approach doesn't work at all.

"One vendor, one tool, everything's sorted."

Tempting, and sometimes true. But that's the question we tackle in Session 34 — with the price that has to be paid first in Session 33.


4. The Official Name

Term Practical meaning
Sidecar A companion container in the same pod, sharing network and volumes, taking over a single responsibility
DaemonSet A Kubernetes workload that guarantees one pod runs on every node
Node agent A log collection process running on the node, outside the application containers
stdout / stderr The two standard streams of a process; the way an application hands its logs to the runtime
Log driver The container runtime mechanism that captures the stdout/stderr streams
Fluent Bit A lightweight forwarder (C, ~1 MB) to install on nodes
Fluentd A plugin-rich aggregator (Ruby) to install centrally
Forwarder / Aggregator Two different roles in a two-tier log pipeline architecture
Buffer Temporary storage while the backend is slow; this is where logs silently disappear
Flush interval How often the buffer is sent to the backend; a trade-off between latency and efficiency
Multiline parser Rules that merge many lines into a single log event
Backpressure When the producer is faster than the consumer; the main source of log loss
Log loss Losing log data without an alarm — the most dangerous failure because it's silent

5. In Our World

In a production system with dozens to hundreds of pods, the log pipeline is infrastructure, not configuration. Signs of a system with a mature log pipeline:

  • Applications write only to stdout; there are no local log files being rotated by the app itself.
  • Every node has one agent (Fluent Bit / Promtail / OTel Collector).
  • A central aggregator does the parsing and routing, not the edge.
  • There's a dashboard for the log pipeline itself: lines/second, drop rate, buffer size, delivery lag.
  • Multiline stack traces are already recognized as a single event.
  • There's a retention policy with a clear business decision behind it.

Signs of an immature system: the logs are "in Kibana", but during the biggest incident, there isn't a single line from the problematic pod. This is the failure that happens most often, and gets noticed least.


6. The Performance Tester's Lens

The numbers to ask for as soon as you see a log pipeline in a design doc:

  1. Log engine throughput — how many lines/second can the pipeline handle at peak load? Test this by sending test traffic while measuring how many lines actually arrive.
  2. Drop rate — ask for an explicit number: what percentage of logs is dropped when the backend is slow. If there's no number, nobody has ever measured it.
  3. Overhead on the node — how much CPU and memory the agent uses, measured under full load, not while idle. An agent that's calm in the lab can turn greedy when the application is screaming.
  4. Buffer depth and drain time — how long can the pipeline hold data when the backend is completely down before it starts dropping?
  5. Test scenario: shut down the log backend (Elastic) for 5 minutes while a load test is running. How many logs are lost? Does the application keep running? Is there an alarm?
  6. Multiline readiness — make sure a stack trace is read as one event, not 40. Test it by throwing nested exceptions.

7. Question for the Next Round

"The logs are now neatly collected in one place. But storing all of it so it can be searched in 1 second — what does that really cost? And why can engineers suddenly not write any logs at all once the disk is full?"

(Continued in Session 33: The Price of an Index — Elasticsearch & Kibana)