Fluentd: Buffer overflow
Fluentd's output buffer is full because the destination cannot accept data fast enough. Logs may be dropped.
Quick fix
Read the commands before running them. Anything that restarts a service, deletes data or changes permissions should be tried on a non-production system first.
# Increase buffer size in fluentd.conf
<buffer>
@type file
path /var/log/fluentd/buffer
total_limit_size 5GB
chunk_limit_size 32MB
overflow_action throw_exception
</buffer>
# Scale destination or add retry
retry_max_times 10
retry_wait 5s
How to diagnose Logging errors
Logging failures are dangerous because they are silent: the application keeps running while its telemetry disappears. The recurring causes are buffer overflow under backpressure (the destination cannot keep up), permissions on log files or directories, and ingestion rate limits at the cloud provider. Monitoring the logging pipeline itself is the only reliable way to notice.
If the quick fix above does not resolve it, work through these steps. They apply to this whole class of error, not just to this one message, which is usually what saves the time.
- Check the collector's own logs first: Fluentd, Logstash and the CloudWatch agent all log their own failures, usually to a separate destination.
- Look for backpressure metrics: buffer queue length, retry counts, and dropped-record counters. A full buffer means the destination, not the collector, is the bottleneck.
- Verify write permissions and disk space on the buffer path. A full disk silently stops most collectors.
- For cloud ingestion, check the API rate limit for the log group or stream and batch more aggressively rather than retrying harder.
- Add a heartbeat log line and alert on its absence. This is the only way to detect a pipeline that has stopped without erroring.
Tools worth reaching for
collector self-logsbuffer/queue metricsdf -haws logs describe-log-streamslogger / fluent-cat for test events
Authoritative references
Primary documentation for this error, worth reading before applying any fix in production.
Related Logging errors
- CloudWatch Logs: Rate exceeded (ThrottlingException)PutLogEvents API calls are being throttled by CloudWatch Logs. The account or log group is…
- journalctl: No journal files were foundjournald stores logs in memory when /var/log/journal does not exist, so everything is…
- Log4j: Appender not foundThe logging configuration references an appender that is not defined. The log4j2.xml or…
- Logback: Failed to create log fileLogback cannot create or write to the log file due to filesystem permissions or the directory…
- logrotate: skipping because parent directory has insecure permissionslogrotate refuses to rotate a file in a directory that is group or world writable, because…
- Logstash: Pipeline errorA Logstash pipeline failed to start or process events, usually due to a bad grok pattern…
- Loki: entry too far behind / out of orderLoki rejected log lines whose timestamps are older than the accepted window for that stream…
- systemd-journald: Suppressed 1234 messages from /system.slice/app.servicejournald rate limits each service, discarding the rest of a burst rather than slowing the…
Browse other categories
- HTTP 494xx client errors, 5xx server errors, redirects, headers and protocol problems.
- JavaScript 42npm resolution, async pitfalls, hydration, memory limits and runtime type…
- Database 41Connections, deadlocks, constraints, replication and memory limits.
- AI 35Rate limits, context windows, GPU memory and model-serving failures.
- Network 35Refused connections, timeouts, resets, MTU problems and port exhaustion.
- Python 35Imports, virtual environments, encoding, concurrency and dependency conflicts.
- Kubernetes 34CrashLoopBackOff, ImagePullBackOff, OOMKilled, RBAC, scheduling and storage.
- Docker 27Daemon connectivity, disk space, image pulls, ports and architecture mismatches.
- System 26Disk space, systemd units, file descriptors, OOM killer and scheduled jobs.
- Cloud 25IAM permissions, quotas, service limits and credential failures.
- Security 25JWT validation, CSRF, OAuth grants, SELinux, SSH host keys and CSP.
- TLS 24Untrusted authorities, expiry, hostname mismatch, chains and cipher negotiation.
Something missing or wrong?
This entry is maintained by hand. If the fix is out of date, incomplete, or you have a better one, email a correction and it will be reviewed.