Data Engineering · Published 9/29/2026 · By

Retries Are Not a Data-Quality Strategy

Reliable ingestion needs clear contracts, idempotent retries, validation, and visible failure paths—not just another retry loop.

A retry loop can turn a temporary network hiccup into a successful delivery. It can also deliver the same event five times, hide a malformed payload behind a pile of logs, or keep retrying data that will never become valid.

Retries solve a transport problem. They do not, by themselves, solve a data-quality problem.

Give an event a stable identity

Suppose a producer sends an equipment reading, but the acknowledgement gets lost. The producer cannot tell whether the receiver stored the event. It sends the reading again.

If the receiver treats every request as new, a perfectly reasonable retry becomes a duplicate record. Give each event a stable identifier and make the write idempotent: processing the same identifier twice should have the same effect as processing it once. A database uniqueness constraint or an idempotency record can enforce that behavior at the storage boundary.

The identifier must belong to the event, not the delivery attempt. Generating a fresh ID each time a producer retries defeats the point.

Validate at the edge of trust

Check the shape and meaning of data before it enters the trusted path. Is the timestamp parseable? Does the measurement have a unit? Is the value within a meaningful domain range? Are required identifiers present?

Validation should distinguish at least three cases: accepted, rejected as invalid, and temporarily unavailable for processing. A bad number will not become good after a five-minute retry. A database timeout might.

That distinction keeps retry policy small and honest: retry transient failures with limits and backoff; route invalid data to a visible rejection path that a person can inspect or correct.

Make failure useful

“Failed to process event” is a status, not an explanation. Record the event ID, source, validation rule, and a safe description of the failure. Give operators a way to see the rejected item and decide whether to repair, replay, or discard it.

Avoid logging sensitive payloads indiscriminately. Observability should reveal what happened without creating a second, less controlled copy of private data.

Test the uncomfortable cases

Send the same event twice. Send it late. Send it out of order. Send a value with the wrong unit. Make the receiver unavailable, then restore it. Confirm that recovery does not silently lose events or double-count them.

The small ingestion demo shows these ideas with synthetic temperature readings. It is deliberately a browser-only teaching example, not a benchmark or production service.

Before adding another retry, ask what the system should do when the data is invalid, duplicated, delayed, or impossible to deliver. A dependable pipeline has an answer for each—and makes that answer visible.

Discuss a system challenge ↗