// BLOG
Improving Messaging Safety with Time-to-Live Values (TTLs) and Expiring Messages
Applies to Akka.NET and any other message-driven technology.
There are many different strategies for retrying the “failed” delivery of messages over the network: exponential backoff, retry loops, circuit breakers, etc.
What these have in common is that they are “sender-side” retry strategies - it is the message sender’s responsibility to ensure that the receiver processes it1.

Sender-side retry strategies are the most popular resiliency strategies for message-oriented technologies like Akka.NET, HTTP APIs, and RPC (remote procedure call) systems alike. They’re popular because they are simple and require the least amount of infrastructure to implement. The recipient can often remain dumb and stateless, and the client / sender only needs to retain some transient request state in order to retry the message n times before dead-lettering it.
But there are two really interesting and potentially harmful problems this “simple” re-delivery strategy imposes on the recipient:
- Idempotency - how do we know if we’ve already processed this message previously? This is important for preventing data corruption and duplication.
- Wasted work - how do we know if this request has already been timed out by the client, so we should skip processing it because the client’s not going to receive our response anyway? This is important for conserving resources and high availability.
That second problem is easy to miss because cancellation on the client doesn’t cooperatively cancel anything on the server. When an Ask or an HttpClient call times out, the caller throws an exception and moves on - usually by retrying - but the original request is still sitting in the recipient’s queue. Telling the server to stop would require a whole other network round trip, and even if we sent a “cancel” message, FIFO ordering means it would land behind the request it’s trying to cancel. So the recipient is going to do all of that work anyway, for a caller who is no longer listening, and every retry adds another copy to the pile.
We’ve seen this take down production systems. One of our customers runs a large financial application that performs a high volume of calculations. Every so often, one of the pieces of data a calculation needed to retrieve was much larger than the others, and while the calculator worked through it, everything queued up behind it timed out and got retried. The retries piled on top of the originals and the calculator spent all of its time reprocessing requests nobody was waiting for, over and over, until the only way out was to shut the system down and reboot it. Once they started discarding expired calculation requests, the system stayed stable - the load spikes still happened and callers still saw the occasional timeout, but the calculator recovered on its own instead of needing a restart.
In this post we’re going to focus on item number 2: preventing the receiver from doing useless work by encoding “time-to-live” (TTL) information into the messages we send to it.
Time-to-Live
Time-to-Live (TTL) is a straightforward software engineering concept: how long is this data valid for? Or another way of putting it: when does this data expire?
Many networking protocols, database platforms, and distributed systems use it as a way of expressing how perishable a piece of data is. In Akka.Discovery.Redis, for instance, we use Redis’ built in TTL support to make sure that records from dead nodes that left abruptly get purged automatically - each live node refreshes its record every 30 seconds, and Redis expires any record that hasn’t been refreshed within the TTL (2 minutes by default):
/// <summary>
/// Refresh the LastUpdate value and update Redis with new TTL
/// </summary>
public async Task UpdateAsync(CancellationToken token = default)
{
if (_entity is null || _myKey is null)
throw new InvalidOperationException("Invalid update operation, client has not been initialized");
token.ThrowIfCancellationRequested();
_entity = _entity.Update();
await _database.StringSetAsync(_myKey,
_entity.ToBytes(), _settings.Ttl);
}
This keeps the Akka.Cluster member Address records limited to nodes that are still alive and refreshing their entries - a good example of TTLs enforcing data “freshness.”
The Problem: Retry Storms
To see why this matters for messages, imagine the following:

- Multiple senders all trying to get their messages processed by a single recipient;
- Acknowledgements from the recipient are not coming back on-time, so all senders retry up to
ntimes; and - The recipient is backed up or blocked by some shared resource - a pegged CPU, a slow network, or other sources of busyness.
When the recipient wakes up and begins catching up on its backlog, its mailbox (queue) will look something like this:

Remember: in Akka.NET actors process their mailboxes in FIFO order, so the oldest messages at the top of the queue get processed first. These are all expired and have been timed out by the sender already, so processing them is pure waste. We want the recipient to discard them QUICKLY and move onto the green messages that still have time left before their deadlines.
That’s exactly what expiring messages let us do.
Using Time-to-Live to Create Expiring Messages
In Akka.NET or any other messaging technology, it is not difficult to implement a Time-to-Live value that the recipient can use to detect and discard messages that have already been expired or timed out by the sender. Some messaging tools even have this built in - NServiceBus has its [TimeToBeReceived] attribute and RabbitMQ supports per-message expiration - but it’s easy to roll your own for anything that doesn’t.
My preferred approach is to define something like a Deadline type that can be used to express when a message is considered to be “overdue” or “expired:”
// Deadline struct in C# computed from the current time and the timeout
// value. The deadline is used to determine if a request has timed out.
public readonly struct Deadline
{
public Deadline(TimeSpan timeout)
: this(DateTime.UtcNow + timeout)
{
}
// used by serializers to rehydrate the deadline from the wire
public Deadline(DateTime deadlineTime)
{
DeadlineTime = deadlineTime;
}
public DateTime DeadlineTime { get; }
public bool IsOverdue => DeadlineTime < DateTime.UtcNow;
}
Then we just incorporate this data structure into our message types that are going over the wire:
public interface IWithDeadline
{
Deadline Deadline { get; }
}
public sealed record Request(string Payload, Deadline Deadline) : IWithDeadline;
An important but subtle detail about generating these Deadlines on the sender side:
DeadlineTime must be an absolute time, not a TimeSpan - what we’re really doing here is having the receiver predict when the client’s CancellationToken is going to fire. The receiver needs to know when the message expires relative to the moment the client sent it, not when the receiver received it. Therefore, the clock starts ticking upon send, not upon receive.
So when we send the message from the sender / client to the recipient, we should do as follows:
var timeout = TimeSpan.FromSeconds(5);
using var cts = new CancellationTokenSource(timeout);
var request = new Request(payload, new Deadline(timeout));
// could be an HTTP client, IActorRef.Ask, etc...
return await client.SendAsync(request, cts.Token);
On the recipient side, the key is to discard expired messages instead of processing them. In an Akka.NET actor, that’s just one extra Receive handler:
public sealed class RequestHandler : ReceiveActor
{
public RequestHandler()
{
// expired - the sender already gave up, so don't bother
Receive<Request>(r => r.Deadline.IsOverdue, r =>
{
// TOTALLY OPTIONAL BUT NOT A BAD IDEA
// let the sender know we skipped an expired message
Sender.Tell(new RequestTimedOut(r));
});
Receive<Request>(r =>
{
// real work here
});
}
}
If you’re processing requests with Akka.Streams instead, the same check fits in a single Where or DivertTo stage.
You can take this one step further and pass the remaining time into the recipient’s own I/O calls as a CancellationToken (i.e. new CancellationTokenSource(req.Deadline.DeadlineTime - DateTime.UtcNow)), so work that is already in progress gets cut short once the deadline passes - not just work that is still waiting in the queue.
This is what lets the system heal from pauses and retry storms without any complicated infrastructure - a cheap “has this request expired?” check is all our actors need.
Conclusion
Not all distributed systems techniques are complicated - this one is simple and effective. If you rely on sender-side retry strategies, your recipients need a way to tell work that still matters apart from work nobody is waiting for anymore, and putting a Deadline on every request is the cheapest way to give them one.
Expiring messages only solve half of the problem we started with, though. They stop the recipient from doing work for callers who have already given up, but they don’t stop it from processing the same live request twice when a retry arrives before the deadline. That’s the job of idempotency, which we’ve written about before in “Why You Should Try to Avoid Exactly Once Message Delivery.”
If you’d like to see a working example of this, you can see my original AkkaStreams.Demo.HttpClient that I created many years ago which uses this technique to demo bearer-token refresh using TTL deadlines as a key ingredient.
And if you want some help applying these techniques to your own Akka.NET applications, that’s what our Akka.NET Support plans and consulting services are for.
-
The alternative to sender-side strategies is consumer-side (pull) strategies, such as having the sender write messages into Kafka or a durable queue and letting the consumer be responsible for guaranteeing processing. ↩
Observe and Monitor Your Akka.NET Applications with Phobos
Phobos automatically instruments your Akka.NET applications with OpenTelemetry — traces, metrics, and logs with built-in dashboards.
Enjoyed this post? Subscribe to our newsletter for more insights on distributed systems, Akka.NET, and .NET + AI.
// COMMENTS