In this post, I show how my Redis based LockService 2.0 handles Redis restarts and detects deadlocks across multiple application instances.
Intro
When you develop backend applications that deal with data and run as multiple instances, you often need some sort of coordination between those instances. Without it, you get concurrency issues where multiple instances try to access and modify the same data at the same time.
For example: imagine you own a webshop with the following business rule:
It is only possible to order products which can be shipped from the nearest warehouse.
Without proper concurrency control across instances, a clever user might be able to order any product from anywhere in the world using the following carefully timed requests:
This is a typical check-then-act race condition, where the code:
- reads some data (check), then
- evaluates it, and then
- acts upon it (act)
Without concurrency control, the data being read might already be outdated/stale the moment it leaves the database.
To solve this, it is common to use some sort of locking mechanism, which prevents multiple threads from accessing/modifying the same data at the same time. This locking mechanism needs to work across instances. That means when a thread inside instance #3 has acquired a lock, no other thread (locally or on other instances) must be able to acquire the same lock.
What is needed here is a distributed lock. A database row lock (SELECT ... FOR UPDATE) might get you some of the way,
but it quickly gets complicated when your data is spread over multiple tables, and it cannot lock something that
does not exist yet.
Redis based Lock Service 1.0
Redis is an excellent, high-performance, in-memory key/value store, which is quite easy to use from a developer's standpoint. Its performance and simplicity make it a great candidate for implementing distributed locks.
I have worked on Redis based distributed lock implementations in the past; let's call that LockService 1.0.
But it always bugged me that I had no good solution for the case where Redis restarts and potentially loses its data.
Problem 1 - Redis might lose data
Since Redis is primarily an in-memory key/value store, it loses its data when restarted (there is an option to persist data to disk, but it is best-effort without any guarantees). This might lead to the following failure scenario:
While working on the kotlin-redis client library, I realized that it is possible to query Redis' uptime (how long it has been up and running). This led me to the solution for this problem.
Transactions must use a single connection
Assuming there is no TCP proxy between the application and Redis, we can assume that when Redis
restarts, the TCP connection is also reset/lost. So when a transaction runs over a single connection
(which kotlin-redis supports), there cannot be a restart between the WATCH, GET, MULTI and EXEC
commands (see Redis transactions). Thus our
transactions behave correctly ✅: when we reach the EXEC step, our decision to write (act) is guaranteed to be based
on up-to-date data.
Detect Redis restarts between transactions
To solve the problem illustrated in the drawing above, we can use Redis' uptime to detect whether Redis has restarted recently. If so, we simply hold off acquiring the lock(s) in question until Redis has been up for at least the lease time of the lock. This gives the thread owning the lock enough headroom to run its automatic lease renewal logic, which restores the lock entry in Redis and prevents any other thread from acquiring it.
This results in the following updated flow:
As shown above, this solves the restart problem.
Problem 2 - preventing deadlocks
Another issue in my LockService 1.0 implementation from the JavaZone talk was the lack of deadlock detection. Imagine:
- Thread-1 owns lock-A and wants to acquire lock-B
- Thread-2 owns lock-B and wants to acquire lock-A
Or even worse:
- Thread-1 owns lock-A and wants to acquire lock-B
- Thread-2 owns lock-B and wants to acquire lock-C
- Thread-3 owns lock-C and wants to acquire lock-A
We want to detect these kinds of scenarios so that we can break them up. Even though acquiring locks should always be accompanied by a reasonable timeout, a deadlock still blocks the threads involved until that timeout runs out, and probably renders the frontend unresponsive. Deadlocks like these reveal coding errors which should be fixed.
To detect them, I opted for a solution where each thread waiting for locks creates an entry in a Redis hash describing which locks it owns and which locks it is currently waiting for.
When a lock cannot be acquired and deadlock detection is enabled, LockService 2.0
fetches this entire hash (using the HGETALL command) and quickly checks whether there is a cyclic
path. If it finds one, one of the threads in the cycle gives up, so the others can continue.
With this approach, deadlocks can now be detected across JVMs ✅
RedisLockService 2.0
Armed with these new solutions, it was time to work on an improved version of my RedisLockService.
The new version supports:
- Acquiring several locks atomically: either all of them, or none.
- Reentrant locks per thread: a thread that already holds a lock can acquire it again.
- Automatic renewal of the lease of held locks while the protected code runs.
- Detection of deadlocks between threads (across JVMs).
- A single time source for all JVMs: Redis' clock (
TIME) is used, so the JVMs' clocks need not be in sync. - Protection against Redis restarts / data loss.
- Plain Redis 6.2 or newer - no Lua scripts are used.
The code is released under the MIT license as part of the kotlin-redis client library. You can find the documentation here.
Happy locking.