Stalwart 0.16.23: mail delivery stops while messages remain queued

Issue Description

After upgrading from Stalwart 0.16.20 to 0.16.23, mail delivery initially works normally, but after some time delivery appears to stop completely.
The important observation is that this affects both local and remote delivery.

Messages continue to be accepted and placed into the queue, but no corresponding delivery.attempt-start events appear afterwards. The queue grows to approximately 100 messages.

After downgrading back to 0.16.20, the queued messages are processed and delivered normally.

Expected Behavior

Messages placed in the delivery queue should eventually be processed and delivered.

Actual Behavior

After running 0.16.23 for some time:
SMTP submissions are still accepted.
Incoming SMTP messages are still accepted and placed in the queue.
Outgoing messages are still accepted and placed in the queue.
However, delivery attempts stop being started.
Both local and remote delivery appear to be affected.
The queue grows while no corresponding delivery.attempt-start events are generated.
The problem does not appear to be related to a particular remote mail provider.

Relevant Log Output

Immediately after starting 0.16.23, mail delivery works normally.
For example, an outgoing message was successfully delivered to Gmail:

2026-09-22T21:25:38Z INFO Queued message submission for delivery
…
queueId = 330998405339808768
…
2026-09-22T21:25:38Z INFO Delivery attempt started
…
2026-09-22T21:25:39Z INFO Message delivered
…
code = 250
…
2026-09-22T21:25:39Z INFO Delivery completed

Another remote delivery also completed successfully at 21:45 UTC:
2026-09-22T21:45:01Z INFO Delivery attempt started
queueId = 331000843276913152
queueName = “remote”
…
2026-09-22T21:45:06Z INFO Message delivered
hostname = “mx.spamexperts.com”
code = 250
…
2026-09-22T21:45:06Z INFO Delivery completed
2026-09-22T21:45:06Z INFO Delivery attempt ended

Later, however, messages continued to enter the queue without being processed.
For example, an incoming message was accepted at 04:47 UTC:

2026-09-23T04:47:19Z INFO Queued message for delivery
queueId = 331053968713071616
queueName = not yet assigned in this event
from = “<>”
to = [“[email protected]”]
nextRetry = 2026-09-23T04:47:19Z

There was then no delivery attempt for this message while 0.16.23 was running.
At 09:17 UTC I downgraded back to 0.16.20. Immediately afterwards the queued message was processed:

2026-09-23T09:17:14Z INFO Delivery attempt started
queueId = 331053968713071616
queueName = “local”
…
2026-09-23T09:17:14Z INFO Delivery completed
…
code = 250

The actual local delivery itself took only milliseconds:
delivery.attempt-end … elapsed = 35ms

The queued message had therefore remained unprocessed until the downgrade/restart.

Stalwart Version

v0.16.x

Installation Method

Binary (Linux)

Database Backend

RocksDB

Blob Storage

RocksDB

Search Engine

Internal

Directory Backend

Internal

Additional Context

The issue affects both directions:
Incoming → local mailbox

Messages are accepted by the SMTP listener and queued, but are not delivered to the local mailbox while the problem is occurring.

Authenticated submission → remote recipient

Messages submitted by authenticated users are accepted and queued, but are not delivered to remote recipients.

This makes a problem with a specific remote MX, SMTP/TLS, DNS, DANE, or MTA-STS unlikely.

There is also a startup warning:

Configuration build warning
source = “DnsResolver”
reason = “The configured DNS resolver cannot validate DNSSEC.
DANE has been disabled to avoid deferring mail.
Ensure the resolver is DNSSEC-capable and reachable over TCP.”

However, outbound delivery works normally immediately after startup, including successful SMTP/TLS delivery to Gmail and other remote MX servers. Therefore I am not currently assuming that this DNSSEC/DANE warning is the cause.
Recovery
Downgrading from 0.16.23 back to 0.16.20 immediately caused the queued messages to start processing again. The accumulated queue subsequently drained normally.
No configuration changes were made between the two versions.

I have reviewed the documentation and FAQ and confirm that my issue is NOT addressed there.

on

I have searched this support forum (open and closed topics) and confirm this is not a duplicate.

on

I understand that topics in this category are triaged by a bot first but a human reply will follow up. If I’d prefer a human-only reply, I’ll add the no-ai tag to my topic.

on

This isn’t documented in changelog form here. Let me search prior forum topics and known issue trackers regarding this specific bug.

The DNSSEC warning isn’t the cause. It means the validating resolver failed its startup check, so on your server DANE is off and every MX and address lookup goes through the plain resolver; the DNSSEC changes in 0.16.21 to 0.16.23 never come into play. The queue manager is also unchanged between 0.16.20 and 0.16.23, so I don’t have a suspect yet and need to catch it in the act.

Next time it stops, before you restart, grep the log for queue.back-pressure, queue.locked, server.thread-error and any “Failed to read queue.” or “Failed to lock event.” errors. Then take the hour before the stall and list the delivery.attempt-start spans that never got a delivery.attempt-end, with the last event logged in each one; if workers are hanging, that shows where. Also look at the queue in the web UI while it’s stuck: are the nextRetry times in the past, and does it show the queue as paused?

It would also help to know how long it runs before stopping, whether it’s a single node, what threadsPerNode your virtual queues use, and what the system resolver points at (systemd-resolved on 127.0.0.53?).