A correction to my own report, and an updated patch.
*The mechanism I described cannot happen.* I wrote that iter.get_document() throws outside the try block and that the exception then propagates to optimize_box_do()'s generic catch (Xapian::Error &e). That is wrong on both counts.
fts_flatcurve_xapian_optimize_rebuild() is called from *inside* catch (Xapian::InvalidOperationError &e), and the generic catch (Xapian::Error &e) is a sibling handler of the same try block — in C++ an exception thrown inside one handler is never caught by another handler of that same try. And there is no try/catch further up the stack, so an escaping exception would have reached std::terminate, not produced the rc=0 I actually observed.
*What really happens:* the throw comes from replace_document(), which was already inside the try. A Xapian::Document obtained from an MSet is a lazy proxy, so the termlist is read when the document is used, not when it is fetched. The pre-existing catch (Xapian::Error &e) inside the rebuild loop catches it and returns -1, and optimize_box_do() then logs "Optimize failed" and does return 0.
The observable behaviour and the conclusion are unchanged — a failed rebuild is reported to the caller as success, shards are never merged and accumulate on every delivery. Only the explanation of how the exception travels was wrong.
*Consequences for the patch.* Moving get_document() inside the try is inert at runtime, since that call does not throw here. I kept it as a guard in case a future Xapian reads eagerly, but the DocNotFoundError handler is the entire fix.
The updated patch also differs from the first one in two ways that came out of review:
- the per-document warning is capped and an aggregate count is emitted at the end — a badly damaged index can hold thousands of unreadable documents, and one log line each would bury the output exactly when it is being read;
- updates is reset before continue, because commit() is inside the same try: a throw can arrive with documents counted but not committed, and without the reset a persistently failing commit would leave every later document uncommitted.
- The attached patch is against current main .
Reproduction, verification numbers and the note about return 0 in optimize_box_do() all stand as reported.
Updated patch attached.
On Tue, Sep 22, 2026 at 2:07 PM Ihor Rusyn <0k.a.b.a0@gmail.com> wrote:
Hello,
Summary
fts_flatcurve_xapian_optimize_rebuild()aborts on the first document whose termlist is missing, which makes shard compaction impossible for the affected folder from that point on. The failure is silent — the operation reports success — so shards accumulate without bound. We have seen folders reach several thousand shards, with an index orders of magnitude larger than the messages it covers.Present in 2.4.4, 2.4.5 and current main (checked at 732917749a).
Symptom
doveadm fts optimize -u <user> Error: fts-flatcurve: Optimize failed: DocNotFoundError: No termlist for document 257091
exit code 0, shard count unchanged
In production this is logged by indexer-worker and repeats every 20-30 minutes, once per delivery, each run leaving another shard behind. Affected folders also keep an
optimizedirectory containing an empty Xapian database.Mechanism
- Long-lived, high-churn folders end up with shards holding disjoint UID ranges.
db->compact()withDBCOMPACT_NO_RENUMBERthrowsInvalidOperationError, which the code expects and handles.- It falls back to
fts_flatcurve_xapian_optimize_rebuild(), which walks the MSet and copies each document into a fresh database.iter.get_document()sits **outside** the try block. When a document in the MSet has no termlist it throwsDocNotFoundError, which the local handler never sees.- The exception propagates to
fts_flatcurve_xapian_optimize_box_do(), is caught bycatch (Xapian::Error &e), setsfailed = TRUE— and that path doesreturn 0.The caller therefore sees success. Old shards are never deleted, the merge never happens, and every subsequent message adds another shard. The folder can never recover on its own.
Reproduction
Reproducible from scratch in about five minutes on a scratch mailbox, with generated messages only. Three conditions have to coincide; raising uidnext alone does not reproduce it.
U=test@example.com
1. push uidnext past 2^31
doveadm mailbox update -u $U --min-next-uid 2277107224 <(227)%20710-7224> INBOX
2. + 3. churn: grow, index into many shards, expunge half.
the expunged messages leave their documents behind in older shards.
for round in $(seq 1 10); do n=$(doveadm -f flow mailbox status -u $U messages INBOX | tail -1 | sed 's/.*messages=//') [ "$n" -gt 400 ] && n=400 doveadm copy -u $U INBOX mailbox INBOX 1:$n doveadm -o fts_flatcurve_rotate_count=25 index -u $U INBOX doveadm expunge -u $U mailbox INBOX 1:$((n/2)) done
doveadm fts optimize -u $U
Rounds 1-9 compact cleanly to 2 shards. The tenth breaks: 1,838 live messages, 84 shards, uidnext 2277110749 <(227)%20711-0749> — and optimize then fails as above and never merges again.
Suggested fix
Move
get_document()inside the try, and skip an unreadable document instead of aborting the rebuild. A document whose termlist is gone belongs to a message that no longer exists, so there is nothing to carry into the optimized database; the rebuild exists precisely to salvage databases the native compact refused.Patch attached (
fts-flatcurve-optimize-rebuild.diff).Other
Xapian::Errortypes keep aborting the rebuild, so genuine corruption is still not silently swallowed.Separately, the
return 0onfailedinfts_flatcurve_xapian_optimize_box_do()looks questionable regardless of this fix: a failed optimize is reported to the caller as success, which is why this went unnoticed for so long.Notes
Affected folders are long-lived and high-churn: most of the UIDs they have issued belong to messages that were deleted years ago. Folders that compact normally stay at 9-11 shards, so the ones in this state stand out immediately by shard count alone — which is also the cheapest way to find them:
find <index root> -type d -path '*/fts-flatcurve/*' -prune -printf '%h\n'
| sort | uniq -c | sort -rn-- Best regards, Ihor Rusyn
-- Best regards, Ihor Rusyn
A correction to my own report, and an updated patch.
The mechanism I described cannot happen. I wrote that iter.get_document() throws outside the try block and that the exception then propagates to optimize_box_do()'s generic catch (Xapian::Error &e). That is wrong on both counts.
fts_flatcurve_xapian_optimize_rebuild() is called from inside catch (Xapian::InvalidOperationError &e), and the generic catch (Xapian::Error &e) is a sibling handler of the same try block -- in C++ an exception thrown inside one handler is never caught by another handler of that same try. And there is no try/catch further up the stack, so an escaping exception would have reached std::terminate, not produced the rc=0 I actually observed.
What really happens: the throw comes from replace_document(), which was already inside the try. A Xapian::Document obtained from an MSet is a lazy proxy, so the termlist is read when the document is used, not when it is fetched. The pre-existing catch (Xapian::Error &e) inside the rebuild loop catches it and returns -1, and optimize_box_do() then logs "Optimize failed" and does return 0.
The observable behaviour and the conclusion are unchanged -- a failed rebuild is reported to the caller as success, shards are never merged and accumulate on every delivery. Only the explanation of how the exception travels was wrong.
Consequences for the patch. Moving get_document() inside the try is inert at runtime, since that call does not throw here. I kept it as a guard in case a future Xapian reads eagerly, but the DocNotFoundError handler is the entire fix.
The updated patch also differs from the first one in two ways that came out of review:
o the per-document warning is capped and an aggregate count is emitted
at the end -- a badly damaged index can hold thousands of unreadable
documents, and one log line each would bury the output exactly when it
is being read;
o updates is reset before continue, because commit() is inside the same
try: a throw can arrive with documents counted but not committed, and
without the reset a persistently failing commit would leave every
later document uncommitted.
o The attached patch is against current main .
Reproduction, verification numbers and the note about return 0 in optimize_box_do() all stand as reported.
Updated patch attached.
On Tue, Sep 22, 2026 at 2:07PM Ihor Rusyn <[1]0k.a.b.a0@gmail.com> wrote:
Hello,
## Summary
`fts_flatcurve_xapian_optimize_rebuild()` aborts on the first document
whose
termlist is missing, which makes shard compaction impossible for the
affected
folder from that point on. The failure is silent -- the operation
reports success
-- so shards accumulate without bound. We have seen folders reach
several
thousand shards, with an index orders of magnitude larger than the
messages it
covers.
Present in 2.4.4, 2.4.5 and current main (checked at 732917749a).
## Symptom
doveadm fts optimize -u <user>
Error: fts-flatcurve: Optimize failed: DocNotFoundError: No termlist for
document 257091
# exit code 0, shard count unchanged
In production this is logged by indexer-worker and repeats every 20-30
minutes,
once per delivery, each run leaving another shard behind. Affected
folders also
keep an `optimize` directory containing an empty Xapian database.
## Mechanism
1. Long-lived, high-churn folders end up with shards holding disjoint
UID
ranges. `db->compact()` with `DBCOMPACT_NO_RENUMBER` throws
`InvalidOperationError`, which the code expects and handles.
2. It falls back to `fts_flatcurve_xapian_optimize_rebuild()`, which
walks the
MSet and copies each document into a fresh database.
3. `iter.get_document()` sits **outside** the try block. When a document
in the
MSet has no termlist it throws `DocNotFoundError`, which the local
handler
never sees.
4. The exception propagates to `fts_flatcurve_xapian_optimize_box_do()`,
is
caught by `catch (Xapian::Error &e)`, sets `failed = TRUE` -- and that
path
does `return 0`.
The caller therefore sees success. Old shards are never deleted, the
merge never
happens, and every subsequent message adds another shard. The folder can
never
recover on its own.
## Reproduction
Reproducible from scratch in about five minutes on a scratch mailbox,
with
generated messages only. Three conditions have to coincide; raising
uidnext
alone does not reproduce it.
U=[2]test@example.com
# 1. push uidnext past 2^31
doveadm mailbox update -u $U --min-next-uid [3]2277107224 INBOX
# 2. + 3. churn: grow, index into many shards, expunge half.
# the expunged messages leave their documents behind in older shards.
for round in $(seq 1 10); do
n=$(doveadm -f flow mailbox status -u $U messages INBOX | tail -1 | sed
's/.*messages=//')
[ "$n" -gt 400 ] && n=400
doveadm copy -u $U INBOX mailbox INBOX 1:$n
doveadm -o fts_flatcurve_rotate_count=25 index -u $U INBOX
doveadm expunge -u $U mailbox INBOX 1:$((n/2))
done
doveadm fts optimize -u $U
Rounds 1-9 compact cleanly to 2 shards. The tenth breaks: 1,838 live
messages,
84 shards, uidnext [4]2277110749 -- and optimize then fails as above and
never
merges again.
## Suggested fix
Move `get_document()` inside the try, and skip an unreadable document
instead of
aborting the rebuild. A document whose termlist is gone belongs to a
message
that no longer exists, so there is nothing to carry into the optimized
database;
the rebuild exists precisely to salvage databases the native compact
refused.
Patch attached (`fts-flatcurve-optimize-rebuild.diff`).
Other `Xapian::Error` types keep aborting the rebuild, so genuine
corruption is
still not silently swallowed.
Separately, the `return 0` on `failed` in
`fts_flatcurve_xapian_optimize_box_do()`
looks questionable regardless of this fix: a failed optimize is reported
to the
caller as success, which is why this went unnoticed for so long.
## Notes
Affected folders are long-lived and high-churn: most of the UIDs they
have
issued belong to messages that were deleted years ago. Folders that
compact
normally stay at 9-11 shards, so the ones in this state stand out
immediately by
shard count alone -- which is also the cheapest way to find them:
find <index root> -type d -path '*/fts-flatcurve/*' -prune -printf
'%h\n' \
| sort | uniq -c | sort -rn
--
Best regards,
Ihor Rusyn
-- Best regards, Ihor Rusyn
References
Visible links
- mailto:0k.a.b.a0@gmail.com
- mailto:test@example.com
- file:///tmp/tmplhro4bel/tel:(227)%20710-7224
- file:///tmp/tmplhro4bel/tel:(227)%20711-0749