fix(ai): 修复本地检索进度事件导致 ai.sqlite3 无限膨胀 - #162
Merged
2977094657 merged 2 commits intoSep 27, 2026
Merged
Conversation
本地检索每次进度变化都会向 events 表追加一份约 216 KB 的完整任务快照, 且没有 unique_key、没有保留策略,任务运行期间以约 3 行/秒持续写入, 使 output/local_search/ai.sqlite3 增长到 144 GB。 - 进度事件改用 replace=True 原地替换,每个任务只保留最新一行;自增 id 仍 递增,依赖 Last-Event-ID 重连的 SSE 依然能收到最新进度。 - 提醒类通知保留 INSERT OR IGNORE 去重语义,重复 key 不会重置 delivered, 避免已确认的提醒被重新投递。 - 进度事件剔除可从 records 重建的大字段(config/coverage/segments/ read_starts),完整快照仍以 records 表为准;前端改为合并增量字段, 不再整体替换任务对象。 - 新增启动维护:折叠旧版遗留的重复快照、按 24 小时 TTL 回收事件、空闲页 超过 64 MB 时执行 VACUUM,并在后台线程对 summary 与 search 两个 AI 库 执行,避免大型遗留库拖慢启动;初始化不再同步建 events 索引,杜绝阻塞启动。 - 新增 tests/test_ai_storage_retention.py 覆盖原地替换、非 replace 去重、 遗留快照折叠、TTL 与 VACUUM;tests/test_ai_message_pages.py 同步更新断言。
升级前已膨胀的库只靠 TTL 回收会残留很久,直接删除重建又会丢失本地检索配置 (records 中的 active 索引指针),导致需要重新整理/重新向量化。 - 库超过 512 MB 时改为重建:把 records 与未投递提醒复制到新库, 丢弃可再生的 events;耗时与库体积无关。 - 先原子替换主库、成功后再清理旧 WAL/SHM;替换失败自动重试, 仍失败则保留原库并告警,应用照常运行,下次启动再试。 - 保留 sqlite_sequence,重建后事件 id 继续递增,SSE 不断档、前端无感。 - maintain() 在超大库上走重建分支,不再原地删除/VACUUM。 - tests/test_ai_storage_retention.py 增加超大库重建用例。
Member
|
正在查看相关问题 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
问题
启用本地检索并建立索引时,
output/local_search/ai.sqlite3会在任务运行期间持续写入。本机实测约 2.9–3.5 行/秒,48 小时增长到了 144 GB(
events表 699,215 行,占 99.9%)。根因是每次进度变化都向
events表追加一份约 216 KB 的完整任务快照(含 599 个群的
config/coverage/segments),既无unique_key去重,也无 TTL 或条数保留策略;
delivered字段仅用于通知,进度事件永不清理。以相邻两条记录
id=348397 / 348398为例:31 个字段中 28 个完全相同,单条 216 KB 里只有约 2 KB 是新信息,冗余率约 99%。
根因
LocalSearch.update()每次都会调用AIStore.event()无条件INSERT新行。events表只增不减,没有unique_key,也没有任何保留策略。records= 任务配置/历史,events= 进度快照),真实数据在
databases/与local_search/indexes/,因此进度快照完全可以丢弃。改动
8312046a52aecaad5820e1.
8312046阻止无限增长replace=True原地替换,每个任务只保留最新一行;INSERT OR REPLACE会删除旧行并写入新行,自增id仍单调递增,依赖
Last-Event-ID重连的 SSE 依然能收到最新进度。INSERT OR IGNORE去重语义,重复 key 不会重置delivered,避免已确认的提醒被重新投递。
records重建的大字段(
config/coverage/segments/read_starts),完整快照仍以records为准;前端改为合并增量字段,不再整体替换任务对象。空闲页超过 64 MB 时才执行
VACUUM,并在后台线程对 summary 与 search两个 AI 库执行,避免大型遗留库拖慢启动。
events同步建索引,避免在遗留巨型库上阻塞后端启动。2.
a52aeca自动重建已膨胀的旧库records与未投递提醒复制到新库,丢弃可再生的
events;耗时与库体积无关,不会把 WAL 放大到库体积。保留原库并告警,应用照常运行,下次启动再试。
sqlite_sequence,重建后事件 id 继续递增,SSE 不断档、前端无感。影响与兼容性
(
records中的active索引指针),无需重新整理/重新向量化。databases/、local_search/indexes/,events只是可再生产的过程快照。测试
tests/test_ai_storage_retention.py:原地替换、非 replace 去重、遗留快照折叠、TTL、
VACUUM、超大库重建。tests/test_ai_message_pages.py:改为运行中捕获事件,验证“读取中实时进度”仍在,同时断言库内每个任务只剩一行且不含大字段。
records与未投递提醒完好;相关 AI、前端、桌面用例通过。
复现
output/local_search/ai.sqlite3与-wal的合计大小;数据安全
删除该文件不影响任何用户数据。仅当在应用外手动删除
local_search/ai.sqlite3时会丢失本地检索配置、需要重新整理;本 PR 的自动重建会保留配置与未投递提醒。