feat(scanner): report and parallelize the reconcile walk

A regular Refresh walked every configured root sequentially, so a full
25 TB drive delayed the roots behind it, and the dialog sat on "Checking
for changes..." at 0 % for the whole walk with no way to tell it was
working. On a cold external drive the walk itself dominates the cost, so
the fix is to overlap the drives and to show what the walk is doing.

Backend (py/services/model_scanner.py):
* Extract the per-root walk into the synchronous _walk_root_for_reconcile()
  worker and merge its results on the event loop afterwards, in configured
  root order: which business path wins a file reachable through several
  roots must not depend on the order the workers happened to finish in.
* Group roots by device (_root_device_key: drive letter on Windows, st_dev
  on POSIX) and run one worker per device. Roots sharing a device stay
  sequential, so directory claims and the overlap dedup (#871, #1041) keep
  their configured-order semantics; different devices run in parallel.
* Track walk progress per root (_ReconcileWalkTracker), weighted by each
  root's cached entry count because the real file count is only known once
  the walk ends. The bar splits into walk (0-50 %) and new-file (50-99 %)
  phases so it never jumps backwards, and the walk broadcasts files seen,
  active roots and the ETA counters.
* Replace the Windows case-insensitive fallback -- a scan of every cached
  path per miss, i.e. O(files x cached) -- with a lazily built lower-cased
  index (_CachedPathLookups, also now guarding the realpath alias index for
  worker threads), and lift the os.name == "nt" gate into the module-level
  _CASE_INSENSITIVE_PATHS so the branch is testable off Windows.
* Excluded-model membership is a set lookup instead of a list scan.

Frontend:
* render the walk phase as "Checking for changes... <roots> (N files)" with
  the ETA, and reset the ETA tracker when the stage changes: the per-file
  rate of counting files says nothing about processing them.
* add common.scanProgress.walkFiles; the other locales keep the sanctioned
  [TODO: Translate] placeholder until the translation pass.

Verified: no-change reconcile over 10k files/500 dirs 101 ms and 10k/5000
dirs 230 ms (was 93/229 ms, within noise); a two-device sandbox walk runs
both roots concurrently, names them in the progress messages and finishes
with added=15, removed=0; 3665 passed, 7 skipped; frontend 1456 passed,
vue widgets 96 passed.
This commit is contained in:
Will Miao
2026-10-06 12:40:57 +08:00
parent 8cd53c20f8
commit 14e660e876
16 changed files with 1059 additions and 187 deletions
@@ -1,7 +1,7 @@
# Reconcile 的 Windows 大小写回退分支 - 待验证清单
# Reconcile 的 Windows 大小写回退分支
> **状态**: 待 Windows 环境验证 | **创建日期**: 2026-09-11
> **相关文件**: `py/services/model_scanner.py` (`ModelScanner._reconcile_cache`)
> **状态**: 已改为 O(1) 索引(2026-10-06),不再需要 Windows 机器验证 | **创建日期**: 2026-09-11
> **相关文件**: `py/services/model_scanner.py`(`_walk_root_for_reconcile` / `_CachedPathLookups`)
> **相关历史**: #871 (`76ee59cd`, 路径重叠去重)、#1108 (按文件夹扫描的需求)
---
@@ -9,84 +9,70 @@
## 背景
Refresh 按钮走的是 `_reconcile_cache()`(快速增量对账)。2026-09-11 做了一轮性能优化,把两处"预防性"的
realpath 全量遍历改成按需触发(详见下方"已完成")。优化后,一次零变更 Refresh 在 5 万文件库上从
~1400 ms 降到 ~120 ms。
清理过程中发现**唯一一处遗留的可疑点**:Windows 专属的大小写不敏感回退分支。它无法在 Linux 上验证,
因此单独记录,留待 Windows 机器上确认。
realpath 全量遍历改成按需触发。清理时留下的唯一可疑点,是 Windows 专属的大小写不敏感回退分支:它排在
精确匹配和 realpath 别名匹配之后,只有**未命中**的文件才会走到,但一旦走到就是 O(文件数 × 缓存条目数)。
---
## 待验证分支(现状)
## 现状(已改造)
`py/services/model_scanner.py` 中 `_reconcile_cache()` 的 walk 循环内:
分支语义保持不变,查找改为 O(1):
```python
# Try case-insensitive match on Windows
if os.name == 'nt':
lower_path = file_path.lower()
matched = False
for cached_path in cached_paths: # 每个未命中文件都全量扫一遍缓存
if cached_path.lower() == lower_path:
found_paths.add(cached_path)
matched = True
break
if matched:
continue
if _CASE_INSENSITIVE_PATHS: # 模块级常量,默认 os.name == "nt"
cached_case_match = lookups.match_casefold_path(file_path)
```
它排在精确匹配(`file_path in cached_paths`)和 realpath 别名匹配之后,只有**未命中**的文件才会走到。
`lookups` 是 `_CachedPathLookups`:小写索引 `{lower(cached_path): cached_path}` 在第一次未命中时构建一次
(与 realpath 别名索引同样懒构建,并用 `threading.Lock` 护住——walk 现在跑在工作线程里),之后每次未命中
只做一次字典查询。
### 为什么可疑
1. **可能不可达**:Windows 上 `os.path.realpath()` 会返回磁盘上的真实大小写,因此"缓存路径大小写与磁盘
不一致"的情形,理论上已经被上一步的 realpath 别名匹配覆盖。若如此,这段就是纯冗余代码。
2. **一旦可达就是 O(N×M)**:每个未命中文件都要遍历全部 `cached_paths` 做小写比较。若某种路径写法让
整个库都变成"未命中"(例如缓存里的盘符/大小写形式与 walk 结果系统性不一致),一次 Refresh 会退化
成 文件数 × 缓存条目数 次字符串比较,比真实 IO 还贵。
3. **没有测试覆盖**:`tests/services/test_model_scanner.py` 没有任何针对该分支的用例(它在 Linux 上
被 `os.name == 'nt'` 短路,无法覆盖)。
`_CASE_INSENSITIVE_PATHS` 提成模块级常量的原因:这条分支在 Linux 上原本被 `os.name == 'nt'` 短路、没有任何
测试覆盖;现在测试可以 monkeypatch 该常量,在 Linux 上真实执行这条分支。
---
## 待办
- [ ] **验证可达性**:在 Windows 上构造"缓存路径与磁盘真实大小写不一致"的场景,确认 realpath 别名匹配
是否已经命中,即上面的 `if os.name == 'nt'` 分支是否还有进入的必要。
- [ ] **若不可达 / 冗余**:删除该分支,并在删除处留注释说明 realpath 已覆盖大小写归一(附验证记录)。
- [ ] **若可达**:保留语义但改成 O(1)——预先构建一次 `lower_path -> cached_path` 映射(与
`cached_real_paths` 同样按需、懒构建),把内层全量扫描换成一次字典查询。
- [ ] **补一个 Windows-only 的回归测试**(`pytest.mark.skipif(os.name != "nt", ...)`),锁定最终结论。
- [ ] 把验证结论回填到本文件,并同步更新状态行。
- [x] **消除 O(N×M)**:改成懒构建的小写索引,查询降为 O(1)。
- [x] **补回归测试**:`tests/services/test_model_scanner.py::test_reconcile_case_fold_fallback_is_indexed_not_linear`
(缓存路径与磁盘仅大小写不同、realpath 别名不命中 → 断言条目保留、不重新处理、索引只构建一次)。
- [x] 回填本文件。
- [ ] (可选,纯代码瘦身)在 Windows 上确认 realpath 别名匹配是否已覆盖全部情形;若确认该分支不可达,
可以整体删掉这一层。它现在的成本已经可以忽略,删除不再是性能问题。
---
## 验证方法(Windows)
## 验证方法(可选,Windows)
1. **构造不一致的大小写**:让缓存里的 `file_path` 与磁盘实际路径大小写不同(例如改过盘符/目录大小写,
或从另一台机器迁移了 `settings.json` 与持久化缓存),然后在 UI 点 Refresh。
2. **看后端日志判据**:
- 若 realpath 已覆盖 → 日志应显示 `Cache reconciliation completed in X seconds. Added 0, removed 0 models.`,
且**没有** `Found N new files to process` / `Processing <path>`。
- 若回退分支在起作用 → 同样应该是 `Added 0, removed 0`(因为 `found_paths` 被补上),这是"分支可达"
的证据;反之若出现大量 `Processing ...` 并重新 hash,说明连回退分支也没命中,问题更严重
(缓存路径被当成了新文件 + 旧条目被删)。
3. **跑测试**:`python -m pytest tests/services/test_model_scanner.py -k reconcile`(该文件在 Windows 上会
真实执行 `os.name == 'nt'` 分支)。
4. **量化**:如果需要,可在 `_reconcile_cache` 里临时插桩统计该分支的进入次数与内层迭代次数,确认是否为 0。
1. 构造"缓存路径与磁盘真实大小写不一致"的场景(改过盘符大小写、迁移过 `settings.json` 与持久化缓存),点 Refresh。
2. 日志判据:`Cache reconciliation completed in X seconds. Added 0, removed 0 models.`,且**没有**
`Found N new files to process` / `Processing <path>`。
3. 跑测试:`python -m pytest tests/services/test_model_scanner.py -k reconcile`。
---
## 已完成(本轮优化,供对照)
## 已完成(历史,供对照)
同一次清理里已经落地并验证的部分(Linux,5 万文件库):
2026-09-11 那轮清理里在 Linux(5 万文件库)验证过的部分:
- `cached_real_paths` 别名映射改为**首次未命中时**懒构建(原来每次 Refresh 都对全部缓存条目算一次 realpath)。
- 每个文件的 `realpath` 移到精确命中检查**之后**(原来对每个文件都算,命中即丢弃)。
- `get_model_roots()` 在新增文件处理阶段只快照一次(原来每个新文件重读一次)。
- 全量去重 pass 加了 O(1) 前置判断(`cached_size_before != len(cached_paths) or total_added > 0`),
零变更且缓存干净时跳过;快照本身含重复路径时仍会自愈。
- 全量去重 pass 加了 O(1) 前置判断(`cached_size_before != len(cached_paths) or total_added > 0`)。
结果:零变更 Refresh 5 万文件 **~1400 ms → ~120 ms**;根目录顺序/符号链接别名翻转场景仍是
`re-processed=0`(不重新读 metadata、不重新 hash)。测试:`tests/services/test_model_scanner.py`
47 项、全量后端 2567 项全部通过。
`re-processed=0`。
---
## 2026-10-06 追加:walk 阶段的并发与进度
同一次改动还做了两件事(与大小写分支无关,但都动到了同一段 walk 循环,故一并记录):
- walk 循环抽成同步函数 `_walk_root_for_reconcile()`,并按设备分组
(`_root_device_key()` / `_group_roots_by_device()`)在工作线程里并行执行:同一设备内的 root 仍按配置顺序
串行(保证目录认领与去重的确定性),不同设备之间才并行。结果回到事件循环后按**配置的 root 顺序**合并,
因此"同一个文件可达多条业务路径时谁胜出"与并发完成顺序无关。
- walk 阶段按 root 广播进度:`_ReconcileWalkTracker` 以「该 root 的缓存条目数」为权重估算进度(真实文件数
只有走完才知道),进度条把 walk 记为 0-50%,新增文件处理阶段顺延为 50-99%,两段之间不会回退。