The Refresh menu can scope a scan to one model root, but the folder a user is
looking at lives in the sidebar's unified tree, which merges every root into one
relative-path namespace. "Scan this folder" therefore addresses the folder, not a
root: the backend walks that relative path under every root that holds it, which
is also what makes the action safe while another drive is switched off.
Backend:
* GET /scan accepts `folder=<rel>` (alone or with `roots=`), rejected together
with full_rebuild=true like the roots parameter. Validation reuses
normalize_relative_folder(), extracted to module scope from ModelMoveService so
the folder operations and the scan endpoint reject the same input (absolute
paths, drive letters, `..` climbing) instead of each carrying its own copy.
* The reconcile summary carries `scope_label` (the folder) for a folder scope, so
the result toast names the folder the user clicked instead of the roots it
happens to live under; the completed WS payload carries it too.
* `folder` is a scope prefix exactly like a root: only that subtree is re-read or
pruned, and an unreachable root keeps the entries that fall inside it.
Frontend:
* The sidebar folder context menu gains "Scan this folder" above "Check for
updates in this folder" (they share the refresh divider); the entry is gated by
the same supportsFolderManagement flag as the other folder operations.
* SidebarManager.scanFolder() resolves the node through the existing
_resolveFolderCandidates() before doing anything: a folder no root holds any
more explains itself ("no longer exists on disk") instead of scanning nothing,
and an unresolvable multi-root node is refused rather than guessed.
* PageControls.refreshModels() and BaseModelApi.refreshModels() forward the folder
scope, and _showRefreshSummary() prefers scope_label over the walked roots.
Verified live on the three-root sandbox with one drive switched off:
GET /scan?folder=pack000 walks drive-G and drive-Y, reports scope_label=pack000,
keeps drive-Z's 6 entries under that folder (kept_unreachable=6) and leaves all
420 models cached. 3704 passed, 7 skipped; frontend 1495 passed (150 files); vue
widgets 96 passed. The 2 new sidebar keys are [TODO: Translate] placeholders
pending the feature owner's go-ahead.
Refreshing had no way to say "scan only this drive": a user with three external
drives had to spin all of them up for every refresh, and switching a drive off
made the next refresh treat its whole library as deleted (rows pruned from the
memory cache and the SQLite cache, preview_url stripped on the next scroll).
Backend (py/services/model_scanner.py, py/config.py):
* ReconcileScope(roots, folder) + _reconcile_cache(scope=...): files inside the
scope reconcile normally, everything outside is neither re-read nor removed.
The folder half is plumbing for the sidebar entry in the next change.
* Path-level pruning guard: cached entries under a path this walk could not
read are kept and reported instead of removed. Sources: a configured root that
is not reachable (drive switched off while LM runs), a directory os.walk
failed to enter (permissions / I/O error / Windows junction to an offline
drive), and a known first-level symlink whose target is gone
(Config.iter_path_mappings()).
* The recorded folder list is unioned instead of replaced whenever the scan did
not verify every root, so a scoped scan cannot empty the sidebar.
* _reconcile_cache returns a summary (added / removed / repaired /
scanned_roots / skipped_roots / unavailable_paths / kept_unreachable),
exposed as ModelScanner.last_reconcile_summary, returned by
BaseModelService.scan_models() and broadcast in the completed WS payload.
* _root_display_labels(): set-aware labels ("G: loras", "usb/loras") grown
leftwards with real parent segments until unique, shared by the walk-progress
line and the roots API.
* GET /scan accepts repeated `roots` (400 for unknown roots, 400 combined with
full_rebuild=true); GET /roots gains root_details (label / reachable / cached
count) while `roots` stays a plain path list for existing callers.
* serve_preview: a 404 no longer clears the cached preview_url when the file's
own directory is unreachable - browsing the grid with a drive off used to
strip preview references from the persistent cache.
Frontend:
* Refresh ▾ gains a "Scan one folder" section listing the page's roots with
their cached counts; offline roots stay clickable and explain themselves; rows
are wired by delegation (new static/js/components/controls/ScanScopeMenu.js).
* A scoped scan reports "Scanned <root>: N new, M removed"; a scan that kept
entries reports "<N> models kept: <paths> not reachable".
* registerAPI() now injects the two cross-page passthroughs (fetchModelRoots and
an argument-forwarding refreshModels) so a page facade cannot drop them: the
first version rendered an empty menu and would have run a full refresh.
* createToastElement whitelists toast types, so a wrong `type` argument degrades
to the info style instead of rendering an unstyled box.
Verified in a sandbox instance with three roots: a scoped scan walks only the
requested root (progress roots=0/1, 240 files); a full refresh with one root
offline reports kept_unreachable=60 and leaves all 420 models cached; /roots
reports the offline root with its cached count. 3699 passed, 7 skipped;
frontend 1488 passed (148 files); vue widgets 96 passed.
The recorded folder list is a union over the model roots keyed by relative
path, but remove_known_folder() dropped an entry unconditionally. Deleting
<rootA>/test in the sidebar therefore hid a "test" that <rootB> still held:
the node disappeared from the next tree load and came back after the next
scan, which reads as "the delete did not work". The same call purged cache
entries by relative folder, so removing an empty <rootA>/test2 also evicted
the model cards of <rootB>/test2 until the next scan, and
rename_known_folder() re-keyed both the recorded folder and the "folder"
field of models that never moved.
* _folders_present_on_disk() answers "which of these relative folders does
some root still hold?" with stats off the event loop (model roots can live
on slow network shares) and reports nothing for stand-in scanners without
roots, which preserves their previous behaviour.
* remove_known_folder(folder, absolute_path=None) keeps the entries another
root still owns and purges by the removed directory's absolute path when
the caller knows it. Without a path the relative-folder filter stays as the
documented fallback: it prunes the entry (disk-verified either way) but
cannot tell same-named copies apart.
* rename_known_folder() re-adds the survivors of the old name and only
touches cache entries whose file_path sits inside the renamed directory.
Tests: the two delete tests now remove the directory from disk first, which
is what the contract always assumed, plus three new scanner tests (a twin
keeps the entry, the entry goes once no root holds it, the twin's cards
survive the purge), one for the legacy fallback and one for a rename that
leaves a twin behind.
A regular Refresh walked every configured root sequentially, so a full
25 TB drive delayed the roots behind it, and the dialog sat on "Checking
for changes..." at 0 % for the whole walk with no way to tell it was
working. On a cold external drive the walk itself dominates the cost, so
the fix is to overlap the drives and to show what the walk is doing.
Backend (py/services/model_scanner.py):
* Extract the per-root walk into the synchronous _walk_root_for_reconcile()
worker and merge its results on the event loop afterwards, in configured
root order: which business path wins a file reachable through several
roots must not depend on the order the workers happened to finish in.
* Group roots by device (_root_device_key: drive letter on Windows, st_dev
on POSIX) and run one worker per device. Roots sharing a device stay
sequential, so directory claims and the overlap dedup (#871, #1041) keep
their configured-order semantics; different devices run in parallel.
* Track walk progress per root (_ReconcileWalkTracker), weighted by each
root's cached entry count because the real file count is only known once
the walk ends. The bar splits into walk (0-50 %) and new-file (50-99 %)
phases so it never jumps backwards, and the walk broadcasts files seen,
active roots and the ETA counters.
* Replace the Windows case-insensitive fallback -- a scan of every cached
path per miss, i.e. O(files x cached) -- with a lazily built lower-cased
index (_CachedPathLookups, also now guarding the realpath alias index for
worker threads), and lift the os.name == "nt" gate into the module-level
_CASE_INSENSITIVE_PATHS so the branch is testable off Windows.
* Excluded-model membership is a set lookup instead of a list scan.
Frontend:
* render the walk phase as "Checking for changes... <roots> (N files)" with
the ETA, and reset the ETA tracker when the stage changes: the per-file
rate of counting files says nothing about processing them.
* add common.scanProgress.walkFiles; the other locales keep the sanctioned
[TODO: Translate] placeholder until the translation pass.
Verified: no-change reconcile over 10k files/500 dirs 101 ms and 10k/5000
dirs 230 ms (was 93/229 ms, within noise); a two-device sandbox walk runs
both roots concurrently, names them in the progress messages and finishes
with added=15, removed=0; 3665 passed, 7 skipped; frontend 1456 passed,
vue widgets 96 passed.
A repository is not a model identity: collection repos on Hugging Face
and ModelScope host many unrelated models, which were wrongly shown as
versions of each other.
- Hugging Face models no longer auto-group (the Hub exposes no
site-native model id)
- ModelScope models group by the site's native published-model id
(MuseInfo modelVersion.modelId), extracted during enrichment and
persisted on the sidecar as source_model_id/source_version_id;
unenriched models stay standalone instead of collapsing a whole repo
into one group
- TensorArt grouping unchanged (its URL id is already model-level)
- Frontend group-key derivation mirrors the new backend semantics
Follows the folder create/delete work: a typo'd directory could be
removed but not corrected, and for a folder holding models the only fix
was to move every model out by hand.
Adds POST /api/lm/{prefix}/rename-folder. Unlike the delete path this one
deliberately works on folders that hold models — a rename keeps every
file, so nothing is cascaded over: the directory is renamed on disk and
the scanner re-keys the records that pointed at the old prefix (recorded
folder list, cache file_path/folder/preview_url, hash and autov3 index
paths, excluded-model paths, and the metadata sidecars that travelled
with the directory). Ancestors are never touched, and only the leaf name
is accepted so a rename can never escape its parent.
Library roots, top-level symlinks and folders holding a staged delete are
refused; the last because a staging manifest records absolute
original/staged paths, so moving it would break undo and purge. A name
collision is a 409 target_exists conflict.
The sidebar reuses the inline-row idiom from folder creation: prefilled
with the current name, inserted in place of the node with that node
hidden while editing, Enter confirms and Escape/blur cancels. The
persisted selection and the expanded set are re-keyed across the rename
so the user keeps their place in the refreshed tree.
Folders created from the sidebar had no in-app way back out: the only
removal path was to leave ComfyUI, delete the directory by hand and
rescan. A typo'd folder also polluted the move/download destination
picker permanently, since it reads the same all_folders source.
Adds POST /api/lm/{prefix}/delete-folder, restricted to directories
whose subtree holds no model weight files — a folder-level cascade would
bypass the per-model lifecycle bookkeeping (metadata sidecars, previews,
cache entries, pending-delete staging, recipe references). The service
walks the directory itself instead of trusting the possibly stale cache,
reports what it would remove (models / files / subfolders / symlinks),
and refuses library roots, top-level symlinks (shutil.rmtree rejects
those) and folders holding a staged delete, whose manifest would be
invalidated by the move. Symbolic links inside the subtree are counted
but never followed.
ModelScanner.remove_known_folder mirrors add_known_folder: the removed
subtree leaves all_folders while ancestors are kept (every recorded
ancestor exists on disk in its own right), stale cache entries under the
prefix are purged and the folder list recomputed. The handler broadcasts
models_changed so destination pickers drop the folder too.
The sidebar entry is a destructive context-menu item. The modal opens in
a confirm state for model-free folders and an explanatory one when the
subtree still holds models, decided from the models-only set that already
dims empty nodes; a stale tree is caught by the 409 not_empty/busy
conflict. Truly empty folders get the existing 20s undo affordance,
implemented by re-creating the directory.
Empty folders (tracked in the scan-recorded all_folders list, same source
the move/download destination picker uses) can now be surfaced in the
folder sidebar via a view-options toggle, dimmed when their subtree holds
no models. Folders can be created directly from the sidebar through a new
POST /api/lm/{prefix}/create-folder endpoint with library-root
containment checks; the scanner records the new directory incrementally
so the tree reflects it without a rescan.
The sidebar header moves its view toggles (tree/list, recursive, empty
folders) into a "..." menu to fit the new create-folder button.
A model file could only ever be linked to huggingface.co: `set_hf_url`
validated the URL with a huggingface-only regex, the agent fetched the card
from a hardcoded HF URL, and the readme processor built every relative image
path off `https://huggingface.co/{repo}/resolve/main`. ModelScope publishes the
same model-card convention (README.md + YAML frontmatter, often carrying
`base_model:` and `trigger_words:`) behind a public, key-less API, so the
enrichment pipeline could already serve it - it was the plumbing that was
HF-shaped, not the idea.
Make the external source a first-class, provider-driven concept:
- New `py/services/model_sources/` registry. A `ModelSource` owns URL
recognition (lenient for stored values, strict for user input), the
canonical page URL, model-card fetching, the asset base URL and the
capability flags. `HuggingFaceSource` is the previous logic relocated;
`ModelScopeSource` reads `/models/{o}/{n}/resolve/{master|main}/README.md`
and falls back to `/api/v1/models/{o}/{n}/repo`. `TensorArtSource` is
link-only on purpose: tensor.art answers plain HTTP clients with a
Cloudflare challenge and its internal API (ap-east-1.tensorart.cloud /
cn.tensorart.net) rejects every /v1/model/* route with "invalid
authorization header", so it declares supports_enrichment=False rather than
failing silently later.
- Metadata gains `source_platform` + `source_url`; `hf_url` stays as a
read/write alias, written only for Hugging Face, so existing sidecars,
cached rows and third-party consumers keep working. Normalisation runs at
the scanner, the persistent cache (both directions, plus two new columns
behind an ALTER migration) and the linking handler - which is what stops a
user who switches sources from leaving a stale `hf_url` on a ModelScope
model.
- The agent pipeline keys off the provider instead of `hf_url`: the fast-fail
gate now explains *why* a model is skipped (no source / unknown source /
source without a reachable card), the prompt context exposes
source_url/source_id/source_label/asset_base_url while still filling the
legacy hf_url/repo aliases, and the four README image extractors take a
base_url (defaulting to HF) so relative paths resolve against the right
site. Version grouping generalises to hf: / ms: / ta: keys.
- `POST /api/lm/set-hf-url` keeps its path and its legacy payload keys but
accepts `source_url`, validates against every provider and returns the
platform. `GET /api/lm/model-sources` lets the UI render the supported-site
list from the server.
- Frontend: a `modelSourceHelpers` mirror of the registry drives the link
dialog, the card/modal globe (branded "View on ModelScope/TensorArt"), the
version-group key and the enrichment gate; the versions tab no longer sends
ms:/ta: keys to the CivitAI API.
TensorArt stays in the list because provenance is worth keeping even when the
card is unreadable - the dialog says so plainly ("Sites that don't expose one
(currently TensorArt) can only be linked") and the context menu disables
enrichment with a matching tooltip, instead of the user getting
"Unsupported URL".
Verified against the real ModelScope API: jj3550945163/Krea-2-LORA returns a
1882-byte card whose frontmatter carries base_model/tags/trigger_words, and
relative images resolve to .../resolve/master/....
Tests: backend 2815 passed; frontend 1130 JS + 91 Vue passed; pytest
tests/i18n and a Jinja compile pass over templates/. The nine locales carry
[TODO: Translate] for the new strings, completed in the next commit.
The include_empty folder tree (download/move modals) walked every model
root synchronously on the event loop via get_all_folders(). On network
(NAS) roots this froze the whole server for the duration of the walk —
blocking WebSocket progress, aria2 RPC and the download queue — and the
5s TTL re-triggered the walk on nearly every modal interaction.
The scanners already visit every directory during cache scans, so record
the full directory list (including empty folders) there instead:
- _gather_model_data/_reconcile_cache collect directories during the
existing walks; reconcile refreshes and persists the list even when no
model files changed.
- ModelCache gains an all_folders field (None = never recorded).
- PersistentModelCache stores the list in a new folders table, with a
cache_meta flag distinguishing 'recorded empty' from legacy snapshots.
- get_all_folders() is now a pure in-memory read. A legacy snapshot
triggers a one-shot backfill walk in a worker thread (never on the
event loop) that records and persists the list.
- Moves add the destination folder (and parents) incrementally instead
of invalidating a TTL cache.
A no-change Refresh still computed os.path.realpath for every model file
in the library and for every cached entry. Both values are only ever
consulted when a discovered file is missing from the cache, so on a
50k-file library they cost ~1.3s and ~0.6s while being used zero times.
- Compute the per-file realpath only after the exact cache match fails
- Build the physical-path alias map lazily on the first miss; the
cross-run alias guard (overlapping roots / symlink layout changes)
still keeps the cached entry instead of a delete + re-add, which would
re-read metadata and re-hash the whole library
- Snapshot get_model_roots() once for the new-file pass instead of
re-reading it for every added file
- Run the duplicate-path integrity pass only when the snapshot already
contained duplicates or files were appended; a clean, unchanged cache
has nothing to clean. Duplicates can only be introduced by external
code rewriting raw_data or by this pass's own appends.
Zero-change reconcile drops from ~1400ms to ~120ms on 50k files, and an
alias flip still re-processes 0 files (#1108 investigation).
Broadcast typed scan_progress messages over /ws/fetch-progress from the
manual refresh/rebuild paths of ModelScanner and RecipeScanner, and
render percent, processed/total, current file name and an EMA-smoothed
ETA in the loading overlay. Hardcoded refresh strings move to i18n
(common.scanProgress); WS connection failure falls back to the previous
static loading behavior.
Bulk delete merged staged batches by physically moving each loser's
files into the winner's batch dir with os.rename. Cross-volume bulks
(winner and loser on different filesystems) always hit EXDEV, forcing a
rollback and degrading to the batch_ids array with per-batch undo.
Merge is now manifest-only: loser entries are appended to the winner's
manifest with their staged paths unchanged, so staged files keep living
in each model's own .lm-pending-delete/<batch_id> dir (no data IO, no
EXDEV). Loser dirs are recorded in the winner manifest's merged_sources
and each loser manifest is stamped merged_into so its own purge timer, a
post-restart sweep or a direct undo call no-op. A cross-volume bulk is
one undoable batch again, and undo/purge clean up the loser dirs once
the merged batch settles.
- Remove delete_undo_enabled setting (backend default, frontend state,
settings modal UI, 10 locales); staged deletes with 30s undo are now
the only delete path and stale settings keys are silently ignored
- Remove the 1500ms delete-button arm delay (armDeleteButton) from all
delete modals; misclicks are recoverable via the undo toast
- Delete modal always shows the recoverable warning
- Log the first staged file path in staging log lines for easier support
Fix ~790 basedpyright errors across the test suite:
- Type stub subclasses of real production classes with super().__init__()
- Add missing generic type arguments and Dict[str, Any] annotations
- Add None guards before subscript/member access
- Adapt tests to production API changes (removed dead handlers,
PersistentModelCache.get_default, _i18n_filter_added location)
- Add PersistentModelCache.update_single_model() for lightweight targeted
SQL update (single row + incremental tag/hash deltas, no full table scan)
- Add ModelScanner.sync_cache_from_metadata() with compare-first logic:
skips entirely when cache is already in sync; when stale, updates the
entry in-place (O(1) instead of O(n) remove+append), incrementally
adjusts tag counts/hash index/version index, and resorts only when
sort-relevant fields changed
- Wire sync_cache_from_metadata() into BaseModelService.get_model_metadata()
via fire-and-forget asyncio.create_task — disk I/O is already paid for
- Include identity re-validation guard against concurrent cache replacement
- Add 16 tests covering _cache_entries_differ, sync_cache_from_metadata
(no-change, in-place, fallback, conditional resort), and
update_single_model (insert, tag delta, hash delta)
- _resolve_commercial_bits() no longer has Sell-implies-Image
cascading; each CommercialUse value sets only its own bit,
matching CivitAI's modern array-format API.
- Keep filter tag label as 'Allow Selling' for brevity; add
title/tooltip 'Allow selling generated images' on hover.
- Same tooltip treatment for 'No Credit Required'.
- Add i18n keys for both tooltips across all 10 locales.
Duplicate filename detection is only relevant for LoRAs, which use
basename-only syntax (<lora:name:strength>). Checkpoints and diffusion
models reference files via relative paths with extensions, so filename
conflicts there are false positives — there is no resolution ambiguity.
Both _log_duplicate_filename_summary() and DoctorHandler's
_check_filename_conflicts() now skip scanners with model_type != 'lora'.
Detects when multiple model files share the same basename (causing
ambiguity in LoRA resolution), logs warnings during scanning, and
provides a "Resolve Conflicts" button in the Doctor panel. Resolution
renames duplicates with hash-prefixed unique filenames, migrates all
sidecar and preview files, and updates the cache and frontend scroller
in-place so the model modal immediately reflects the new filename.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Remove backward compatibility code for `model_type` in `ModelScanner._build_cache_entry()`
- Update `CheckpointScanner` to only handle `sub_type` in `adjust_metadata()` and `adjust_cached_entry()`
- Delete deprecated aliases `resolve_civitai_model_type` and `normalize_civitai_model_type` from `model_query.py`
- Update frontend components (`RecipeModal.js`, `ModelCard.js`, etc.) to use `sub_type` instead of `model_type`
- Update API response format to return only `sub_type`, removing `model_type` from service responses
- Revise technical documentation to mark Phase 5 as completed and remove outdated TODO items
All cleanup tasks for the model type refactoring are now complete, ensuring consistent use of `sub_type` across the codebase.
Add license resolution utilities and integrate license information into model metadata processing. The changes include:
- Add `resolve_license_payload` function to extract license data from Civitai model responses
- Integrate license information into model metadata in CivitaiClient and MetadataSyncService
- Add license flags support in model scanning and caching
- Implement CommercialUseLevel enum for standardized license classification
- Update model scanner to handle unknown fields when extracting metadata values
This ensures proper license attribution and compliance when working with Civitai models.